TLDR - Quantized and calibrated MiniMax-M3-GPTQ checkpoint paired with 2xGB10 deployment package. Optimized fp8, nvfp4, KVarN, and EAGLE-3 verify kernels integrated with b12x and vllm. Vision is not tested or validated but should work in theory.
Container, build steps, recipes, and checkpoints:
This package is not stable. All three recipes push the absolute limit of the available system memory. I run my nodes headless and clear the page caches before each cluster launch. Expect, errors, OOMs, etc before you find the sweet spot for your configuration. Bring your agent.
Credit to @PILCOTHINK for both introducing KVarN as a potential optimization and validating the implementation directly: Serving Qwen3.5-397B-A17B at 1M Tokens on 2× DGX Spark — MiniMax M3 Is Next
His calculations for KV-cache size potential were spot on.
fp8 w/EAGLE-3 - 131k ctx
nvfp4 w/EAGLE-3 - 196k ctx
KVaRN - 262k ctx, up to 370k ctx observed
I spent a lot of time producing the quantization checkpoint. The calibration dataset included simulated agentic trajectories using the MiniMax-M3 api rendered through pi and OpenCode for SWE and terminal tasks. Math, retail, and general chat were also included domains. With that said, I cannot make any concrete claims to actual quality of the quant - what I do know is that I did my due diligence, for what its worth, to do the GPTQ properly without cutting corners. How well this worked, or whether it worked at all, I don’t know. On tool-eval-bench hardmode in a best of 5 with fp8 KV-cache, the quant scored on par with the MiniMax-M3 api. Admittedly, both the quant and the api had mediocre scores, averaging around 77/100.
I recommend using adaptive thinking (included in the recipe kwargs, along with important reasoning parser/chat template fixes in the mods). Throughput benchmarks are below for nvfp4-EAGLE3 and KVaRN. fp8-EAGLE-3 has comparable throughput to nvfp4-EAGLE-3. Prompt processing is low due to lowering max_num_batched_tokens to 1024 to increase KV-cache size. But more aggressive batching is theoretically possible if that is priority for your workload. Adding the EAGLE-3 drafter to the KVarN configuration hurts available KV-cache more than the other KV cache quants. So the recipe leaves it out, but EAGLE-3 is still compatible with it. Something else worth noting is that the EAGLE-3 drafter results in diminishing returns at longer context lengths. My hypothesis for this is that the target verify step is more costly, because the target model needs to produce top-k block scores for a longer context length for each drafted token, hence hurting throughput at longer context lengths.
nvfp4-EAGLE-3
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) |
|:------------------------------|----------------:|----------------:|-------------:|-----------------:|-----------------:|-----------------:|
| Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 | 1376.00 ± 12.12 | | 2080.75 ± 13.04 | 1488.49 ± 13.04 | 2145.45 ± 14.66 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 | 35.49 ± 4.95 | 36.79 ± 4.96 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d4096 | 1182.70 ± 2.50 | | 4056.37 ± 7.35 | 3464.12 ± 7.35 | 4119.00 ± 3.24 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d4096 | 25.85 ± 2.23 | 27.33 ± 2.36 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d4096 | 964.46 ± 13.72 | | 2716.15 ± 29.93 | 2123.89 ± 29.93 | 2781.90 ± 30.29 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d4096 | 31.87 ± 1.12 | 32.93 ± 1.16 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d8192 | 1076.15 ± 3.53 | | 8205.60 ± 24.95 | 7613.34 ± 24.95 | 8262.40 ± 19.94 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d8192 | 27.01 ± 1.61 | 30.49 ± 2.26 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d8192 | 952.11 ± 17.91 | | 2744.04 ± 40.55 | 2151.78 ± 40.55 | 2812.80 ± 41.01 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d8192 | 30.82 ± 3.51 | 31.59 ± 3.76 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d16384 | 1016.11 ± 0.79 | | 16717.49 ± 12.49 | 16125.23 ± 12.49 | 16772.56 ± 12.27 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d16384 | 29.25 ± 1.88 | 30.72 ± 4.18 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d16384 | 886.72 ± 7.38 | | 2902.05 ± 19.33 | 2309.79 ± 19.33 | 2959.50 ± 16.64 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d16384 | 32.84 ± 2.91 | 33.70 ± 3.33 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d32768 | 959.78 ± 2.51 | | 34734.21 ± 89.70 | 34141.95 ± 89.70 | 34791.61 ± 88.23 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d32768 | 27.87 ± 1.92 | 30.61 ± 3.01 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d32768 | 782.12 ± 3.49 | | 3210.84 ± 11.66 | 2618.58 ± 11.66 | 3270.24 ± 18.36 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d32768 | 24.35 ± 2.42 | 26.67 ± 2.62 | | | |
KVarN
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) |
|:------------------------------|----------------:|----------------:|-------------:|------------------:|------------------:|------------------:| | Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 | 1767.04 ± 48.35 | | 1439.90 ± 32.36 | 1159.88 ± 32.36 | 1483.13 ± 13.15 | | Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 | 22.65 ± 0.01 | 23.00 ± 0.00 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d4096 | 1332.75 ± 1.34 | | 3353.85 ± 2.86 | 3073.84 ± 2.86 | 3419.90 ± 4.94 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d4096 | 22.60 ± 0.02 | 23.00 ± 0.00 | | | | | Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d4096 | 937.57 ± 8.59 | | 2464.57 ± 20.15 | 2184.55 ± 20.15 | 2510.73 ± 7.26 | | Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d4096 | 22.39 ± 0.19 | 23.00 ± 0.00 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d8192 | 1147.31 ± 7.91 | | 7421.13 ± 48.78 | 7141.11 ± 48.78 | 7488.76 ± 52.95 | | Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d8192 | 22.30 ± 0.06 | 23.00 ± 0.00 | | | | | Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d8192 | 913.49 ± 9.47 | | 2522.22 ± 23.41 | 2242.20 ± 23.41 | 2568.48 ± 31.48 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d8192 | 22.24 ± 0.24 | 23.00 ± 0.00 | | | | | Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d16384 | 1053.77 ± 6.43 | | 15829.54 ± 95.34 | 15549.53 ± 95.34 | 15896.51 ± 94.64 | | Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d16384 | 22.15 ± 0.01 | 23.00 ± 0.00 | | | | | Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d16384 | 890.17 ± 4.65 | | 2580.77 ± 12.02 | 2300.75 ± 12.02 | 2642.97 ± 13.78 |
| Sebesky/MiniMax-M3-W4A16-GPTQ | tg32 @ d16384 | 22.21 ± 0.01 | 23.00 ± 0.00 | | | | | Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_pp @ d32768 | 1001.66 ± 4.03 | | 32995.13 ± 131.94 | 32715.11 ± 131.94 | 33045.46 ± 114.71 | | Sebesky/MiniMax-M3-W4A16-GPTQ | ctx_tg @ d32768 | 21.78 ± 0.03 | 22.00 ± 0.00 | | | |
| Sebesky/MiniMax-M3-W4A16-GPTQ | pp2048 @ d32768 | 832.69 ± 3.10 | | 2739.56 ± 9.15 | 2459.54 ± 9.15 | 2786.99 ± 3.53 |
Hoping this serves as a baseline for more experimentation, or even real workloads on the 2xGB10.