MiniMax-M3-MXFP4 on 4× GB10, TP=4, EAGLE3, 262k context, ~35 tok/s

Sharing a recipe for MiniMax-M3-MXFP4 across 4× DGX Spark at TP=4 on patched vLLM nightly, using olka-fi/MiniMax-M3-MXFP4 · Hugging Face .

Taking no credit, just having Claude glue things together until something sticks :)

Haven’t seen an MXFP4 version yet on here, so thought I’d try it out based on the RTX PRO 6000 version from 0xSero and the MXFP8 work on 8x Sparks by @ciprianveg

Performance seems compelling!

Setup: 4× Sparks with vLLM nightly pinned at 93d8f834, TP=4, bf16 KV, EAGLE3 k=2, 262k ctx.

Files and build info here: GitHub - BokuNoGF/minimax-m3-mxfp4-4x-gb10: MiniMax-M3 (MXFP4, MSA) serving on 4x GB 10 (DGX Spark), patched vLLM image + launch recipe · GitHub

Throughput (prefill=2048, decode=512, EAGLE k=2, 3 runs averaged via sparkrun benchmark):

depth conc pp tok/s tg tok/s ttfr ms
0 1 2020 34.8 1015
0 2 2047 50.3 1652
0 5 2154 70.1 3536
4096 1 1726 34.2 1188
16384 1 1552 33.2 1331
32768 1 1558 30.8 1316
65536 1 1383 28.6 1482
65536 5 1358 49.7 5232

Single-stream decode holds ~35 → 29 tok/s from 0 → 64k depth; ~70 tok/s aggregate at concurrency 5.

Draft tokens at 2 had the best performance, saw decreases across the board at 1 and 3.

Gotchas:

  • No fp8 KV. --kv-cache-dtype fp8 causes a crash on startup. Same issue as the folks running MXFP8 ran into, so reduces the available KV cache to about 800k.
  • Pinned vllm nightly image since patches need to be applied directly to it.
  • MXFP4 quant that I can’t speak to the quality of, but so far seems to be pretty solid anecdotally

Credits: @ciprianveg for the sm121 image approach + MSA/sparse-attn Triton path, and 0xSero/minimax-m3-sm120 for the MXFP4 clamp fix.

7 Likes