Sharing a recipe for MiniMax-M3-MXFP4 across 4× DGX Spark at TP=4 on patched vLLM nightly, using olka-fi/MiniMax-M3-MXFP4 · Hugging Face .
Taking no credit, just having Claude glue things together until something sticks :)
Haven’t seen an MXFP4 version yet on here, so thought I’d try it out based on the RTX PRO 6000 version from 0xSero and the MXFP8 work on 8x Sparks by @ciprianveg
Performance seems compelling!
Setup: 4× Sparks with vLLM nightly pinned at 93d8f834, TP=4, bf16 KV, EAGLE3 k=2, 262k ctx.
Files and build info here: GitHub - BokuNoGF/minimax-m3-mxfp4-4x-gb10: MiniMax-M3 (MXFP4, MSA) serving on 4x GB 10 (DGX Spark), patched vLLM image + launch recipe · GitHub
Throughput (prefill=2048, decode=512, EAGLE k=2, 3 runs averaged via sparkrun benchmark):
| depth | conc | pp tok/s | tg tok/s | ttfr ms |
|---|---|---|---|---|
| 0 | 1 | 2020 | 34.8 | 1015 |
| 0 | 2 | 2047 | 50.3 | 1652 |
| 0 | 5 | 2154 | 70.1 | 3536 |
| 4096 | 1 | 1726 | 34.2 | 1188 |
| 16384 | 1 | 1552 | 33.2 | 1331 |
| 32768 | 1 | 1558 | 30.8 | 1316 |
| 65536 | 1 | 1383 | 28.6 | 1482 |
| 65536 | 5 | 1358 | 49.7 | 5232 |
Single-stream decode holds ~35 → 29 tok/s from 0 → 64k depth; ~70 tok/s aggregate at concurrency 5.
Draft tokens at 2 had the best performance, saw decreases across the board at 1 and 3.
Gotchas:
- No fp8 KV. --kv-cache-dtype fp8 causes a crash on startup. Same issue as the folks running MXFP8 ran into, so reduces the available KV cache to about 800k.
- Pinned vllm nightly image since patches need to be applied directly to it.
- MXFP4 quant that I can’t speak to the quality of, but so far seems to be pretty solid anecdotally
Credits: @ciprianveg for the sm121 image approach + MSA/sparse-attn Triton path, and 0xSero/minimax-m3-sm120 for the MXFP4 clamp fix.