Sometimes I get ~70 tok/sec other times I get ~3 tok/sec, is there any way to make this more stable?
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| MiniMax-M3 (428B MoE + vision) at ~14–15 tok/s on 2× DGX Spark — EAGLE3 speculative decoding is the unlock | 2 | 786 | July 4, 2026 | |
| Minimax3 on 2 nodes decode ~10.7 tok/s, 4bits | 26 | 1617 | June 21, 2026 | |
| MiniMax M3 : NVFP4 for Quad DGX Spark | 116 | 8514 | June 25, 2026 | |
| MiniMax M3 NVFP4 and NVFP4 REAP 50 for 4x & 2x DGX Sparks | 53 | 4297 | July 2, 2026 | |
| Minimax-m3 internal server error | 8 | 725 | June 24, 2026 | |
| MiniMax-M3-AWQ on 4× GB10, fp8 KV, 262k context, adaptive reasoning, ~30 tok/s | 17 | 1264 | July 21, 2026 | |
| Getting "Too Many Requests" always in MiniMax M3 | 3 | 273 | August 10, 2026 | |
| Many models are very very slow to start answer or thinking | 0 | 612 | April 26, 2026 | |
| Minimax-m2.7 | 2 | 174 | July 28, 2026 | |
| MiniMax M2.5 released (not available on HuggingFace as of now) -- is DGX Spark ready? | 92 | 6921 | April 12, 2026 |