I kind of did…. check the other thread.
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| vLLM on GB10: gpt-oss-120b MXFP4 slower than SGLang/llama.cpp... what’s missing? | 143 | 8058 | February 24, 2026 | |
| Marlin Fix: NVFP4 Actually Works on SM121 (DGX Spark) | 15 | 3025 | April 12, 2026 | |
| Help: Running NVFP4 model on 2x DGX Spark with vLLM + Ray (multi-node) | 18 | 2802 | December 25, 2025 | |
| PSA: State of FP4/NVFP4 Support for DGX Spark in VLLM | 234 | 13931 | May 15, 2026 | |
| Two-Spark cluster with vLLM using tensor-parallel-size 2 causes one node to drop while the other's GPU goes 100% forever | 36 | 2044 | February 13, 2026 | |
| New bleeding-edge vLLM Docker Image: avarok/vllm-nvfp4-gb10-sm120 | 32 | 3371 | December 17, 2025 | |
| GLM-4.7-Flash-NVFP4 was just released, but for Transformers 5.0 + vLLM 0.14...? | 89 | 4826 | February 13, 2026 | |
| vLLM containers | 45 | 2646 | July 22, 2026 | |
| Llama.cpp experimental native mxfp4 support for blackwell PR | 12 | 1796 | January 7, 2026 | |
| vLLM 0.17.0 MXFP4 Patches for DGX Spark: Qwen3.5-35B-A3B 70 tok/s, gpt-oss-120b 80 tok/s (TP=2) | 32 | 2670 | April 13, 2026 |