BTW, I’ve posted in the other thread, but I incorporated @mjpansa’s improvements (the essential one, at least, the other one is properly solved by using Transformers 5 version) into spark-vllm-docker as a mod. Mods can now be used on a single node too, not just in cluster: Make GLM-4.7-Flash go BRRRRR - #3 by eugr
eugr
85
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Help: Running NVFP4 model on 2x DGX Spark with vLLM + Ray (multi-node) | 18 | 2845 | December 25, 2025 | |
| Two-Spark cluster with vLLM using tensor-parallel-size 2 causes one node to drop while the other's GPU goes 100% forever | 36 | 2163 | February 13, 2026 | |
| vLLM on GB10: gpt-oss-120b MXFP4 slower than SGLang/llama.cpp... what’s missing? | 143 | 8182 | February 24, 2026 | |
| New bleeding-edge vLLM Docker Image: avarok/vllm-nvfp4-gb10-sm120 | 32 | 3412 | December 17, 2025 | |
| PSA: State of FP4/NVFP4 Support for DGX Spark in VLLM | 234 | 14176 | May 15, 2026 | |
| We unlocked NVFP4 on the DGX Spark: 20% faster than AWQ! | 144 | 9811 | March 14, 2026 | |
| Make GLM-4.7-Flash go BRRRRR | 18 | 2815 | March 25, 2026 | |
| vLLM containers | 45 | 2746 | July 22, 2026 | |
| Install and Use vLLM for Inference on two Sparks does not work | 159 | 5995 | December 9, 2025 | |
| Gemma 4 Models - which vLLM version? Any PRs spotted? | 177 | 13024 | April 16, 2026 |