Not sure if it would be helpful for you, but you could give this a shot for improving performance:
christopher_owen
146
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| vLLM container out of date for new models | 9 | 2004 | October 31, 2025 | |
| vLLM containers | 45 | 2636 | July 22, 2026 | |
| Install and Use vLLM for Inference on two Sparks does not work | 159 | 5878 | December 9, 2025 | |
| vLLM on GB10: gpt-oss-120b MXFP4 slower than SGLang/llama.cpp... what’s missing? | 143 | 8047 | February 24, 2026 | |
| I'd like to learn how to use the latest vLLM on DGX Spark | 9 | 2484 | November 29, 2025 | |
| GLM-4.7-Flash-NVFP4 was just released, but for Transformers 5.0 + vLLM 0.14...? | 89 | 4824 | February 13, 2026 | |
| New pre-built vLLM Docker Images for NVIDIA DGX Spark | 73 | 9692 | March 27, 2026 | |
| VLLM -- the $150M train wreck? | 24 | 1591 | February 27, 2026 | |
| HOW-TO: Run Qwen3-Coder-Next on Spark | 92 | 10952 | March 24, 2026 | |
| Running Mistral Small 4 119B NVFP4 on NVIDIA DGX Spark (GB10) | 66 | 5546 | July 22, 2026 |