# Running Existing vLLM and SGLang Setups with sparkrun

**URL:** <https://forums.developer.nvidia.com/t/running-existing-vllm-and-sglang-setups-with-sparkrun/381051>\
**Category:** DGX Spark / GB10\
**Tags:** spark\
**Created:** [August 24, 2026, 5:52am UTC](https://forums.developer.nvidia.com/t/running-existing-vllm-and-sglang-setups-with-sparkrun/381051 "2026-08-24T05:52:58Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 25, 2026, 7:17am UTC](https://forums.developer.nvidia.com/t/running-existing-vllm-and-sglang-setups-with-sparkrun/381051/4 "2026-08-25T07:17:16Z")

</div>

Here are our results for that spark-arena recipe:

Conc 1: TTFT=330ms [OK] TPS=18.4 [OK]  
Conc 2: TTFT=520ms [OK] TPS=18.2 [OK]  
Conc 4: TTFT=533ms [OK] TPS=17.4 [OK]  
Conc 8: TTFT=628ms [OK] TPS=15.4 [OK]  
Conc 16: TTFT=808ms [OK] TPS=12.6 [X]

The results are taken with [CordatusAI LLM Benchmark Tool](https://github.com/CordatusAI/llm-benchmark) (128 tokens in, 128 tokens out. Pass criteria at a concurrency: TTFT \< 1000ms AND TPS \>= 15 tok/s)

By the way, we have a [Comprehensive Qwen3.8-27B Study](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102) where all the setups are documented with sparkrun recipes and launch commands. We evaluated various quantizations, speculative decoding methods, scalings, and serving images — results go up to 77 tok/s. There we had similar runs to this recipe. This recipe uses the dgx-vllm-eugr-nightly:latest for the FP8 model at TP=1, whereas we used the vllm/vllm-openai:qwen38 and got almost exactly the same performance. We also have different quantizations and tensor parallelism configurations with the eugr image there.

If you are not happy with the FP8 results, you can try the other configurations from our post, such as the NVFP4 version with SGLang.

---

_[View the full topic](https://forums.developer.nvidia.com/t/running-existing-vllm-and-sglang-setups-with-sparkrun/381051)._
