# Qwen3.5-122B-A10B on single Spark: up to 51 tok/s (v2.1 — patches + quick-start + benchmark)

**URL:** <https://forums.developer.nvidia.com/t/qwen3-5-122b-a10b-on-single-spark-up-to-51-tok-s-v2-1-patches-quick-start-benchmark/365639>\
**Category:** DGX Spark / GB10\
**Tags:** performance-tuning, llm, performance, docker, cuda\
**Created:** [April 5, 2026, 8:02am UTC](https://forums.developer.nvidia.com/t/qwen3-5-122b-a10b-on-single-spark-up-to-51-tok-s-v2-1-patches-quick-start-benchmark/365639 "2026-04-05T08:02:00Z")\
**Posts on this page:** 1\
**Showing post:** 396

<div class="post-metadata">

**Author:** ![norman.2](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/norman.2/32/468361_2.png) [@norman.2](https://forums.developer.nvidia.com/u/norman.2)\
**Post date:** [May 9, 2026, 9:35am UTC](https://forums.developer.nvidia.com/t/qwen3-5-122b-a10b-on-single-spark-up-to-51-tok-s-v2-1-patches-quick-start-benchmark/365639/396 "2026-05-09T09:35:02Z")

</div>

Will you be looking into [Atlas, pure Rust inference — Intelligence, on your terms](https://atlasinference.io/#models) ? As its for Spark and RTX specifically and a lot smaller, things like call overhead could be more easily addressed / not become an issue :)

> [@Atlas: Open-source inference engine for DGX Spark \<2minute cold start, 100+ tok/s on Qwen3.6-35B-FP8, 13+ supported models](https://forums.developer.nvidia.com/t/atlas-open-source-inference-engine-for-dgx-spark-2minute-cold-start-100-tok-s-on-qwen3-6-35b-fp8-13-supported-models/369263):
>
> We’ve just open-sourced Atlas, an LLM inference engine purpose-built (but not limited to) GB10-class hardware, Repository is here: [https://github.com/Avarok-Cybersecurity/atlas](https://github.com/Avarok-Cybersecurity/atlas) and we need the communities help to keep elevating it for developers. Many DGX Spark owners and Blackwell developers read this forum, we wanted to make sure this landed in front of you directly rather than relying on it filtering over from elsewhere. Earlier we shared some informal benchmarks showing Qwen3.5-35B run…

---

_[View the full topic](https://forums.developer.nvidia.com/t/qwen3-5-122b-a10b-on-single-spark-up-to-51-tok-s-v2-1-patches-quick-start-benchmark/365639)._
