Will you be looking into Atlas, pure Rust inference — Intelligence, on your terms ? As its for Spark and RTX specifically and a lot smaller, things like call overhead could be more easily addressed / not become an issue :)
norman.2
396
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Qwen/Qwen3.5-122B-A10B - Alibaba/Qwen thought about us... :-D | 340 | 19573 | March 24, 2026 | |
| Qwen/Qwen3.6-35B-A3B (and FP8) has landed | 309 | 34439 | June 22, 2026 | |
| Qwen3.5-122B-A10B NVFP4 Quantized for DGX Spark — 234GB → 75GB, Runs on 128GB | 44 | 13488 | April 9, 2026 | |
| Does Qwen3.5-35B-A3B on GB10 leave a lot of performance on the table? | 40 | 7391 | March 16, 2026 | |
| Qwen3.5-35B-A3B optimizations on single Spark | 48 | 4320 | May 22, 2026 | |
| What's the best speed we can get with Qwen 3.6 27B without quantizing? | 64 | 27694 | July 6, 2026 | |
| Qwen3.5-122B-A10B on single Spark: 15 → 21.5 tok/s with hybrid GPTQ-INT4 + FP8 dense layers (https://github.com/rmstxrx/vllm-hybrid-quant) | 9 | 1121 | March 20, 2026 | |
| Qwen3.5-397B-A17B run in dual spark! but I have a concern | 236 | 11260 | June 6, 2026 | |
| Qwen3.5-397B-A17B + DGX Spark (duo) | 62 | 7408 | June 14, 2026 | |
| HOW-TO: Run Qwen3-Coder-Next on Spark | 92 | 12186 | March 24, 2026 |