I wanted to share my current experience on a single DGX Spark / GB10, because after several tests I think I finally have a setup that is genuinely usable for real work.
I am currently running DeepSeek V4 Flash with ds4 by antirez, using the q2-imatrix quant:
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf
The model is served through ds4-server with CUDA on port 30007. My current launch configuration is:
/home/athena/ds4/ds4-server \
--cuda \
-m /home/athena/ds4/ds4flash.gguf \
-c 131072 \
-n 2200 \
-t 10 \
--host 0.0.0.0 \
--port 30007 \
--kv-disk-dir /home/athena/ds4-kv \
--kv-disk-space-mb 8192
I rebuilt ds4 with:
make cuda-spark
and verified that CUDA is actually being used: ds4-server appears as a compute process on the GB10, with GPU utilization reaching around 94% during generation.
For speed testing I used llama-benchy against the OpenAI-compatible endpoint. With the previous non-imatrix quant I was seeing roughly:
pp2048 → tg128: ~29.6 tok/s
pp4096 → tg128: ~27.6 tok/s
pp7000 → tg128: ~27.5 tok/s
With the new q2-imatrix quant I got:
pp2048 → tg128: 30.60 ± 1.48 tok/s
pp4096 → tg128: 29.83 ± 1.33 tok/s
pp7000 → tg128: 28.40 ± 0.85 tok/s
So at least in my setup the imatrix quant did not slow things down. It actually improved generation speed slightly in this benchmark.
More importantly, the quality seems better in real use. I tested it on a long translation / editing workflow with tool calls, file reads/writes, checklist validation, and a context growing beyond 45K tokens. In that scenario the runtime obviously feels slower than the short benchmark, because the model spends a lot of time in tool/thinking cycles, but it completed the task correctly and produced a noticeably better result than the previous quant.
My impression so far:
-
q2-imatrix is probably the best default choice for a single 128 GB Spark
-
the non-imatrix quant is already very good and feels very fast
-
imatrix feels more robust for longer reasoning/editing/tool workflows
-
full residency is clearly preferable if you want maximum speed
-
SSD streaming is interesting, but I would only use it if I needed to free memory for other workloads
For my use case, this has crossed the line from “interesting experiment” to “actually useful local workstation model.” A single Spark with ds4 and DeepSeek V4 Flash q2-imatrix is now good enough for serious local work, not just demos.
Huge thanks to antirez and everyone contributing to ds4. This project is making the Spark much more useful than I initially expected.