Single-Spark setups — which models do you actually run for coding, and how? (sharing mine + a test prompt)

I run a single DGX Spark serving models headless with sparkrun into the OpenCode harness. I’d like to hear how others have set theirs up, especially on one node (so much of the recipe ecosystem assumes 2-Spark tensor-parallel):

  • Which models you’ve found genuinely useful solo, and for what (autonomous coding / agentic / vision / general)?
  • Your serving stack — vLLM vs llama.cpp, quant choices, sampling settings?
  • Gotchas you hit and fixed? Mine was ensuring it reconnects to wi-fi reliably.

To give before I ask, here’s what’s worked for me as an unsupervised coding agent on a single Spark:

  • Qwen3.6-35B-A3B — PrismaQuant 4.75-bit (vLLM) — my daily driver; fastest and most capable small model I’ve found for the Spark.
  • DeepSeek-V4-Flash (Q2 GGUF, llama.cpp) — strong coder; Q2 is the only quant that fits one Spark, but it holds up.
  • Holo-3.1-35B-A3B (NVFP4) — best agentic tool-user of its size (figures out gcc, runs compound shell commands, reads docs); needs scaffolding on hard problems but gets there.
  • Gemma-4-26B-A4B — also robust on my coding tasks.

Sampling tip that mattered: set "repetition_penalty": 1.02. Without it, models loop forever emitting repeating digits (2.6666…, 2.0000…) while debugging a bignum library. 1.02 is gentle enough not to harm normal code tokens but reliably breaks the digit loop.

My test prompt — one concrete bar I use to judge “can it actually code unsupervised?”:

Please write a C99 program calculating 100 digits of Pi (don't hardcode). Use C:\soft\w64devkit to compile it.

I only count it a pass if the model can, on its own: plan a real algorithm (bignum/spigot, not hardcoded digits), work around build/system quirks (find gcc inside w64devkit, set PATH), debug its own code when it’s wrong, and stay persistent without giving up. It’s a surprisingly good filter — plenty of capable-looking models stall on the build environment or loop instead of debugging.

Questions:

  1. On a single Spark, what’s your go-to for autonomous coding — model, quant, engine?
  2. Anyone running larger MoE (Step-3.7-Flash, GLM, MiniMax) solo via GGUF/llama.cpp — worth it over the 35B class?
  3. vLLM vs llama.cpp for solo serving — where’s your line?
  4. Recipe/sampling settings you consider essential?

Happy to share more detail on any of these. Thanks!

I guess it very much depends on your tasks, complexity, codebade type, prompt and info injection quality. Better information and plan - more chances that smaller models can deliver. However, really aggressive quant like q2xxs with DwarfStar really struggle with multistep reasoning and long task execution. Quite okay for oneff well defined task, chat, but hard quantization degrades deep reasoning. If it work for you - then great. Otherwise only one model really stands up for a single spark: qwen 3.5 122b. Not slow (26 t/s and steady pp up until 256k), 3nough knowledge, extremely good at tool calling. Qwen 3.6 27b is better at actually writing the code but lacks knowledge and must have extremely well defined information for execution. Best as subagent, but slow-ish on one spark (20-21 t/s). I did not find other models useful at all for my work, very shallow.

Vllm beats llamacpp every day if you have a proper quant for it, like Nvfp4 done properly or int4 autoround. But if you have a well done quant for llamacpp with mtp then it’s close. Llamacpp simply does not have any of the vllm optimizations community had built into vllm. In theory llamacpp is faster for 1-2 seqs, but lacking sm121 kernels it’s not. But if you built from source (I did) and match with proper quant llamacpp is pretty much on oar for smaller models with more standard arch.

I had bad experience with qwen3.5 - it failed task without being explicitly told to use git and plan, and even with git was slow. So it can’t work unattended. Right now I tested Step-3.7 in IQ3_M, which a goes at 22-23 t/s, but it ended up entering complex loops, which repetition_penalty doesn’t break

122b? Sure you used recommended quant?
Honestly models that can fit on one spark can be great for certain tasks but not for more complex multistep ones. Either model is not super capable being too small or it’s so quantized it looses reasoning ability. I would not trust full automation that can lead to any destruction actions even any model that can run on two, including deepseek V4 flash. Maybe mimo v2. 5 is better, still haven’t made it boot. People do run autonomous agents successfully even using much smaller models but it depends on tasks, instruction quality and a lot depends on harness.

scottgl/MiniMax-M2.7-REAP-172B-A10B-NVFP4-GB10 at about 2000pp and 40tg, you can probably push it harder with b12x

Reap is bad. Period.

can you elaborate please?

You need question answer chat bot it can work. Multistep - not, it’s half brained. I tested minimax m3 reap 50, it’s hilariously bad and universal outcry confirms it. But if you enjoy model talking like a drunk dude then it’s a win

50% reap is a bit too much. the one I shared is about 25% only… but I see your point there…

It’s similar to q2 compression. It can’t hold a line of thought, if you give it 500 tokens it will reply fine. More, or multistep - starts forgetting where it started. I posted an open letter written by minimax m3 reap50 to the m3 thread here on forum. 25 certainly won’t be as bad but not good either.

My own approach is multi agent orchestration instead of chasing bigger and bigger models. One planner and dispatcher, another executor. But it’s hard on one box and no other inference. I use two sparks for main model, big brain ds4f and MacBook to run fast hands 122b and meticulous auditor 27b. But yes it’s essentially 3 box setup on ram and 4 on perf.

For one there are not much options. But you can try Mistral 4 Small 119B A6. 5B. It’s fast, very well rounded but not a great coder if you are after coding mostly.

Or use mixed setup: deepseek V4 pro from deepseek inference - cheap and fast - to plan and verify and 122b or 27b to implement.

so homegrown FUGU. thanks for your answers. i guess i will have to get to 4x GX-10s

That’s exactly what I have done. Even further - I am stopping using opencode myself, instead I talk to main brain agent in hermes and he prompts either opencode or goose via cli to execute a task and validates it. It’s insanely efficient and cool. Main agent runs with 1m context, loads massive wiki on boot, has skills and memories and formulated tasks to executor better than I do. So human is not even orchestrator or architect anymore - switching to visionary and qa role :)

I told my agent yesterday that he’s not only better at coding but also at planning and architecting. Then I watch his thoughts, he talks to himself well at least that human is humble enough to recognize the most obvious thing :D

have you considered HHEM models for grounding-validator? I use it with my self hosted searXNG + firecrawl. searXNG does web_search, firecrawl + HHEM does web_fetch + validation. it validates all the sources, all quotations, everything to make sure that the llm that does the scraping + fetching doesn’t hallucinate anything. I also use executor.sh MCP to do programmatic tool calling and I am thinking to also use the HHEM there to validate the tool calls and the structured JSON the model actually receives. small models like deep seek v4 flash are really decent at tool calling, if you can lower the hallucination rate it actually can become a good model

You must be coming from 8xH200 background if you are calling 300B parameters deepseek v4 flash a small model :D

initially I wrote shi** but it blurred it, so I corrected to small :P

Good for you.

Returning to the OT..

albond/Qwen3.5-122B-A10B. Has been extremely good for me. Mainly running in Agent Zero.

DeepSeek-V4-Flash (Q2 GGUF, llama.cpp) Also very hard worker..

Oddly, I just read about Holo-3.1-35B-A3B today, and it is on my list to try. Would be interested in your findings etc..

Qwen3.6-27B — PrismaAURA 5.5 currently trying to get it up and running. Could be a winner..

Anyway.. hoping for more contributions for this thread.

  • Qwen3.6-35B-A3B-FP8 (vLLM) — daily for hermes, auto tasks, simple coding etc. Fallback to CC/CLI (Opus) on hard coding tasks.

I currently use at home and at work Qwen3.5-122B-A10B-int4-autoround with vLLM (eugr/spark-vllm-docker). Until one week ago also with dflash, but it currently is not working.