Nemotron3 super 120Gb on ollama v0.30.x-v0.31.2 is broken

Looks like it might be related to this issue posted against llama.cpp - Nemotron-3-Super 120B on GB10 — llama.cpp sm_121 build + Ollama GGUF incompatibility fix . Since Ollama is using llama.cpp since v.0.30, and this has broken my setup.

The fix was to switch back to v0.24.0, the latest working version for Ollama. But I might invesitigate further with llama.cpp, since it should give a perf boost.

Symptom:
Parser aborts the SSE stream mid-response → client sees no finish_reason.

Not: your config, the model, temperature, context size, or multiple users. Confirmed server-side parser regression. Temp 0.3 only reduced the rate; stock model failed identically.

Fix: downgraded spark to Ollama 0.24.0 (last known-good). Verified 20/20 multi-tool requests clean, zero parse errors. v0.31.2-rc1 does NOT fix it.

im trying to avoid Ollama period. I just dont trust it with my business or my DGX spark

Ok. What is the reason for not trusting them? I don’t use their agent, only use it to host the llm locally. For agents I use Pi or Hermes. And Claude Code.
It’s also interesting to me that both ollama and vllm are now a wrapper on top of llama.cpp. I am wondering if it is worth it now to keep that extra layer.

More details.
Tested on DGX Spark (GB10, Ubuntu 24.04.4 LTS, aarch64, kernel 6.17.0-1021-nvidia, 128 GB unified memory, CUDA 13.0, Driver 580.159.03), Ollama 0.24.0, nemotron-3-super-512k (nemotron_h_moe, 123.6B-A12B MoE, Q4_K_M ~87 GB, 524288 context), July 2026.

well im running a legit business and llama holds back my dgx sparks full potential its also vurnurable to malware and bugs