# Very poor performance with Ollama on DGX Spark – looking for help

**URL:** <https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456>\
**Category:** DGX Spark / GB10 Projects\
**Created:** [December 3, 2025, 5:01pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456 "2025-12-03T17:01:01Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![deeduckme](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@deeduckme](https://forums.developer.nvidia.com/u/deeduckme)\
**Post date:** [December 3, 2025, 5:01pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/1 "2025-12-03T17:01:01Z")

</div>

Hi everyone,

I installed Ollama on my **DGX Spark** to run a **20B ChatGPT OSS model** , and the performance is honestly terrible.

From my Mac, I run a small Python script that reads about ten lines from an Excel file and sends each line to Ollama using:

```auto
http://<DGX_IP>:11434/api/generate

```

Everything runs through a Docker container, with **only one Ollama instance** active.

### What I’m seeing:

- For each generation, GPU usage jumps to **around 89%**.

- Despite this high usage, **latency is very bad**.

- Processing just 10 lines → 10 requests takes far longer than expected.

- The 20B model performs nowhere near what I’d expect on hardware like a DGX Spark.

### My question:

Am I missing something in the Ollama or container configuration?  
Has anyone else experienced similar behavior on DGX systems or other GPU platforms?

Thanks in advance for any insights or feedback.

---

<div class="post-metadata">

**Author:** ![ibrunton\_smith](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@ibrunton\_smith](https://forums.developer.nvidia.com/u/ibrunton_smith)\
**Post date:** [December 3, 2025, 7:41pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/3 "2025-12-03T19:41:49Z")

</div>

No answer from me, but I experienced similarly poor performance using Ollama with oss-got-20b and 120b on the spark. Switching to lm studio was much faster.

I also tried sglang but couldn’t beat the performance from lm studio,

interested to know if there is an obvious explanation.

---

<div class="post-metadata">

**Author:** ![raphael.amorim](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/raphael.amorim/32/446733_2.png) [@raphael.amorim](https://forums.developer.nvidia.com/u/raphael.amorim)\
**Post date:** [December 4, 2025, 4:58am UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/4 "2025-12-04T04:58:22Z")

</div>

@ibrunton_smith@deeduckme Ollama image is slow, please try this: [GDX Spark is extremely slow on a short LLM test - #5 by cosinus](https://forums.developer.nvidia.com/t/gdx-spark-is-extremely-slow-on-a-short-llm-test/350703/5)

---

<div class="post-metadata">

**Author:** ![deeduckme](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@deeduckme](https://forums.developer.nvidia.com/u/deeduckme)\
**Post date:** [December 6, 2025, 10:01pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/6 "2025-12-06T22:01:10Z")

</div>

I posted another thread performance - much better with llama.cpp

> [@Building llama.cpp container images for Spark/GB10](https://forums.developer.nvidia.com/t/building-llama-cpp-container-images-for-spark-gb10/353664/4):
>
> answering to my-self ;) docker run -d –name llama-spark-mistral32 –gpus all -p 3010:8080 -v /home/user/models:/models llama.cpp:server-spark –host 0.0.0.0 –port 8080 -m /models/mistral-small-3.2-24b-ud-q4\_k\_xl.gguf –ctx-size 4096 –threads -1 –n-gpu-layers 16 –flash-attn auto better by limiting the gpu layer to 16… around 27-30% of GPU use ;)

---

<div class="post-metadata">

**Author:** ![eugr](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/eugr/32/449615_2.png) [@eugr](https://forums.developer.nvidia.com/u/eugr)\
**Post date:** [December 11, 2025, 7:40pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/7 "2025-12-11T19:40:05Z")

</div>

There is very little reason to use Ollama these days.

Llama.cpp is faster, has a decent built-in webui, has very active development, and just introduced first-party model switching on demand: [New in llama.cpp: Model Management](https://huggingface.co/blog/ggml-org/model-management-in-llamacpp)

llama-swap still allows more granular control and can control vllm and other inference engines, but it’s great to have this functionality built-in.

---

<div class="post-metadata">

**Author:** ![raphael.amorim](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/raphael.amorim/32/446733_2.png) [@raphael.amorim](https://forums.developer.nvidia.com/u/raphael.amorim)\
**Post date:** [December 12, 2025, 1:18am UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/8 "2025-12-12T01:18:05Z")

</div>

Ollama is a no-go nowadays for me. Completely pointless.

---

<div class="post-metadata">

**Author:** ![vmm1234](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@vmm1234](https://forums.developer.nvidia.com/u/vmm1234)\
**Post date:** [January 18, 2026, 5:48am UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/9 "2026-01-18T05:48:52Z")

</div>

Just to share my experience, Ollama’s docker image have serious performance issues when running on DGX Spark. After re-installing it locally instead of the docker image the performance becomes normal.

So if your Ollama is a docker image I think this is the case. Also somehow even now webUI have weird issues when using images. I’m still looking for solutions :)

---

<div class="post-metadata">

**Author:** ![aweb](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@aweb](https://forums.developer.nvidia.com/u/aweb)\
**Post date:** [January 20, 2026, 10:27pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/10 "2026-01-20T22:27:09Z")

</div>

I’m very curious about this. I’m trying to do OpenAI (or Ollama) remote LLM testing. Currently testing /embeddings and the llama.cpp responses are different from an OpenAI response (JSON formatting). So it’s not really “OpenAI compatible” yet.  
I would try Ollama if it can perform at least close to the llama.cpp server. What’s the best way to install Ollama locally (not docker)?

Thanks.

---

<div class="post-metadata">

**Author:** ![system](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/system/32/68080_2.png) [@system](https://forums.developer.nvidia.com/u/system)\
**Post date:** [February 3, 2026, 10:27pm UTC](https://forums.developer.nvidia.com/t/very-poor-performance-with-ollama-on-dgx-spark-looking-for-help/353456/11 "2026-02-03T22:27:42Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
