# Help, team spark!

**URL:** <https://forums.developer.nvidia.com/t/help-team-spark/370983>\
**Category:** DGX Spark / GB10\
**Created:** [May 22, 2026, 1:04am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983 "2026-05-22T01:04:07Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![jc2375](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jc2375/32/522115_2.png) [@jc2375](https://forums.developer.nvidia.com/u/jc2375)\
**Post date:** [May 22, 2026, 1:04am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/1 "2026-05-22T01:04:07Z")

</div>

Dear community of developers,  
I am but a simple doctor who has been enthusiastically using a DGX cluster of 2 to run Qwen 397b as an inference model for my clinical notes. Believe it or not, it’s the first and only local model that I have found that effortlessly generates a medical note from my conversation, a system prompt, and a couple of small search tools. This has been a huge boon to my patient care since I can just have a conversation, human being to human being, and let a robot tale care of 90% of modern medicine’s woes, aka note writing and filling in billing codes.

Now, all was well…until I decided to update the docker repository that @eugr_nv has kindly created. Luck would have it, it has broken the memory allocation for this model and I’m unable to run it.

I am wondering how I can roll back to an earlier version of the repository and vLLM build?

By the way, I know there is awesome software engineering talent here, fantastic CAD wizards….but I’ll nominate myself to be the team doc :D!

Appreciate any help you can give me.

---

<div class="post-metadata">

**Author:** ![j0n](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@j0n](https://forums.developer.nvidia.com/u/j0n)\
**Post date:** [May 22, 2026, 1:10am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/2 "2026-05-22T01:10:57Z")

</div>

@jc2375 Fantastic use of the Sparks. We may be in similar boats. I’ve been unable to run vLLM since I installed NVIDIA Sync updates today. (It hangs both Sparks brutally.) In my case, when I check system logs:

```auto
 journalctl -b -1 -k | grep -iE 'memory'

```

I see this error:

```auto
NVRM: nvCheckOkFailedNoLog: Check failed: Out of memory [NV_ERR_NO_MEMORY] (0x00000051) returned from _memdescAllocInternal(pMemDesc) @ mem_desc.c:1359

```

---

<div class="post-metadata">

**Author:** ![josephbreda](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@josephbreda](https://forums.developer.nvidia.com/u/josephbreda)\
**Post date:** [May 22, 2026, 1:20am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/3 "2026-05-22T01:20:00Z")

</div>

Just check out a prior tag from the GIT repo. Also, try the build-and-copy.sh with a vllm-ref of a few iterations back.

---

<div class="post-metadata">

**Author:** ![jc2375](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jc2375/32/522115_2.png) [@jc2375](https://forums.developer.nvidia.com/u/jc2375)\
**Post date:** [May 22, 2026, 3:02am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/4 "2026-05-22T03:02:00Z")

</div>

I think it worked. Had to go bsck and change to qwen3\_coder in the recipe vs the xml, tool calls are more reliable for me with that template. But it is working again!!

---

<div class="post-metadata">

**Author:** ![eugr\_nv](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/eugr_nv/32/508945_2.png) [@eugr\_nv](https://forums.developer.nvidia.com/u/eugr_nv)\
**Post date:** [May 22, 2026, 5:19am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/5 "2026-05-22T05:19:43Z")

</div>

Yes, there seems to be an issue with the most recent vLLM build and how it calculates available memory. I’ll see if I can find the cause if it’s not fixed in the next nightly build.

---

<div class="post-metadata">

**Author:** ![j0n](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@j0n](https://forums.developer.nvidia.com/u/j0n)\
**Post date:** [May 22, 2026, 7:03am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/6 "2026-05-22T07:03:17Z")

</div>

Following up on potentially the root cause of @jc2375’s issue, regarding the `NV_ERR_NO_MEMORY` kernel log entry: (I apologise for the AI-generated phrasing, but I’ve reviewed it for accuracy):

I’ve observed it on a docker-pinned vLLM build (`0.19.1rc1.dev1315+g102aaddf8.d20260519`, image built three days ago via `eugr/spark-vllm-docker`), so it’s not unique to the most recent vLLM nightly.

Today’s occurrence was non-fatal — single dmesg entry at the exact moment the MTP drafter model finished loading and shared-weight remapping began (`mtp.py:484` + `llm_base_proposer.py:1392/1448`). Run proceeded to serve traffic normally and is still serving. dmesg only shows the bare allocation failure (`_memdescAllocInternal @ mem_desc.c:1359`) — no `os_acquire_rwlock_read` cascade like last night.

Setup: 2x DGX Spark, driver 580.159.03, kernel 6.17.0-1018-nvidia, NCCL 2.30.6, DeepSeek-V4-Flash, TP=2 across nodes, `--gpu-memory-utilization 0.85`, `--kv-cache-dtype fp8`.

So I think there are two distinct phenomena: (1) `_memdescAllocInternal` returns NO\_MEMORY more often than vLLM expects on Spark UMA, which is mostly cosmetic, and (2) sometimes that failure triggers an RM lock cascade (matching open-gpu-kernel-modules#968), which is fatal. Today I hit (1) only; yesterday I hit (1)→(2). Whether vLLM’s memory accounting is the cause or a contributing factor to (1), the underlying recoverability of (1) might also be different across vLLM versions.

Happy to grab any additional logs that would help.

---

<div class="post-metadata">

**Author:** ![system](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/system/32/68080_2.png) [@system](https://forums.developer.nvidia.com/u/system)\
**Post date:** [June 5, 2026, 7:03am UTC](https://forums.developer.nvidia.com/t/help-team-spark/370983/7 "2026-06-05T07:03:45Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
