Spark `hangs` - requires a hard-reset (physically unplugging)

We’re running custom PyTorch RL training code on the DGX Spark, and we’ve hit a reliability issue around out-of-memory events. When the model OOMs, the entire system becomes unresponsive—SSH freezes and the box effectively hangs—so recovery requires a hard power cycle (we’ve literally had to unplug the power cable).

On our H100 systems we OOM frequently as well, but it’s typically a process-level failure and the machine stays reachable. On the DGX Spark, because memory is unified, an OOM appears to exhaust system memory and takes down the whole host, leaving us with no clean way to recover remotely.

This is becoming super-annoying, as we run everything remotely…

Some answers in the forum include disabling your swap so the process crashes rather than the system, or to tune your process to match memory constraints (keeping cpu and gpu constraints in mind).

I bought a remote KVM to see if I could alt+ctnl+del when it happens, but I haven’t tried it yet. I selected this one: Comet Pro (GL-RM10) Remote KVM over Wi-Fi — GL.iNet EU

I also researched ‘smart plugs’ for power cycling but didn’t go that route.

Fingers crossed that the next DGXos release gives us some relief.

Sorry to hear you are running into this problem. This is a known issue which we are actively working to fix. The next major Spark OS release should have better handling of OOM and other issues.

Hi @aniculescu, is there a timeline when when a patch/fix/solution might be released. I have two Sparks and this keeps happening to me. Thank you!

The behavior difference versus H100 comes down to memory architecture. On H100, GPU memory is discrete — an OOM is contained to the CUDA context and the host stays up. On GB10, CPU and GPU share the same physical memory pool, so an OOM event can starve the kernel itself before the OOM killer gets a chance to act.

While waiting for the OS-level fix, the earlyoom approach documented here is a practical mitigation: https://forums.developer.nvidia.com/t/mitigating-oom-system-freezes-on-uma-based-single-board-computers/362769

Do we have a timeline for this, since this makes the dgx unusable when we only have the ssh access, even if I can physically reach the device, I have to hard reset when OOM.
For LLama.cpp, it’s fine that if we already know what settings will not cause OOM and quite stable.
But for embedding model like using transformers library, flag embedding( already tried different versions of the library, CUDA and pytorch), I don’t know why, it cause OOM and seems the VRAM leak unusually, for the same script on a X86 machine with CUDA(GPU), it never happen and only use like 20% of VRAM compared to dgx spark also no memory leak, but in dgx spark, the starting VRAM is already much large compared to on a X86 machine, and on X86 it doesn’t leak but on spark it leaks.
So for now, I can only run llama.cpp and some OCR on the spark, was planned to run embedding model on it, but it is not possible and require hard reset when it hangs.

Adding a data point that may be relevant to the acknowledged UMA-OOM freeze bug — a variant without memory pressure.

Same idle-hard-lockup signature described in this thread (SSH dead, no panic, no clean shutdown, hard reset required) — but on my most recent occurrence (2026-08-10, 20:34 CEST), memory usage w

nvidia-bug-report.log.gz (15.6 MB)

as only ~3-4% at the moment of the freeze, not elevated/spiking like the OOM-triggered cases discussed above.

System: ASUS Ascent GX10 (GB10), DGX OS 7.5.0, driver 580.173.02, kernel 6.17.0-1029-nvidia.

Independent vitals logger (fsync’d outside journald, sampling every 3s specifically to catch the moment before these freezes). Last sample before this one:

load=0.11 0.11 0.09 mem=3951/123394MB (96.8% free)

temps=[40-44C] gpu=0%|41C 0 GPU processes, disk I/O flat, dmesg clean

A second, unrelated application on the box independently logged its own unclean-exit within 15 seconds of that sample, also confirming ~122GB free at that instant. journald shows the identical gap — zero panic/OOM/hung_task/Xid signals.

From nvidia-bug-report.sh, run right after the reboot:

- No Xid errors anywhere in the report.

- ECC fully unsupported — ECCSupported: 0, every ECC query (single-bit, double-bit, aggregate) returns N/A. No ECC telemetry path exists at all on this GPU, not just unexposed via nvidia-smi.

- All Clocks Event Reasons “Not Active” — no thermal slowdown, no power braking, no sync boost logged around the incident.

- GPU at P8 (idle), 208 MHz, ~5W draw — matches the vitals logger’s idle reading.

- Only NVRM lines present are boot-time cosmetic warnings (nv_acpi_evaluate_dsm_method failed, RmFetchGspRmImages: no GSP-RM logs) — identical on every boot, unrelated to this freeze.

This is the 4th hard lockup on this box since Aug 6. Three of the four clearly fit the OOM pattern (elevated memory beforehand); this one didn’t — zero memory pressure, zero ECC/Xid/thermal signal anywhere. Wondering if @aniculescu’s bookkeeping-starvation explanation can also trigger below the memory thresholds discussed so far, or if this is a distinct failure mode sharing the same symptom.

Full nvidia-bug-report.log.gz attached.

Driver updates released in May and July have mechanisms to stop unit from freezing due to being OOM by stopping processes that take too much memory.
@luis.poveda9321 Can you tell me what workload you were running when your unit crashed? The nvidia-bug-report only shows dmesg for the current boot. For messages from previous boots, you will need to look at kern.log or run journalctl -k -b -1 -e

kernlog-crashboot-2026-08-10.txt (3.4 KB)

Thanks for looking into this — to answer both of your questions directly:

Workload at crash time: none. The system was fully idle when it froze — ~3.5 GB of 123 GB used (≈97% memory free), no vLLM/inference server, no Docker containers, zero GPU compute processes, load average ~0.1, GPU at P8 / idle clocks / ~4 W, thermals 40–45 °C. Two independent out-of-band logs (a vitals logger fsync’d every 3 s, plus an app-level heartbeat) both confirm that idle state to within ~15 s of the freeze.

So the safeguard you mentioned — stopping processes that consume too much memory — wouldn’t engage here: there was no memory pressure and no runaway process to stop. This is a different failure mode from the OOM freezes elsewhere in this thread. I do also hit the classic OOM variant under heavy load (a 27B model at 256K context genuinely exhausted the unified-memory pool and took the host down — that one fits your description exactly), but the incident in my bug report is the opposite: a silent hard lockup with the box doing essentially nothing.

Worth flagging: this unit is already on 580.173.02 (Linux release date 06/29/2026), i.e. within the May–July window whose OOM-handling improvements you referenced — and the idle freeze still occurred. The improvement doesn’t cover this case because there’s no allocation for it to catch.

Previous-boot logs: checked, and clean. I keep persistent journald plus rotated kern.log, so history survives the hard resets. I’ve attached the journalctl -k -b -1 extract for the crash boot (kernlog-crashboot-2026-08-10.txt). Key points from it:

- Scanning the entire crash boot: zero kernel panic, OOM-killer, hung_task, soft-lockup, RCU-stall, Xid, or NVRM-error entries.

- The kernel log ends at 20:34:14 on a routine NVRM: nvidia_ctl_close (a userspace process closing its GPU fd), with the very last line truncated mid-write — consistent with the machine freezing as it was written.

- No systemd-shutdown / SIGTERM / poweroff / reboot sequence follows; the log simply stops.

- The next entry is a cold boot ~100 s later (Booting Linux on physical CPU 0, 20:35:54), after I applied a manual hard reset.

That’s the signature of a genuine silent hard freeze rather than a software crash or a clean reboot. The nvidia-bug-report.log.gz on my earlier post corroborates the idle state (Clocks Event Reasons all “Not Active”, GPU at P8, ECCSupported: 0 — no ECC telemetry path on this SKU).

System, for reference:

- ASUS Ascent GX10 (NVIDIA GB10 Superchip, 128 GB unified memory)

- DGX OS 7.5.0 · kernel 6.17.0-1029-nvidia (aarch64) · Ubuntu 24.04.4

- Driver 580.173.02 · CUDA 13.0

hung_task_panic / softlockup_panic / panic_on_rcu_stall are now armed and kdump is active, so if the next freeze leaves any kernel code running we should capture a panic + core dump. Glad to re-run nvidia-bug-report.sh right after the next occurrence, or enable any additional telemetry you’d suggest.

Can you run the DGX Spark Fieldiag? Get the Right Support for Your DGX Spark — DGX Spark User Guide
Please make sure to install the CX7 prerequisites beforehand.

Happy to run the Fieldiag. One blocker: this unit is managed fully remotely with no BMC/KVM, so I can’t disable Secure Boot (needed for the diag’s kernel-module tests) without physical access — I’m arranging that now.

In the meantime, is there a Secure-Boot-compatible way to run it, or would the CPU/SSD/memory subset be enough for now?

Fieldiag requires the normal driver to be unloaded and its own drivers to run. This unfortunately needs Secure Boot disabled to run any of the tests

Quick update on the idle (non-OOM) variant. It recurred three more times on Aug 12, all at idle — ~97% memory free, no GPU workload, no panic/OOM/Xid, hard reset required each time.

New finding: my out-of-band vitals logger now captures the CPU idle-state trajectory, and on all three freezes the SoC was descending hard into its deepest idle state (LPI-3) right up to the moment it locked, with PCIe ASPM set to default. This looks consistent with a wake-from-deep-idle lockup (a CPU/PCIe idle power-state transition failing to wake) rather than anything workload- or memory-driven — it only ever happens at idle, never under load.

Is there a known GB10 cpuidle/PSCI/ASPM wake issue matching an idle-only silent lockup with no kernel fault logged? We’ll run the Fieldiag as suggested once we have physical access to disable Secure Boot.

Regards,

Luis Poveda

Update on the “idle hard-lockups” I reported earlier on this box (ASUS Ascent GX10 / GB10, on WiFi): they turned out to be WiFi isolation, not host freezes. A CPU-local vitals logger (fsync every ~3 s, independent of the network) kept writing normally – healthy load, nominal temps – straight through every event, so the host was alive the whole time; it had just dropped off the network over WiFi (MediaTek MT7925, mt7925e), which over SSH looked identical to a lockup.

The mechanism: on an AP-driven band-steer roam (2.4 → 5 GHz) the WPA 4-way handshake times out, wpa_supplicant misreads the EAPOL timeout as a wrong pre-shared key (“4-Way Handshake failed - pre-shared key may be incorrect” → WRONG_KEY), and NetworkManager then parks the interface in “failed (reason ‘no-secrets’)”. On a headless box with no agent to answer the “new key” prompt this is a terminal state – zero further reconnect attempts, the network never comes back on its own, and a full power-cycle is the only recovery. The PSK is correct; it’s a false wrong-key on a roam. This looks fixable in the driver/supplicant stack (treat an EAPOL/handshake timeout on a roam as a retry, not a credential failure, and/or have the mt7925e path not drop the handshake frames during a band switch) – would be great to see it addressed in an upcoming DGX OS / driver update. A verbatim journald excerpt is attached and pasted below. (The one genuine host death I captured was under load, not idle – a deep-context vLLM run where the local logger stopped dead mid-run – which is the UMA-OOM-class issue this thread is really about.)

On our side, the workarounds we’re considering meanwhile: (1) use wired Ethernet instead of WiFi – eliminates the roam entirely and is our preferred fix; (2) if WiFi must stay, pin the client to a single BSSID / disable band-steering so the failing 2.4<->5 GHz roam never happens; (3) configure NetworkManager to auto-reconnect (connection.autoconnect-retries=0 for unlimited retries + a dispatcher/nmcli con up fallback) so the interface recovers itself instead of parking in “failed” and requiring a power-cycle. If it would help others hitting this, I can report back on which of these holds up.

Evidence (verbatim journald excerpt, also attached):

===============================================================================
 WiFi isolation evidence - gx10-8be4 (ASUS Ascent GX10 / GB10)
 2026-08-16 - MediaTek MT7925 (mt7925e), NetworkManager + wpa_supplicant
 Source: systemd-journald. All timestamps CEST. Lines verbatim except the
 hostname 'gx10-8be4' elided for width.
===============================================================================

### 1) BASELINE - a normal band-steer roam that SUCCEEDED (16:23:57-58) ###
  Aug 16 16:23:57 wpa_supplicant[2019]: wlP9s9: WNM: Disassociation Imminent - Disassociation Timer 255
  Aug 16 16:23:57 wpa_supplicant[2019]: wlP9s9: SME: Trying to authenticate with 50:6f:0c:4f:cc:c4 (SSID='Losquefrao' freq=2412 MHz)
  Aug 16 16:23:58 wpa_supplicant[2019]: wlP9s9: Associated with 50:6f:0c:4f:cc:c4
  Aug 16 16:23:58 wpa_supplicant[2019]: wlP9s9: WPA: Key negotiation completed with 50:6f:0c:4f:cc:c4 [PTK=CCMP GTK=CCMP]
  Aug 16 16:23:58 wpa_supplicant[2019]: wlP9s9: CTRL-EVENT-CONNECTED - Connection to 50:6f:0c:4f:cc:c4 completed [id=0 id_str=]
  Aug 16 16:23:58 NetworkManager[2014]: <info>  [1786890238.7271] dhcp4 (wlP9s9): state changed new lease, address=192.168.1.160, acd pending
  Aug 16 16:23:58 NetworkManager[2014]: <info>  [1786890238.7272] dhcp4 (wlP9s9): state changed new lease, address=192.168.1.160

### 2) THE FAILING ROAM - kernel: 2.4GHz(cc:c4) -> 5GHz(cc:c0), assoc OK (16:29:04) ###
  Aug 16 16:29:04 kernel: wlP9s9: disconnect from AP 50:6f:0c:4f:cc:c4 for new auth to 50:6f:0c:4f:cc:c0
  Aug 16 16:29:04 kernel: wlP9s9: authenticate with 50:6f:0c:4f:cc:c0 (local address=50:2e:91:d4:5a:ec)
  Aug 16 16:29:04 kernel: wlP9s9: associate with 50:6f:0c:4f:cc:c0 (try 1/3)
  Aug 16 16:29:04 kernel: wlP9s9: RX ReassocResp from 50:6f:0c:4f:cc:c0 (capab=0x11 status=0 aid=10)
  Aug 16 16:29:04 kernel: wlP9s9: associated

### 3) HANDSHAKE TIMES OUT -> false WRONG_KEY -> NM parks 'failed(no-secrets)' ###
  Aug 16 16:29:14 wpa_supplicant[2019]: wlP9s9: Authentication with 50:6f:0c:4f:cc:c0 timed out.
  Aug 16 16:29:14 kernel: wlP9s9: deauthenticating from 50:6f:0c:4f:cc:c0 by local choice (Reason: 3=DEAUTH_LEAVING)
  Aug 16 16:29:14 wpa_supplicant[2019]: wlP9s9: WPA: 4-Way Handshake failed - pre-shared key may be incorrect
  Aug 16 16:29:14 wpa_supplicant[2019]: wlP9s9: CTRL-EVENT-SSID-TEMP-DISABLED id=0 ssid="Losquefrao" auth_failures=1 duration=10 reason=WRONG_KEY
  Aug 16 16:29:14 NetworkManager[2014]: <info>  [1786890554.8229] device (wlP9s9): Activation: (wifi) disconnected during association, asking for new key
  Aug 16 16:29:14 NetworkManager[2014]: <info>  [1786890554.8230] device (wlP9s9): state change: activated -> need-auth (reason 'supplicant-disconnect', sys-iface-state: 'managed')
  Aug 16 16:31:14 NetworkManager[2014]: <warn>  [1786890674.8761] device (wlP9s9): no secrets: No agents were available for this request.
  Aug 16 16:31:14 NetworkManager[2014]: <info>  [1786890674.8762] device (wlP9s9): state change: need-auth -> failed (reason 'no-secrets', sys-iface-state: 'managed')

### 4) NEVER RECOVERS - zero further reconnect attempts until manual power-cycle ###
  wpa_supplicant 'Trying to associate' attempts, 16:31:15 -> boot end: 0
  (0 = the supplicant/NM made NO further attempt to reconnect for the rest of the boot;
   the interface stayed in 'failed' until the box was power-cycled at ~18:00)

### 5) HOST WAS ALIVE THE WHOLE TIME (network-agnostic local vitals logger) ###
  crash-hunter fsync's every ~3s, independent of the network.
  Samples logged during the outage window 16:29:00 -> 17:58:55: 1473
  first sample after WiFi died:
    2026-08-16 16:29:01 load=0.24 0.23 0.27 mem=81057/123394MB temps=[acpitz=48 acpitz=44 acpitz=44 acpitz=45 acpitz=44 acpitz=48 acpitz=46 ] gpu=0| [N/A]| 45| 11.35 gpu_procs=[380234| 68940;]
  last sample before the manual power-cycle:
    2026-08-16 17:58:55 load=0.52 0.43 0.28 mem=81091/123394MB temps=[acpitz=48 acpitz=44 acpitz=44 acpitz=44 acpitz=44 acpitz=48 acpitz=45 ] gpu=0| [N/A]| 45| 11.27 gpu_procs=[380234| 68940;]
  (healthy idle host throughout: load ~0.2-0.5, temps ~44-48C, GPU idle)

### 6) NO KERNEL FAULT (it was not a freeze) ###
  panic>oom>hung_task|soft lockup|rcu|Xid|NVRM-error entries in crash boot: 0
  (0 = no software/kernel fault of any kind; the host simply lost WiFi and stayed up)

===============================================================================
 SUMMARY: AP band-steer roam -> WPA 4-way handshake timeout -> wpa_supplicant
 misreads it as WRONG_KEY -> NetworkManager parks the interface in
 'failed (reason no-secrets)' -> headless box, no agent to supply a key ->
 ZERO further reconnect attempts -> network never recovers on its own ->
 manual power-cycle required. The host stayed ALIVE (local logger continuous,
 no kernel fault) the entire time - this is network isolation, not a lockup.
===============================================================================

Regards,
Luis Poveda

(attachments)

wifi-isolation-evidence-2026-08-16.txt (4.82 KB)

@luis.poveda9321 This looks to be another appearance of an issue reported here: [solved.. disappointing GB10 wifi card/fw/drivers] WIFI Authentication Loop - Post Setup. We are actively investigating this issue and I will post an update there once I have one. I will mark this thread as closed since no one else has reported an OOM hang after the July Update