Adding a data point that may be relevant to the acknowledged UMA-OOM freeze bug — a variant without memory pressure.
Same idle-hard-lockup signature described in this thread (SSH dead, no panic, no clean shutdown, hard reset required) — but on my most recent occurrence (2026-08-10, 20:34 CEST), memory usage w
nvidia-bug-report.log.gz (15.6 MB)
as only ~3-4% at the moment of the freeze, not elevated/spiking like the OOM-triggered cases discussed above.
System: ASUS Ascent GX10 (GB10), DGX OS 7.5.0, driver 580.173.02, kernel 6.17.0-1029-nvidia.
Independent vitals logger (fsync’d outside journald, sampling every 3s specifically to catch the moment before these freezes). Last sample before this one:
load=0.11 0.11 0.09 mem=3951/123394MB (96.8% free)
temps=[40-44C] gpu=0%|41C 0 GPU processes, disk I/O flat, dmesg clean
A second, unrelated application on the box independently logged its own unclean-exit within 15 seconds of that sample, also confirming ~122GB free at that instant. journald shows the identical gap — zero panic/OOM/hung_task/Xid signals.
From nvidia-bug-report.sh, run right after the reboot:
- No Xid errors anywhere in the report.
- ECC fully unsupported — ECCSupported: 0, every ECC query (single-bit, double-bit, aggregate) returns N/A. No ECC telemetry path exists at all on this GPU, not just unexposed via nvidia-smi.
- All Clocks Event Reasons “Not Active” — no thermal slowdown, no power braking, no sync boost logged around the incident.
- GPU at P8 (idle), 208 MHz, ~5W draw — matches the vitals logger’s idle reading.
- Only NVRM lines present are boot-time cosmetic warnings (nv_acpi_evaluate_dsm_method failed, RmFetchGspRmImages: no GSP-RM logs) — identical on every boot, unrelated to this freeze.
This is the 4th hard lockup on this box since Aug 6. Three of the four clearly fit the OOM pattern (elevated memory beforehand); this one didn’t — zero memory pressure, zero ECC/Xid/thermal signal anywhere. Wondering if @aniculescu’s bookkeeping-starvation explanation can also trigger below the memory thresholds discussed so far, or if this is a distinct failure mode sharing the same symptom.
Full nvidia-bug-report.log.gz attached.