DGX Spark: log-less hard power-off under load — thermal sensor fault, firmware up-to-date, matches RMA'd case 377365

DGX Spark Log-less Hard Power-Off — RMA Evidence Summary

  • Device: NVIDIA DGX Spark (GB10 Grace-Blackwell)
  • Date: 2026-10-07
  • Purpose: Evidence package for NVIDIA support to request RMA

0. Summary

This unit exhibits a thermal sensor mapping/calibration defect that causes the Embedded Controller (EC) to falsely trigger an over-temperature condition and perform a log-less hard power-off. The symptom matches the already-RMA-approved case reported by “digiegg” on the NVIDIA Developer Forum (thread 377365). The unit has been updated to the latest official firmware (EC 3.5.8 / SoC 2.155.11) yet the fault persists. After installing an external 50 W TEC cooling module, a GPU full-load test still trips the 87 °C cutoff, further confirming that the readings are anomalous rather than genuinely thermal.


1. Device Information

Item Value
Model NVIDIA_DGX_Spark / P4242
BIOS (SBIOS) 5.36_0ACUM018 (2025-08-06)
Kernel 6.11.0-1014-nvidia
Driver 580.82.09 (open kernel)
EC firmware 0x03000508 (= latest official 3.5.8)
SoC firmware 0x02009b0b (= latest official 2.155.11)
USB-C PD firmware 0x00000516 (= official 0.5.22)

2. Core Symptom: Log-less Hard Power-off (10 occurrences)

10 occurrences in a single day (2026-10-06), with the interval shrinking from 18 hours down to 47 seconds:

# Power-off time Uptime
1 07:04:11 18h17m
2 07:32:03 19m24s
3 08:01:15 10m27s
4 08:04:31 2m49s
5 08:05:43 47s
6 08:06:59 50s
7 08:54:37 41m17s
8 09:49:29 12m46s
9 10:45:35 43m25s
10 12:10:26 14m57s

Identical signature for every occurrence (journalctl -b -1):

  • Journal truncated mid-line after ordinary application logs
  • No systemd shutdown sequence, no kernel panic/oops, no OOM, no thermal critical, no Xid, no MCE/AER
  • /sys/fs/pstore empty; no vmcore from kdump (crashkernel=1G-:0M renders kdump effectively disabled)
  • last -x marks every session as crash

This is the signature of an EC-level power cut that occurs faster than the kernel can log — identical to threads 373251 / 377365.


3. Sensor Reading Anomalies (core evidence)

3.1 Physically impossible instantaneous jumps

07:27:50 → 07:28:00, acpitz zone0 fell from 84 °C to 67 °C in 10 seconds (a 17 °C drop), which thermal mass does not permit:

代码块

07:27:50  z0=84  z4=84
07:28:00  z0=67  z4=67   ← 10 s, -17 °C
07:28:31  z0=82  z4=82

3.2 Only three discrete values across seven zones + perfect synchronization

Under GPU full load, the seven thermal zones report only three distinct values, with zone0/zone5 perfectly synchronized and zone1–4 perfectly synchronized:

代码块

zone0=87  zone1=66  zone2=66  zone3=67  zone4=66  zone5=87  zone6=71
   ▲───────────────── 87 °C band (zone0/zone5)
              ▲──────────── 66 °C band (zone1–4)

Seven physically distinct locations cannot collapse to three tidy bands — this is the classic signature of a sensor mapping/calibration defect (digiegg’s words: “Two zones trading values … it reads as a sensor handoff or a mapping/calibration problem”).

3.3 87 °C reported on the first sample (physically impossible)

Approximately 2 seconds after GPU load starts (first sample), zone0/zone5 report 87 °C, while the GPU die itself reports only 76 °C via nvidia-smi. Thermal mass does not permit a 38 °C rise in 2 seconds.


4. Firmware Already Latest, Fault Persists

fwupdmgr get-history confirms the EC was upgraded 0x02004e18 → 0x03000508 (latest official 3.5.8) on 2026-10-06, and the SoC is already 0x02009b0b (latest official 2.155.11). The 10 hard power-offs occurred after the firmware was updated, eliminating “stale firmware” as a variable.


5. GPU Thermally Locked Down + Cooling Verification

5.1 The GPU was never actually running at full load

Condition GPU clock GPU power Note
No external cooling 617 MHz 15 W thermally locked, false full-load
50 W TEC cooling added 2171 MHz 83.6 W true full-load (against 2200 cap)

5.2 50 W cooling still trips the 87 °C cutoff (reproducible)

代码块

10:13:55  zone=87 66 66 67 66 87 71   GPU=76°C|83.60W|93%|2171MHz
10:17:20  zone=87 66 66 67 66 87 71   GPU=76°C|82.75W|93%|2171MHz

Two identical runs. Note: the GPU die is 76 °C and zones 1–4 are 66 °C, yet zone0/zone5 report 87 °C — two “hot zones” 11 °C hotter than the GPU die itself, reached instantly. This points to sensor over-reporting rather than genuine cooling insufficiency.


6. Comparison with RMA-approved digiegg case (377365)

Dimension digiegg This unit
SBIOS 5.36_0ACUM018 ✅ Same
EC / SoC 0x03000508 / 0x02009b0b ✅ Same
Log-less hard power-off journal truncated, no vmcore/pstore ✅ Same
Sensor zone handoff/sync Yes ✅ Yes
Error code reading MODS-020000600139 = “temperature limits exceeded or the thermal sensor is broken or miscalibrated” Consistent with this unit’s symptoms
Outcome RMA approved Pending

7. RMA Request

This unit exhibits a reproducible, firmware-independent thermal sensor fault causing log-less hard power-offs. I request authorization to RMA this unit.

Attachments:

  1. journalctl --list-boots output (crash timeline)
  2. sudo journalctl -b -1 > journal_prev_boot.txt
  3. sudo journalctl -k -b -1 > kernel_prev_boot.txt
  4. nvidia-smi -q > nvidia-smi-q.txt
  5. Thermal zone readings: for z in /sys/class/thermal/thermal_zone*; do echo "=== $z ==="; cat $z/type $z/temp; done
  6. gpu-stress-monitor.log (two 87 °C cutoffs + three-band readings)

Serial number available on request.

@moderators Could this be moved to the DGX Spark / GB10 category?