DGX Spark Log-less Hard Power-Off — RMA Evidence Summary
- Device: NVIDIA DGX Spark (GB10 Grace-Blackwell)
- Date: 2026-10-07
- Purpose: Evidence package for NVIDIA support to request RMA
0. Summary
This unit exhibits a thermal sensor mapping/calibration defect that causes the Embedded Controller (EC) to falsely trigger an over-temperature condition and perform a log-less hard power-off. The symptom matches the already-RMA-approved case reported by “digiegg” on the NVIDIA Developer Forum (thread 377365). The unit has been updated to the latest official firmware (EC 3.5.8 / SoC 2.155.11) yet the fault persists. After installing an external 50 W TEC cooling module, a GPU full-load test still trips the 87 °C cutoff, further confirming that the readings are anomalous rather than genuinely thermal.
1. Device Information
| Item | Value |
|---|---|
| Model | NVIDIA_DGX_Spark / P4242 |
| BIOS (SBIOS) | 5.36_0ACUM018 (2025-08-06) |
| Kernel | 6.11.0-1014-nvidia |
| Driver | 580.82.09 (open kernel) |
| EC firmware | 0x03000508 (= latest official 3.5.8) |
| SoC firmware | 0x02009b0b (= latest official 2.155.11) |
| USB-C PD firmware | 0x00000516 (= official 0.5.22) |
2. Core Symptom: Log-less Hard Power-off (10 occurrences)
10 occurrences in a single day (2026-10-06), with the interval shrinking from 18 hours down to 47 seconds:
| # | Power-off time | Uptime |
|---|---|---|
| 1 | 07:04:11 | 18h17m |
| 2 | 07:32:03 | 19m24s |
| 3 | 08:01:15 | 10m27s |
| 4 | 08:04:31 | 2m49s |
| 5 | 08:05:43 | 47s |
| 6 | 08:06:59 | 50s |
| 7 | 08:54:37 | 41m17s |
| 8 | 09:49:29 | 12m46s |
| 9 | 10:45:35 | 43m25s |
| 10 | 12:10:26 | 14m57s |
Identical signature for every occurrence (journalctl -b -1):
- Journal truncated mid-line after ordinary application logs
- No systemd shutdown sequence, no kernel panic/oops, no OOM, no thermal critical, no Xid, no MCE/AER
/sys/fs/pstoreempty; no vmcore from kdump (crashkernel=1G-:0Mrenders kdump effectively disabled)last -xmarks every session ascrash
This is the signature of an EC-level power cut that occurs faster than the kernel can log — identical to threads 373251 / 377365.
3. Sensor Reading Anomalies (core evidence)
3.1 Physically impossible instantaneous jumps
07:27:50 → 07:28:00, acpitz zone0 fell from 84 °C to 67 °C in 10 seconds (a 17 °C drop), which thermal mass does not permit:
代码块
07:27:50 z0=84 z4=84
07:28:00 z0=67 z4=67 ← 10 s, -17 °C
07:28:31 z0=82 z4=82
3.2 Only three discrete values across seven zones + perfect synchronization
Under GPU full load, the seven thermal zones report only three distinct values, with zone0/zone5 perfectly synchronized and zone1–4 perfectly synchronized:
代码块
zone0=87 zone1=66 zone2=66 zone3=67 zone4=66 zone5=87 zone6=71
▲───────────────── 87 °C band (zone0/zone5)
▲──────────── 66 °C band (zone1–4)
Seven physically distinct locations cannot collapse to three tidy bands — this is the classic signature of a sensor mapping/calibration defect (digiegg’s words: “Two zones trading values … it reads as a sensor handoff or a mapping/calibration problem”).
3.3 87 °C reported on the first sample (physically impossible)
Approximately 2 seconds after GPU load starts (first sample), zone0/zone5 report 87 °C, while the GPU die itself reports only 76 °C via nvidia-smi. Thermal mass does not permit a 38 °C rise in 2 seconds.
4. Firmware Already Latest, Fault Persists
fwupdmgr get-history confirms the EC was upgraded 0x02004e18 → 0x03000508 (latest official 3.5.8) on 2026-10-06, and the SoC is already 0x02009b0b (latest official 2.155.11). The 10 hard power-offs occurred after the firmware was updated, eliminating “stale firmware” as a variable.
5. GPU Thermally Locked Down + Cooling Verification
5.1 The GPU was never actually running at full load
| Condition | GPU clock | GPU power | Note |
|---|---|---|---|
| No external cooling | 617 MHz | 15 W | thermally locked, false full-load |
| 50 W TEC cooling added | 2171 MHz | 83.6 W | true full-load (against 2200 cap) |
5.2 50 W cooling still trips the 87 °C cutoff (reproducible)
代码块
10:13:55 zone=87 66 66 67 66 87 71 GPU=76°C|83.60W|93%|2171MHz
10:17:20 zone=87 66 66 67 66 87 71 GPU=76°C|82.75W|93%|2171MHz
Two identical runs. Note: the GPU die is 76 °C and zones 1–4 are 66 °C, yet zone0/zone5 report 87 °C — two “hot zones” 11 °C hotter than the GPU die itself, reached instantly. This points to sensor over-reporting rather than genuine cooling insufficiency.
6. Comparison with RMA-approved digiegg case (377365)
| Dimension | digiegg | This unit |
|---|---|---|
| SBIOS | 5.36_0ACUM018 |
✅ Same |
| EC / SoC | 0x03000508 / 0x02009b0b |
✅ Same |
| Log-less hard power-off | journal truncated, no vmcore/pstore | ✅ Same |
| Sensor zone handoff/sync | Yes | ✅ Yes |
| Error code reading | MODS-020000600139 = “temperature limits exceeded or the thermal sensor is broken or miscalibrated” |
Consistent with this unit’s symptoms |
| Outcome | RMA approved | Pending |
7. RMA Request
This unit exhibits a reproducible, firmware-independent thermal sensor fault causing log-less hard power-offs. I request authorization to RMA this unit.
Attachments:
journalctl --list-bootsoutput (crash timeline)sudo journalctl -b -1 > journal_prev_boot.txtsudo journalctl -k -b -1 > kernel_prev_boot.txtnvidia-smi -q > nvidia-smi-q.txt- Thermal zone readings:
for z in /sys/class/thermal/thermal_zone*; do echo "=== $z ==="; cat $z/type $z/temp; done gpu-stress-monitor.log(two 87 °C cutoffs + three-band readings)
Serial number available on request.