for some reason, customer support told me i need to post in the forum to get an RMA, which seems silly for a unit like this, but here it is:
Unit:
- NVIDIA DGX Spark
- BIOS 5.36_0ACUM018 (08/06/2025), DGX OS / Ubuntu 24.04.4, kernel 6.17.0-1021-nvidia
Diagnostic result (sudo ./partnerdiag --field): Final Result: FAIL
- PowerStress FAILED at ~8:08, error code MODS-020000600139 — “Acceptable temperature limits exceeded or the thermal sensor is broken or miscalibrated.”
- All other subtests PASS (GpuStress, C2CStress, CpuStress1, CpuStress2).
- Reproducible: 4 separate runs, all fail PowerStress at the same ~8-minute mark. Full fieldiag logs available.
Real-world symptom: the unit hard-powers-off / freezes under sustained high-power GPU inference, recoverable only by a cold power cycle — no kernel panic, no OOM.
It’s isolated to this unit, not my environment: my identical second DGX Spark passes the same PowerStress test, but this unit keeps failing PowerStress
under the same (better) cooling, even at modest external temps. So it’s a unit-level thermal-sensor / power fault, not ambient.
I am happy to supply logs.
Testing GpuStress OK [ 3:19s ]
Testing C2CStress OK [ 0:06s ]
Testing CpuStress1 OK [ 0:08s ]
Testing CpuStress2 OK [ 10:03s ]
Testing PowerStress FAILED [ 8:10s ]
Exit Code | Virtual Id | Test | Subtest | Component | Component Id | Notes
MODS-000000000000 | GpuStress | custommods | | GPU | | OK
MODS-000000000000 | C2CStress | custommods | | C2C | | OK
MODS-000000000000 | CpuStress1 | custommods | | CPU | | OK
DGX-000000000000 | CpuStress2 | cpustress | | CPU | | OK
MODS-020000600139 | PowerStress | custommods | | Power | | Acceptable temperature limits exceeded or the thermal sensor is broken or miscalibrated
####### #### ######## ###
####### ###### ######## ###
## ##
## ##
####### ######## ## ###
####### ######## ## ###
## ##
## ##
## ##
Final Result: FAIL
End time: Sun, 14 Jun 2026 16:14:59 [ 21:48s elapsed ]
Copying logs…