My DGX Spark GPU fails to initialize after a GSP firmware failure. Requesting
developer team review and RMA validation. NVIDIA Customer Care (case #260620-000330) directed me here for confirmation before RMA.
HARDWARE: DGX Spark (Founders Edition), Serial 1983925007599, Under Warranty
(exp 06/03/2027), purchased directly from NVIDIA.
SYSTEM: DGX OS 7.5.0, kernel 6.17.0-1021-nvidia, driver 580.159.03.
SYMPTOM: GPU fails to initialize. nvidia-smi returns “No devices were found” on
every boot. The GPU worked normally for 40+ hours before the failure.
NVIDIA kernel modules load fine (nvidia, nvidia_uvm, nvidia_modeset, nvidia_drm)
This is specifically a GSP firmware bootstrap failure (SEC2 secure boot timeout)
TRIGGER: The GPU wedged during a large LLM load that drove unified memory + swap
into thrashing; the OS became unresponsive and required a forced power-off. The
GSP failure began after that.
ALREADY TRIED (all failed):
Full power-disconnect cold boot, 10+ minutes, multiple times
ADDITIONAL: The monitor shows no output past the firmware splash, so I cannot
enter UEFI locally to disable Secure Boot for field diagnostics. SSH still works
and nvidia-bug-report.log.gz is available.
This matches thread 373394 (identical OS/kernel/driver and error signature),
where the conclusion was an RMA after System Recovery also failed.
Could the DGX Spark team confirm whether this requires an RMA? I’ll relay the
confirmation to Customer Care (case #260620-000330). Thank you.
I ran the field diagnostic as requested. Result: partnerdiag cannot run because its MODS diagnostic driver is blocked by Secure Boot.
The MODS driver compiles and installs fine, but fails to load:
insmod: ERROR: could not insert module mods.ko: Key was rejected by service
This is Secure Boot rejecting the unsigned diagnostic module. The documented fix is to disable Secure Boot in UEFI (mokutil confirms: SecureBoot enabled). But the GPU failure also eliminated all video output — the monitor shows nothing past the firmware splash — so I cannot enter UEFI setup to disable Secure Boot. There is no BMC/remote console on DGX Spark.
So the hardware failure blocks the standard diagnostic on two fronts:
No video output, so UEFI is unreachable to disable Secure Boot for fieldiag
This isn’t a driver/setup issue — the diagnostic driver builds correctly; the hardware fault itself prevents running the diagnostic. The signature is identical to thread 373394, which concluded in an RMA.
Attached: full fieldiag logs (driver.log shows “Key was rejected by service”). Could you confirm RMA, or advise how to run fieldiag without UEFI access?
Serial 1983925007599, Under Warranty (06/03/2027), Customer Care case #260620-000330.
Thanks aniculescu. I can run sudo systemctl reboot --firmware-setup over SSH, but
I want to flag the risk before I do: the GPU failure has eliminated all video
output — the monitor is black past the firmware splash. So even if the system
boots into the UEFI menu, I won’t be able to see or navigate it to change Secure
Boot. And if it halts at the (invisible) UEFI menu, SSH won’t come back and I’ll
have to force power-off again (which is what triggered this GSP state originally).
Is there a way to disable Secure Boot or run fieldiag without a working display
(e.g. via mokutil/efibootmgr from the CLI, or a fieldiag option that skips the
MODS signature check)? If you still want me to try the firmware-setup reboot
given the no-display constraint, I’ll do it — just confirming you’re aware SSH
may not return.
One more consideration before any further reboots or recovery steps: this unit stored sensitive customer data (a data-sovereignty product for law firms), and I haven’t been able to securely wipe the SSD yet. If a step risks locking me out of SSH (e.g. halting at an invisible UEFI menu), I’d lose the ability to sanitize the drive before return. Could you advise whether I may remove and retain the SSD for the RMA, returning only the chassis? That would let me proceed with recovery steps without risking unsanitized data leaving my custody.
You need Secure Boot disabled to run the fieldiag. Please try to access BIOS with the command I sent. You cannot retain any components to RMA the device.
I set the MOK with mokutil --disable-validation as a first step, but I need to flag a hard constraint, and I’ve now physically verified the display failure.
DISPLAY IS DEAD (verified): I connected a monitor directly and tried multiple times. The first screen appears briefly, then it immediately goes to a black screen — I cannot even reach the login/password prompt, let alone UEFI or the MOK Manager. Same result on every attempt, across reconnections. The GPU failure took the display output with it.
So disabling Secure Boot to run fieldiag is not physically possible in this state:
GPU failed → nvidia-smi returns “No devices found”, and the GSP firmware fails to initialize (SEC2 secure boot timeout, RmInitAdapter failed 0x62:0x65:2028)
Display output is dead → I cannot see UEFI or the MOK Manager, so I cannot disable Secure Boot or confirm the MOK enrollment even with physical access
After sudo reboot, the MOK Manager requires keyboard confirmation on a screen I cannot see; if it halts there, SSH won’t return and the unit is bricked
I confirmed the fieldiag itself is blocked by Secure Boot — the MODS diagnostic driver builds and installs fine but fails to load:
insmod: ERROR: could not insert module mods.ko: Key was rejected by service
(full fieldiag logs are attached to an earlier post). This is a Secure Boot rejection, and the only documented fix — disabling Secure Boot in UEFI — is unreachable because the display is dead.
The hardware failure blocks both nvidia-smi and the field diagnostic, and the dead display blocks the one workaround. This is the same dead-end as thread 373394, which concluded in an RMA after System Recovery also failed.
Could we proceed with the RMA on this basis? Serial 1983925007599, Under Warranty (exp 06/03/2027), Customer Care case #260620-000330.
Update with physical access today (day 7 of this issue).
To recap: I’ve followed every step NVIDIA directed — opened an enterprise case (01201344) and a DGX Spark ticket (#260620-000330), was routed to this forum, posted here, and worked through each instruction. I ran fieldiag, hit the Secure Boot block, set the MOK, and connected a monitor — I’ve completed everything asked except the final step (disabling Secure Boot), and that one is not something I can do, because the hardware failure makes it physically impossible:
nvidia-smi: “No devices were found” (GPU dead; GSP firmware fails to init, RmInitAdapter 0x62:0x65:2028)
Monitor physically connected (DRM port reports “connected”), but the screen stays black after the initial splash, and keyboard input produces no response — I cannot reach the login prompt, UEFI, or the MOK Manager
So Secure Boot cannot be disabled, and fieldiag’s MODS driver stays rejected by Secure Boot (“Key was rejected by service”)
The GPU failure has killed both compute and display, and the dead display blocks the only documented workaround. This is the same dead-end as thread 373394, which concluded in an RMA.
It’s now been 7 days, with no moderator reply for the last two, and the downtime is blocking active development that’s time-critical for us. Could we please proceed to RMA as quickly as possible? nvidia-bug-report and full fieldiag logs are already attached. Serial 1983925007599, Under Warranty (06/03/2027), case #260620-000330. Thank you for any speed you can offer.