Not just myself, but quite a number of users have experienced hardware-level shutdowns while running ComfyUI or the community benchmark tool llama-benchy. Furthermore, I understand that these users all passed the NVIDIA field diagnostics test provided at NVIDIA DGX Spark Field Diagnostics | NVIDIA. In other words, the field diagnostics indicate no problems.
However, shutdowns ultimately occur under high load, and the most commonly cited workaround is sudo nvidia-smi -lgc min,max. Typically, setting the max to 2300 or below appears to eliminate these shutdown experiences.
The question is: are these shutdowns caused by firmware issues, given that the device cannot handle the initial factory settings (min:max=2418:3003)? Or does this require RMA? We purchased this product based on advertised performance claims, but if we need to use sudo nvidia-smi -lgc min,max for normal operation, then we essentially purchased a misrepresented product.
In my case, with llama-benchy, I experience no shutdowns even under high load conditions with Qwen3.5-397B-A17B-int4-AutoRound (dual), Qwen3.5-122B-A10B-int4-AutoRound (single), and gpt-oss-120b (single) models at --depth values of 262144, 131072, 65536, 32768 and --concurrency ranging from 10 to 100.
However, with ComfyUI, the Wan2.2 i2v default template causes guaranteed shutdowns. The qwen image edit 2512 sometimes shuts down but mostly succeeds.
According to “nvtop”, “lm-sensor” logs, the basic average temperature and load are similar between these two workloads, so why does ComfyUI consistently cause shutdowns?
If this were a firmware issue, all devices should experience the same problem. However, other devices with the same kernel and firmware versions (updated at the same time) do not experience any hardware-level shutdowns with the workflows mentioned above.
Naturally, I tried complete power disconnection and reconnection from the wall outlet and device, but the symptoms persisted. I also changed the power strip, but the issue remained the same.
It would be easier to just send it for RMA, but I heard the shocking news that the vendor requires approximately 5 weeks or more for RMA processing. I use this device for my livelihood and research—what am I supposed to do if it disappears for 5 weeks or more? For the time being, I’m continuing to use the device since it works when I apply sudo nvidia-smi -lgc min,max.
Therefore, I ask NVIDIA staff: For devices that experience shutdowns under high load but operate stably when sudo nvidia-smi -lgc min,max is configured:
- Is this simply a unit-to-unit variation?
- Or does this require inspection and RMA?
- Is this a firmware issue that NVIDIA is aware of and working to resolve?
- Or is this a symptom caused by a design flaw?
I look forward to your response.