Unexpected Shutdown During ComfyUI Inference on DGX Spark (Occurs on Two Units)

, ,

Hello,

I would like to discuss an unexpected shutdown issue that occurs during ComfyUI inference.

I own two NVIDIA DGX Spark systems, and the same issue is now occurring on both devices. Initially, the problem appeared on only one unit, but it is currently happening on both.

I followed the field diagnostics guide provided here:

I ran the full diagnostic test three times on each device, and all tests passed successfully.

I also reviewed and referenced the following forum post:

As mentioned in that post, when I apply a power limit, the shutdown issue does not occur. But fundamentally, I can’t understand why it can’t even sustain the amount of power required to hold the stock/default clock speeds.

For context, I am able to run the recently released Qwen 3.5 122B-A10B model with 250 concurrent requests for game translation workloads. Even under heavy load with extremely high cooling fan activity, the system does not shut down.

However, when I start inference in ComfyUI, the system shuts down after only 1–2 steps. Occasionally it completes without issue, but in most cases the power is abruptly cut off.

Is this a ComfyUI-specific issue?
Is this a hardware issue with my units?
Could it be variability between individual devices?

What makes this more confusing is that my wife also owns a DGX Spark, and her system does not shut down during ComfyUI inference, regardless of power limits.

Given that both of my units now exhibit the same behavior while passing diagnostics, I am uncertain whether this is a firmware, power delivery, thermal protection, or workload-related issue—or whether I should proceed with an RMA request.

I would greatly appreciate any insight into what might be causing this behavior.

I also have similar issue, but not NVIDIA FE. it’s MSI EdgeXpert.
Unexpected shutdown occurred during llama-benchy.

It’s not a comfy issue it’s the power draw when the GPU ramps up. It’s not thermal related. I recall someone mentioning that different models have different power configurations which would make sense in your case with two different ones. Throttling the clock speed has negligible effect to performance and most servers are doing this already since they run 24/7. I’m sure Nvidia will fix it now that they read my post.

users experiencing the same issue and an Nvidia moderator participated in the conversation:

The Nvidia moderator response was that this behavior only occurs in certain units and that an RMA is recommended. Therefore, I have already applied for an RMA before my warranty expires.

More importantly, the DGX Spark does not appear to be designed like a typical server-based system. Even in actual GPU server environments, clock speeds are lowered primarily due to environmental considerations (such as thermals, power limits, or stability margins), not because the system stops operating under high load.

Based on the moderator’s response, it is also difficult to conclude that different models have fundamentally different power configurations.

Most importantly, the fact that the system cannot sustain the default factory settings despite no overclocking being applied suggests that there may be an underlying issue. I am not insisting that everyone must proceed with an RMA, but if you are experiencing the same symptoms and still have warranty coverage, I would recommend submitting your field diagnostic log files first and then considering an RMA.

As you can see from the post I shared earlier, most of the users who participated in the discussion passed the field diagnostics, yet their systems only operate normally after lowering the clock speed.

I have the same issue. The solution @jas.burton pointed out works for me. I recall undervolting or underclocking systems for many years to deal with stability issues on certain temparature, voltage and clock related workloads. For me personal it is a developement unit running any experimental code so unexpected behavour can occur. I think in my dgx spark the NV ComfyUi playbook only worked with tweaks.