To NVIDIA Staff: Is This a Hardware Issue Requiring Repeated Shutdowns and RMA Under High Load?

Not just myself, but quite a number of users have experienced hardware-level shutdowns while running ComfyUI or the community benchmark tool llama-benchy. Furthermore, I understand that these users all passed the NVIDIA field diagnostics test provided at NVIDIA DGX Spark Field Diagnostics | NVIDIA. In other words, the field diagnostics indicate no problems.

However, shutdowns ultimately occur under high load, and the most commonly cited workaround is sudo nvidia-smi -lgc min,max. Typically, setting the max to 2300 or below appears to eliminate these shutdown experiences.

The question is: are these shutdowns caused by firmware issues, given that the device cannot handle the initial factory settings (min:max=2418:3003)? Or does this require RMA? We purchased this product based on advertised performance claims, but if we need to use sudo nvidia-smi -lgc min,max for normal operation, then we essentially purchased a misrepresented product.

In my case, with llama-benchy, I experience no shutdowns even under high load conditions with Qwen3.5-397B-A17B-int4-AutoRound (dual), Qwen3.5-122B-A10B-int4-AutoRound (single), and gpt-oss-120b (single) models at --depth values of 262144, 131072, 65536, 32768 and --concurrency ranging from 10 to 100.

However, with ComfyUI, the Wan2.2 i2v default template causes guaranteed shutdowns. The qwen image edit 2512 sometimes shuts down but mostly succeeds.

According to “nvtop”, “lm-sensor” logs, the basic average temperature and load are similar between these two workloads, so why does ComfyUI consistently cause shutdowns?

If this were a firmware issue, all devices should experience the same problem. However, other devices with the same kernel and firmware versions (updated at the same time) do not experience any hardware-level shutdowns with the workflows mentioned above.

Naturally, I tried complete power disconnection and reconnection from the wall outlet and device, but the symptoms persisted. I also changed the power strip, but the issue remained the same.

It would be easier to just send it for RMA, but I heard the shocking news that the vendor requires approximately 5 weeks or more for RMA processing. I use this device for my livelihood and research—what am I supposed to do if it disappears for 5 weeks or more? For the time being, I’m continuing to use the device since it works when I apply sudo nvidia-smi -lgc min,max.

Therefore, I ask NVIDIA staff: For devices that experience shutdowns under high load but operate stably when sudo nvidia-smi -lgc min,max is configured:

  1. Is this simply a unit-to-unit variation?
  2. Or does this require inspection and RMA?
  3. Is this a firmware issue that NVIDIA is aware of and working to resolve?
  4. Or is this a symptom caused by a design flaw?

I look forward to your response.

Can second this. I am in the exact same boat. I have two DGX Spark FEs. One will consistently shut down after reaching over 65K context on llama-benchy, whether it is the head node or worker node. It will also fail under real-world workflows that involve large context processing. The units are running exactly the same firmware, and are both completely up to date. The other unit (purchased a few months later) works fine.

If I lower the maximum GPU clock to 2200 it runs fine under stress indefinitely. The unit passes the field diagnostics without lowering the clock but failure condition can be consistently reproduced under real-world loads.

Here is what I don’t want: Have the unit in limbo for more than a month for one of three things:

  1. Somebody at NVIDIA runs the same field diagnostics and concludes “it’s fine” and ships it back.
  2. NVIDIA eventually swaps it out with another unit with the same problem (seems widespread)
  3. I actually get a working unit or this one is fixed.

Either way I’m without the item for a month or more, and I am rolling the dice that I get outcome #3.

Crazy idea: Cross ship me a unit that works and I’ll send the defective one back to you. If you want to hold a credit card for a short period of time until you receive shipment, that’s fine. This is how customer-centric companies, especially those that deal in the enterprise space, deal with hardware issues.

failure condition can be consistently reproduced under real-world loads.

Hi,

It’d help us triage this if you can share the minimal repro steps. What are the exact llama-benchy version and parameters used?

Can you please share the passing Field Diag logs from the unit that consistently shuts down?

==========================================
llama-banchy

I cloned the main branch of llama-benchy from “GitHub - eugr/llama-benchy: llama-benchy - llama-bench style benchmarking tool for all backends · GitHub” just 2 days ago and used it.

The execution flags are as follows. ( However, I did not experience shutdown symptoms with llama-benchy, but many users have reported shutdown issues during this bench test.)


llama-benchy \
  --base-url http://localhost:8000/v1 \
  --model /workspace/Model/gpt-oss-120b \
  --depth 0 16384 32768, 65536, 131072, 262144 \
  --concurrency 20

llama-benchy \
  --base-url http://localhost:8000/v1 \
  --model /workspace/Model/Qwen3.5-122B-A10B-int4-AutoRound \
  --depth 0 16384 32768, 65536, 131072, 262144 \
  --concurrency 20

llama-benchy \
  --base-url http://localhost:8000/v1 \
  --model /workspace/Model/Qwen3.5-122B-A10B-int4-AutoRound \
  --depth 0 16384 32768, 65536, 131072, 262144 \
  --concurrency 20

==========================================
ComfyUI wan 2.2 i2v

video_wan2_2_14B_i2v.txt (19.8 KB)

I have attached the ComfyUI Workflow as a txt file. You need to convert it to JSON for actual application and testing.

==========================================
NVIDIA DGX Spark Field Diagnostics

NVIDIA DGX Spark Field Diagnostics
I have sent the complete field diagnostic log via DM. and I’m showing the run.log results.

Command Line: onediagfield.r9.257.3 --keep_disp_enabled --force_product=spark --run_spec=spec_dgx_spark_field_level2.json


******************************************************************
*                                                                *
*                      DGX FIELD DIAGNOSTIC                      *
*                                                                *
******************************************************************

Version                  r9.257.3
Python                   /opt/nvidia/dgx-spark-fieldiag/dgx/python/python-3.10.7-glibc-2.17-aarch64/bin/python3
Build Date               Fri, 19 Dec 2025
Start time               Tue, 03 Mar 2026 18:20:33
Est. Completion time     Tue, 03 Mar 2026 21:23:03
Logs                     /opt/nvidia/dgx-spark-fieldiag/dgx/logs-20260303-182013
Inforom logs             /opt/nvidia/dgx-spark-fieldiag/dgx/logs-20260303-182013/inforom
Product                  NVIDIA_DGX_Spark
Product Version          A.7
Family                   DGX Spark
SKU                      0000
Serial Number            xxxxxxxxxxxxxxxxx

Testing GpuStress OK [ 3:19s ]
Testing C2CStress OK [ 0:06s ]
Testing CpuStress1 OK [ 0:08s ]
Testing CpuStress2 OK [ 10:21s ]
Testing PowerStress OK [ 8:11s ]
Testing ThermalStress OK [ 0:42s ]
Testing FioSSD OK [ 0:12s ]
Testing MemStress OK [ 7:12s ]

Exit Code         | Virtual Id    | Test       | Subtest | Component | Component Id | Notes
===========================================================================================
MODS-000000000000 | GpuStress     | custommods |         | GPU       |              | OK
MODS-000000000000 | C2CStress     | custommods |         | C2C       |              | OK
MODS-000000000000 | CpuStress1    | custommods |         | CPU       |              | OK
DGX-000000000000  | CpuStress2    | cpustress  |         | CPU       |              | OK
MODS-000000000000 | PowerStress   | custommods |         | Power     |              | OK
MODS-000000000000 | ThermalStress | custommods |         | Thermal   |              | OK
DGX-000000000000  | FioSSD        | custom     |         | SSD       |              | OK
DGX-000000000000  | MemStress     | custom     |         | Memory    |              | OK

#######     ####     ######    ###### 
########   ######   ########  ########
##    ##  ##    ##  ##     #  ##     #
##    ##  ##    ##   ###       ###    
########  ########    ####      ####  
#######   ########      ###       ### 
##        ##    ##  #     ##  #     ##
##        ##    ##  ########  ########
##        ##    ##   ######    ###### 

Final Result: PASS
End time: Tue, 03 Mar 2026 18:50:48 [ 30:15s elapsed ]

“The diagnostic log shows all tests passed (Final Result: PASS)”

In this case, I would appreciate it if you could let me know as soon as possible whether this requires an RMA, whether it is an issue caused by firmware, or whether it is a limitation of the design. If an RMA is necessary, I need to proceed with it as soon as possible to resolve the problem.

Thanks for the additional context, gpieceoffice.

I did not experience shutdown symptoms with llama-benchy,

What are the exact workloads associated with your unit to shut down?

Thanks for your patience. Gathering this information now helps the engineering team prioritize and narrow down the issue, and reduces back-and-forth later.

This occurs when running image generation models like Wan 2.2 i2v or Qwen Image series in ComfyUI. In particular, this happens even with the default Wan 2.2 i2v workflow provided by ComfyUI (I also uploaded my workflow earlier).

Image generation and video generation models use FP16 precision.
In case you need it, I’m uploading the workflow that causes the problem again.
video_wan2_2_14B_i2v.txt (19.8 KB)

Additionally, this same workflow runs indefinitely on two other devices without any issues, but on the problematic device, it causes guaranteed shutdowns during operation.

Let me reiterate: even the problematic device passes the Field Diagnostics no matter how many times I run the test.

Also, the shutdown does not occur during the llama-benchy test.

I will upload my info shortly

The only time I encountered the system’s power/frequency being locked at a threshold was right after running DGX Spark Field Diagnostics, so I have reason to suspect that DGX Spark Field Diagnostics is one of the culprits causing the issue.

Hey, I have exactly the same problem.

llama-benchy runs smoothly.
ComfyUI image generation - Crash

With “nvidia-smi -lgc 0.2300” it runs stable.

RMA or is there a fix coming?

The DM I personally received was asking me to request an RMA.

Looking at this alone, it seems that anyone with similar symptoms as mine would be considered to have a defective product.

I would like to ask once again:

Many people appear to be experiencing these shutdown issues. Is the only solution to proceed with an RMA due to a hardware defect, or could this potentially be related to firmware, drivers, or even the kernel?

If it is indeed a hardware defect, the failure rate seems quite high, which makes me wonder whether there could be a design-related issue as well. I have already gone through the RMA process once, and the unit I received as a replacement is now experiencing problems again—so this would be my second RMA.

I’m also very concerned about the possibility that the same issue could occur again

It would be greatly appreciated if representatives from NVIDIA could clearly address this issue and confirm whether it has been identified and included in the list of known issues.

Yeah, not good. Your scenario is basically, why I have not requested a return yet. If the answer is a firmware update that reduces clocks to avoid this it’s not ideal, but at least it is addressing the issue.

field diagnostic logs and reproduction steps provided via DM

Thanks. We do need the entire field diag *.tgz log bundle. A single log misses many contexts of the entire log bundle. And it can be attached via DM.

Ok, can you confirm where I should grab that? I zipped the logs directory under the field diag directory and it was apparently too large to attach.

Per NVIDIA DGX Spark Field Diagnostics | NVIDIA

Logs

Logs are saved under:

/opt/nvidia/dgx-spark-fieldiag/logs-<timestamp>/

Key files:

  • output.log, run.log
  • Per-test logs in subfolders
  • summary.json for RMA review

In my case, I have already provided the log files via DM. Is this issue something that requires mandatory inspection and RMA (Return Merchandise Authorization), or can it be resolved through a future software update? I would appreciate any information, whether it’s in a DM or a post.

Update:

We have not been able to reproduce the issue internally using multiple Founders Edition and partner SKU units while repeatedly running the workloads described in this thread. The GPU PD check tool also returns a pass across a range of SKUs.

My understanding is that specific units within your Sparks cluster show this behavior. All units are running the same OS, firmware, software stack, and workload, yet only certain nodes consistently crash. Based on that pattern, we recommend RMA for the units that consistently fail. If this were a software or firmware issue, it would likely affect all units similarly.

Unfortunately, this doesn’t seem to be a rare occurrence. Can you advise what one can typically expect via the RMA process?

I had the same issue and received a replacement device via RMA yesterday.

First, let me tell you the results: the issue was resolved after the device was replaced via RMA. The NVIDIA Test Tool run by the distributor gave a PASS result, so even though it wasn’t originally eligible for RMA, I was able to get a replacement.

As a result, the replacement device is fine. Below is the relevant article.
MSI EdgeXpert Suddenly Power-Off During llama-benchy – Possible PD Firmware Issue? - DGX Spark / GB10 User Forum / DGX Spark / GB10 - NVIDIA Developer Forums

I would appreciate it if NVIDIA officials could officially investigate this issue and provide a response. Even the distributor officials were completely clueless about how to address this issue.

The frustrating part about the RMA process was that, although the test tool returned a PASS result, in my experience, the device would still unexpectedly shut down. This clearly requires further investigation.

What are your thoughts on NVIDIA collecting and investigating the devices returned due to this issue?