Tue Jan 20 04:49:39 2026
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 |
±----------------------------------------±-----------------------±---------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GB10 On | 0000000F:01:00.0 On | N/A |
| N/A 50C P0 12W / N/A | Not Supported | 0% Default |
| | | N/A |
±----------------------------------------±-----------------------±---------------------+
±----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 4135 G /usr/lib/xorg/Xorg 400MiB |
| 0 N/A N/A 4316 G /usr/bin/gnome-shell 224MiB |
| 0 N/A N/A 5126 G /usr/bin/wezterm-gui 32MiB |
| 0 N/A N/A 7600 G …rack-uuid=3190708988185955192 87MiB |
±----------------------------------------------------------------------------------------+
When performing image generation tasks using Qwen image, GPU utilization reaches up to 96% and the temperature reaches 80 degrees Celsius. Once the temperature reaches 80 degrees Celsius, the power is shut down with a near 100% probability after about 10-15 seconds. Power consumption is approximately 90 watts. Memory consumption is approximately 36 GB.
In addition to Qwen image, when GPU utilization approaches 96%, including gpu-burn, the temperature remains at around 78-80 degrees Celsius for about 10 seconds before the power is shut down.
I would like to know if this behavior is normal.
(I wanted to upload a bug report file, but the attachment wasn’t working properly.)
I suspect my product is an MSI product, so it may not be eligible for official support on this forum. I’ve already contacted the manufacturer, but is there any way to inspect this issue down on the dgx OS? (command or tools)
hi the snippet you shared is mostly early-boot bringup messages. then "systemd-journald … journal … corrupted or uncleanly shut down at 04:44:12
" which appears after an unplanned reset.
Since you seem to have an MSI version, you’ll need to discuss the hardware triage with MSI. You can install and run the Spark Field Diagnostics tool, which should give us more insights. Instructions are here NVIDIA DGX Spark Field Diagnostics | NVIDIA
Similar issue so I am plonking it here.
My findings after two weeks of using the DGX is that if I go anywhere near to close to max the machine resources, especially on utilising all CPU cores in parallel, the freezes and then after a period just shuts down in a non graceful manner!. Only a hard reboot will fix it. Additionally, when connecting using Cursor/ssh, it is common for the connect to freeze up and I have to again hard reboot the device to get it back working. It gets super hot also which I’m sure will not be great for the machine long term. Not great reliability for the price! (and I have 4 of them waiting to build into a cluster!)
thanks for your help. I got Field diagnostics and it failed with GPU Stress test. like below [cmd-3] Error Code = 020000021139 (Acceptable temperature limits exceeded or the thermal sensor is broken or miscalibrated)
only getting back to this as was travelling - it turns out the problem is that at least one of the 4 x devices I purchased are running too hot on idle and more so under load. I am having a pretty poor experience trying to get these taken back under warranty by the reseller in Spain where I purchased them (€20k spent) … is there a way I can escalate this via NVidia before resorting to having to send legal letters?
Piling up here. Purchased DGX on day 1. Was crashing sometimes. Now it crashes all the time. Running gpu_burn is enough to show that this is simply a thermal issue. The unit heats then dies (non gracefully). Seems like a bad thermal design.
developer@sparky
:
~/github/gpu-burn
$ ./gpu_burn 300
Using compare file: compare.fatbin
Burning for 300 seconds.
GPU 0: NVIDIA GB10 (UUID: GPU-cf0ef592-0423-ee7e-c6ee-65276769d5a0)
cuInit returned 0 (no error)
Initialized device 0 with 124610 MB of memory (11610 MB available, using 10449 MB of it), using FLOATS
Results are 268435456 bytes each, thus performing 38 iterations
10.7% proc’d: 532 (18765 Gflop/s) errors: 0 temps: 83 C
Summary at: Mon Jun 15 09:27:07 AM CEST 2026
21.0% proc’d: 1026 (16842 Gflop/s) errors: 0 temps: 78 C
Summary at: Mon Jun 15 09:27:38 AM CEST 2026
31.3% proc’d: 1520 (17605 Gflop/s) errors: 0 temps: 85 C
Summary at: Mon Jun 15 09:28:09 AM CEST 2026
38.7% proc’d: 1862 (17127 Gflop/s) errors: 0 temps: 85 C
It would probably be quicker to remove the case, add big fans and they will most likely work fine. Identical experience with NVIDIA here in terms of (lack of) support.
“remove the case, add big fans” … true … thinking about a harness that includes fans and possibly exposing the base of the device … clearly looking at some of the other OEM design decisions the original is not the best design for thermal release…