Shutdown under high utilization during image generation using qwen

Tue Jan 20 04:49:39 2026
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05              Driver Version: 580.95.05      CUDA Version: 13.0     |
±----------------------------------------±-----------------------±---------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GB10                    On  |   0000000F:01:00.0  On |                  N/A |
| N/A   50C    P0             12W /  N/A  | Not Supported          |      0%      Default |
|                                         |                        |                  N/A |
±----------------------------------------±-----------------------±---------------------+

±----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|    0   N/A  N/A            4135      G   /usr/lib/xorg/Xorg                      400MiB |
|    0   N/A  N/A            4316      G   /usr/bin/gnome-shell                    224MiB |
|    0   N/A  N/A            5126      G   /usr/bin/wezterm-gui                     32MiB |
|    0   N/A  N/A            7600      G   …rack-uuid=3190708988185955192         87MiB |
±----------------------------------------------------------------------------------------+

When performing image generation tasks using Qwen image, GPU utilization reaches up to 96% and the temperature reaches 80 degrees Celsius. Once the temperature reaches 80 degrees Celsius, the power is shut down with a near 100% probability after about 10-15 seconds. Power consumption is approximately 90 watts. Memory consumption is approximately 36 GB.

In addition to Qwen image, when GPU utilization approaches 96%, including gpu-burn, the temperature remains at around 78-80 degrees Celsius for about 10 seconds before the power is shut down.

I would like to know if this behavior is normal.

(I wanted to upload a bug report file, but the attachment wasn’t working properly.)

can you please share more info? what exactly do you mean “power is shutdown”? meaning the device turns off? power drops? or something else?

Thank you for your attention. The device suddenly turned off.

When I looked at journalctl -k -b -1 it ended with this message.

1월 20 04:42:04 edgexpert-8de4 kernel: nvidia 000f:01:00.0: Adding to iommu group 20
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: Core ver 2.22
1월 20 04:42:04 edgexpert-8de4 kernel: NET: Registered PF_BLUETOOTH protocol family
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: HCI device and connection manager initialized
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: HCI socket layer initialized
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: L2CAP socket layer initialized
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: SCO socket layer initialized
1월 20 04:42:04 edgexpert-8de4 kernel: RPC: Registered rdma transport module.
1월 20 04:42:04 edgexpert-8de4 kernel: RPC: Registered rdma backchannel transport module.
1월 20 04:42:04 edgexpert-8de4 kernel: nvidia 000f:01:00.0: vgaarb: VGA decodes changed: olddecodes=io+mem,decodes=none:owns=none
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: hci0: HW/SW Version: 0x00000000, Build Time: 20250721233113
1월 20 04:42:04 edgexpert-8de4 kernel: usbcore: registered new interface driver btusb
1월 20 04:42:04 edgexpert-8de4 kernel: mt7925e 0009:01:00.0: HW/SW Version: 0x8a108a10, Build Time: 20251015212927a
1월 20 04:42:04 edgexpert-8de4 kernel: input: NVIDIA HDMI/DP,pcm=3 as /devices/platform/NVDA2014:00/sound/card0/input10
1월 20 04:42:04 edgexpert-8de4 kernel: input: NVIDIA HDMI/DP,pcm=7 as /devices/platform/NVDA2014:00/sound/card0/input11
1월 20 04:42:04 edgexpert-8de4 kernel: input: NVIDIA HDMI/DP,pcm=8 as /devices/platform/NVDA2014:00/sound/card0/input12
1월 20 04:42:04 edgexpert-8de4 kernel: input: NVIDIA HDMI/DP,pcm=9 as /devices/platform/NVDA2014:00/sound/card0/input13
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.533:4): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“brave” pid=1718 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.533:5): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“ch-run” pid=1725 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.533:6): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“element-desktop” pid=1731 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:7): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“firefox” pid=1735 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:8): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“ch-checkns” pid=1724 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:9): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“buildah” pid=1719 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:10): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“Discord” pid=1714 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:11): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“1password” pid=1713 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:12): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“vscode” pid=1727 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: audit: type=1400 audit(1768851724.534:13): apparmor=“STATUS” operation=“profile_load” profile=“unconfined” name=“crun” pid=1728 comm=“apparmor_parser”
1월 20 04:42:04 edgexpert-8de4 kernel: mt7925e 0009:01:00.0: WM Firmware Version: ____000000, Build Time: 20251015213023
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: BNEP (Ethernet Emulation) ver 1.3
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: BNEP filters: protocol multicast
1월 20 04:42:04 edgexpert-8de4 kernel: Bluetooth: BNEP socket layer initialized
1월 20 04:42:04 edgexpert-8de4 kernel: NET: Registered PF_QIPCRTR protocol family
1월 20 04:42:04 edgexpert-8de4 kernel: loop25: detected capacity change from 0 to 8
1월 20 04:42:05 edgexpert-8de4 kernel: mt7925e 0009:01:00.0 wlP9s9: renamed from wlan0
1월 20 04:42:05 edgexpert-8de4 kernel: mlx5_core 0002:01:00.0 enP2p1s0f0np0: Link down
1월 20 04:42:05 edgexpert-8de4 kernel: mlx5_core 0002:01:00.1 enP2p1s0f1np1: Link down
1월 20 04:42:05 edgexpert-8de4 kernel: enP7s7: 0xffff800085900000, fc:9d:05:13:8d:e4, IRQ 337
1월 20 04:42:05 edgexpert-8de4 kernel: mlx5_core 0000:01:00.0 enp1s0f0np0: Link down
1월 20 04:42:05 edgexpert-8de4 kernel: xhci_endpoint_init rsv 0x801
1월 20 04:42:05 edgexpert-8de4 kernel: xhci_endpoint_init rsv 0x801
1월 20 04:42:05 edgexpert-8de4 kernel: Bluetooth: hci0: Device setup in 1598965 usecs
1월 20 04:42:05 edgexpert-8de4 kernel: Bluetooth: hci0: HCI Enhanced Setup Synchronous Connection command is advertised, but not supported.
1월 20 04:42:05 edgexpert-8de4 kernel: Bluetooth: hci0: AOSP extensions version v1.00
1월 20 04:42:05 edgexpert-8de4 kernel: Bluetooth: hci0: AOSP quality report is supported
1월 20 04:42:05 edgexpert-8de4 kernel: Bluetooth: MGMT ver 1.23
1월 20 04:42:05 edgexpert-8de4 kernel: NET: Registered PF_ALG protocol family
1월 20 04:42:05 edgexpert-8de4 kernel: mlx5_core 0000:01:00.1 enp1s0f1np1: Link down
1월 20 04:42:06 edgexpert-8de4 kernel: evm: overlay not supported
1월 20 04:42:07 edgexpert-8de4 kernel: Initializing XFRM netlink socket
1월 20 04:42:07 edgexpert-8de4 kernel: bridge: filtering via arp/ip/ip6tables is no longer available by default. Update your scripts to load br_netfilter if you need this.
1월 20 04:42:07 edgexpert-8de4 kernel: NVRM: loading NVIDIA UNIX Open Kernel Module for aarch64 580.95.05 Release Build (dvs-builder@U22-I3-AF08-06-3) Tue Sep 23 09:46:53 UTC 2025
1월 20 04:42:07 edgexpert-8de4 kernel: nvidia-modeset: Loading NVIDIA UNIX Open Kernel Mode Setting Driver for aarch64 580.95.05 Release Build (dvs-builder@U22-I3-AF08-06-3) Tue Sep 23 09:36:29 UTC 2025
1월 20 04:42:09 edgexpert-8de4 kernel: [drm] [nvidia-drm] [GPU ID 0x000f0100] Loading driver
1월 20 04:42:09 edgexpert-8de4 kernel: [drm] Initialized nvidia-drm 0.0.0 for 000f:01:00.0 on minor 1
1월 20 04:42:12 edgexpert-8de4 kernel: rfkill: input handler disabled
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: authenticate with 80:ca:4b:25:b9:8f (local address=50:bb:b5:d9:7a:10)
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: send auth to 80:ca:4b:25:b9:8f (try 1/3)
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: send auth to 80:ca:4b:25:b9:8f (try 2/3)
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: authenticated
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: associate with 80:ca:4b:25:b9:8f (try 1/3)
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: RX AssocResp from 80:ca:4b:25:b9:8f (capab=0x1511 status=0 aid=3)
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: AP has invalid WMM params (AIFSN=1 for ACI 3), will use 2
1월 20 04:42:12 edgexpert-8de4 kernel: wlP9s9: associated
1월 20 04:42:13 edgexpert-8de4 kernel: wlP9s9: Limiting TX power to 20 (23 - 3) dBm as advertised by 80:ca:4b:25:b9:8f
1월 20 04:44:12 edgexpert-8de4 systemd-journald[830]: File /var/log/journal/295f5139615f4bbaa29921a29574c7a3/user-1000.journal corrupted or uncleanly shut down, renaming and replacing.
1월 20 04:44:13 edgexpert-8de4 kernel: rfkill: input handler enabled
1월 20 04:44:15 edgexpert-8de4 kernel: kauditd_printk_skb: 217 callbacks suppressed

I suspect my product is an MSI product, so it may not be eligible for official support on this forum. I’ve already contacted the manufacturer, but is there any way to inspect this issue down on the dgx OS? (command or tools)

hi the snippet you shared is mostly early-boot bringup messages. then "systemd-journald … journal … corrupted or uncleanly shut down at 04:44:12
" which appears after an unplanned reset.

Since you seem to have an MSI version, you’ll need to discuss the hardware triage with MSI. You can install and run the Spark Field Diagnostics tool, which should give us more insights. Instructions are here NVIDIA DGX Spark Field Diagnostics | NVIDIA

Similar issue so I am plonking it here.
My findings after two weeks of using the DGX is that if I go anywhere near to close to max the machine resources, especially on utilising all CPU cores in parallel, the freezes and then after a period just shuts down in a non graceful manner!. Only a hard reboot will fix it. Additionally, when connecting using Cursor/ssh, it is common for the connect to freeze up and I have to again hard reboot the device to get it back working. It gets super hot also which I’m sure will not be great for the machine long term. Not great reliability for the price! (and I have 4 of them waiting to build into a cluster!)

Hi Allen, certainly not the behavior we are expecting. Do you have a minimal repro that i can take to engineering to review?

I will make a mock repo that reproduces and then share.

thanks for your help. I got Field diagnostics and it failed with GPU Stress test. like below
[cmd-3] Error Code = 020000021139 (Acceptable temperature limits exceeded or the thermal sensor is broken or miscalibrated)

I will ask manufacturer

Can you please DM me your Diagnostics logs? I’d like to share with quality team. Thanks.

Hi,

only getting back to this as was travelling - it turns out the problem is that at least one of the 4 x devices I purchased are running too hot on idle and more so under load. I am having a pretty poor experience trying to get these taken back under warranty by the reseller in Spain where I purchased them (€20k spent) … is there a way I can escalate this via NVidia before resorting to having to send legal letters?

Thanks,

Allen.

Piling up here. Purchased DGX on day 1. Was crashing sometimes. Now it crashes all the time. Running gpu_burn is enough to show that this is simply a thermal issue. The unit heats then dies (non gracefully). Seems like a bad thermal design.

developer@sparky

:

~/github/gpu-burn

$ ./gpu_burn 300
Using compare file: compare.fatbin
Burning for 300 seconds.
GPU 0: NVIDIA GB10 (UUID: GPU-cf0ef592-0423-ee7e-c6ee-65276769d5a0)
cuInit returned 0 (no error)
Initialized device 0 with 124610 MB of memory (11610 MB available, using 10449 MB of it), using FLOATS
Results are 268435456 bytes each, thus performing 38 iterations
10.7% proc’d: 532 (18765 Gflop/s) errors: 0 temps: 83 C
Summary at: Mon Jun 15 09:27:07 AM CEST 2026

21.0% proc’d: 1026 (16842 Gflop/s) errors: 0 temps: 78 C
Summary at: Mon Jun 15 09:27:38 AM CEST 2026

31.3% proc’d: 1520 (17605 Gflop/s) errors: 0 temps: 85 C
Summary at: Mon Jun 15 09:28:09 AM CEST 2026

38.7% proc’d: 1862 (17127 Gflop/s) errors: 0 temps: 85 C

It would probably be quicker to remove the case, add big fans and they will most likely work fine. Identical experience with NVIDIA here in terms of (lack of) support.

“remove the case, add big fans” … true … thinking about a harness that includes fans and possibly exposing the base of the device … clearly looking at some of the other OEM design decisions the original is not the best design for thermal release…