DGX Spark: Boot To Idle then doing nothing & idle, the unit is unusually warm (45C)

I noticed consistently in the past 2 weeks, when I turn on my DGX Spark then left alone doing nothing (not logged on via desktop) for 2 hours, touching the unit’s top is unusually warm. After noticing it, I log on and temperature readout (via nvidia-smi) is 45C. It wasn’t like that before. Last Feb and months prior, power on boot to idle then left it alone doing nothing, the temperature is 37C. Touching the unit’s top is not warm.

I think a software update (via DGX Dashboard) this month or last month, something has changed. I recall after the much touted software update in Jan which reduced the average power of DGX Spark (by auto-power down the ConnectX), the unit did not have this problem.

Pls advise how to resolve this. The unit has no docker container running and ollama has no model loaded. Dashboard reports 0% GPU utilization and 5.55 GB of System Memory used.

Pls advise how to fix this.

Without knowing the ambient temperature in both scenarios, that different could mean nothing. 45C at idle isn’t warm. Could be colder? yes. But doesn’t scream ’ issues’ to me.

For comparison – Ambient here is 25C and my DGX Spark at idle with model loaded is CPU 40C and GPU 38C, 11.8w. A higher ambient might track higher. I have mine standing on its end with a USB fan. Performing a quantisation run it got up to 95C – hence I got the fan. No thermal shutdowns, power problems or anything like that since I have had it thought.

I have the same problem with temperatures up to 50C. What I noticed is that, as soon as I attach a monitor to the Spark and log in with my user, even if the machine is then left to do nothing, temperatures decrease.

I ran a cron job to monitor temperatures for a couple of days, as suggested in this thread. And I also annotated the log when I rebooted the machine, or stopped the gdm and gnome-remote-desktop services.
For what I observed, stopping the above services did not improve idle temperatures.

I understand that these temperatures are not high per se. But i have 3 other spark machines in my office. I set up all of them and gave them to colleagues, and none of them gets as hot at idle as mine. I did not run the cron jobs on those other machines, but when I come in the office in the morning after a night where all machines were left idle, mine is the only one that is hot to the touch (and around 50C).

Hi azenuser, I have the same observation on my unit. After I do local login, the issue goes away. If I don’t do local login (ie do remote login via ssh), this issue appears.

I’m using an Asus GX10 and I’ve noticed the exact same thing. The temperature rises even at idle when I’m only logged in via SSH, but it drops back down once I log in locally.

Did anyone solved this in any way? Having similar issue with a Gigabyte AI TOP. I have 2 pcs (same mfg date) and just one has this issue. The other one stays at 38-39 degrees while the hot one gets to 50 degrees. The fan does not start on the hot one unless under load.

For anyone experiencing higher temps at idle (before logging in), please follow these steps to collect temperature data while the unit is heating before login.

  1. add a script in your home directory (change path to your username)
    nano /home/yourusername/monitor_temps.sh

  2. Content is

    monitor_temps.sh (5.7 KB)

  3. make it executable (change path to your username)
    chmod +x /home/yourusername/monitor_temps.sh

  4. open and edit cron job config to execute this script every 10 min

  • open crontab
    crontab -e

  • add this single line at the bottom (change path to your username)
    */10 * * * * /home/yourusername/monitor_temps.sh >> /home/yourusername/temperature_monitor.log 2>&1

  1. log out and wait 20 more minutes after spark getting hot
  2. log in and collect /home/yourusername/temperature_monitor.log and provide to us
  3. remove the single line in step4 to restore the change

Hi @aniculescu. I’ve collected logs using your monitor_temps.sh script from two idle DGX-Sparks over 24 hours. One is running very hot to touch (around 52 degrees C), the other is running slightly warm to the touch (around 44 degrees C). The hot Spark is running the most recent software version (7.5.0) I’ve installed, while the cooler Spark is running a much older version (7.2.3) [versions from running cat /etc/dgx-release]. What’s the best way for me to send you these logs for diagnosis?

Edit: Looking at the conversation in this other thread, neither Spark is connected to a monitor and it appears that gnome-remote-desktop is disabled on the hot spark and running on the cooler spark. Both sparks are showing as in p8 GPU power mode in nvitop and both are drawing about 4W in the GPU box of the nvitop stats.

Edit2: I’ve sent you the logs via a forum DM.

nvidia-bug-report-72d1.log.gz (417.1 KB)

temperature_monitor_72d1.log (83.6 KB)

temperature_monitor_12ce.log (83.5 KB)

nvidia-bug-report-12ce.log.gz (355.6 KB)

Hi @aniculescu . I have included logs for two identical DGX Spark computers. When I say identical, I mean in terms of hardware, operating environment, software, everything has been controlled for. As far as I can ascertain anyway. One (72d1) idles significantly warmer than the other (12ce) however, to the tune of 10+ degrees C. I’m hoping you (or anyone else) might see something I didn’t. I’m 99.9% certain this is a hardware issue of some sort. I’ve heard some of these computers left the factory with less than expertly applied thermal compound. Regardless, thanks for the help, looking forward to hearing from you.

EDIT: Now that I’ve slept and looked at this with fresh eyes it seems that the hot computer might have a firmware issue. Looks like the GPU got pinned to P0 intermittently on account of some failure at that level. Not that I’m an expert or anything. Continuing to work on it.

Hi, I have noticed the same issue with you as well.
My ambient temperature is around 21C. SSH into the Spark, ran sensors and nvidia-smi, and got the GPU temp at 50C, and GPU utilization is 0%. But when I logged in with gnome-remote-desktop, GPU utilization jumped to 10%, and the temperature started dropping down to ~30C. Then, when I stopped the remote desktop login and went back to SSH, the temperature started to climb slowly back up to ~50C.
I think this might be related to the fan profile.
I’m actually really worried if 50C idle is safe for the Spark long-term since I’m intending to let it run 24/7 headless. So I have two options, and I would love to have anyone’s advice on which is the best course of action:

  • Leave the Spark running headless 24/7 with 50C idle
  • Plug a monitor in and have the monitor stay on 24/7 so that the fan is keeping the idle temp down to 30C

I have now captured a clean A/B result that matches the local-login/display-dependent behavior described here.

I have two identical ASUS Ascent GX10 units with matching kernel 6.17.0-1029-nvidia, driver 580.173.02, DGX OTA 7.5.0, BIOS and EC/UEFI/PD firmware. One unit can sit at GPU 55–58°C / ACPI 58.5–61.7°C while P8, 0% utilization and around 3.8–4.0 W; the identical control stays at 34–35°C.

The affected unit cooled to GPU 36–37°C after HDMI initialization and a local X11 login, even without keyboard or mouse activity. I changed only GNOME’s blank-screen timeout from five minutes to Never. I then physically unplugged HDMI while keeping the same user session active and IdleHint=no. Over the next five minutes GPU temperature rose from 36°C to 45°C and max ACPI from 39.8°C to 47.9°C, while power remained only 3.48–3.71 W.

The control unit remains cold while fully headless at the GDM greeter, so this is not a universal headless-mode behavior. It appears to be a display/EC state interaction affecting one unit, possibly combined with a fan or EC hardware difference.

@aniculescu @NVES, could you advise whether NVIDIA/ASUS engineering is tracking this and whether the affected unit should be RMA’d? I have sanitized CSV telemetry and private nvidia-bug-report.sh archives from both units.