Quick Test Results Regarding DGX SPARK Temperatures

Since there has been renewed discussion regarding DGX SPARK temperatures lately, I decided to run some quick tests.

TL;DR :

  1. There is no significant difference in clock speed between idle and load states.
  2. Therefore, monitoring the clock speed alone is insufficient; you must also verify the power consumption.
  3. The exhaust temperature can reach at least 60°C under load.

Note: There appears to be an unexplained phenomenon where temperatures rise to a certain level if no one is logged in locally, and only begin to drop once a local login occurs. The cause remains unclear.

I am using an Asus GX10. I am not utilizing the X-7 connect, HDMI, or wired LAN. Instead, I am using Wi-Fi, Bluetooth, and all available USB-C ports (for power, DP-Alt mode, an external SSD, and an RF wireless dongle).

I used the following script to monitor time, power draw, temperature, and clock speed every 5 seconds via nvidia-smi:

#!/usr/bin/env bash

while true; do
  timestamp="$(date '+%Y-%m-%d %H:%M:%S')"
  read -r power temp clk <<< "$(nvidia-smi --query-gpu=power.draw,temperature.gpu,clocks.current.graphics --format=csv,noheader,nounits 2>/dev/null | head -n1 | awk -F', ' '{print $1" "$2" "$3}')"
  printf "%s | POWER %4sW | TEMP %2sC | CLOCK %4sMHZ\n" "$timestamp" "$power" "$temp" "$clk"
  sleep 5
done

Logs:

1. Model loading via llama.cpp from an idle state

2026-05-27 08:38:57 | POWER 10.46W | TEMP 45C | CLOCK 2411MHZ
2026-05-27 08:39:02 | POWER 10.41W | TEMP 46C | CLOCK 2411MHZ
2026-05-27 08:39:07 | POWER 10.40W | TEMP 46C | CLOCK 2411MHZ

2. Inference after model loading is complete (Long Context)

2026-05-27 08:41:24 | POWER 10.16W | TEMP 45C | CLOCK 2411MHZ
2026-05-27 08:41:29 | POWER 10.14W | TEMP 45C | CLOCK 2411MHZ
2026-05-27 08:41:34 | POWER 67.24W | TEMP 59C | CLOCK 2392MHZ
2026-05-27 08:41:39 | POWER 67.46W | TEMP 61C | CLOCK 2392MHZ
2026-05-27 08:41:44 | POWER 70.07W | TEMP 63C | CLOCK 2392MHZ
2026-05-27 08:41:49 | POWER 71.12W | TEMP 63C | CLOCK 2392MHZ
2026-05-27 08:41:54 | POWER 68.85W | TEMP 64C | CLOCK 2392MHZ
2026-05-27 08:41:59 | POWER 73.95W | TEMP 68C | CLOCK 2385MHZ
2026-05-27 08:42:04 | POWER 75.65W | TEMP 69C | CLOCK 2385MHZ
2026-05-27 08:42:09 | POWER 76.26W | TEMP 70C | CLOCK 2385MHZ
2026-05-27 08:42:14 | POWER 77.47W | TEMP 70C | CLOCK 2385MHZ
2026-05-27 08:42:19 | POWER 78.37W | TEMP 73C | CLOCK 2379MHZ
2026-05-27 08:42:24 | POWER 78.99W | TEMP 72C | CLOCK 2379MHZ
2026-05-27 08:42:29 | POWER 80.08W | TEMP 72C | CLOCK 2379MHZ
2026-05-27 08:42:34 | POWER 76.79W | TEMP 73C | CLOCK 2385MHZ
2026-05-27 08:42:39 | POWER 82.61W | TEMP 75C | CLOCK 2379MHZ

3. Second inference following the first round (Short Context)

2026-05-27 08:45:51 | POWER 11.06W | TEMP 51C | CLOCK 2411MHZ
2026-05-27 08:45:56 | POWER 11.06W | TEMP 51C | CLOCK 2411MHZ
2026-05-27 08:46:01 | POWER 10.99W | TEMP 50C | CLOCK 2411MHZ
2026-05-27 08:46:07 | POWER 57.94W | TEMP 57C | CLOCK 2476MHZ
2026-05-27 08:46:12 | POWER 34.60W | TEMP 57C | CLOCK 2476MHZ
2026-05-27 08:46:17 | POWER 34.94W | TEMP 58C | CLOCK 2476MHZ
2026-05-27 08:46:22 | POWER 29.81W | TEMP 58C | CLOCK 2405MHZ
2026-05-27 08:46:27 | POWER 29.89W | TEMP 58C | CLOCK 2405MHZ

As shown in the logs, when the GPU is engaged in inference, power consumption spikes, which naturally leads to a rapid increase in temperature. Even after computation stops, power consumption remains around 10W with clock speeds staying near 2400MHz. Therefore, if you feel that performance/computation speed is not meeting expectations, I recommend checking whether the actual power draw increases correctly during tasks, rather than looking solely at the clock speed.

Regarding the exhaust temperature, I took a thermal image of the exhaust after some usage. Given that this was captured shortly after the fans ramped up for a few seconds, it can be inferred that the exhaust temperature reaches approximately 60°C.

Lastly, regarding the phenomenon where temperatures rise while idling at the login screen: please refer to the following posts for more information. This may not affect those who use local logins, but I recommend that users working exclusively via remote access look into this.

does anyone try to figure out the difference of llama.cpp and vllm trigger fan speedup machanism?
because I was using llama.cpp for models and the fan speedup like aircraft. then I use vllm the fan keep quiet till machine so hot and crash.

I’m currently using llama.cpp, but it was pretty much the same when I previously used vLLM via Docker to run Gemma 4. That said, since versions can vary, it might not be exactly the same vLLM setup. Also, I’ve lowered the clock speed to 2000MHz as shown below; this has resolved the overheating issue with almost no noticeable drop in performance.

sudo nvidia-smi -lgc 0,2000