What are normal temps under load? Is 94.6c too hot?

I was running some benchmarks locally and noticed the temps much higher than usual. It’s a Founders Edition Spark. Ambient temp is about 25 degrees C.

danny@toad:~$ paste <(ls /sys/class/thermal/thermal_zone*/type) <(cat /sys/class/thermal/thermal_zone*/temp)
/sys/class/thermal/thermal_zone0/type   94600
/sys/class/thermal/thermal_zone1/type   68800
/sys/class/thermal/thermal_zone2/type   68800
/sys/class/thermal/thermal_zone3/type   69200
/sys/class/thermal/thermal_zone4/type   68700
/sys/class/thermal/thermal_zone5/type   94600
/sys/class/thermal/thermal_zone6/type   71600

They’ve been fluctuating a bit, but spending a lot of time between 90 and 95.

The Spark is raised up on a metal monitor riser (so there is lots of clear air on all sides) and has a desk fan blowing air across it!

i’d say thats too hot you need a external cooling fan or else your likely to get a thermal shutdown

That’s pretty standard maxes for the CPU temps under benchmark load. GPU max temps should be ~10 C under that (Founder Edition).

It feels hot to me, and I already have a desk fan on a high setting blowing over it. Although it didn’t shut down overnight at these temps. If they are not normal, I don’t understand why the fan isn’t ramping up higher.

Yeah, GPU is usually around 10 degrees lower. What’s odd though is that the CPU usage is almost 0%, it’s all GPU work. Seems odd that the GPU is the “coolest” part 🙃

I wish NVIDIA would tell us what those thermal zones are! There are seven zones in the sysfs but all are generic acpitz types.

In Danny’s case zone 0/5 are about 20 degrees hotter than the rest. Are those zone readings from the GPU or CPU+GPU? Who knows! Paging @eugr_nv 🙂

There’s tegrastats utility but won’t reveal more than what can be read in sysfs. I copied the binary from a Jetson Orin Nano. Here’s sample output on an idling Spark:

07-19-2026 08:24:21 RAM 1431/124610MB (lfb 91x4MB) SWAP 0/16384MB (cached 0MB) CPU [0%@2132,0%@2262,0%@2106,0%@2132,0%@2132,0%@3146,0%@3068,0%@3250,0%@3432,0%@7436,0%@2314,0%@2340,0%@2106,0%@2054,0%@2106,0%@3094,0%@3042,0%@3094,0%@3510,0%@3588] acpitz@34.8C acpitz@36.9C acpitz@36.9C acpitz@34.8C acpitz@34.8C acpitz@35.8C acpitz@33.8C
07-19-2026 08:24:22 RAM 1432/124610MB (lfb 91x4MB) SWAP 0/16384MB (cached 0MB) CPU [0%@2314,0%@2236,0%@2262,0%@2106,0%@2080,0%@3068,0%@3094,0%@3146,0%@3406,0%@6162,0%@2340,0%@2340,0%@2340,0%@2262,0%@2080,0%@3094,0%@3094,0%@3146,0%@3588,0%@3510] acpitz@34.8C acpitz@37.1C acpitz@37.1C acpitz@34.8C acpitz@34.8C acpitz@35.8C acpitz@33.9C

Have you tried capping the GPU clock? Lots of opinions out there that you will lose only a bit of performance by applying this.

See this article: Your DGX Spark Is Cooking Itself | Wild Pines AI

I hadn’t yet, but I was aware of it and might do it. Although I was hoping to see a response from nvidia about what temps are normal and when we should be concerned.

My Spark crashed and powered off again this morning while benchmarking Laguna s.2. So I’ve decided to give the throttling a go. I’ve set it 300-2200.

The temps are currently around 78 and 68 which seems better but since they fluctuate a bit it’s hard to be sure.

I’m still not sure if it’ll really solve the issue if the fan curves are the same, and therefore the fans will just run more slowly at the lower temps and let the heat build up, but we’ll see. It only took a few hours before it crashed earlier (although I’m also not sure it wasn’t fluke and might’ve got through a re-run without).

ASUS Ascent GX10 Cooling Stand for 140mm Intake Fan by srinathh | Download free STL model | Printables.com Never seend temps like yours. Attached is simplest and fastest solution to print I found.

Too hot IMO. I have the GX10 and have reached low 90s during benchmarking too. But I’ve successfully lowered the temps by ~20*C with the addition of a Noctua 3000 rpm PPC. Not sure if this would help others, but I can upload it to github including the temperature control software too

Yeah, I had intended to print something like this from a contributor here that sucked air through it, but hadn’t gotten around to it yet (and since it has so much clear air on every side, I kept telling myself I shouldn’t need anything).

Since restricting the GPU to 2200mhz it hasn’t gone above 80c (while running the same benchmark where it crashed this morning), but I don’t know if that’s because of the change or just variance (for ex. some caches are shared so it might not have done the same startup compilation work this time as it did the first time).

Mine started this way - only a few failures, so capped the clocks, then a few more failures. It was progressive and makes me believe the early builds used a thermal paste which degrades with use. The forum is seeing more of these reports as later buyers have accumulated hours. Heavy prefill is where I see the most thermal buildup.

What was the outcome - are you still using it, or did you get it swapped? Or did you re-paste it? (I’m not sure how comfortable I’d be doing that, at least not unless it was out of warranty and really bad!)

My opinion is that warranty shouldn’t be rejected as this was only a handful of screws and ultimately repaired an OEM issue, but I won’t be holding my breathe either.

With the cabinet door closed the units idle mid 40s. Single user sessions generally hang out in the 50s/60s with a heavy prefill maybe getting to 70. 20 concurrent requests all performing prefill can briefly spike 82 but will primarily hang out in the 70s. Heat is generated no matter what. It seems like the workload drives the level of cooling solution. Since going this route I haven’t had a single shutdown and I’ve been running 24/7 for weeks with this style of heavy prefill.

Is 94.6c too hot?

Yes. Put a fan direct in front of it. Be extremely cautious of plastic wires and combustable materials around the unit.

I already had a fairly powerful desk fan (Meaco Sefte) blowing over it when these temps were seen. Things have been better (mostly below 80, but sometimes hitting 83) since I added the gpu clock limit, although I’m also not certain that the workload was the same (because of caches).

I was somewhat hoping nvidia would confirm what are expected/safe temps.. seems like a reasonable request, particularly given we can’t do anything to control the fans. It doesn’t seem normal to have to babysit a machine like this.

https://forums.developer.nvidia.com/t/cooler-gb10-temps-almost-no-performance-lost/372662

Give that a test :)

I’m currently running with clocks set to 300-2200. But I’m trying to understand if I should be concerned with the temps I’m seeing (for ex. some people have mentioned a bad thermal paste job).

I don’t want to hide the issue and then go out of warranty if there’s something wrong that should be RMA’d 🙃

Definitely understand your point. If you read about GB10s temp issues in general (not just nVidia DGX Spark), you’ll see people reporting sub-par thermal paste jobs. And that’s actually true to (maybe) 90% of the products that use thermal paste out there in the industry, at least consumer products.

But on the other side, even with a good/decent thermal paste application, you’ll see throttling at stock clocks speeds and ~double the power consumption, so lowering to 2000 or 2100 will save you watts, temperature and headaches.

I’ve done sweeps from 1500-2400mhz and found the sweet spot to be 2000, but also 2100Mhz is still very decent efficiency-wise, at 2150Mhz the diminishing returns in terms of Wattage per Tok/s peformance (for both decode and generation) are noticeable already. The optimization engineer in me tells me “never run above 2100” :)

My theory:

My guess is that the last few hundred Mhz were required to reach that 1PFlops of the advertised performance: Check this post GB10 really does hit ~1 PFLOP NVFP4 (2:4 sparse) — measured, with an open-source tool to reproduce it

When you downclock to the ideal Power/Performance figures, the below-Petaflop drop is quite big , actually but since inference is more memory bandwidth-bound than everything else, you will barely (if at all) see a drop when using your LLM. Look at this test from that same thread above:

“780 TeraFlops sounds less cool than 1 PetaFlop. so let’s make this super hot to the silicon limit to hit the number” :)

I could believe that - although I don’t understand why they wouldn’t just ramp the fans up more. It seems like the fans are barely doing anything, and I’d rather a bit more noise than too much heat (or having to drop the clocks).

Ah, that’s useful to know. At 2200 I’ve only seen it go a little over 80 once, but I suspect 2000 isn’t going to change much for me (things are already slow enough that most of what I use it for will be async/background work anyway), so I might do that.

Although my initial concern remains - if the paste is bad and it might continue to get worse over time (even when clocked down to 2000), I’d rather have nvidia fix that in warranty than have to do it myself 🙃