Understanding Optimal GPU Temperature and Default GPU Fan Curve (NVIDIA RTX 6000 Ada)

Hi Lukas, thanks for all these good questions…

the actively cooled workstation GPUs are aimed to be used in (under-the-desk) workstations, and as that, NOISE=fan rpm needs to be a limiting factor.
Your (not unique) usecase of likely favoring lower temps over more noise we currently DO NOT COVER, so we do not have a way for you to allow to ‘just run the fans faster’ = more noisy…
Your best approach for such configs is to provide better (cold) airflow TO the GPU, so it would stay cooler… (for the server version of our GPUs = passively cooled, all such control is in the central element of the server chassis cooling then…).
With the noise being a critical parameter for actively cooled cards, each product has its own value for the (noise limited) max. rpm, which we show as a percentage of the vendor specified absolute max rpm of only the fan. So yes, like 60% is correct (a).

(b) basically NO, assuming we are not talking like thousands of GPUs per month? Work on the chassis cooling, which will impact the GPU cooling logic…
(c) between some 2 generations of our GPUs we changed the temp mechanisms (and readings) substantially, I would assume this is a glitch resulting from that change. Unless it still repro’s with current drivers (580/590 branch), I don’t see a chance to fix/backport this (only for older drivers, your 535 branch..) I’m sorry…
(d) might also be a result of old driver, changed temp logic and readings, but not sure… again, we won’t add any other but security fixes to old r535, so for a test, if you could check how all this looks and works with a r580/590 driver?

(e) one intention of the new ‘logic’ was to have homogenous (delta)temp values for all our different SKUs, and only store a GPU specific absolute temp with the GPU. That’s why we now show t_limit, as the delta you are running from the SW slowdown temp, which varies per GPU… t_avg being the average of all internal GPU temp sensors (likely shows as current temp)…

The target temp is a range, we want the GPUs to stay within over a period of time, short spikes of individual sensors spiking higher are fine….
This in your case reads to: you run at 88°, which is 4° away from SW slowdown, which for this sku would then be 92°. HW slowdown would be -2°, so 2° above that 92°, with a HW SHUTdown kicking in at -7°. We no longer show you fixed values (that would be different per GPU), but only your delta to that, and your absolute average temp at any time…
[server temp and fan control have much easier logic now, focusing only on that t_limit value, and keeping it positive…] make (some) sense?
(f) as we don’t offer you any modification of temp and rpm, I wonder how you even change target temp ?? (though I’m not really familiar what tolls we offer on Linux actually…).

The way we qualify and certify the (active cooled) products, we guarantee proper function, and accept RMA for failures, (with a given fan-inlet air temp, the chassis vendor needs to guarantee to the GPU, for any situation and thermal profile the OEM certifies its chassis for… )
As that, we can’t really offer for users to modify any of the qualified and hence guaranteed presets..

Lukas, can you tell (DM) me, what type of appliance and usecase you run, and roughly the number of GPUs…? I’m still collecting ‘needs’, and reporting them to product owners and engineering leads…

many thanks

-Frank