Critical Issue:
Temperature thresholds are completely wrong:
GPU Shutdown T.Limit Temp : -5 C
GPU Slowdown T.Limit Temp : -2 C
GPU Max Operating T.Limit : 0 C
Because of these wrong thresholds, the driver believes the GPU
is always overheating, even at idle (actual temp: 32~38°C).
This permanently activates SW Power Cap, and causes abnormal
power readings far exceeding the rated limit:
Power Draw (idle) : ~1200W
Power Draw (loaded) : ~1219W
Current Power Limit : 600W
pviol : 100% (always)
GPU clock is throttled to ~510MHz instead of 3090MHz.
Actual GPU performance is reduced to about 1/6.
GPU is completely unusable for production workloads.
nvidia-smi dmon output during full workload (SM=100%):
Hi, has this been solved in some way?
The VBIOS for your GPU would need to be provided by the distribution partner we sold the GPU to, so most likely like PNY or TDSynnex…
Though I doubt a VBIOS update would fix this, since the temp readings are NOT a bug/mistake; as outlined in the pointer Markus gave…
I’m still interested in the performance/clock degradation you described, it sounds similar to this posting:
Any chance, you could send me a full nvidia-smi log of the system, more logs and tests maybe..? many thanks
-Frank
At 100% CUDA load, the card reports 600 W, asserts clocks_event_reasons.sw_power_cap, and throttles to 652–675 MHz at 35–37 °C.
Thermal and hardware slowdown flags remain inactive; no Xid, ECC, or PCIe
errors occur. Suspected onboard power-sensor/PMIC telemetry failure.
Also, I noticed that the idle power is surprisingly high 200-300 W without any workloads, which is also mentioned in other threads.