RTX PRO 6000 Blackwell (Workstation Edition, 96 GB) — “SW Power Cap” pins core to ~580–650 MHz under load at a FALSE-looking 600 W / 35 °C. Driver-, kernel-, and BIOS-independent. Fix, or RMA?
SUMMARY
Under any sustained 100% compute load, my RTX PRO 6000 Blackwell Workstation
Edition throttles its core (SM) clock to ~580–650 MHz (max 3090 MHz) with
clock event reason “SW Power Cap” active continuously. nvidia-smi reports the
board at ~600 W, but GPU temperature stays only ~34–36 °C — which is not
plausible for genuine 600 W dissipation on this cooler. Result: roughly
~6–11 TFLOP/s FP32 instead of the ~110+ this SKU should deliver (~1/10–1/20).
The card ran at full performance when new and has since regressed on an
UNCHANGED driver stack. I have ruled out driver, kernel, motherboard BIOS,
power limit, and cable/sense-pin derating. This looks like on-board power
telemetry reading false-high and continuously tripping SW Power Cap.
I am asking whether this matches the known power-sensor degradation pattern
(threads below), and whether any fix exists short of RMA.
SYSTEM
- GPU: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB
- Device / board: 10DE:2BB1
- VBIOS: 98.02.81.00.07
- OS: Ubuntu 24.04 LTS
- Driver: NVIDIA open kernel modules (reproduced on three branches — see below)
- Power: 12V-2x6 / CEM5 path; power limit reports Default = Max = Enforced = 600 W
(not derated to 450/300 W via sense pins)
THE SYMPTOM (measured under 100% compute)
Steady state under matmul / training:
utilization : 100%
SM clock : ~577–650 MHz (max 3090 MHz)
memory clock: ~13365 MHz (OK)
power draw : ~600 W (pegged at limit)
temperature : 34–36 °C
throttle : SW Power Cap ACTIVE continuously
Key observation: in idle gaps BETWEEN kernels (util drops), SM clock jumps to
~2600 MHz with no throttle reason — then collapses back to ~600 MHz the
instant sustained compute resumes. The GPU CAN clock high; the SW Power Cap
under load is what holds it down.
The “600 W @ ~35 °C” combination is the tell. Real 600 W should heat the
board well above this within seconds. Power reading appears false-high.
I can attach full:
nvidia-smi -q -d PERFORMANCE,POWER,CLOCK,TEMPERATURE
under load on request (happy to paste now if preferred).
WHAT I HAVE RULED OUT
-
DRIVER — identical behaviour on three open branches:
- 570.211.01-open : ~577 MHz / SW Power Cap / ~6 TFLOP/s
- 580.105.08-open : ~607 MHz / SW Power Cap / ~11 TFLOP/s
- 595.71.05-open : ~585–645 MHz / SW Power Cap / ~9 TFLOP/s
(Blackwell requires open modules; expected.)
-
KERNEL — reproduced across multiple 6.8.x kernels.
-
MOTHERBOARD BIOS — updated to latest; no change.
-
POWER LIMIT / CLOCK LOCKS — already at max (600 W).
nvidia-smi -pl 600 and --lock-gpu-clocks have no lasting effect under load
(SW Power Cap overrides). -
CABLE / SENSE — power limit remains 600 W (not sense-derated to 450/300 W),
so the cable is advertising a 600 W-capable budget. -
ERROR FLAGS — no Xid, no ECC, no row-remap; HW thermal slowdown inactive.
This is SW Power Cap, not HW thermal / HW power brake. -
VBIOS — already on 98.02.81.00.07 (past the early MIG update). I am not
aware of any newer Workstation Edition VBIOS that addresses this, and
partner-only flashes for MIG do not appear to fix the throttle pattern
reported by others on the same VBIOS.
HISTORY (points to in-field degradation, not software)
- Installed ~October 2025; full performance initially (580-open).
- Then ran ~8 months on a single unchanged 580.105.08-open stack.
- Over that period, performance regressed to the throttled state above
with NO driver / kernel / hardware change during that window. - Regression on a fixed software stack is consistent with hardware
degradation of power sensing (also consistent with elevated idle power
reports from other users with the same fingerprint).
RELATED THREADS
-
“Blackwell Pro 6000 MHz degrading”
Blackwell Pro 6000 MHz degrading
Multiple users, same fingerprint; several report wall/PDU draw far below
nvidia-smi power; issue follows the physical card; VBIOS flash did not fix
at least one case. -
“Wrong temperature thresholds / SW Power Cap / 510 MHz”
RTX Pro 6000 Blackwell - Wrong temperature thresholds causes SW Power Cap, GPU throttled to 510MHz (1/6 performance) -
“5090 founders edition suddenly decreased performance”
5090 founders edition suddenly decreased performance in Ubuntu 24.04.3 nvidia driver 575/580
Same family pattern; support language quoted there: perceived false power
limit / hardware sensing error when SW Power Cap is Active at low temp
with clocks stuck near base.
Note: I am aware T.Limit deltas (e.g. Shutdown/Slowdown constants) are
by design and not absolute °C thresholds. My concern is specifically the
false-high power + permanent SW Power Cap under load, not a misread of T.Limit.
QUESTIONS
-
Does this match the on-board power-sensor / false-telemetry degradation
described in the threads above? -
Is there ANY Workstation Edition VBIOS / firmware / driver fix, or is
RMA the only resolution path? -
Is there a software-only way to positively confirm false power telemetry
short of a wall/PDU meter? (I can still arrange external power measurement
if required for RMA.) -
What exact logs / tests do you need to authorize RMA?
Happy to post full nvidia-smi -q under load, dmesg | grep -i nvrm, and
pcie generation/width immediately if helpful.
Thank you for any guidance.