PRO 6000 Blackwell WS: SW Power Cap ~600 MHz @ 600 W / 35°C — fix or RMA?

RTX PRO 6000 Blackwell (Workstation Edition, 96 GB) — “SW Power Cap” pins core to ~580–650 MHz under load at a FALSE-looking 600 W / 35 °C. Driver-, kernel-, and BIOS-independent. Fix, or RMA?


SUMMARY

Under any sustained 100% compute load, my RTX PRO 6000 Blackwell Workstation
Edition throttles its core (SM) clock to ~580–650 MHz (max 3090 MHz) with
clock event reason “SW Power Cap” active continuously. nvidia-smi reports the
board at ~600 W, but GPU temperature stays only ~34–36 °C — which is not
plausible for genuine 600 W dissipation on this cooler. Result: roughly
~6–11 TFLOP/s FP32 instead of the ~110+ this SKU should deliver (~1/10–1/20).

The card ran at full performance when new and has since regressed on an
UNCHANGED driver stack. I have ruled out driver, kernel, motherboard BIOS,
power limit, and cable/sense-pin derating. This looks like on-board power
telemetry reading false-high and continuously tripping SW Power Cap.

I am asking whether this matches the known power-sensor degradation pattern
(threads below), and whether any fix exists short of RMA.


SYSTEM

  • GPU: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB
  • Device / board: 10DE:2BB1
  • VBIOS: 98.02.81.00.07
  • OS: Ubuntu 24.04 LTS
  • Driver: NVIDIA open kernel modules (reproduced on three branches — see below)
  • Power: 12V-2x6 / CEM5 path; power limit reports Default = Max = Enforced = 600 W
    (not derated to 450/300 W via sense pins)

THE SYMPTOM (measured under 100% compute)

Steady state under matmul / training:

utilization : 100%
SM clock : ~577–650 MHz (max 3090 MHz)
memory clock: ~13365 MHz (OK)
power draw : ~600 W (pegged at limit)
temperature : 34–36 °C
throttle : SW Power Cap ACTIVE continuously

Key observation: in idle gaps BETWEEN kernels (util drops), SM clock jumps to
~2600 MHz with no throttle reason — then collapses back to ~600 MHz the
instant sustained compute resumes. The GPU CAN clock high; the SW Power Cap
under load is what holds it down.

The “600 W @ ~35 °C” combination is the tell. Real 600 W should heat the
board well above this within seconds. Power reading appears false-high.

I can attach full:
nvidia-smi -q -d PERFORMANCE,POWER,CLOCK,TEMPERATURE
under load on request (happy to paste now if preferred).


WHAT I HAVE RULED OUT

  1. DRIVER — identical behaviour on three open branches:

    • 570.211.01-open : ~577 MHz / SW Power Cap / ~6 TFLOP/s
    • 580.105.08-open : ~607 MHz / SW Power Cap / ~11 TFLOP/s
    • 595.71.05-open : ~585–645 MHz / SW Power Cap / ~9 TFLOP/s
      (Blackwell requires open modules; expected.)
  2. KERNEL — reproduced across multiple 6.8.x kernels.

  3. MOTHERBOARD BIOS — updated to latest; no change.

  4. POWER LIMIT / CLOCK LOCKS — already at max (600 W).
    nvidia-smi -pl 600 and --lock-gpu-clocks have no lasting effect under load
    (SW Power Cap overrides).

  5. CABLE / SENSE — power limit remains 600 W (not sense-derated to 450/300 W),
    so the cable is advertising a 600 W-capable budget.

  6. ERROR FLAGS — no Xid, no ECC, no row-remap; HW thermal slowdown inactive.
    This is SW Power Cap, not HW thermal / HW power brake.

  7. VBIOS — already on 98.02.81.00.07 (past the early MIG update). I am not
    aware of any newer Workstation Edition VBIOS that addresses this, and
    partner-only flashes for MIG do not appear to fix the throttle pattern
    reported by others on the same VBIOS.


HISTORY (points to in-field degradation, not software)

  • Installed ~October 2025; full performance initially (580-open).
  • Then ran ~8 months on a single unchanged 580.105.08-open stack.
  • Over that period, performance regressed to the throttled state above
    with NO driver / kernel / hardware change during that window.
  • Regression on a fixed software stack is consistent with hardware
    degradation of power sensing (also consistent with elevated idle power
    reports from other users with the same fingerprint).

RELATED THREADS

Note: I am aware T.Limit deltas (e.g. Shutdown/Slowdown constants) are
by design and not absolute °C thresholds. My concern is specifically the
false-high power + permanent SW Power Cap under load, not a misread of T.Limit.


QUESTIONS

  1. Does this match the on-board power-sensor / false-telemetry degradation
    described in the threads above?

  2. Is there ANY Workstation Edition VBIOS / firmware / driver fix, or is
    RMA the only resolution path?

  3. Is there a software-only way to positively confirm false power telemetry
    short of a wall/PDU meter? (I can still arrange external power measurement
    if required for RMA.)

  4. What exact logs / tests do you need to authorize RMA?

Happy to post full nvidia-smi -q under load, dmesg | grep -i nvrm, and
pcie generation/width immediately if helpful.

Thank you for any guidance.

1 Like