Blackwell Pro 6000 MHz degrading

After 6 months of 80%+ utilization of these workstation cards, I notice they are beginning to fail. Specifically the max Mhz at 600W is slowly decreasing (if not under load, nvidia smi reports 2500MHz+). I’ve sent 2 back to the vendor for investigation but also notice 2 more starting to exhibit issues - These show up noticeably in training graphs as all the the “healthy” GPUs are waiting for the failing GPU to sync gradients.

Anyone else seeing this? Temps seem fine, no errors on the bus, just Mhz slowly degrading to 500MHz worst case under load (600W)

Further debugging: I’ve moved these over to another server and they still throttle - this pretty much rules out PSU/motherboard issues. This is very similar to the issue reported here:

I’m currently at 800Mhz and 1400Mhz but there is a distinct downward trend (I log the power curves through wandb). The other healthy gpus average about 2500 Mhz

I also am getting max 270watt out of 600w as HW Power brake is Active. Did you solved it ? Drivers ? or GPU BIOS firmware ?

Not confirmed, but the vendor said updating VBIOS has fixed the issue - I’ll check later this week to see if its true. They also said one card failed to take the VBIOS and basically bricked so they have to send it back to the manufacturer. Sounds like updating the VBIOS is dicey / should not be done by the enduser.

I have found the issue what was causing it jn my case.

ain my case new soldered risers with rtx6000 were having this problem while 5090 gpus on the same risers had no issues with gull 600w power.

Anyway will swap risers and now it works as intended. Maybe this is vbios issue but overall i was able to fix this myself. Thanks

Hi - I’m having the exact same issue on 2 cards on 2 different machines.
All was fine - and now they both go at 620 MHz instead of 2500Mhz.

Were you able to fix this?
When I look at logs I can see that there is SW Power Cap : Active.

Should I try to update the VBIOS? Is that the fix?

No, the VBIOS flash didn’t fix the issue. I’m currently working with NVIDIA support to debug further. Happy to share updates as I learn more.

Same issue for me as well. Clock at 600Mhz when power hit 600 watts.
Temperature sits at 35-40 degrees C.
It was working fine for the previous 2 months, beats my 5090 in inference/training speed reaching in the 75 degree C.
Tried on 3 different machine, Windows and Linux, same story.

Yeah that’s really odd 🤔

If it’s dropping to 600 MHz while pulling that much power and temps are still low, it doesn’t sound like thermal throttling at all. More like some kind of power/firmware or driver issue.

The fact it happens across different machines and OS makes it even more suspicious, might be worth checking BIOS/firmware updates or even considering a hardware fault.

Hello!
I am currently experiencing similar issues with a 6000 PRO Blackwell
In stress test the video card stays at around 700 MHz but shows a power draw of 600W
I check the power consumption on grid and that power draw is false
Also the fans stay at 30% and the temps are around 45 Celsius

Did you fixed your problem, or got any resolution from your vendor / Nvidia?

Hi everyone,

I’m running two identical RTX PRO 6000 Blackwell Workstation Edition cards in a dual-GPU setup (Gigabyte TRX50 AI TOP + Threadripper PRO, Ubuntu 24.04, driver 590.48.01, VBIOS 98.02.81.00.07 on both).

One of the cards started showing clear degradation:

  • In P8 idle (0 % GPU utilization, memory clock 405 MHz) it constantly draws ~75–80 W.

  • The second identical card in the exact same conditions draws only ~9–10 W.

I did a full physical GPU swap between the two PCIe slots.
The issue follows the physical card — high idle power and degraded performance stayed with the same GPU. The healthy card now idles normally (~9.6 W) in the previously problematic slot.

D2D / VRAM bandwidth (Vast.ai test container + CUDA bandwidthTest --memory=pinned --mode=range):

  • Affected card: very low and unstable — drops to ~300 GB/s, mostly circulating around 700–1000 GB/s

  • Healthy card: stable ~1.35–1.40 TB/s (1400 GB/s)

H2D/D2H PCIe bandwidth is normal and identical on both (~56 GB/s). No ECC errors, no XID errors, no retired pages.

Has anyone else seen this exact pattern (high P8 idle power + severe D2D/VRAM bandwidth drop on one card only) on RTX PRO 6000 Blackwell?

Quick update on the problematic RTX PRO 6000 Blackwell (the one that stays at ~75–80 W in P8 idle while the identical second card is at ~9–10 W).

I just confirmed a very important new symptom:

When the affected card is under real client load:

  • nvidia-smi / nvtop reports ~600–620 W (P1 state, memory clock 13365 MHz)

  • However, measuring the entire server at the wall/PDU shows only ~250 W total

After subtracting CPU + system (~100 W) and the healthy card (~10 W), the problematic card is physically drawing only ~140–150 W, while the driver thinks it is pulling 600+ W.

This is a clear faulty / miscalibrated power sensor (PMIC) on the GPU board.

Summary of all confirmed symptoms (unchanged):

  • High idle power in P8 (~75–80 W vs ~9–10 W on healthy card)

  • Severe and unstable D2D/VRAM bandwidth (drops to ~300 GB/s, mostly 700–1000 GB/s instead of 1.4 TB/s)

  • Problem follows the physical card after both PCIe slot swap and power cable swap

  • Clean ECC, no XID errors, no retired pages

Has anyone else seen this exact combination on RTX PRO 6000 Blackwell — especially the huge discrepancy between reported power (600+ W) and actual measured power (~150 W)?

Any similar experiences or advice before the RMA process would be greatly appreciated.

Thanks!

Update / Additional information:

I have two identical RTX PRO 6000 Blackwell Workstation Edition cards, but I bought them one month apart.

The first card (the one that is now problematic) already showed higher idle power right from the beginning:

  • Bad card in P8 idle: 20–30 W

  • Good card in P8 idle: 6–8 W

At that time I didn’t pay much attention to it and thought it was normal variation or that the system was simply using one card more often. Only after putting both cards into the same server did I start monitoring them closely.

Now it’s clear that this was already an early sign of degradation on that card. Over time the problem worsened significantly — the D2D/VRAM bandwidth dropped from ~1.4 TB/s down to 300–1000 GB/s (performance loss of 50–75 % in some runs).

All other diagnostics (full physical GPU + power cable swap, clean reboots, driver resets, etc.) confirm the issue follows the physical card that had elevated idle power from day one.

Has anyone else noticed elevated idle power (even 20–30 W) on a new Blackwell PRO 6000 that later turned into severe degradation?

Not sure if I have a bad card as well, It’s idling around 12s (sometime 11W), where as my RTX 5090 FE idles around 6W.

I was only able to get around 2000MHz at 380W (840mW), will monitor if it continues to degrade. Temp has been good at 62C max (90% fan at that temp, manual fan profile)

Same issue on my side.

Running a hashcat benchmark shows one GPU dropping to 40% clock rate related to remaining GPUs.

Every 1.0s: nvidia-smi --query-gpu=gpu_name,clocks.sm,clocks.mem,temperature.gpu,power.draw,clocks_throttle_reasons.sw_thermal_slowdown,clocks_throttle_reasons.sw_power_cap,pcie.link.width.current --f… genoa4: Wed Jul 8 14:41:01 2026

name, clocks.current.sm [MHz], clocks.current.memory [MHz], temperature.gpu, power.draw [W], clocks_event_reasons.sw_thermal_slowdown, clocks_event_reasons.sw_power_cap, pcie.link.width.current
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 2760 MHz, 13365 MHz, 58, 599.90 W, Not Active, Not Active, 16
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 1237 MHz, 13365 MHz, 43, 599.99 W, Not Active, Active, 16
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 2760 MHz, 13365 MHz, 59, 599.99 W, Not Active, Not Active, 16
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 2790 MHz, 13365 MHz, 58, 599.98 W, Not Active, Not Active, 16

Has anyone received feedback regarding a RMA process?