HW Power Brake Slowdown on a pair of RTX Ada R6000

I have RTX 6000 Ada Generation dual-GPU server (Supermicro X12DAI-N6, dual Xeon Gold 5320, EVGA 1600W PSU) presented with severe LLM token generation performance degradation on dense models — approximately 3 t/s single GPU and 7 t/s dual GPU against an expected ~33 t/s single GPU for Qwen3.5-27B Q8_0. Additionally, running any GPU heavy like Alphafold 3 have the GPU stuck in P2 mode instead of P0

Here is system spec

Dual Intel Xeon Gold 5320
Supermicro X12DAI-N6
2TB RAM
Dual RTX Ada R6000
1600Watt EVGA PSU
X710 10G NIC

Ubuntu 24.04 Server LTS
580.126.09 Server Driver

Here is what I tried

  1. Upgrading system BIOS and BMC firmware to latest
  2. Removed T400 ( used to have a seprate T400 on the server)
  3. Moved X710 to a different CPU’s PCI-E slot
  4. Checked, reseated 12VHPWR cable into GPU
  5. Checked, switched GPU PCI-E power plug on the PSU side
  6. Changed BIOS setting Hardware P-States to Out-of-Band from Disabled

Nothing worked. Anyone experienced this type of HW power braking before?

Driver Version : 580.126.09
CUDA Version : 13.0

Attached GPUs : 2
GPU 00000000:B1:00.0
Performance State : P2
Clocks Event Reasons
Idle : Not Active
Applications Clocks Setting : Not Active
SW Power Cap : Active
HW Slowdown : Active
HW Thermal Slowdown : Not Active
HW Power Brake Slowdown : Active
Sync Boost : Not Active
SW Thermal Slowdown : Not Active
Display Clock Setting : Not Active
Clocks Event Reasons Counters
SW Power Capping : 218684551062 us
Sync Boost : 0 us
SW Thermal Slowdown : 0 us
HW Thermal Slowdown : 0 us
HW Power Braking : 265955594550 us
Sparse Operation Mode : N/A

GPU 00000000:CA:00.0
Performance State : P2
Clocks Event Reasons
Idle : Not Active
Applications Clocks Setting : Not Active
SW Power Cap : Active
HW Slowdown : Active
HW Thermal Slowdown : Not Active
HW Power Brake Slowdown : Active
Sync Boost : Not Active
SW Thermal Slowdown : Not Active
Display Clock Setting : Not Active
Clocks Event Reasons Counters
SW Power Capping : 265702379036 us
Sync Boost : 0 us
SW Thermal Slowdown : 0 us
HW Thermal Slowdown : 0 us
HW Power Braking : 265954937816 us
Sparse Operation Mode : N/A

Not sure if Robert’s reply in this thread is relevant.

You can find various reports of power brake issues on these forums, rs277 has pointed out one of them. (Here is another example. And there are others as well.)

While you could possibly go to extraordinary measures (taping off a specific PCIE pin), barring that sort of extreme (which I definitely do not recommend), I don’t think there is anything you can do to fix this yourself. Given that both GPUs exhibit the same report, I think this is unlikely to be a GPU defect or other one-off defect, I would suspect a systemic issue.

My suggestion, not unlike that other thread, is to contact SuperMicro (SMC) support. It will matter a lot whether you bought the GPUs in the server configured that way from SMC, in which case they absolutely should fix the issue for you, or else if you added the GPUs yourself, then I’m not sure of the outcome. But you should discuss it with SMC, in my opinion. I don’t happen to know if that server was intended for use with that GPU by SMC or not. If it was not, then they may not be able to help much. If you bought it with the GPUs installed by SMC, then it was obviously intended for use that way.

After rereading, I see that you mention a EVGA PSU. I don’t have any knowledge to suggest that is relevant to the issue, but it suggests to me you may have assembled pieces yourself. If that is the case, I’m not sure what SuperMicro’s position might be. And if this is a private label server from some other server outfit that happens to use a SMC motherboard but you bought the system elsewhere, then replace all my usages of SMC here with whatever packaged server outfit you bought the box from.

Microway are an NVIDIA-Certified Workstation Partner, per NVIDIA’s website. And as far as I recall, they have held that status for many years. Nothing wrong with your choice of system integrator.

My expectations would be that if you contact Microway regarding this issue, they will resolve this with their component suppliers (SuperMicro, EVGA, and NVIDIA) and then get back to you. In particular, I would expect them to have a designated support contact at NVIDIA. In other words, provided that your support contract with Microway is still in force, you should not need to do anything after reporting the issue to Microway.

Given that, I am a bit surprised to read that “Microway … recommended [you] check with NVIDIA.” The issue at hand is not one that can typically be resolved by end users. All the component vendors involved generally supply high-quality parts, and Microway has been supplying HPC workstations for over three decades, so you should be in experienced hands. My best guess is that this is a resolvable configuration issue, but it may require the combined technical expertise from multiple vendors to resolve it.

[Later:]

Is that a single 1600W PSU supplying the whole system? If so, that seems under-dimensioned to me. You have 2 CPU @ 185W each, 2 GPUs @ 300W each, NIC/SSD/assorted motherboard components @ 50W, and 2 TB (!) of DRAM at 0.35W per GB of DDR4. That sums to a nominal load of 1740W. I do not see how even an EVGA SuperNOVA 1600 T2 (80PLUS Titanium compliant) would support that kind of load reliably, especially since modern CPUs and GPUs often have short-time power spikes due to dynamic clocking. The sum of the nominal load of all system components should be well under the nominal rating of the PSU.