my guess would also be changes in the system. I cannot explain any of this conclusively, I’m just offering opinions.
Power brake originated in server platforms as a desire to allow the server (perhaps in the context of a larger scope, i.e. a cluster) to deliver best possible performance when the total aggregate power available is less than the total aggregate power that the system (server or cluster, depending on how power delivery system is designed) could draw at peak demand.
The basic mechanism is that the the power delivery system (e.g. rack PDU or other cluster-level power indication) and/or server PSU would signal to the motherboard of one or more systems that the power was at or exceeding an aggregate level that the system could safely deliver. The system then asserts power brake, and the GPUs reduce their power draw, preventing sustained operation in an overload state.
The idea for the benefit, if we look at it at the single server level, is the possibility that I could design a system lets say with 4 GPUs, each of which can draw 300W, but I only have at most 1kW available for GPU power in my server. Knowing that GPUs don’t always run at their peak load, we could allow the GPUs to run normally until the aggregate demand was at or above 1kW, as determined by the PSU. The PSU then signals the motherboard (e.g. the server BMC), and the next step in the chain is that power brake is asserted. The GPUs reduce their power consumption.
You can imagine then that this scheme would allow two or possibly even 3 GPUs to run unconstrained, if the 4th was lightly loaded. Given that in a datacenter the loading across processors is not always equal, this is considered a useful system design strategy, in a certain light. Of course you could just prefer a system that did not do this and delivered 1.2kW to the GPUs. But sometimes server manufacturers make design tradeoffs that seem appropriate (to them, at least).
I don’t know if typical workstation platforms do this or not. But the RTX A5000 GPU is designed to be suitable for server applications as well as workstation applications (unlike, eg. a RTX A500 which would not have any typical server footprint, differentiating here between servers and rack workstations, but we are starting to get into the weeds now.)
And we should not conflate uncertainty in application description with uncertainty of diagnosis here. The fact remains that regardless of this discussion, the GPU is declaring that it believes the power brake input is asserted. And when the GPU believes that, it is completely normal for it to restrict clocks and power consumption. That part of the description here leaves little doubt, in my view.
Furthermore, without better description of the provenance of the system, I wouldn’t bet that this is anything other than a motherboard asserting a signal it should not (especially, for example, if power brake is asserted all the time, even when the GPUs are idle). I’m not suggesting there is an actual properly functioning power brake regime here, along the lines of the description I have given. Even if it were an actual properly designed power brake setup, and the system as a whole were behaving “as intended”, there isn’t anything that can be done at the GPU level. The issue must be clarified by working at (and understanding) the system level behavior.