I am reporting a critical system stability issue affecting dual NVIDIA GeForce RTX 5090 GPUs setup. This appears to be related to the known GSP firmware bug affecting Blackwell (RTX 50 series) GPUs. Please help to identify the issue.
The GSP timeout is obviously caused by the GPU falling off the bus in the first place: how can it not time-out if the GPU is no longer there? ;-)
Now regarding falling off the bus, in the words of an NV eng:
So check these 3 things first and definitely the most recent driver (595.71 currently).
However if you browse this forum a bit, you will find that falling off the bus is something that Blackwell just tends to do for many ppl…
I think also encountering this error or a related error with a 5070 Ti on an Asus TRX50 board. My system runs Windows, however symptoms are effectively identical:
the OS still responds to VNC connections (but show a black screen)
the GPU fully freezes showing a static image
the OS is unresponsive to any keyboard/mouse inputs
I’m unable to restore operation, requiring a power cycle to restore the system.
If I have telemetry/monitoring software displaying system metrics across the moment of failure, there are some patterns for things like PCIe bus interface speeds, bus utilisation, GPU +12V supply, power and fan speed. The crash/failure moment is otherwise ‘silent’, and happens under a variety of conditions where the GPU has a steady, medium workload (in my case NVENC/NVDEC). I’ve seen discussions where CUDA workloads have been involved.
My crashes have occurred between 4-6 days. The system previously benchmarked OK running the ‘usual’ industry benchmark tools for several days, so perhaps I didn’t leave the system benchmarking long enough to reveal the fault earlier.
A lack of relevant event logs or other telemetry suggests to me that the system is not kernel panicking, and the GPU drivers are not catching the fault - or perhaps they are causing the fault? Maybe it’s a GSP issue occurring so quickly that TDR isn’t able to handle it?
The fault has followed the 50 series GPU across two very different spec test systems using various releases of Studio drivers over the last few months. In the TRX50 system, after I installed an 30-series RTX GPU for testing, the fault has not reoccurred.
GPU and CPU thermals are fine, the PSU is intentionally massively overspecced for the system, all firmwares, drivers and BIOSes are latest stable versions.