Filing a heavily instrumented data point for the Blackwell "GPU has fallen off the bus" family, in the hope it can be attached to **internal bug 6426268** (acknowledged by @amrits in [thread 376750](https://forums.developer.nvidia.com/t/deterministic-xid-120-gsp-kernel-panic-on-rtx-pro-6000-blackwell-under-sustained-cublas-workload-595-84/376750)). Community tracking: [open-gpu-kernel-modules#1151](https://github.com/NVIDIA/open-gpu-kernel-modules/issues/1151) — I've posted the full dataset there. ## Summary One PNY RTX 5080 (GB203, VBIOS `98.03.6C.00.41`) has produced **29 identical faults since 2026-07-30**: instant, precursor-free device death. `Xid 79 → Xid 154`, preceded by hundreds of `_issueRpcAndWait: rpcSendMessage failed with status 0xf for fn 76`; on 595.84 also `tmrGetTimeEx: Consistently Bad TimeLo value ffffffff` and `_kgspRpcRecvPoll: GSP RM heartbeat timed out`. **Zero PCIe AER events in any boot, ever.** With the device dead, `lspci -vv` still reports `Speed 32GT/s, Width x16` with both functions enumerated. What makes this report worth attaching to the bug: **1. Same card, both operating systems.** Identical failure under Linux open kernel modules **580.173.02, 595.84, 610.43.02** and under **Windows 11, driver 610.88** — including a fresh Windows install that failed 71 minutes after first boot. On Windows the failure surfaces as bugcheck `0x133`, or as LiveKernelReport `0x141 VIDEO_ENGINE_TIMEOUT_DETECTED` with the host surviving (`nvidia-smi: GPU is lost`, card still enumerated `CM_PROB_NONE` at Gen5 x16). Cross-OS reproduction on one unit points below the driver boundary — GSP firmware or silicon. **2. Power excluded by measurement, not inference.** A Corsair HX1000i logged at 2 s intervals ran *through* three faults (the host survived one, so sampling continued past the death-second): ``` 20:05:11 +12V 12.047 V 38.25 A 456 W (GPU at 100% util, 308 W, 65 °C, P0) 20:05:13 +12V 12.094 V 6.75 A 80 W (GPU gone from the bus) ``` 31 A of load disappeared and the rail **rose**. No OCP trip, 52% of PSU rating at peak, wall-side input stable. The load vanished; the supply never sagged. **3. Full platform elimination.** The card was moved between two complete systems (Z390/i9-9900K/DDR4/Gen3 → X870/Ryzen 9900X/DDR5-JEDEC/Gen5, different PSUs). The fault followed the card unchanged. Thermal (65–73 °C at death, throttle clean), ASPM (disabled, verified at register level), memory pressure, VRAM pressure (one fault at 1.4 GB used during model load), and workload type (CUDA inference, D3D gaming, near-idle) are all excluded — details and tables in the GitHub post. **4. Trigger-envelope note.** Unlike most #1151 reports (death at idle, P8, \~50 W), this unit dies at **full boost — 2850–2955 MHz, 310–360 W, parked in P0** with no P-state/mem-clock transitions for long stretches beforehand. Combined with the thread's idle-death captures, the failure spans both ends of the DVFS range on GB20x. **5. MTBF is degrading monotonically** on this unit — from half-day intervals in late July to 71–115 minutes now. Same shrinking-MTBF pattern at least three #1151 reporters describe. ## Data Four instrumented captures — 2 s GPU + PSU CSVs, gap-free through the death-second, plus symbolized `!analyze -v` outputs — are published (CC0): https://github.com/rudutoitnuhome/rtx5080-gb203-xid79-telemetry Available on request: * `nvidia-bug-report.log.gz` and journal extracts from the Linux installs (\~40 captured events) * the minidump and the 10.5 GB `0x141` live kernel dump, via a private channel (memory contents) ## Questions 1. Can this report be linked to bug 6426268, and is there any driver/GSP-firmware build in which a fix is expected to land? 2. Is there an engineering VBIOS or GSP firmware build for GB203 you would want tested on a unit that reproduces this every 1–2 hours? Given the reproduction rate, this card is an unusually fast test vehicle. 3. For a unit on a degrading-MTBF curve, is there any diagnostic you'd want captured while a fault is live (the host survives some faults, so the dead card can be probed before reboot)?