Dual-GPU (RTX 4090 + RTX 5060 Ti): Xid 158 NV_UFLUSH_FB_FLUSH → Xid 175 GSP RPC timeout → black screen on near-idle display GPU (595.71.05 open)

Summary

On a dual-GPU Ubuntu desktop, the RTX 4090 (display GPU) repeatedly hangs with:

  1. GSP CrashCat / GSP task watchdog timeout
  2. Xid 158 – timeout waiting for NV_UFLUSH_FB_FLUSH
  3. Xid 175 – GSP RPC timeout (GSP_RM_CONTROL)
  4. Xid 154 – GPU Reset Required
  5. nvidia-modeset: Failed detecting connected display devices → black screen

Recovery requires reboot. Soft reset does not restore the display.

This is not a CUDA OOM. It also occurs when heavy FLUX/ComfyUI containers are not running, including after ~35–40 minutes of near-idle desktop use.

System

Item Value
OS Ubuntu 22.04.5 LTS
Kernel 6.8.0-124-generic
Desktop GNOME on X11
CPU AMD Ryzen 9 9950X
Motherboard ASUS ProArt X870E-CREATOR WIFI
BIOS 1715 (09/19/2025)
Driver 595.71.05NVIDIA UNIX Open Kernel Module
GSP Firmware 595.71.05 (both GPUs)
Persistence Mode Disabled
Cmdline extras none (quiet splash only; no pcie_aspm=off, no GSP disable)

GPUs

GPU Bus UUID VBIOS Role
GeForce RTX 4090 0000:01:00.0 GPU-c3aff235-d99c-e258-1fd4-a828b1540f35 95.02.3C.00.60 Primary display (HDMI-1-1, 2560x1440) + sometimes CUDA workloads
GeForce RTX 5060 Ti 0000:09:00.0 GPU-89e50e82-df56-fb9f-05fb-5f1323ba61a6 98.06.39.40.F8 Secondary / compute (often vLLM); display usually inactive

nvidia-smi after reboot typically shows:

  • GPU0 (4090): display_active = Enabled
  • GPU1 (5060 Ti): display_active = Disabled

Symptoms

  • Screen goes black / unrecoverable display loss
  • Machine often still partially alive (SSH / some processes), but GPU0 needs reboot
  • last session shows crash for the graphical login
  • No CUDA out of memory messages correlated with these hangs

Recent incidents (JST)

A) 2026-07-21 ~12:42

  • GPU0 GSP heartbeat timeout → Xid 175Xid 154 → display detection failures
  • CUDA workloads (FLUX/vLLM containers) may have been present
  • Reboot required

B) 2026-07-22 ~07:11 (containers stopped)

  • Boot 06:34 → hang ~07:11
  • Cascade: Xid 158 (NV_UFLUSH_FB_FLUSH) → Xid 154 → GSP heartbeat / Xid 175 (process names in log included nvidia-smi, electron)
  • FLUX / vLLM / image-analyzer containers were not running
  • Session ended as crash; reboot ~08:52

C) 2026-07-22 ~14:40 (best instrumented; near-idle)

Custom 5s polling showed GPU0 for many minutes before the hang as approximately:

  • P-state: P5
  • util.gpu: ~27–31%
  • memory.used: fixed ~9080 MiB (of 24564)
  • power: ~17 W
  • temp: ~52 °C
  • PCIe: gen 2 x8 (idle-ish link)

No OS suspend around the event.

Timeline (kernel, JST):

  • 14:40:06 – GSP-CrashCat / GSP task watchdog timeout @ pc:0x1ad4644, partition:2#0, task:3
  • 14:40:31Xid 158 NV_UFLUSH_FB_FLUSH = 0x2
  • 14:41:16Xid 175 GSP RPC timeout (75s) for GSP_RM_CONTROL seq 293677
    also: Memory Subsystem Error / kgmmuInvalidateTlb failed
    then Xid 154 GPU Reset Required
    another CrashCat with same pc / buildId
  • 14:43:17 – repeated Failed detecting connected display devices

GSP bin buildId: