RTX 5090 (open 610.57.04): display enumeration fails ~26s into `nvidia-drm` load after warm reboot, no auto-retry — worse on DP 2.1 than DP 1.4

Summary

After a warm reboot (systemctl reboot), the driver fails to enumerate my only display and
the screen stays black.

Never happens on a cold boot / full power-off; a full power-off is the
only reliable recovery. Present since at least driver 590.48.01 (Feb, kernel 6.19); still present
on 610.57.04 (kernel 7.2.3).

Hardware / software

  • GPU: NVIDIA GeForce RTX 5090 (GB202), VBIOS 98.02.2E.40.AC, board MSI [1462:5302]
  • Motherboard: ASUS ProArt X870E-CREATOR WIFI (Ryzen, AMD iGPU present but unused — single monitor on the 5090)
  • Monitor: Samsung Odyssey Neo G9 57" (G95NC), 7680x2160 @ 120 Hz, single DisplayPort cable to the 5090
  • Driver: nvidia-open 610.57.04 (also reproduced on 590.48.01)
  • Kernel: 7.2.3 (CachyOS); nvidia_drm.modeset=1, fbdev=1
  • Wayland (KDE Plasma / kwin_wayland)

Repro

  1. Working desktop, monitor’s OSD set to DisplayPort 2.1.
  2. systemctl reboot.
  3. Boot completes but the monitor gets no signal.

What I see in the log — two failure sub-modes, same root cause

At nvidia-drm load, on a bad boot, enumeration comes up empty:

[drm] [nvidia-drm] [GPU ID 0x00000100] Loading driver
   ...~26s gap, no other GPU activity...
[drm] [nvidia-drm] [GPU ID 0x00000100] Failed to get dynamic displays during device registration.
nvidia 0000:01:00.0: [drm] Cannot find any crtc or sizes

(a good boot instead says, within ~1s: Encoder with NvKmsKapiDisplay 0x00000001 already exists)

Capture 1 (Feb, driver 590.48.01): after this, kwin_wayland keeps retrying against a GPU with
no CRTCs for about a minute, and the GSP firmware then crashes:

NVRM: Xid 109 ... CTX SWITCH TIMEOUT
NVRM: Xid 119 ... Timeout after 6s of waiting for RPC response from GPU0 GSP! Expected function 10 (FREE)
NVRM: Xid 120 ... GSP task exception: supervisor timer interrupt
nvidia-modeset: WARNING: GPU:0: Failed to determine which devices were hotplugged: 0x62
nvidia-modeset: WARNING: GPU:0: Failed to allocate display resource.

This repeats for minutes; kwin_wayland / the login manager crash-loop.

Capture 2 (today, driver 610.57.04, kernel 7.2.3): same Failed to get dynamic displays /
Cannot find any crtc or sizes after the same ~26s stall — but no Xid, no GSP crash this
time. plasma-login-greeter/plasma-login-wallpaper immediately log There are no outputs - creating placeholder screen and stay there. A few minutes later, with the physical monitor never
touched, cat /sys/class/drm/card1-DP-4/status reports connected, enabled, dpms: On, and a
full EDID mode list — so the DP AUX channel does recover on its own, the driver just never
re-probes/rebinds a CRTC to it afterwards. Power-cycling the monitor at that point produced a
visible mouse cursor on an otherwise still-black screen — some partial output state gets
re-established (enough for a cursor layer) but never a full modeset.

So the actual bug looks narrower than my original report suggested: nvidia_drm gives up on
display enumeration after ~26s during device registration
(likely a DP AUX/DPCD read racing the
monitor’s own link re-training, worse at UHBR13.5/DP2.1 than at HBR3/DP1.4) and never retries.
Whether that then cascades into a GSP crash seems to depend on how hard the compositor hammers it
afterwards. simpledrm is not involved — it initializes identically on every boot, good or bad,
and nvidia-drmdrmfb always takes over fb0 cleanly either way.

Workarounds

  • Full power-off (not reboot) — always works.
  • Setting the monitor’s OSD to DisplayPort 1.4 (HBR3) instead of 2.1 (UHBR): in my testing the
    failure did not occur at DP 1.4. Same monitor, same cable, only the OSD link-rate cap changed.
  • Does NOT help: disabling the AMD iGPU in BIOS; forcing the compositor onto the NVIDIA node (KWIN_DRM_DEVICES); disabling UEFI Fast Boot.

Attached

  • nvidia-bug-report_BAD.log.gz — captured over SSH ~1 min into the failed state (today, no Xid)
  • nvidia-bug-report_GOOD.log.gz — same machine, healthy boot, for comparison
  • badboot_kmsg.txt, badboot_full.txt, badboot_dp4.txt — full journal + DP-4 sysfs state
  • (from the original report, Feb, 590.48.01) journal_gpu_focused_20min.log — has the Xid
    109/119/120 GSP-crash variant

Happy to test patches / debug builds.

nvidia-bug-report_GOOD.log.gz (546.3 KB)

nvidia-bug-report_BAD.log.gz (510.4 KB)

badboot_logs.tar.gz (77.5 KB)

journal_gpu_focused_20min.log (76.7 KB)

Ran into exactly this same issue this morning after installing CachyOS. My hardware is almost the same as yours, in particular same monitor (G95NC) and an RTX 5090. Same driver.

HW Diffs:

  • Motherboard: ASUS TUF GAMING X870-PLUS WIFI motherboard
  • GPU: Founders Edition RTX 5090

Nvidia bug “6732678” created for same. We will keep posted about updates.