Ubuntu 24.04: suspend/display power transition causes soft lockup in nvidia-modeset nvWriteGpEntry, system requires hard power-off

The system was unresponsive during the failure, so I could not run nvidia-bug-report.sh while the machine was stuck. The attached nvidia-bug-report.log.gz was collected after forced power-off and reboot. The actual failure is visible in previous-boot-kernel.log, where nvidia-modeset repeatedly soft-locks in nvWriteGpEntry and the KMS thread is blocked.

previous-boot-kernel.log (1.2 MB)

nvidia-bug-report.log.gz (461.2 KB)

syslog-snippet.txt (4.6 KB)

I have exactly the same problem. RTX 5060 with a Ryzen 7 Linux Mint system.
After some time the system soft-crashes (either directly after waking up from standby or sometimes after the system was running for several hours).

It still reacts to the REISUB sequence.
See the syslog snippet (it is nearly identical to xianchao’s log).
I already tried everything: Using different driver versions (open driver 580 and 595), different kernel versions (6.8, 6.17, etc), downgrading the PCI to Gen4 and Gen2.
The problem persists.

Same soft lockup problem here — adding my data point and a captured stack trace that identifies the culprit as DIFR (Display Idle Frame Refresh).

System:

  • Desktop: MSI B650M Project Zero (MS-7E09), BIOS 1.H5, Ryzen (AM5)
  • GPU: RTX 5070 Ti (Blackwell, so the open kernel module is mandatory)
  • Ubuntu 24.04, kernel 6.17.0-1028-oem, GNOME Wayland session
  • Driver: nvidia-driver-595-open (595.71.05, DKMS), Secure Boot on
  • Sleep mode: S3 “deep”, NVreg_PreserveVideoMemoryAllocations=1, NVreg_UseKernelSuspendNotifiers=1

Symptom: Suspend and resume themselves complete perfectly (deep S3 first try, no NVRM errors). Then 20–60 seconds after resume the desktop freezes hard and requires a power-off. Frequency is roughly 1 in every 3–5 resumes; the other resumes are completely clean. Most freezes leave nothing in the journal (journald dies before the watchdog line flushes), but on my latest hit the system stayed half-alive for ~6 minutes and flushed the full trace:

watchdog: BUG: soft lockup - CPU#2 stuck for 365s! [nvidia-modeset/:1002]
RIP: 0010:nvWriteGpEntry+0x100/0x370 [nvidia_modeset]
Call Trace:
  nvPushKickoff
  PrefetchHelperSurfaceEvo
  nvDIFRPrefetchSurfaces
  DifrPrefetchEventDeferredWork
  nvkms_kthread_q_callback
  _main_loop

accompanied by:

INFO: task KMS thread:3440 blocked on a semaphore likely last held by task nvidia-modeset/:1002

So the culprit is the DIFR prefetch: after an S3 resume, GSP raises a DIFR prefetch event (DifrPrefetchEventDeferredWork), and nvWriteGpEntry spins forever waiting for pushbuffer space on a channel the GPU never services. The nvkms kthread soft-locks, the KMS thread blocks on its semaphore, and the whole display stack is dead. DIFR runs fine during normal desktop use — it’s only its first engagement after S3 resume that hits stale state, which explains the variable 20–60s delay (first idle moment after resume) and why silent freezes with no logged stack are the same bug.

What I’ve tried — none of it fixed the lockup:

  1. Driver upgrade 580.159.03 → 595.71.05 (open): reduced frequency (580 locked up on nearly every resume, 595 is ~1 in 3–5) but did not eliminate it. Identical stack on both versions.
  2. Xorg → Wayland: fixed an unrelated EnterVT session-death race; no effect on this lockup.
  3. NVreg_UseKernelSuspendNotifiers=1 (new in 595) instead of the nvidia-suspend/resume services: made suspend fast and clean (also eliminated an NVRM mmuWalkSparsify/VA-space assert storm during eviction) — but the post-resume DIFR lockup still occurs.
  4. DIFR idle-threshold regkeys — NVreg_RegistryDwords="RMLpwrMsDifrCgIdleThresholdUs=0xFFFFFFFF;RMLpwrMsDifrSwAsrIdleThresholdUs=0xFFFFFFFF" (found in GSP firmware strings; the idea was to push DIFR idle-entry out to ~71 min so it never engages). Verified active in /proc/driver/nvidia/params at the time of the latest freeze — it locked up anyway. Since the hang enters via a GSP-originated prefetch event, these idle thresholds apparently don’t gate the post-resume prefetch path.

Also verified: there is no user-facing way to disable DIFR in 595.71.05 — no nvidia_modeset module parameter, no nvkms config-file key, and RmForceEnableDIFR only forces it on. The Ubuntu/.run DKMS package ships the nvkms core as a precompiled blob (nv-modeset-kernel.o_binary), so patching it out means building the full open-gpu-kernel-modules tree from source (nvDIFRAllocate() in src/nvidia-modeset/src/nvkms-difr.c stubbed to return NULL), which is my next step if switching to s2idle doesn’t avoid it.

NVIDIA folks: this reproduces across Blackwell, Ada, Ampere and Pascal per this thread and GitHub issues #1205/#1177. Could we get either a fix for the post-S3 DIFR prefetch race or a supported regkey/module param to disable DIFR entirely?

Same issue with 610.43.03.

System info:

  • Gigabyte X670 GAMING X AX V2 mobo
  • Ryzen 7 7800 X3D CPU
  • RTX 4070 Ti Super
  • Driver 580.142 - 610.43.03
  • Rocky Linux 10.2 with kernel 6.12 with these options:
    • options nvidia-drm modeset=1 fbdev=1
    • options nvidia NVreg_PreserveVideoMemoryAllocations=1 NVreg_TemporaryFilePath=/var/tmp
  • Also tried openSUSE Tumbleweed Slowroll with kernel 7.1 and NVIDIA driver packaged for openSUSE with default settings, same issue.

After resume the system is responsive for a few seconds and then the graphics freeze and the CPU fan spins up. The console is frozen but you can log in via SSH.

While the console is frozen “soft lockup” is logged to /var/log/messages, about every 26 seconds:

kernel: watchdog: BUG: soft lockup - CPU#8 stuck for 26s! [nvidia-modeset/:943]
[...]
kernel: watchdog: BUG: soft lockup - CPU#8 stuck for 2191s! [nvidia-modeset/:943]

Followed by this (also every 26 seconds):

kernel: CPU: 8 UID: 0 PID: 943 Comm: nvidia-modeset/ Tainted: G           OE      ------  ---  6.12.0-211.40.1.el10_2.x86_64 #1 PREEMPT(voluntary)
[...]
kernel: RIP: 0010:nvWriteGpEntry+0xfc/0x370 [nvidia_modeset]
kernel: Code: d2 0f 1f 00 80 0b 10 48 83 c4 28 44 89 f8 5b 5d 41 5c 41 5d 41 5e 41 5f c3 0f 1f 44 00 00 48 8b 83 10 01 00 00 8b 38 c1 ef 12 <01> ff 39 fd 74 a2 48 8b 7c 24 08 44 89 f2 43 8d 04 12 48 03 54 24
kernel: RSP: 0018:ffffcf908133fd18 EFLAGS: 00000212
kernel: RAX: ffff8b17f89c4000 RBX: ffff8b17fbb9e030 RCX: ffffcf908292e3a0
kernel: RDX: ffff8b17f89d8000 RSI: 00000000000003ac RDI: 0000000000000002
kernel: RBP: 0000000000000004 R08: 0000000000000000 R09: 0000000000000001
kernel: R10: 0000000000000002 R11: 0000000000000010 R12: ffffcf90809280d8
kernel: R13: 00000000000003ac R14: 0000000000000344 R15: 0000000000000000
kernel: FS:  0000000000000000(0000) GS:ffff8b1ede600000(0000) knlGS:0000000000000000
kernel: CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
kernel: CR2: 00007fccc89e1000 CR3: 00000001d7a24000 CR4: 0000000000f50ef0
kernel: PKRU: 55555554
kernel: Call Trace:
kernel: <IRQ>
kernel: ? srso_alias_return_thunk+0x5/0xfbef5
kernel: ? show_trace_log_lvl+0x255/0x2f0
kernel: ? show_trace_log_lvl+0x255/0x2f0
kernel: ? nvPushKickoff+0x28/0x50 [nvidia_modeset]
kernel: ? watchdog_timer_fn.cold+0x3d/0xa0
kernel: ? __pfx_watchdog_timer_fn+0x10/0x10
kernel: ? __hrtimer_run_queues+0x139/0x2a0
kernel: ? hrtimer_interrupt+0xff/0x230
kernel: ? __sysvec_apic_timer_interrupt+0x52/0x100
kernel: ? sysvec_apic_timer_interrupt+0x6c/0x90
kernel: </IRQ>
kernel: <TASK>
kernel: ? asm_sysvec_apic_timer_interrupt+0x1a/0x20
kernel: ? nvWriteGpEntry+0xfc/0x370 [nvidia_modeset]
kernel: nvPushKickoff+0x28/0x50 [nvidia_modeset]
kernel: PrefetchHelperSurfaceEvo+0x45c/0x650 [nvidia_modeset]
kernel: nvDIFRPrefetchSurfaces+0xb1/0x1f0 [nvidia_modeset]
kernel: DifrPrefetchEventDeferredWork+0x16/0x30 [nvidia_modeset]
kernel: nvkms_kthread_q_callback+0xbc/0x150 [nvidia_modeset]
kernel: _main_loop+0x8f/0x150 [nvidia_modeset]
kernel: ? __pfx__main_loop+0x10/0x10 [nvidia_modeset]
kernel: kthread+0xfa/0x240
kernel: ? __pfx_kthread+0x10/0x10
kernel: ret_from_fork+0x31/0x50
kernel: ? __pfx_kthread+0x10/0x10
kernel: ret_from_fork_asm+0x1a/0x30
kernel: </TASK>

The system can not be rebooted via SSH. Tried REISUB, didn’t work (might be because of keyboard/KVM). Power off is required.

I have tried these driver versions, they all have the same issue:

580.142
595.71.05
595.80
595.84
610.43.03

The last driver that doesn’t have this issue is 580.126.18.

Samma issue with 610.57.04. Option NVreg_UseKernelSuspendNotifiers=1 and “setenforce 0” before suspend.

aug 14 09:28:52 hostname kernel: nvidia-modeset: Loading NVIDIA UNIX Open Kernel Mode Setting Driver for x86_64  610.57.04  Release Build  (dvs-builder@U22-I3-AF05-32-1)  Wed Jul 29 02:31:30 UTC 2026
aug 15 11:19:28 hostname kernel: watchdog: BUG: soft lockup - CPU#5 stuck for 26s! [nvidia-modeset/:920]

Having the same issue, confirmed to have occurred using both 595 and 610 driver branches.

Specs:
Lenovo Legion Slim 5 16ARP9
GPU: NVIDIA GeForce RTX 4060 Max-Q / Mobile (no iGPU, NVIDIA is the sole display controller)
Kernel: Linux 7.1.8-1-cachyos
OS: CachyOS
Desktop Environment: Gnome 50 (Wayland)

Freezing is most commonly triggered by the display waking up after blanking out due to inactivity, but will on occasion occur briefly after waking from suspend or during normal usage. Bug is triggered basically every day I use the device.

Same root cause as diagnosed above (DIFR prefetch deadlock). More info on my response to this thread.

nvidia-bug-report.log.gz (115.4 KB)

Confirming the same DIFR / nvWriteGpEntry deadlock on another configuration:

Fedora 44 KDE / KWin Wayland, Intel UHD 770 + RTX 4060 Laptop, NVIDIA open driver 610.57.04, kernel 7.1.8, S3 deep.

The freeze typically occurs after 2–3 suspend/resume cycles followed by normal desktop use. Ctrl+Alt+F3 is unresponsive, but Magic SysRq still works.

SysRq w captured kwin_wayland blocked in:

nvkms_ioctl_from_kapi_try_pmlock -> ApplyModeSetConfig -> nv_drm_atomic_apply_modeset_config

SysRq l simultaneously captured nvidia-modeset/ in:

nvWriteGpEntry -> nvPushKickoff -> PrefetchHelperSurfaceEvo -> nvDIFRPrefetchSurfaces -> DifrPrefetchEventDeferredWork

This appears to be the same root cause discussed here. I added the full system details and sanitized SysRq trace to NVIDIA/open-gpu-kernel-modules issue #1289:

I also have a full nvidia-bug-report.log.gz available privately if NVIDIA needs it.

I’ve been running 595.99.02 for a week now without issues after resume, looks promising.

Option NVreg_UseKernelSuspendNotifiers=1 and “setenforce 0” before suspend.

Two weeks without issues, 595.99.02 seems to have resolved the nvWriteGpEntry problem, at least when related to suspend/resume.

The release notes for 595.99.02 say “Fixed a bug that caused resume from suspend to fail if the system was previously hibernated”, I have never hibernated my machine, only suspended.