RTX 2070 SUPER fails to suspend | 595.71.05 | Linux 7.0.0-28

System

  • GPU: NVIDIA GeForce RTX 2070 SUPER (TU104)
  • Driver: 595.71.05-0ubuntu0.24.04.1 (nvidia-driver-595-open installed via Driver Manager)
  • Kernel: 7.0.0-28-generic
  • Distribution: Linux Mint 22.3 (based on Ubuntu 24.04 LTS)
  • Desktop: Cinnamon 6.6.9 (Xorg)

nvidia-bug-report.log.gz (419.9 KB)

Summary

After upgrading to NVIDIA driver 595.71.05, suspend no longer works.

The suspend sequence is aborted by the NVIDIA kernel driver with:

NVRM: GPU 0000:01:00.0: PreserveVideoMemoryAllocations module parameter is set. System Power Management attempted without driver procfs suspend interface. Please refer to the 'Configuring Power Management Support' section in the driver README.

nvidia 0000:01:00.0: PM: pci_pm_suspend(): nv_pmops_suspend [nvidia] returns -5
nvidia 0000:01:00.0: PM: dpm_run_callback(): pci_pm_suspend returns -5
nvidia 0000:01:00.0: PM: failed to suspend async: error -5

PM: Some devices failed to suspend, or early wake event detected

The machine never actually enters suspend but only blanks the screens for a moment. Shortly after, it resumes normal operation - i.e. prompts for login.

With the help of ChatGPT I’ve extensively tried to find issues in my configuration - see below.

Kernel messages

$ dmesg

[14492.588432] PM: suspend entry (deep)
[14492.601799] Filesystems sync: 0.013 seconds
[14493.061746] Freezing user space processes
[14493.063339] Freezing user space processes completed (elapsed 0.001 seconds)
[14493.063341] OOM killer disabled.
[14493.063342] Freezing remaining freezable tasks
[14493.064331] Freezing remaining freezable tasks completed (elapsed 0.000 seconds)
[14493.064370] printk: Suspending console(s) (use no_console_suspend to debug)
[14493.080638] serial 00:00: disabled
[14493.087525] sd 3:0:0:0: [sda] Synchronizing SCSI cache
[14493.094890] ata4.00: Entering standby power mode
[14493.101841] xhci_hcd 0000:01:00.2: xHC error in resume, USBSTS 0x401, Reinit
[14493.101843] usb usb3: root hub lost power or was reset
[14493.101844] usb usb4: root hub lost power or was reset
[14493.210454] NVRM: GPU 0000:01:00.0: PreserveVideoMemoryAllocations module parameter is set. System Power Management attempted without driver procfs suspend interface. Please refer to the 'Configuring Power Management Support' section in the driver README.
[14493.210460] nvidia 0000:01:00.0: PM: pci_pm_suspend(): nv_pmops_suspend [nvidia] returns -5
[14493.210673] nvidia 0000:01:00.0: PM: dpm_run_callback(): pci_pm_suspend returns -5
[14493.210679] nvidia 0000:01:00.0: PM: failed to suspend async: error -5
[14493.914314] PM: Some devices failed to suspend, or early wake event detected
[14493.918717] serial 00:00: activated
[14493.940942] nvme nvme1: D3 entry latency set to 8 seconds
[14493.946467] nvme nvme0: 7/0/0 default/read/poll queues
[14493.959994] nvme nvme1: 20/0/0 default/read/poll queues
[14494.069835] OOM killer enabled.
[14494.069842] Restarting tasks: Starting
[14494.071065] Restarting tasks: Done
[14494.071068] efivarfs: resyncing variable state
[14494.088579] efivarfs: finished resyncing variable state
[14494.088594] random: crng reseeded on system resumption
[14494.088685] PM: suspend exit
[14494.088714] PM: suspend entry (s2idle)
[14494.091883] snd_hda_codec_nvhdmi hdaudioC2D0: HDMI: invalid ELD data byte 78
[14494.093550] Filesystems sync: 0.004 seconds
[14494.254300] ata1: SATA link down (SStatus 4 SControl 300)
[14494.254575] ata4: SATA link up 6.0 Gbps (SStatus 133 SControl 300)
[14494.254980] ata6: SATA link down (SStatus 4 SControl 300)
[14494.255030] ata2: SATA link down (SStatus 4 SControl 300)
[14494.255089] ata3: SATA link down (SStatus 4 SControl 300)
[14494.255231] ata5: SATA link down (SStatus 4 SControl 300)
[14494.256257] sd 3:0:0:0: [sda] Starting disk
[14494.257967] ata4.00: configured for UDMA/133
[14494.258065] ata4.00: Entering active power mode
[14494.831823] Freezing user space processes
[14494.834138] Freezing user space processes completed (elapsed 0.002 seconds)
[14494.834141] OOM killer disabled.
[14494.834142] Freezing remaining freezable tasks
[14494.835791] Freezing remaining freezable tasks completed (elapsed 0.001 seconds)
[14494.835795] printk: Suspending console(s) (use no_console_suspend to debug)
[14497.270692] igc 0000:04:00.0 enp4s0: NIC Link is Up 1000 Mbps Full Duplex, Flow Control: RX
[14497.279981] serial 00:00: disabled
[14497.284765] sd 3:0:0:0: [sda] Synchronizing SCSI cache
[14497.285237] ata4.00: Entering standby power mode
[14497.331417] NVRM: GPU 0000:01:00.0: PreserveVideoMemoryAllocations module parameter is set. System Power Management attempted without driver procfs suspend interface. Please refer to the 'Configuring Power Management Support' section in the driver README.
[14497.331420] nvidia 0000:01:00.0: PM: pci_pm_suspend(): nv_pmops_suspend [nvidia] returns -5
[14497.331559] nvidia 0000:01:00.0: PM: dpm_run_callback(): pci_pm_suspend returns -5
[14497.331562] nvidia 0000:01:00.0: PM: failed to suspend async: error -5
[14497.934299] PM: Some devices failed to suspend, or early wake event detected
[14497.938779] serial 00:00: activated
[14497.960043] nvme nvme1: D3 entry latency set to 8 seconds
[14497.965623] nvme nvme0: 7/0/0 default/read/poll queues
[14497.978278] nvme nvme1: 20/0/0 default/read/poll queues
[14498.089032] OOM killer enabled.
[14498.089039] Restarting tasks: Starting
[14498.090110] Restarting tasks: Done
[14498.090114] efivarfs: resyncing variable state
[14498.108624] efivarfs: finished resyncing variable state
[14498.108641] random: crng reseeded on system resumption
[14498.108751] PM: suspend exit

Relevant configuration

The driver reports:

$ cat /proc/driver/nvidia/params | grep Preserve
PreserveVideoMemoryAllocations: 1

$ cat /proc/driver/nvidia/params | grep TemporaryFilePath
TemporaryFilePath: "/var"

The procfs interface exists:

$ ls -l /proc/driver/nvidia/suspend
-rw-r--r-- 1 root root 0 ... /proc/driver/nvidia/suspend

$ cat /proc/driver/nvidia/suspend
suspend hibernate resume

The NVIDIA systemd units are installed and enabled:

$ systemctl status nvidia-suspend.service nvidia-hibernate.service nvidia-resume.service

nvidia-suspend.service    enabled
nvidia-hibernate.service  enabled
nvidia-resume.service     enabled

The dependency graph also appears correct:

$ systemctl show systemd-suspend.service -p Wants -p Requires

Requires=sleep.target system.slice
Wants=nvidia-suspend.service nvidia-resume.service

and

$ systemctl show -p WantedBy nvidia-suspend.service

WantedBy=systemd-suspend.service

The system is configured to use deep sleep:

$ cat /sys/power/mem_sleep
s2idle [deep]

Additional observations

The procfs interface appears to be functional.

Executing

sudo sh -c 'echo suspend > /proc/driver/nvidia/suspend'

is accepted by the driver and immediately blanks both displays. The machine does not recover from this state until it is rebooted, indicating that the procfs interface exists and the driver reacts to it.

This makes the kernel message

“System Power Management attempted without driver procfs suspend interface”

somewhat unexpected.

Hibernation

Hibernate works correctly on this system.

Only suspend fails.

Since both suspend and hibernate are expected to use /proc/driver/nvidia/suspend when PreserveVideoMemoryAllocations=1 is enabled, I am unsure why only the suspend path reports that the procfs interface was not used.

Notes

The driver was installed using Mint’s Driver Manager tool.
This is not a manual .run installation.

I have tried different drivers, open and proprietary. With 535 and 550 I am able to suspend, hibernate and resume my system. With 570, 580, 590 & 595, I believe, suspension does not work.

Also, up until kernel 6.8.0-85 I have not encountered this issue. It was with its successor (6.8.0-87 or something) that I started having this problem (and still have with 7.0.0).

You’ll find nvidia-bug-report.log.gz attached to this post.

Thanks in advance!
Cheers