Random crashes on wayland since driver 580 on Linux

General information:

Distribution: Arch Linux

Kernel: 6.16.7.arch1-1

Driver 580.82.09-1, nvidia-dkms package (not the nvidia-open driver)

GPU: RTX 3070 Laptop (110w)


Since driver 580, I’ve hit a regression that after about 1 hour of activity on any Wayland-based interface (tested with hyprland and sway), my laptop will freeze without further logs to the point I have to force poweroff by holding the power button. Nothing on journalctl -b -1

The 1 hour mark is an average, but this error sometimes reproduces at the 20 minutes mark, or at the 2 hour mark, and it does not even needs to involve GPU load with 3D games. Just normal web browsing and routine tasks may trigger it. Disabling the nvidia-powerd.service will make the system more stable and not crash under the 10 minutes mark, but it will eventually crash.

Important: It does not reproduces on X11even if I deliberately put load on the system and spend hours logged. Some of the pictures I’ve got.

hyprland, yesterday night after 1:45h of use:

sway, past week after 2h of use(same artifacts but since there is some personal data at the left I had to blur it with before posting here).

wayland periodically entirely crashes since installing 580.82.09, every application quits and I need to restart the whole session, its odd.
This is the error:

kwin_wayland[1683]: kwin_scene_opengl: Invalid framebuffer status: "GL_FRAMEBUFFER_INCOMPLETE_ATTACHMENT"
systemd-coredump[14730]: [🡕] Process 1683 (kwin_wayland) of user 1000 dumped core.

This is with kde plasma

General information:

Distribution: Ubuntu 24.04.3 (Wayland)

Kernel: Linux 6.11.0-29-generic

Driver 580.65.06, nvidia-dkms package (not the nvidia-open driver)

GPU: RTX 3050 Laptop

Power profile: NVIDIA On-Demand

Use the system for anything other than games: no crash

Launch any game: the game crash after 10 to 20min

Re-launch a game that has already crashed: no crash

This happen only with driver version 580, tried to completely wipe the driver and all related config files and then re-install the driver, but the same issue persist

And adding more context here:

  • Doesn’t matter if using nvidiaor nvidia-open here, both crash
  • This does not reproduces with nouveau + zink

So literally the community is making better drivers than the company that is supposed to fix those issues.;

$ memtest_vulkan
https://github.com/GpuZelenograd/memtest_vulkan v0.5.0 by GpuZelenograd
To finish testing use Ctrl+C

1: Bus=0x01:00 DevId=0x249D   8GB NVIDIA GeForce RTX 3070 Laptop GPU (NVK GA104)
Standard 5-minute test of 1: Bus=0x01:00 DevId=0x249D   8GB NVIDIA GeForce RTX 3070 Laptop GPU (NVK GA104)
      1 iteration. Passed  0.0761 seconds  written:    1.0GB  39.1GB/sec        checked:    2.0GB  39.6GB/sec
     68 iteration. Passed  1.0099 seconds  written:   67.0GB 187.4GB/sec        checked:  134.0GB 205.4GB/sec
    435 iteration. Passed  5.0012 seconds  written:  367.0GB 214.2GB/sec        checked:  734.0GB 223.2GB/sec
   2617 iteration. Passed 30.0129 seconds  written: 2182.0GB 212.1GB/sec        checked: 4364.0GB 221.3GB/sec
   4782 iteration. Passed 30.0015 seconds  written: 2165.0GB 210.8GB/sec        checked: 4330.0GB 219.5GB/sec
   6936 iteration. Passed 30.0069 seconds  written: 2154.0GB 209.6GB/sec        checked: 4308.0GB 218.3GB/sec
   9083 iteration. Passed 30.0128 seconds  written: 2147.0GB 208.9GB/sec        checked: 4294.0GB 217.6GB/sec
  11224 iteration. Passed 30.0088 seconds  written: 2141.0GB 209.0GB/sec        checked: 4282.0GB 216.7GB/sec
  13359 iteration. Passed 30.0042 seconds  written: 2135.0GB 208.7GB/sec        checked: 4270.0GB 216.0GB/sec
  15490 iteration. Passed 30.0100 seconds  written: 2131.0GB 208.4GB/sec        checked: 4262.0GB 215.4GB/sec
  17617 iteration. Passed 30.0089 seconds  written: 2127.0GB 208.3GB/sec        checked: 4254.0GB 214.9GB/sec
  19740 iteration. Passed 30.0093 seconds  written: 2123.0GB 207.9GB/sec        checked: 4246.0GB 214.5GB/sec
Standard 5-minute test PASSed! Just press Ctrl+C unless you plan long test run.
Extended endless test started; testing more than 2 hours is usually unneeded
use Ctrl+C to stop it when you decide it's enough

memtest_vulkan indicates that there is no memory issues with my GPU as well, and no crash was triggered so far since I started using nouveau on my laptop.

Are you still experiencing these issues?

Yes.

Drivers 590 is still didn’t fix a thing. Random freezing without any extra dmesg log, with disk activity led turned fixed on (no blinking) .

The weird artifacts are not happening anymore, but now, the computer is just completely frozen and a long power button press is needed.

I have tested out the GPU load with software, no issues on multiple benchmarks, used memtest.efi and made a long smartctl test to make sure that it isn’t something with memory and disks, and also, used stress-ng for the CPU. All of them passed.

The only thing left is the nvidia-open driver that after version 570 just became utterly crap.

Edit: Another thing. It’s been 2 months that I’m using the RTX3070 exclusively and I have put my laptop in Ultimate(Dedicated) mode because I thought that Optimus could be part of this equation here but, in ultimate mode it is still happening :(

Have you reported the issue here?

github.com/NVIDIA/open-gpu-kernel-modules

Not yet because I have ordered a new battery for my laptop.

upower --dump is showing a 62% health metric, and since I have tested every component that could be a source of error, I want to play safe before reporting.

If this behavior continues after the battery replacement, I’ll retest everything and sure report there

Have you tried cleaning the GPU fans, replacing the thermal paste, and raising the laptop’s back with attachable legs?

Have you tried using a version of nvidia-open-dkmsthat doesn’t contain the 0003-Revert-a-change-related-to-the-display-stack-in-580..patch?

For example the one in the Express Repository?

Did the cleaning already and also, replaced the CPU and GPU dyes with PTM7950 and the GPU vRAM with some Upsiren thermal putty. Temperatures are great and thermal dissipation and management is better than what I’ve had with thermal paste which was really dried out when I replaced 10 months ago. I try to run maintenance once each year on my laptops.

I’ll try the dkms after the battery replacement, to isolate all the variables.

Regards,

Temperatures monitoring aren’t reliable, because CPUs and GPUs under performance tend to clock till reaching 95ºC under any circumstance. Also they tend to exclude the VRAM itself.

Only choosing the appropriate thermal solution, and mounting it properly does it.

The effects you are seeing could be very likely due to VRAM overheating or harm.

Also if the laptop has a charger, that can also be the culprit. Chargers lose conductivity over time, hence Watts, and may not provide enough power.

The effects you are seeing could be very likely due to VRAM overheating or harm.

While this is a valid investigation path, this issue is happening even when Nvidia is under no load at all, and the laptop is running on battery for like 10 minutes only, with Firefox and 3 tabs…

Also, I have put this gpu under stress and sometimes I have hours of reliable use or 4h gaming sessions without any issue so, I do believe it is temperature related.

It happened on both situations: while on battery and while charging as well.

The only way to tell is by keeping swapping software and hardware components, and seeing if it still happens. Till you find the specific one that changes behavior.

I wish you good luck 🍀

Hello all. I am having a very similar issue as nwildner has reported. This is happening with a system that has a 5060 GPU installed. I have an identical system that has a 3060 installed which has no problems whatsoever. The lockup issue happens when I plug a monitor into the computer via an HDMI connector (which is then a dual monitor setup). I receive a hard lock at random time intervals. This happens under Fedora (version does not matter and has occurred using 41 through 44) KDE Plasma (aka Wayland) environment. Using anything that does not produce an image onto the screen works perfectly while utilizing 100% of the GPU power and all the vRAM for days on end. The moment that I end this process to free the full potential of the video card to use for personal gaming the game will work beautifully for random periods of time (usually minutes, but sometimes an hour or so) and then hard lock the system. If it is played without the HDMI monitor connected, it works perfectly. Under Windows 11, this problem has been recently resolved due to some driver related update (it was occurring under the Windows environment as well until about a month or two ago) playing the same game. The driver version that this has been happening under has included the versions from 580 through to 595.71.05-1 (have not attempted nouveau as those drivers will not allow cuda interfacing for the non-graphical computations that the computer does when not being used for gaming).

I am however not having the problems with lockups during usage of normal web browsing or other application usage with anything that is not utilizing the GPU. Examples include Spotify, Steam interface, FireFox, Google Chrome, various system tools, and Konsole. Sound works through the HDMI connector on the nVidia card as well when the HDMI connector is connected.

The fact that Windows 11 working flawlessly with no setup/hardware changes and in Wayland there is hard lockups seems to suggest that this is related to something in the driver or interaction between the driver and kernel or some power setting being called from the video driver that is causing a kernel panic.