610 release feedback & discussion

the 5070ti will in general absolutely crush the 9070XT in compute performance. Theres also the fact that ROCM kinda sucks and Cuda is king. The 9070XT is about the range of an RTX4070 in compute (heavily depends on the workload, sometimes 4070 crushes it), that puts it pretty well behind the 5070ti.

Hi @matusmeister

Can you please help to capture nvidia bug report immediately after game crash and upload here.

Capcom explicitly disables raytracing due to a lot of potential bugs using it on Linux with RE Engine be it AMD or Nvidia. I would say stop using that launch option if you dont want to experience crashing and try tinkering with your mods as Im not experiencing this issue running the game on 595 or 610.

using raytracing in RE Engine games on linux is dragons territory

@wanboyao

You wrote:

Just add execution permission to pahole.sh; it’s a known bug.

I assume you mean ‘chmod +x’, as described here:

https://gist.github.com/SeriousPassenger/73392e80c64cc1b63839658cc17926e1

sudo chmod +x /usr/src/nvidia-595.71.05/pahole.sh
sudo chmod +x /var/lib/dkms/nvidia/595.71.05/build/pahole.sh 2>/dev/null || true

In my case, driver 595.58.03 is now installed.
The problem is that there is no ‘pahole.sh’ file in the ‘/usr/src/nvidia-595.58.03’ directory.
There is also no ‘/var/lib/dkms/nvidia/595.58.03/build’ directory.

So how do I fix this?

I noticed that the ‘/var/lib/dkms/nvidia/610.43.02/build’ directory is created every time the 610.43.02 driver is installed.

It is deleted when the driver installs successfully, and remains if the driver fails to install.
In that case, it contains the ‘pahole.sh’ file, whose permissions can be changed using the ‘chmod +x’ command.
But that doesn’t help at all, because when you run the driver installation, the installer deletes that directory and creates a new one, in which, after a failed installation, we’ll find the ‘pahole.sh’ file again without ‘x’ permissions.
And so on and so forth.

Maybe I’m doing something wrong.
Could you point me toward the correct solution?

Best regards

P.S.
I’m installing the driver in the console, i.e., when the GRUB menu appears, I press the “E” key, add a “3” after the space at the end of the line starting with “linux,” and press “Ctrl+X.”

@amrits, could we ask for a bug tracking ID for this issue?

Ray tracing in The Witcher 3 has been crashing the game since the 590 branch. I can confirm that reverting to the 580.xx driver fixes the problem.

The build folder in /var/ is temporary, deleted after successful KM compilation is expected. And that build folder is not reused for recompilation after fail. It is recreated and copied from /usr/src/nvidia-<your-version>/ if KM compilation initated, so just chmod +x here.

As for 595.58, the pahole.sh file is introduced after this version I think, maybe 595.71. You can confirm it by bash run_file.run -x, extract the driver files to see if it exists

@wanboyao

Here’s what happened in my case:

  1. I had kernel 6.12.90-1 installed.
  2. I installed driver 610.43.02.
    Everything worked fine.
  3. A new kernel, 6.12.90-2, was released, so I installed it.
    The NVIDIA modules won’t compile.

From what you wrote, I understand that instead of reverting to the previous driver 595.58.03, I should find the file ‘pahole.sh’ in the directory ‘/usr/src/nvidia-610.43.02/’, which should be there (because ver. 610.xx.xx already includes it), run the command ‘chmod +x pahole.sh’, and retry installing the new kernel.

Am I thinking correctly?

When I find some time, I’ll look through the previous driver versions to see starting from which version the ‘pahole.sh’ file is included.

Best regards.

Yes, that’s correct. I also encounter this issue on Debian testing and recent LTS 580 release. Method I mention above solved my issue.

I don’t have my linux install anymore so I was wondering if there’s any improvement with 5374195 in this driver? (Low framerate compensation/VRR flicker)

Has anyone been able to test?

I’m still experiencing issues with VRR flicker.

In Hyprland, when I fullscreen an application with VRR on fullscreen-only-mode (=2), every window keeps flickering like crazy. In another wayland compositor - mangowc - it was even worse. I was getting heavy flickering randomly on the browser for no obvious reason, based on inactivity.

I’d really like to encourage NVIDIA to choose the open source route. I believe there are so many passionate customers that would be willing to actually help NVIDIA with these issues. It’s absolutely unfathomable for me why everything is still closed source and a buggy mess.

VRR has not been fixed.

I managed to install driver 610.43.02 on kernel 6.12.90-2

After successfully installing the driver, I checked and the ‘pahole.sh’ file is in the ‘/usr/src/nvidia-610.43.02/’ directory, but its permissions are ‘-rw-r–r–’, meaning the ‘x’ permission is missing.
I ran:

chmod +x /usr/src/nvidia-610.43.02/pahole.sh

This means that when driver 610.43.02 (newer than 595.58.03 – see below) is installed on a Debian system, the compilation of the NVIDIA driver modules will fail when it’s time to update the kernel – the cause is the lack of the executable attribute for the ‘pahole.sh’ file.

Therefore, as soon as we correctly install a driver newer than 595.58.03, the first step we must take is:

chmod +x /usr/src/nvidia-*/pahole.sh

This will likely allow for future kernel updates without issues.

Of course, during driver installation, when asked: “Would you like to register the kernel module sources with dkms?”, answer: “Yes”.

Checking if the ‘pahole.sh’ file is present in the driver directory

According to:

the subsequent driver versions are:
595.58.03
595.71.05
595.80
610.43.02

I am using the ‘no-compat32’ files:
NVIDIA-Linux-x86_64_{nr-ver}–no-compat32.run

I extracted these files (e.g.: ‘./NVIDIA-Linux-x86_64_595.58.03–no-compat32.run -x’)

Version 595.58.03 does not include the ‘pahole.sh’ file.

In version 595.71.05, it is present:
NVIDIA-Linux-x86_64-595.71.05-no-compat32/kernel/pahole.sh
NVIDIA-Linux-x86_64-595.71.05-no-compat32/kernel-open/pahole.sh

In version 595.80 – it is:
NVIDIA-Linux-x86_64-595.80-no-compat32/kernel/pahole.sh
NVIDIA-Linux-x86_64-595.80-no-compat32/kernel-open/pahole.sh

In version 610.43.02 – it is:
NVIDIA-Linux-x86_64-610.43.02-no-compat32/kernel/pahole.sh
NVIDIA-Linux-x86_64-610.43.02-no-compat32/kernel-open/pahole.sh

These are the same files with an MD5 sum of: e84bbf8835bd40a78320bb6b3a7d3c49.

Perhaps my description will be useful to someone.

Best regards

Added on May 30:

In addition to the ‘pahole.sh’ file, there is also the ‘conftest.sh’ file, which probably also needs to be made executable. Therefore, we should issue the command:

chmod +x /usr/src/nvidia-*/*sh

edit: in reply to 610 release feedback & discussion - #22 by Corben78

To create a bug report, I re-installed 610.43.02 and tried to see if the crash happens again at the same situation. It did not. Yet after playing a bit further, it happened at a later point:

[11240.127018] NVRM: GPU at PCI:0000:01:00: GPU-3acf5c71-99a1-b081-aca0-d79b2e9238e8
[11240.127022] NVRM: Xid (PCI:0000:01:00): 31, pid=44497, name=GameThread, channel 0x00000028, intr 00000000. MMU Fault: ENGINE GRAPHICS HUBCLIENT_FE faulted @ 0x1_04e00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_WRITE

MetalEden.log and bug report created by nvidia-bug-report.sh attached.

MetalEden.zip (629.5 KB)

nvidia-bug-report.log.gz (770.9 KB)

just enabled VRR in both hyprland and mangowm and it does not flicker at all on my 5070ti

Link to README: Appendix L. Wayland Known Issues

Workaround HDR command: modprobe -r nvidia_drm ; modprobe nvidia_drm color_pipeline=0

It should still work with multiple X screens and with multiple displays on a single X screen, just not with Xinerama. Can you please attach a bug report log from your configuration?

I’ve seen this opinion on the web, but couldn’t find any specific benchmarks: would you be able to share any, so that we get an idea what order of magnitude the difference actually is?
Thanks! :)

I don’t have much experience running compute workloads on AMD, but I’ve heard that on any card that supports Vulkan, you should use that as the backend: ROCm works ok only on sever Instinct cards.

Thanks for your reply. The language barrier. The issue I am raising is related to consumer GPUs, such as the 5090 and Blackwell 6000 Pro in desktop environments. Moving from the 580 driver on Linux has become a bit of a conundrum, with non-deterministic issues appearing. I haven’t seen any significant problems with server-grade GPUs, working fine.

The latest 610 driver broke my linux machine, and I need to reimage the OS. The last major change I made was installing the driver. I don’t think this is related to the motherboard, thermal issues, BIOS, or anything like that.

There is a well-known issue where the 595 driver breaks Nvidia Isaac Sim on Ubuntu: Installing Isaac Sim 5.1.0 on a Dual Blackwell RTX PRO 6000 Workstation

It’s a bit strange that instead of receiving proper fixes, people often end up finding workarounds to get things working. With all high respect, it is understandable to have problems, this is part of engineering and life, there is nothing perfect, but I don’t see much engagement from Nvidia in terms of providing proper feedback or addressing issues. It is really disappointing to see this.

you wont get much in the way of 5070ti linux benchmarks (and heck even windows ones) but on windows ones ive found it sits around the performance of a 4080 to 4080super in compute. You can see in this phoronix article the 4080 pretty handily beats the 9070xt.

https://www.phoronix.com/review/amd-radeon-rx9070-linux-compute

some windows benchmarks for comparison, not direct but to give an idea of 5070ti performance

ROCM works across more than Instinct cards (they have a table for hardware with official support), it just kinda sucks and would frequently break things with newer versions especially in blender. It unofficially works on most of their hardware RDNA1 and newer. I always kinda feared ROCM updates as it was a coin toss.

Vulkan compute isnt going to save you on performance though
https://www.phoronix.com/review/rocm-71-llama-cpp-vulkan

AMD is still worse ( it trades blows with ROCM on AMD) and Im not aware of anyone actually focusing on using Vulkan backend for serious workloads yet, i could be wrong though. Cuda and ROCM would still be the target for most especially due to ROCM offering some assistance in converting Cuda.

This thread is about 610 release feedback and discussion, not Nvidia Vs. AMD.