API_GPU_ATTACHED_SANITY_CHECK_FAILED with no reliable reproduction steps on Fedora 43 KDE

Sep 02 13:18:12 bobtheskull kernel: [drm:nv_drm_gem_alloc_nvkms_memory_ioctl [nvidia_drm]] *ERROR* [nvidia-drm] [GPU ID 0x00000100] Failed to allocate NVKMS memory for GEM object
Sep 02 13:18:12 bobtheskull kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Sep 02 13:18:12 bobtheskull kernel: NVRM: _scrubWaitAndSave: Timed out when waiting for scrub jobs to finish.
Sep 02 13:18:12 bobtheskull kernel: NVRM: nvCheckOkFailedNoLog: Check failed: Call timed out [NV_ERR_TIMEOUT] (0x00000065) returned from _scrubWaitAndSave(pScrubber, pList, requiredItemsToSave) @ mem_scrub.c:626
Sep 02 13:18:12 bobtheskull kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Sep 02 13:18:12 bobtheskull kernel: NVRM: _scrubWaitAndSave: Timed out when waiting for scrub jobs to finish.
Sep 02 13:18:12 bobtheskull kernel: NVRM: nvCheckOkFailedNoLog: Check failed: Call timed out [NV_ERR_TIMEOUT] (0x00000065) returned from _scrubWaitAndSave(pScrubber, pList, requiredItemsToSave) @ mem_scrub.c:626
Sep 02 13:18:12 bobtheskull kernel: [drm:nv_drm_gem_alloc_nvkms_memory_ioctl [nvidia_drm]] *ERROR* [nvidia-drm] [GPU ID 0x00000100] Failed to allocate NVKMS memory for GEM object
Sep 02 13:18:12 bobtheskull kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Sep 02 13:18:12 bobtheskull kernel: NVRM: _scrubWaitAndSave: Timed out when waiting for scrub jobs to finish.
Sep 02 13:18:12 bobtheskull kernel: NVRM: nvCheckOkFailedNoLog: Check failed: Call timed out [NV_ERR_TIMEOUT] (0x00000065) returned from _scrubWaitAndSave(pScrubber, pList, requiredItemsToSave) @ mem_scrub.c:626
Sep 02 13:18:12 bobtheskull kwin_wayland[2200]: Failed to find a working output layer configuration! Enabled layers:
Sep 02 13:18:12 bobtheskull kwin_wayland[2200]: src KWin::RectF(0,0 3440x1440) -> dst KWin::Rect(0,0 3440x1440)
Sep 02 13:18:12 bobtheskull kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Sep 02 13:18:12 bobtheskull kernel: NVRM: _scrubWaitAndSave: Timed out when waiting for scrub jobs to finish.
Sep 02 13:18:12 bobtheskull kernel: NVRM: nvCheckOkFailedNoLog: Check failed: Call timed out [NV_ERR_TIMEOUT] (0x00000065) returned from _scrubWaitAndSave(pScrubber, pList, requiredItemsToSave) @ mem_scrub.c:626
Sep 02 13:18:12 bobtheskull kernel: [drm:nv_drm_gem_alloc_nvkms_memory_ioctl [nvidia_drm]] *ERROR* [nvidia-drm] [GPU ID 0x00000100] Failed to allocate NVKMS memory for GEM object
Sep 02 13:18:12 bobtheskull kwin_wayland[2200]: Pageflip timed out! This is a bug in the nvidia-drm kernel driver
Sep 02 13:18:12 bobtheskull kwin_wayland[2200]: Please report this at https://forums.developer.nvidia.com/c/gpu-graphics/linux
Sep 02 13:18:12 bobtheskull kwin_wayland[2200]: With the output of 'sudo dmesg' and 'journalctl --user-unit plasma-kwin_wayland --boot 0'
Sep 02 13:18:12 bobtheskull kernel: nvidia-modeset: ERROR: GPU:0: Failed detecting connected display devices
Sep 02 13:18:12 bobtheskull kernel: nvidia-modeset: WARNING: GPU:0: Failure processing EDID for display device DP-0.
Sep 02 13:18:12 bobtheskull kernel: nvidia-modeset: WARNING: GPU:0: Unable to read EDID for display device DP-0
Sep 02 13:18:12 bobtheskull kernel: nvidia-modeset: ERROR: GPU:0: Failed detecting connected display devices

Above is a piece of journalctl following a GPU related hang. This has been a recurring problem for some time but this is the first time I’ve seen the log message relating to the kernel driver bug.

dmesg.txt (116.7 KB)

nvidia-bug-report.log.gz (653.2 KB)

journalctl_plasma-kwin_wayland.txt (17 Bytes)

The problem is intermittant, sometimes as often as 5 minute intervals, other times weeks. I have not found a reliable way to reproduce.

i have exact the same Problem.
In my Case on CachyOS Arch Linux.

I discussed this with ChatGPT. He says my configuration is fine and there’s no hardware problem. Apparently, the server sometimes boots with only 3.2GB of VRAM instead of the full 24GB. He also noticed a timeout issue where it tried to allocate RAM but couldn’t. I went through all the configurations and graphics card outputs with ChatGPT. He thinks it’s very likely a problem caused by a bug in the kernel or graphics driver.