465.24.02 page fault

Hi, I am having a similar issue on a GTX 1060, i7 6700K running Arch Linux with current Kernel 5.11.16, nvidia 465.27-

[   31.430235] BUG: kernel NULL pointer dereference, address: 0000000000000000
[   31.430239] #PF: supervisor read access in kernel mode
[   31.430240] #PF: error_code(0x0000) - not-present page
[   31.430241] PGD 0 P4D 0
[   31.430243] Oops: 0000 [#1] PREEMPT SMP PTI
[   31.430245] CPU: 0 PID: 764 Comm: nv_queue Tainted: P           OE     5.11.16-arch1-1 #1
[   31.430247] Hardware name: System manufacturer System Product Name/Z170 PRO GAMING, BIOS 1104 01/11/2016
[   31.430248] RIP: 0010:_nv032382rm+0x21c/0x500 [nvidia]
[   31.430593] Code: 8b 87 e0 01 00 00 e8 c3 3f 4e d5 45 85 ff 45 89 fc 0f 94 c0 41 f7 d4 41 83 e6 01 75 05 84 45 30 75 2c 48 8b 45 18 48 8b 4d 20 <44> 23 38 44 89 39 44 23 20 48 8b 45 28 44 89 20 5b 41 5c 41 5d 41
[   31.430594] RSP: 0018:ffffaf9f80ed7d98 EFLAGS: 00010246
[   31.430596] RAX: 0000000000000000 RBX: 0000000000000004 RCX: 0000000000000000
[   31.430597] RDX: 0000000000000007 RSI: ffff97dcac260008 RDI: ffff97dcac098008
[   31.430598] RBP: ffff97dd2906af60 R08: 0000000000000001 R09: ffff97dd2906ae68
[   31.430599] R10: ffff97dcac098008 R11: 0000000010100000 R12: 00000000fffffbff
[   31.430600] R13: ffff97dcacdfc010 R14: 0000000000000000 R15: 0000000000000400
[   31.430601] FS:  0000000000000000(0000) GS:ffff97e39e400000(0000) knlGS:0000000000000000
[   31.430603] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   31.430604] CR2: 0000000000000000 CR3: 0000000513610001 CR4: 00000000003706f0
[   31.430605] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[   31.430606] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400
[   31.430607] Call Trace:
[   31.430609]  ? _nv026304rm+0xce/0x760 [nvidia]
[   31.430896]  ? _nv026309rm+0x1b0/0x1d0 [nvidia]
[   31.431181]  ? rm_execute_work_item+0x108/0x120 [nvidia]
[   31.431436]  ? schedule_timeout+0x11c/0x160
[   31.431440]  ? os_execute_work_item+0x46/0x60 [nvidia]
[   31.431633]  ? _main_loop+0x83/0x130 [nvidia]
[   31.431826]  ? nvidia_modeset_resume+0x20/0x20 [nvidia]
[   31.432019]  ? kthread+0x133/0x150
[   31.432022]  ? __kthread_bind_mask+0x60/0x60
[   31.432024]  ? ret_from_fork+0x22/0x30

~

I have two monitors

  • 24" 4K (portrait) connected over HDMI
  • 27" 2560x1440 connected over DP

What I tested so far:

  • Both plugged and I get the crash but still able to SSH.
  • If I unplug the 27" (DP) and only run the 24" (HDMI), then it works fine. Once booted I tried to plug the 27" DP display => driver crash
  • If I unplug the 24" and only run the 27", it crashes at boot and network was unavailable.

Looks like he 27" over DP is the culprit.

I ssh’ed to the box after the crash but nvidia-bug-report hangs, also when running with nvidia-bug-report.sh --safe-mode --extra-system-data

But in case the very limited info collected might help, I have attached it.
nvidia-bug-report.log.gz (1.1 KB)

This is nvidia-bug-report when I could boot with a single monitor (24" HDMI)
nvidia-bug-report-booting.log.gz (387.2 KB)

First time reporting here - let me know what else I could provide

When I downgrade to 460.67 it works fine.