RTX PRO 6000 Blackwell Max-Q hardware problem? Multitude of MMU errors in dmesg

Hello, I did not find a better category to post this topic so I chose “Linux”.

I have a server with two PRO 6000 Blackwell cards, one is old 600W “Workstation” and another is new 300W “Max-Q Workstation”.
The 600W is more than a year old and I have never experienced any problems with it. The 300W is recently purchased and it seems to have a hardware problem, I suspect the problem is with the memory chips or memory controller.
Unfortunately the shop return period has passed already so I can not simply replace the card through the shop and must send it to Nvidia service center for repair or replacement. And I need your assistance in preliminary testing and writing the RMA request.

Problem: I run “llama-server” from llama.cpp software suite and sometimes it crashes on the first run, and sometimes seems to work well but hundreds of Xid 13 errors appear in “dmesg”, like this:

[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5477b0=0x0 0x5477b4=0x20 0x5477a8=0x1c81fb60 0x5477ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x548730=0x0 0x548734=0x20 0x548728=0x1c81fb60 0x54872c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5487b0=0x0 0x5487b4=0x20 0x5487a8=0x1c81fb60 0x5487ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x549730=0x0 0x549734=0x20 0x549728=0x1c81fb60 0x54972c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5497b0=0x0 0x5497b4=0x20 0x5497a8=0x1c81fb60 0x5497ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54a730=0x0 0x54a734=0x20 0x54a728=0x1c81fb60 0x54a72c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54a7b0=0x0 0x54a7b4=0x20 0x54a7a8=0x1c81fb60 0x54a7ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54b730=0x0 0x54b734=0x20 0x54b728=0x1c81fb60 0x54b72c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54b7b0=0x0 0x54b7b4=0x20 0x54b7a8=0x1c81fb60 0x54b7ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54c730=0x0 0x54c734=0x20 0x54c728=0x1c81fb60 0x54c72c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics SM Warp Exception on (GPC 4, TPC 7, SM 1): MMU Fault
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics SM Global Exception on (GPC 4, TPC 7, SM 1): Multiple Warp Errors
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54c7b0=0x8040017 0x54c7b4=0x24 0x54c7a8=0x1c81fb60 0x54c7ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x555730=0x0 0x555734=0x20 0x555728=0x1c81fb60 0x55572c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5557b0=0x0 0x5557b4=0x20 0x5557a8=0x1c81fb60 0x5557ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x556730=0x0 0x556734=0x20 0x556728=0x1c81fb60 0x55672c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5567b0=0x0 0x5567b4=0x20 0x5567a8=0x1c81fb60 0x5567ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x557730=0x0 0x557734=0x20 0x557728=0x1c81fb60 0x55772c=0x1174

Sometimes the application crashes right after the start, but this time with Xid 31:

[Mon Aug 24 21:52:52 2026] NVRM: Xid (PCI:0000:05:00): 31, pid=75397, name=llama-server, channel 0x00000004, intr 00000000. MMU Fault: ENGINE GRAPHICS GPC11 GPCCLIENT_T1_13 faulted @ 0x1c_02c00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ
[Tue Aug 25 11:51:21 2026] NVRM: Xid (PCI:0000:05:00): 31, pid=152066, name=llama-server_26, channel 0x00000004, intr 00000000. MMU Fault: ENGINE GRAPHICS GPC11 GPCCLIENT_T1_7 faulted @ 0x1c_02c00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ
[Tue Aug 25 15:04:49 2026] NVRM: Xid (PCI:0000:05:00): 31, pid=183004, name=llama-server, channel 0x00000004, intr 00000000. MMU Fault: ENGINE GRAPHICS GPC2 GPCCLIENT_T1_0 faulted @ 0x1c_02c00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ

The errors are always about PCI:0000:05:00, which is the current location of Max-Q card, and never about the 600W Workstation model.

What I’ve tried so far:

  1. swapped the 600W and Max-Q to ensure it is not a bad PCIe connection; it did not help, just “NVRM: Xid …” errors started reporting a different PCI location.
  2. ran a memtest to ensure it is not a system RAM problem; no errors found.
  3. updated NVIDIA driver (595->610) and CUDA (13.2->13.3) and recompiled llama.cpp to ensure it is not a software problem; and this also did not help.

nvidia-smi reports 0 ECC errors:

ECC Mode
    Current                                        : Enabled
    Pending                                        : Enabled
ECC Errors
    Volatile   
        SRAM Correctable                           : 0
        SRAM Uncorrectable Parity                  : 0
        SRAM Uncorrectable SEC-DED                 : 0
        DRAM Correctable                           : 0
        DRAM Uncorrectable                         : 0
    Aggregate  
        SRAM Correctable                           : 0
        SRAM Uncorrectable Parity                  : 0
        SRAM Uncorrectable SEC-DED                 : 0
        DRAM Correctable                           : 0
        DRAM Uncorrectable                         : 0
        SRAM Threshold Exceeded                    : No
    Aggregate Uncorrectable SRAM Sources
        SRAM L2                                    : 0
        SRAM SM                                    : 0
        SRAM Microcontroller                       : 0
        SRAM PCIE                                  : 0
        SRAM Other                                 : 0
    Channel Repair Pending                         : No
    TPC Repair Pending                             : No
    Unrepairable Memory                            : No
Retired Pages  
    Single Bit ECC                                 : N/A
    Double Bit ECC                                 : N/A
    Pending Page Blacklist                         : N/A
Remapped Rows  
    Correctable Error                              : 0
    Inactive Correctable Error                     : 0
    Uncorrectable Error                            : 0
    Inactive Uncorrectable Error                   : 0
    Pending                                        : No
    Remapping Failure Occurred                     : No
    Bank Remap Availability Histogram
        Max                                        : 512 bank(s)
        High                                       : 0 bank(s)
        Partial                                    : 0 bank(s)
        Low                                        : 0 bank(s)
        None                                       : 0 bank(s)

Is there any kind of memtest for NVIDIA cards? Google suggests memcheck but it is a software debugging tool, not a hardware memory test like memtest for the system RAM.

What else should I check to make sure that it is a hardware error?

How do I fill a RMA request for this situation? I guess I can not simply write “something is wrong with the card, please check”.

Yes there are “MODS” and “MATS” tools but they are proprietary and I was able to find only leaked versions only for 50*0 GPUs, not for 6000. And I have just found out that this forum actually has a “Hardware” category :D