Hello, I did not find a better category to post this topic so I chose “Linux”.
I have a server with two PRO 6000 Blackwell cards, one is old 600W “Workstation” and another is new 300W “Max-Q Workstation”.
The 600W is more than a year old and I have never experienced any problems with it. The 300W is recently purchased and it seems to have a hardware problem, I suspect the problem is with the memory chips or memory controller.
Unfortunately the shop return period has passed already so I can not simply replace the card through the shop and must send it to Nvidia service center for repair or replacement. And I need your assistance in preliminary testing and writing the RMA request.
Problem: I run “llama-server” from llama.cpp software suite and sometimes it crashes on the first run, and sometimes seems to work well but hundreds of Xid 13 errors appear in “dmesg”, like this:
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5477b0=0x0 0x5477b4=0x20 0x5477a8=0x1c81fb60 0x5477ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x548730=0x0 0x548734=0x20 0x548728=0x1c81fb60 0x54872c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5487b0=0x0 0x5487b4=0x20 0x5487a8=0x1c81fb60 0x5487ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x549730=0x0 0x549734=0x20 0x549728=0x1c81fb60 0x54972c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5497b0=0x0 0x5497b4=0x20 0x5497a8=0x1c81fb60 0x5497ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54a730=0x0 0x54a734=0x20 0x54a728=0x1c81fb60 0x54a72c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54a7b0=0x0 0x54a7b4=0x20 0x54a7a8=0x1c81fb60 0x54a7ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54b730=0x0 0x54b734=0x20 0x54b728=0x1c81fb60 0x54b72c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54b7b0=0x0 0x54b7b4=0x20 0x54b7a8=0x1c81fb60 0x54b7ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54c730=0x0 0x54c734=0x20 0x54c728=0x1c81fb60 0x54c72c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics SM Warp Exception on (GPC 4, TPC 7, SM 1): MMU Fault
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics SM Global Exception on (GPC 4, TPC 7, SM 1): Multiple Warp Errors
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x54c7b0=0x8040017 0x54c7b4=0x24 0x54c7a8=0x1c81fb60 0x54c7ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x555730=0x0 0x555734=0x20 0x555728=0x1c81fb60 0x55572c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5557b0=0x0 0x5557b4=0x20 0x5557a8=0x1c81fb60 0x5557ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x556730=0x0 0x556734=0x20 0x556728=0x1c81fb60 0x55672c=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x5567b0=0x0 0x5567b4=0x20 0x5567a8=0x1c81fb60 0x5567ac=0x1174
[Tue Aug 25 11:24:34 2026] NVRM: Xid (PCI:0000:05:00): 13, Graphics Exception: ESR 0x557730=0x0 0x557734=0x20 0x557728=0x1c81fb60 0x55772c=0x1174
Sometimes the application crashes right after the start, but this time with Xid 31:
[Mon Aug 24 21:52:52 2026] NVRM: Xid (PCI:0000:05:00): 31, pid=75397, name=llama-server, channel 0x00000004, intr 00000000. MMU Fault: ENGINE GRAPHICS GPC11 GPCCLIENT_T1_13 faulted @ 0x1c_02c00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ
[Tue Aug 25 11:51:21 2026] NVRM: Xid (PCI:0000:05:00): 31, pid=152066, name=llama-server_26, channel 0x00000004, intr 00000000. MMU Fault: ENGINE GRAPHICS GPC11 GPCCLIENT_T1_7 faulted @ 0x1c_02c00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ
[Tue Aug 25 15:04:49 2026] NVRM: Xid (PCI:0000:05:00): 31, pid=183004, name=llama-server, channel 0x00000004, intr 00000000. MMU Fault: ENGINE GRAPHICS GPC2 GPCCLIENT_T1_0 faulted @ 0x1c_02c00000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ
The errors are always about PCI:0000:05:00, which is the current location of Max-Q card, and never about the 600W Workstation model.
What I’ve tried so far:
- swapped the 600W and Max-Q to ensure it is not a bad PCIe connection; it did not help, just “NVRM: Xid …” errors started reporting a different PCI location.
- ran a
memtestto ensure it is not a system RAM problem; no errors found. - updated NVIDIA driver (595->610) and CUDA (13.2->13.3) and recompiled
llama.cppto ensure it is not a software problem; and this also did not help.
nvidia-smi reports 0 ECC errors:
ECC Mode
Current : Enabled
Pending : Enabled
ECC Errors
Volatile
SRAM Correctable : 0
SRAM Uncorrectable Parity : 0
SRAM Uncorrectable SEC-DED : 0
DRAM Correctable : 0
DRAM Uncorrectable : 0
Aggregate
SRAM Correctable : 0
SRAM Uncorrectable Parity : 0
SRAM Uncorrectable SEC-DED : 0
DRAM Correctable : 0
DRAM Uncorrectable : 0
SRAM Threshold Exceeded : No
Aggregate Uncorrectable SRAM Sources
SRAM L2 : 0
SRAM SM : 0
SRAM Microcontroller : 0
SRAM PCIE : 0
SRAM Other : 0
Channel Repair Pending : No
TPC Repair Pending : No
Unrepairable Memory : No
Retired Pages
Single Bit ECC : N/A
Double Bit ECC : N/A
Pending Page Blacklist : N/A
Remapped Rows
Correctable Error : 0
Inactive Correctable Error : 0
Uncorrectable Error : 0
Inactive Uncorrectable Error : 0
Pending : No
Remapping Failure Occurred : No
Bank Remap Availability Histogram
Max : 512 bank(s)
High : 0 bank(s)
Partial : 0 bank(s)
Low : 0 bank(s)
None : 0 bank(s)
Is there any kind of memtest for NVIDIA cards? Google suggests memcheck but it is a software debugging tool, not a hardware memory test like memtest for the system RAM.
What else should I check to make sure that it is a hardware error?
How do I fill a RMA request for this situation? I guess I can not simply write “something is wrong with the card, please check”.