Volatile Uncorr. ECC

I am having my Models crashing on my RTX Pro 6000 despite reboots.

Error:

sudo dmesg | grep -i “NVRM: Xid”
[ 2146.678661] NVRM: Xid (PCI:b4a2:00:00): 48, An uncorrectable double bit error (DBE) has been detected on GPU in the framebuffer at physAddr 0x1246514e60 partition 5, subpartition 1.
[ 2146.678670] NVRM: Xid (PCI:b4a2:00:00): 171, GDDR, Uncorrectable DRAM error in FBPA 5 subpartition 1 physAddr 0x1246514e60
[ 2146.701859] NVRM: Xid (PCI:b4a2:00:00): 48, pid=3370, name=nvidia-cuda-mps, channel 0x00000002
[ 2146.723831] NVRM: Xid (PCI:b4a2:00:00): 48, pid=3298, name=VLLM::EngineCor, channel 0x00000003
[ 2146.748707] NVRM: Xid (PCI:b4a2:00:00): 48, pid=3298, name=VLLM::EngineCor, channel 0x00000004
[ 2146.771851] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000005
[ 2146.781888] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000006
[ 2146.784129] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000007
[ 2146.786399] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000008
[ 2146.788598] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000009
[ 2146.790743] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x0000000a
[ 2146.793012] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x0000000b
[ 2146.795256] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x0000000c
[ 2146.801855] NVRM: Xid (PCI:b4a2:00:00): 63, pid=3370, name=cuda-EvtHandlr, Row Remapper: New row (0x0000001246514e60) marked for remapping, reset gpu to activate.
[ 2146.818170] NVRM: Xid (PCI:b4a2:00:00): 154, GPU recovery action changed from 0x0 (None) to 0x4 (Drain and Reset)

technocrat@ubunturtx6000:~/llm4$ sudo dmesg | grep -i “NVRM: Xid”
[ 2146.678661] NVRM: Xid (PCI:b4a2:00:00): 48, An uncorrectable double bit error (DBE) has been detected on GPU in the framebuffer at physAddr 0x1246514e60 partition 5, subpartition 1.
[ 2146.678670] NVRM: Xid (PCI:b4a2:00:00): 171, GDDR, Uncorrectable DRAM error in FBPA 5 subpartition 1 physAddr 0x1246514e60
[ 2146.701859] NVRM: Xid (PCI:b4a2:00:00): 48, pid=3370, name=nvidia-cuda-mps, channel 0x00000002
[ 2146.723831] NVRM: Xid (PCI:b4a2:00:00): 48, pid=3298, name=VLLM::EngineCor, channel 0x00000003
[ 2146.748707] NVRM: Xid (PCI:b4a2:00:00): 48, pid=3298, name=VLLM::EngineCor, channel 0x00000004
[ 2146.771851] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000005
[ 2146.781888] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000006
[ 2146.784129] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000007
[ 2146.786399] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000008
[ 2146.788598] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x00000009
[ 2146.790743] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x0000000a
[ 2146.793012] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x0000000b
[ 2146.795256] NVRM: Xid (PCI:b4a2:00:00): 48, pid=4632, name=VLLM::EngineCor, channel 0x0000000c
[ 2146.801855] NVRM: Xid (PCI:b4a2:00:00): 63, pid=3370, name=cuda-EvtHandlr, Row Remapper: New row (0x0000001246514e60) marked for remapping, reset gpu to activate.
[ 2146.818170] NVRM: Xid (PCI:b4a2:00:00): 154, GPU recovery action changed from 0x0 (None) to 0x4 (Drain and Reset)
technocrat@ubunturtx6000:~/llm4$ nvidia-smi -q -d ECC

==============NVSMI LOG==============

Timestamp : Mon May 25 13:54:09 2026
Driver Version : 580.126.20
CUDA Version : 13.0

Attached GPUs : 1
GPU 0000B4A2:00:00.0
ECC Mode
Current : Enabled
Pending : Enabled
ECC Errors
Volatile
SRAM Correctable : 0
SRAM Uncorrectable Parity : 0
SRAM Uncorrectable SEC-DED : 0
DRAM Correctable : 0
DRAM Uncorrectable : 1
Aggregate
SRAM Correctable : 0
SRAM Uncorrectable Parity : 0
SRAM Uncorrectable SEC-DED : 0
DRAM Correctable : 0
DRAM Uncorrectable : 1215
SRAM Threshold Exceeded : No
Aggregate Uncorrectable SRAM Sources
SRAM L2 : 0
SRAM SM : 0
SRAM Microcontroller : 0
SRAM PCIE : 0
SRAM Other : 0
Channel Repair Pending : No
TPC Repair Pending : No

sudo nvidia-smi -q

==============NVSMI LOG==============

Timestamp : Mon May 25 14:09:05 2026
Driver Version : 580.126.20
CUDA Version : 13.0

Attached GPUs : 1
GPU 0000B4A2:00:00.0
Product Name : NVIDIA RTX PRO 6000 Blackwell Server Edition
Product Brand : NVIDIA
Product Architecture : Blackwell
Display Mode : Requested functionality has been deprecated
Display Attached : Yes
Display Active : Disabled
Persistence Mode : Enabled
Addressing Mode : HMM
MIG Mode
Current : Disabled
Pending : Disabled
Accounting Mode : Disabled
Accounting Mode Buffer Size : 4000
Driver Model
Current : N/A
Pending : N/A
Serial Number : 1323425424869
GPU UUID : GPU-6bb1acee-557f-8d27-1fae-eb6d42e25b67
GPU PDI : 0xcdb0ea35d61fb25e
Minor Number : 0
VBIOS Version : 98.02.8D.00.01
MultiGPU Board : No
Board ID : 0xb4a20000
Board Part Number : 900-2G153-0000-000
GPU Part Number : 2BB5-895-A1
FRU Part Number : N/A
Platform Info
Chassis Serial Number :
Slot Number : 0
Tray Index : 0
Host ID : 1
Peer Type : Direct Connected
Module Id : 1
GPU Fabric GUID : 0x0000000000000000
Inforom Version
Image Version : G153.0210.00.02
OEM Object : 2.1
ECC Object : 7.16
Power Management Object : N/A
Inforom BBX Object Flush
Latest Timestamp : 2026/05/25 13:58:00.419
Latest Duration : 12159 us
GPU Operation Mode
Current : N/A
Pending : N/A
GPU C2C Mode : Disabled
GPU Virtualization Mode
Virtualization Mode : Pass-Through
Host VGPU Mode : N/A
vGPU Heterogeneous Mode : N/A
GPU Recovery Action : Drain and Reset
GSP Firmware Version : 580.126.20
IBMNPU
Relaxed Ordering Mode : N/A
PCI
Bus : 0x00
Device : 0x00
Domain : 0xB4A2
Base Classcode : 0x3
Sub Classcode : 0x2
Device Id : 0x2BB510DE
Bus Id : 0000B4A2:00:00.0
Sub System Id : 0x204E10DE
GPU Link Info
PCIe Generation
Max : 4
Current : 1
Device Current : 1
Device Max : 5
Host Max : N/A
Link Width
Max : 16x
Current : 16x
Bridge Chip
Type : N/A
Firmware : N/A
Replays Since Reset : 0
Replay Number Rollovers : 0
Tx Throughput : 1658 KB/s
Rx Throughput : 1115 KB/s
Atomic Caps Outbound : N/A
Atomic Caps Inbound : FETCHADD_32 FETCHADD_64 SWAP_32 SWAP_64 CAS_32 CAS_64
Fan Speed : N/A
Performance State : P8
Clocks Event Reasons
Idle : Not Active
Applications Clocks Setting : Not Active
SW Power Cap : Not Active
HW Slowdown : Not Active
HW Thermal Slowdown : Not Active
HW Power Brake Slowdown : Not Active
Sync Boost : Not Active
SW Thermal Slowdown : Not Active
Display Clock Setting : Not Active
Clocks Event Reasons Counters
SW Power Capping : 1238520 us
Sync Boost : 0 us
SW Thermal Slowdown : 1238520 us
HW Thermal Slowdown : 0 us
HW Power Braking : 0 us
Sparse Operation Mode : N/A
FB Memory Usage
Total : 97887 MiB
Reserved : 638 MiB
Used : 3 MiB
Free : 97247 MiB
BAR1 Memory Usage
Total : 131072 MiB
Used : 4 MiB
Free : 131068 MiB
Conf Compute Protected Memory Usage
Total : 0 MiB
Used : 0 MiB
Free : 0 MiB
Compute Mode : Default
Utilization
GPU : 0 %
Memory : 0 %
Encoder : 0 %
Decoder : 0 %
JPEG : 0 %
OFA : 0 %
Encoder Stats
Active Sessions : 0
Average FPS : 0
Average Latency : 0
FBC Stats
Active Sessions : 0
Average FPS : 0
Average Latency : 0
DRAM Encryption Mode
Current : Disabled
Pending : Disabled
ECC Mode
Current : Enabled
Pending : Enabled
ECC Errors
Volatile
SRAM Correctable : 0
SRAM Uncorrectable Parity : 0
SRAM Uncorrectable SEC-DED : 0
DRAM Correctable : 0
DRAM Uncorrectable : 1
Aggregate
SRAM Correctable : 0
SRAM Uncorrectable Parity : 0
SRAM Uncorrectable SEC-DED : 0
DRAM Correctable : 0
DRAM Uncorrectable : 1216
SRAM Threshold Exceeded : No
Aggregate Uncorrectable SRAM Sources
SRAM L2 : 0
SRAM SM : 0
SRAM Microcontroller : 0
SRAM PCIE : 0
SRAM Other : 0
Channel Repair Pending : No
TPC Repair Pending : No
Retired Pages
Single Bit ECC : N/A
Double Bit ECC : N/A
Pending Page Blacklist : N/A
Remapped Rows
Correctable Error : 0
Uncorrectable Error : 9
Pending : Yes
Remapping Failure Occurred : No
Bank Remap Availability Histogram
Max : 508 bank(s)
High : 1 bank(s)
Partial : 3 bank(s)
Low : 0 bank(s)
None : 0 bank(s)
Temperature
GPU Current Temp : 35 C
GPU T.Limit Temp : 50 C
GPU Shutdown T.Limit Temp : -5 C
GPU Slowdown T.Limit Temp : -2 C
GPU Max Operating T.Limit Temp : 0 C
GPU Target Temperature : N/A
Memory Current Temp : N/A
Memory Max Operating T.Limit Temp : N/A
GPU Power Readings
Average Power Draw : 38.56 W
Instantaneous Power Draw : 40.25 W
Current Power Limit : 600.00 W
Requested Power Limit : 600.00 W
Default Power Limit : 600.00 W
Min Power Limit : 300.00 W
Max Power Limit : 600.00 W
GPU Memory Power Readings
Average Power Draw : N/A
Instantaneous Power Draw : N/A
Module Power Readings
Average Power Draw : N/A
Instantaneous Power Draw : N/A
Current Power Limit : N/A
Requested Power Limit : N/A
Default Power Limit : N/A
Min Power Limit : N/A
Max Power Limit : N/A
Power Smoothing : N/A
Workload Power Profiles
Requested Profiles : N/A
Enforced Profiles : N/A
Clocks
Graphics : 277 MHz
SM : 277 MHz
Memory : 405 MHz
Video : 600 MHz
Applications Clocks
Graphics : 2430 MHz
Memory : 12481 MHz
Default Applications Clocks
Graphics : 2430 MHz
Memory : 12481 MHz
Deferred Clocks
Memory : N/A
Max Clocks
Graphics : 2430 MHz
SM : 2430 MHz
Memory : 12481 MHz
Video : 2107 MHz
Max Customer Boost Clocks
Graphics : 2430 MHz
Clock Policy
Auto Boost : N/A
Auto Boost Default : N/A
Fabric
State : N/A
Status : N/A
CliqueId : N/A
ClusterUUID : N/A
Health
Summary : N/A
Bandwidth : N/A
Route Recovery in progress : N/A
Route Unhealthy : N/A
Access Timeout Recovery : N/A
Incorrect Configuration : N/A
Partition Assigned : N/A
Processes : None
Capabilities
EGM : disabled