A100 SXM4 GPUs Disappearing from OS and BMC/IPMI
Hello,
We are investigating an issue on an NVIDIA A100 SXM4 system where the node becomes unresponsive and eventually reboots. Before reboot, in the IPMI console the GPUs were visible, After the reboot,
kernel_logs.txt (99.1 KB)
nvidia-bug-report.log.gz.gz (325.6 KB)
one or more GPUs are no longer detected.
System Configuration
-
8 × NVIDIA A100 SXM4 40GB
-
NVSwitch-based platform
-
AMD EPYC system
-
Broadcom PEX880xx PCIe switches
-
Mellanox ConnectX-6 adapters
-
Linux OS
Symptoms
-
The node becomes unresponsive during operation.
-
Prior to the failure, we occasionally observe NVIDIA Xid 79 events:
NVRM: Xid ... 79, GPU has fallen off the bus
-
Following the event, the node may reboot or require a reboot.
-
After reboot, one or more GPUs are no longer detected by the operating system.
-
The missing GPUs are also not visible in the BMC/IPMI inventory.
-
In some cases, portions of the PCIe hierarchy appear to be missing after the event.
Kernel Messages
We have observed messages such as:
pcieport 0000:09:1f.0:
Unable to change power state from D0 to D3hot,
device inaccessible
During boot we also see:
acpi PNP0A08:00: _OSC: platform does not support [AER LTR DPC]
acpi PNP0A08:01: _OSC: platform does not support [AER LTR DPC]
...
PCIe Topology
On a healthy boot, the system enumerates:
-
All 8 A100 GPUs
-
All NVSwitch devices
-
Broadcom PEX880xx PCIe switches
-
PCIe switch management endpoints
-
ConnectX-6 adapters
The issue appears to occur below the NVIDIA driver layer since the affected GPUs are not visible in either the OS or the BMC after the event.
Questions
-
Has anyone observed Xid 79 events followed by GPUs disappearing from both the OS and BMC/IPMI?
-
Are there known issues involving A100 SXM4 systems, NVSwitch, or Broadcom PEX880xx PCIe switches that could cause a GPU or an entire PCIe branch to disappear?
-
Is the following message typically associated with a downstream PCIe device or switch becoming inaccessible?
pcieport 0000:09:1f.0:
Unable to change power state from D0 to D3hot,
device inaccessible
- What additional diagnostics would NVIDIA recommend to determine whether the root cause is GPU, PCIe switch, platform, power delivery, or NVSwitch related?
Any guidance would be appreciated.
Thank you.

