B300 SXM6: FabricManager & NVLSM Link Up / Master, but NVLink P2P Traffic Fails under Load

Hi everyone,

Need some help here with our NVIDIA B300 SXM6 setup.

The issue is: NVLink looks completely fine when idle, but fails as soon as we run any actual load.

What we checked so far:

  • Static / Control Plane Status: Everything looks super clean. nvidia-smi nvlink -e shows healthy physical links (no major errors, link up). FabricManager and NVLSM (OpenSM) start without any issues. The logs literally say:

    “OpenSM Entering MASTER state” “Successfully configured all the available GPUs and NVSwitches to route NVLink traffic.”

  • The Problem: Despite FM and NVLSM claiming everything is successfully routed, the moment we trigger P2P traffic / workloads (like p2pBandwidthLatencyTest or PyTorch/NCCL), NVLink communication breaks / hangs / falls back.

Environment Info:Intel(R) Xeon(R) 6730P

GPU: 8x NVIDIA B300 SXM6 AC

root@amd-5090:~/cuda-samples-master/build/bin/x86_64/linux/release# nvidia-smi
Mon Aug 31 06:12:13 2026
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 610.57.04 KMD Version: 610.57.04 CUDA UMD Version: 13.3 |
±----------------------------------------±-----------------------±---------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA B300 SXM6 AC On | 00000000:1A:00.0 Off | 0 |
| N/A 31C P0 183W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 1 NVIDIA B300 SXM6 AC On | 00000000:40:00.0 Off | 0 |
| N/A 35C P0 184W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 2 NVIDIA B300 SXM6 AC On | 00000000:62:00.0 Off | 0 |
| N/A 34C P0 184W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 3 NVIDIA B300 SXM6 AC On | 00000000:73:00.0 Off | 0 |
| N/A 32C P0 181W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 4 NVIDIA B300 SXM6 AC On | 00000000:9A:00.0 Off | 0 |
| N/A 32C P0 182W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 5 NVIDIA B300 SXM6 AC On | 00000000:BD:00.0 Off | 0 |
| N/A 35C P0 181W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 6 NVIDIA B300 SXM6 AC On | 00000000:DF:00.0 Off | 0 |
| N/A 35C P0 184W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 7 NVIDIA B300 SXM6 AC On | 00000000:F0:00.0 Off | 0 |
| N/A 32C P0 180W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+

±----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
±----------------------------------------------------------------------------------------+

GPU 0: NVIDIA B300 SXM6 AC (UUID: GPU-8337aacd-dd11-a827-ae2c-9da20f8ea472)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 1: NVIDIA B300 SXM6 AC (UUID: GPU-bb5fe741-d156-7a95-8985-53829e581ac3)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 2: NVIDIA B300 SXM6 AC (UUID: GPU-9590ee95-e06d-59e8-15e4-5ba7fa9d47f2)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 3: NVIDIA B300 SXM6 AC (UUID: GPU-6aeb073e-15e8-ef5d-c2ca-023572e0bd61)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 4: NVIDIA B300 SXM6 AC (UUID: GPU-e822c480-886d-55e3-8f9a-4e39f8519279)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 5: NVIDIA B300 SXM6 AC (UUID: GPU-ffab3c6b-ff72-3301-6fdf-ebe8cede2f57)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 6: NVIDIA B300 SXM6 AC (UUID: GPU-6f11100c-e92f-830e-c248-de77bf05ea83)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 7: NVIDIA B300 SXM6 AC (UUID: GPU-71abcb2c-a244-4ca5-e077-2047cfdd5490)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s

root@amd-5090:~/cuda-samples-master/build/bin/x86_64/linux/release# systemctl status nvidia-fabricmanager.service
● nvidia-fabricmanager.service - NVIDIA fabric manager service
Loaded: loaded (/usr/lib/systemd/system/nvidia-fabricmanager.service; disabled; preset: enabled)
Active: active (running) since Mon 2026-08-31 05:54:35 UTC; 22min ago
Process: 89378 ExecStartPre=/usr/bin/nvidia-fabricmanager-start.sh --mode precheck (code=exited, status=0/SUCCESS)
Process: 89605 ExecStart=/usr/bin/nvidia-fabricmanager-start.sh --mode start (code=exited, status=0/SUCCESS)
Main PID: 89757 (nv-fabricmanage)
Tasks: 90 (limit: 629145)
Memory: 36.0M (peak: 42.4M)
CPU: 8.277s
CGroup: /system.slice/nvidia-fabricmanager.service
├─89735 /opt/nvidia/nvlsm/sbin/nvlsm -F /usr/share/nvidia/nvlsm/nvlsm.conf -B --pid_file /var/run/nvidia-fabricmanager/nvlsm.pid -g 0xac3ae20300b1773a
├─89738 osm_crashd
└─89757 /usr/bin/nv-fabricmanager -c /usr/share/nvidia/nvswitch/fabricmanager.cfg -g 0xac3ae20300b1773a

Aug 31 05:54:29 amd-5090 OpenSM[89732]: -I- Configuration loaded
Aug 31 05:54:29 amd-5090 nvidia-fabricmanager-start.sh[89605]: Started “Nvidia NVLink Subnet Manager”
Aug 31 05:54:29 amd-5090 OpenSM[89735]: /var/log/nvlsm.log log file opened
Aug 31 05:54:29 amd-5090 OpenSM[89735]: OpenSM 2025.10.14_42349a2_821fcee_ae214d2
Aug 31 05:54:29 amd-5090 OpenSM[89735]: Entering DISCOVERING state
Aug 31 05:54:32 amd-5090 OpenSM[89735]: Entering MASTER state
Aug 31 05:54:35 amd-5090 nv-fabricmanager[89757]: NodeId 0 partition id 57082 is activated.
Aug 31 05:54:35 amd-5090 nv-fabricmanager[89757]: Successfully configured all the available GPUs and NVSwitches to route NVLink traffic. NVLink Peer-to-Peer support will be enabled onc>
Aug 31 05:54:35 amd-5090 nvidia-fabricmanager-start.sh[89605]: Started “Nvidia Fabric Manager”
Aug 31 05:54:35 amd-5090 systemd[1]: Started nvidia-fabricmanager.service - NVIDIA fabric manager service.
root@amd-5090:~/cuda-samples-master/build/bin/x86_64/linux/release#

We are testing an 8x NVIDIA B300 SXM6 node and hit a critical roadblock: NVLink control plane looks 100% healthy at idle, but P2P traffic fails immediately under load.

  • Control Plane / Idle State (Looks Good):

    • All 8 GPUs are properly enumerated by nvidia-smi.

    • nvidia-smi nvlink -s shows all 12 links per GPU in Active state.

    • FabricManager & NVLSM (OpenSM) start clean with no errors. Logs explicitly confirm:

      “OpenSM Entering MASTER state”

      “Successfully configured all the available GPUs and NVSwitches to route NVLink traffic.”

  • Data Plane / Workloads (Fails Immediately):

    • The moment we trigger high-bandwidth P2P traffic (p2pBandwidthLatencyTest, nccl-tests, or PyTorch training), NVLink communication hangs, drops, or falls back to PCIe / SYS.

Below are our system outputs and runtime log details:

[P2P (Peer-to-Peer) GPU Bandwidth Latency Test]
Device: 0, NVIDIA B300 SXM6 AC, pciBusID: 1a, pciDeviceID: 0, pciDomainID:0
Device: 1, NVIDIA B300 SXM6 AC, pciBusID: 40, pciDeviceID: 0, pciDomainID:0
Device: 2, NVIDIA B300 SXM6 AC, pciBusID: 62, pciDeviceID: 0, pciDomainID:0
Device: 3, NVIDIA B300 SXM6 AC, pciBusID: 73, pciDeviceID: 0, pciDomainID:0
Device: 4, NVIDIA B300 SXM6 AC, pciBusID: 9a, pciDeviceID: 0, pciDomainID:0
Device: 5, NVIDIA B300 SXM6 AC, pciBusID: bd, pciDeviceID: 0, pciDomainID:0
Device: 6, NVIDIA B300 SXM6 AC, pciBusID: df, pciDeviceID: 0, pciDomainID:0
Device: 7, NVIDIA B300 SXM6 AC, pciBusID: f0, pciDeviceID: 0, pciDomainID:0
Device=0 CANNOT Access Peer Device=1
Device=0 CANNOT Access Peer Device=2
Device=0 CANNOT Access Peer Device=3
Device=0 CANNOT Access Peer Device=4
Device=0 CANNOT Access Peer Device=5
Device=0 CANNOT Access Peer Device=6
Device=0 CANNOT Access Peer Device=7
Device=1 CANNOT Access Peer Device=0
Device=1 CANNOT Access Peer Device=2
Device=1 CANNOT Access Peer Device=3
Device=1 CANNOT Access Peer Device=4
Device=1 CANNOT Access Peer Device=5
Device=1 CANNOT Access Peer Device=6
Device=1 CANNOT Access Peer Device=7
Device=2 CANNOT Access Peer Device=0
Device=2 CANNOT Access Peer Device=1
Device=2 CANNOT Access Peer Device=3
Device=2 CANNOT Access Peer Device=4
Device=2 CANNOT Access Peer Device=5
Device=2 CANNOT Access Peer Device=6
Device=2 CANNOT Access Peer Device=7
Device=3 CANNOT Access Peer Device=0
Device=3 CANNOT Access Peer Device=1
Device=3 CANNOT Access Peer Device=2
Device=3 CANNOT Access Peer Device=4
Device=3 CANNOT Access Peer Device=5
Device=3 CANNOT Access Peer Device=6
Device=3 CANNOT Access Peer Device=7
Device=4 CANNOT Access Peer Device=0
Device=4 CANNOT Access Peer Device=1
Device=4 CANNOT Access Peer Device=2
Device=4 CANNOT Access Peer Device=3
Device=4 CANNOT Access Peer Device=5
Device=4 CANNOT Access Peer Device=6
Device=4 CANNOT Access Peer Device=7
Device=5 CANNOT Access Peer Device=0
Device=5 CANNOT Access Peer Device=1
Device=5 CANNOT Access Peer Device=2
Device=5 CANNOT Access Peer Device=3
Device=5 CANNOT Access Peer Device=4
Device=5 CANNOT Access Peer Device=6
Device=5 CANNOT Access Peer Device=7
Device=6 CANNOT Access Peer Device=0
Device=6 CANNOT Access Peer Device=1
Device=6 CANNOT Access Peer Device=2
Device=6 CANNOT Access Peer Device=3
Device=6 CANNOT Access Peer Device=4
Device=6 CANNOT Access Peer Device=5
Device=6 CANNOT Access Peer Device=7
Device=7 CANNOT Access Peer Device=0
Device=7 CANNOT Access Peer Device=1
Device=7 CANNOT Access Peer Device=2
Device=7 CANNOT Access Peer Device=3
Device=7 CANNOT Access Peer Device=4
Device=7 CANNOT Access Peer Device=5
Device=7 CANNOT Access Peer Device=6

***NOTE: In case a device doesn’t have P2P access to other one, it falls back to normal memcopy procedure.
So you can see lesser Bandwidth (GB/s) and unstable Latency (us) in those cases.

P2P Connectivity Matrix
D\D 0 1 2 3 4 5 6 7
0 1 0 0 0 0 0 0 0
1 0 1 0 0 0 0 0 0
2 0 0 1 0 0 0 0 0
3 0 0 0 1 0 0 0 0
4 0 0 0 0 1 0 0 0
5 0 0 0 0 0 1 0 0
6 0 0 0 0 0 0 1 0
7 0 0 0 0 0 0 0 1
Unidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 5826.15 38.08 38.18 38.41 38.09 38.00 38.07 38.03
1 37.79 5896.23 38.21 38.14 37.66 37.81 38.11 38.12
2 37.78 37.65 5896.23 38.18 37.89 37.97 37.87 38.10
3 37.93 37.79 38.01 5896.23 38.06 38.13 38.06 37.91
4 37.85 37.95 37.78 38.02 5987.31 38.49 38.60 38.57
5 38.00 38.05 38.28 38.22 38.23 5919.26 38.38 38.61
6 37.71 38.08 38.22 37.99 38.39 38.55 5918.56 38.60
7 37.73 37.73 38.05 37.91 38.43 38.46 38.51 5941.06
Unidirectional P2P=Enabled Bandwidth (P2P Writes) Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 5896.23 38.10 37.93 38.22 37.98 37.93 38.20 37.92
1 37.99 5896.23 38.00 38.19 37.64 38.13 38.16 38.12
2 37.67 37.82 5896.23 38.01 37.86 38.07 38.02 38.07
3 37.67 37.81 38.00 5896.23 38.08 38.03 38.05 38.04
4 37.83 37.84 38.10 38.10 5940.36 38.44 38.60 38.55
5 37.89 38.19 38.12 38.29 38.27 5896.23 38.53 38.61
6 37.92 37.90 37.99 38.35 38.26 38.61 5941.77 38.54
7 37.82 37.79 37.95 38.28 38.29 38.22 38.40 5941.77
Bidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 6174.36 47.95 47.98 48.03 47.83 47.82 47.78 47.83
1 48.03 6150.06 48.18 48.21 47.76 47.85 47.87 47.81
2 48.20 48.09 6125.57 48.31 47.71 47.90 48.01 47.79
3 47.81 48.17 48.06 6150.06 47.62 47.83 47.98 47.89
4 47.62 47.76 47.91 47.93 6162.19 48.52 48.73 48.52
5 47.94 47.85 48.00 48.04 48.48 6150.06 48.55 48.67
6 47.92 48.00 48.06 47.99 48.51 48.60 6150.06 48.70
7 47.76 47.73 48.09 47.92 48.60 48.51 48.67 6150.06
Bidirectional P2P=Enabled Bandwidth Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 6125.57 48.00 48.23 47.98 47.58 47.67 47.88 47.74
1 47.71 6125.95 48.22 48.23 47.69 47.82 47.87 47.92
2 47.96 48.26 6150.06 48.17 47.76 48.02 47.86 47.97
3 48.16 48.33 48.08 6125.95 47.91 47.88 48.27 47.79
4 47.94 47.70 47.89 48.04 6150.06 48.53 48.61 48.80
5 47.95 47.90 48.04 48.08 48.58 6198.86 48.66 48.68
6 48.00 47.95 47.87 48.11 48.55 48.63 6125.95 48.76
7 47.79 47.76 47.74 48.06 48.68 48.52 48.76 6125.95
P2P=Disabled Latency Matrix (us)
GPU 0 1 2 3 4 5 6 7
0 1.77 18.51 18.66 18.67 19.55 19.48 19.60 18.87
1 18.53 1.40 18.50 18.51 19.34 19.12 19.44 19.40
2 18.55 18.53 1.38 18.51 18.86 18.87 18.92 19.15
3 18.46 18.50 18.43 1.40 18.55 18.54 18.64 18.58
4 18.55 18.52 18.54 18.53 1.77 18.49 18.44 18.52
5 18.54 18.55 18.53 18.46 18.39 1.40 18.50 18.53
6 18.55 18.46 18.54 18.46 18.52 18.52 1.37 18.45
7 18.45 18.45 18.55 18.46 18.47 18.43 18.45 1.37

CPU 0 1 2 3 4 5 6 7
0 2.56 9.05 8.94 8.82 8.03 8.04 8.07 8.00
1 9.01 2.45 8.88 8.83 8.05 7.96 8.01 8.02
2 9.15 8.84 2.44 8.85 8.12 8.06 8.12 8.01
3 8.94 8.78 8.81 2.46 8.08 7.96 7.83 8.06
4 8.38 8.24 8.28 8.17 2.16 7.35 7.49 7.54
5 8.22 8.20 8.20 8.06 7.51 2.18 7.56 7.52
6 8.32 8.27 8.24 8.10 7.42 7.63 2.18 7.59
7 8.29 8.18 8.23 8.18 7.42 7.55 7.65 2.20
P2P=Enabled Latency (P2P Writes) Matrix (us)
GPU 0 1 2 3 4 5 6 7
0 1.76 18.52 18.39 18.55 19.52 19.39 19.48 19.49
1 18.50 1.39 18.48 18.47 19.06 19.22 18.66 18.73
2 18.47 18.51 1.38 18.44 19.05 18.85 18.79 19.13
3 18.30 18.41 18.42 1.39 18.69 18.70 18.56 18.68
4 18.54 18.56 18.46 18.46 1.78 18.53 18.52 18.49
5 18.46 18.55 18.55 18.46 18.53 1.39 18.37 18.44
6 18.52 18.46 18.52 18.45 18.29 18.43 1.37 18.44
7 18.45 18.45 18.55 18.55 18.32 18.37 18.50 1.39

CPU 0 1 2 3 4 5 6 7
0 2.60 8.84 8.85 8.61 8.01 7.94 8.04 7.81
1 8.85 2.43 8.80 8.73 7.96 7.92 8.00 7.95
2 8.75 8.61 2.39 8.54 7.91 7.95 7.99 7.98
3 8.84 8.75 8.80 2.48 7.97 7.89 8.10 7.91
4 8.27 8.30 8.33 8.17 2.20 7.62 7.72 7.63
5 8.27 8.27 8.31 8.23 7.46 2.13 7.54 7.63
6 8.17 8.31 8.27 8.21 7.56 7.52 2.16 7.66
7 8.32 8.17 8.28 8.16 7.63 7.52 7.75 2.15

NOTE: The CUDA Samples are not meant for performance measurements. Results may vary when GPU Boost is enabled.

root@amd-5090:~/nccl-tests-master/build# ./all_gather_perf -g 8 -b 32m -e 4g -f 2

nccl-tests version 2.17.6 nccl-headers=22707 nccl-library=22707

Collective test starting: all_gather_perf

nThread 1 nGpus 8 minBytes 33554432 maxBytes 4294967296 step: 2(factor) warmup iters: 1 iters: 20 agg iters: 1 validation: 1 graph: 0

Using devices

Rank 0 Group 0 Pid 102006 on amd-5090 device 0 [0000:1a:00] NVIDIA B300 SXM6 AC

Rank 1 Group 0 Pid 102006 on amd-5090 device 1 [0000:40:00] NVIDIA B300 SXM6 AC

Rank 2 Group 0 Pid 102006 on amd-5090 device 2 [0000:62:00] NVIDIA B300 SXM6 AC

Rank 3 Group 0 Pid 102006 on amd-5090 device 3 [0000:73:00] NVIDIA B300 SXM6 AC

Rank 4 Group 0 Pid 102006 on amd-5090 device 4 [0000:9a:00] NVIDIA B300 SXM6 AC

Rank 5 Group 0 Pid 102006 on amd-5090 device 5 [0000:bd:00] NVIDIA B300 SXM6 AC

Rank 6 Group 0 Pid 102006 on amd-5090 device 6 [0000:df:00] NVIDIA B300 SXM6 AC

Rank 7 Group 0 Pid 102006 on amd-5090 device 7 [0000:f0:00] NVIDIA B300 SXM6 AC

amd-5090:102006:102006 [0] NCCL INFO Bootstrap: Using enx00e0bc492457:192.168.99.171<0>
amd-5090:102006:102006 [0] NCCL INFO cudaDriverVersion 13030
amd-5090:102006:102006 [0] NCCL INFO NCCL version 2.27.7+cuda13.0
amd-5090:102006:102075 [1] NCCL INFO NET/Plugin: Could not find: libnccl-net.so.
amd-5090:102006:102075 [1] NCCL INFO NET/IB : Using [0]mlx5_2:1/IB [1]mlx5_3:1/IB [2]mlx5_4:1/IB [3]mlx5_5:1/IB [RO]; OOB enx00e0bc492457:192.168.99.171<0>
amd-5090:102006:102075 [1] NCCL INFO Initialized NET plugin IB
amd-5090:102006:102075 [1] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102075 [1] NCCL INFO Using network IB
amd-5090:102006:102081 [7] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102081 [7] NCCL INFO Using network IB
amd-5090:102006:102078 [4] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102078 [4] NCCL INFO Using network IB
amd-5090:102006:102079 [5] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102079 [5] NCCL INFO Using network IB
amd-5090:102006:102074 [0] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102074 [0] NCCL INFO Using network IB
amd-5090:102006:102077 [3] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102077 [3] NCCL INFO Using network IB
amd-5090:102006:102076 [2] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102076 [2] NCCL INFO Using network IB
amd-5090:102006:102080 [6] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102080 [6] NCCL INFO Using network IB
amd-5090:102006:102075 [1] NCCL INFO DMA-BUF is available on GPU device 1
amd-5090:102006:102081 [7] NCCL INFO DMA-BUF is available on GPU device 7
amd-5090:102006:102078 [4] NCCL INFO DMA-BUF is available on GPU device 4
amd-5090:102006:102079 [5] NCCL INFO DMA-BUF is available on GPU device 5
amd-5090:102006:102074 [0] NCCL INFO DMA-BUF is available on GPU device 0
amd-5090:102006:102077 [3] NCCL INFO DMA-BUF is available on GPU device 3
amd-5090:102006:102076 [2] NCCL INFO DMA-BUF is available on GPU device 2
amd-5090:102006:102080 [6] NCCL INFO DMA-BUF is available on GPU device 6
amd-5090:102006:102075 [1] NCCL INFO ncclCommInitRankConfig comm 0x59ec45086370 rank 1 nranks 8 cudaDev 1 nvmlDev 1 busId 40000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102074 [0] NCCL INFO ncclCommInitRankConfig comm 0x59ec44f7e760 rank 0 nranks 8 cudaDev 0 nvmlDev 0 busId 1a000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102076 [2] NCCL INFO ncclCommInitRankConfig comm 0x59ec4518df80 rank 2 nranks 8 cudaDev 2 nvmlDev 2 busId 62000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102079 [5] NCCL INFO ncclCommInitRankConfig comm 0x59ec454a9cd0 rank 5 nranks 8 cudaDev 5 nvmlDev 5 busId bd000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102078 [4] NCCL INFO ncclCommInitRankConfig comm 0x59ec4539d7a0 rank 4 nranks 8 cudaDev 4 nvmlDev 4 busId 9a000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102075 [1] NCCL INFO RAS client listening socket at 127.0.0.1<28028>
amd-5090:102006:102077 [3] NCCL INFO ncclCommInitRankConfig comm 0x59ec45295b90 rank 3 nranks 8 cudaDev 3 nvmlDev 3 busId 73000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102080 [6] NCCL INFO ncclCommInitRankConfig comm 0x59ec455b8980 rank 6 nranks 8 cudaDev 6 nvmlDev 6 busId df000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102081 [7] NCCL INFO ncclCommInitRankConfig comm 0x59ec456c7630 rank 7 nranks 8 cudaDev 7 nvmlDev 7 busId f0000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102076 [2] NCCL INFO Bootstrap timings total 0.001817 (create 0.000035, send 0.000096, recv 0.001232, ring 0.000324, delay 0.000000)
amd-5090:102006:102075 [1] NCCL INFO Bootstrap timings total 0.002059 (create 0.000044, send 0.000119, recv 0.000454, ring 0.001061, delay 0.000001)
amd-5090:102006:102080 [6] NCCL INFO Bootstrap timings total 0.000697 (create 0.000036, send 0.000085, recv 0.000277, ring 0.000210, delay 0.000000)
amd-5090:102006:102081 [7] NCCL INFO Bootstrap timings total 0.000591 (create 0.000052, send 0.000094, recv 0.000167, ring 0.000144, delay 0.000000)
amd-5090:102006:102078 [4] NCCL INFO Bootstrap timings total 0.001270 (create 0.000031, send 0.000090, recv 0.000223, ring 0.000321, delay 0.000000)
amd-5090:102006:102077 [3] NCCL INFO Bootstrap timings total 0.000966 (create 0.000040, send 0.000104, recv 0.000331, ring 0.000312, delay 0.000000)
amd-5090:102006:102079 [5] NCCL INFO Bootstrap timings total 0.001601 (create 0.000036, send 0.000131, recv 0.001078, ring 0.000252, delay 0.000000)
amd-5090:102006:102074 [0] NCCL INFO Bootstrap timings total 0.001927 (create 0.000045, send 0.000134, recv 0.000257, ring 0.000159, delay 0.000000)
amd-5090:102006:102080 [6] NCCL INFO MNNVL busId 0xdf000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102075 [1] NCCL INFO MNNVL busId 0x40000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102076 [2] NCCL INFO MNNVL busId 0x62000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102078 [4] NCCL INFO MNNVL busId 0x9a000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102079 [5] NCCL INFO MNNVL busId 0xbd000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102081 [7] NCCL INFO MNNVL busId 0xf0000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102074 [0] NCCL INFO MNNVL busId 0x1a000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102077 [3] NCCL INFO MNNVL busId 0x73000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80

[2026-08-31 06:21:40] amd-5090:102006:102074 [0] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102074 [0] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102074 [0] NCCL INFO init.cc:816 → 1
amd-5090:102006:102074 [0] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102074 [0] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102081 [7] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102081 [7] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102081 [7] NCCL INFO init.cc:816 → 1
amd-5090:102006:102081 [7] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102081 [7] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102077 [3] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102077 [3] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102077 [3] NCCL INFO init.cc:816 → 1
amd-5090:102006:102077 [3] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102077 [3] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102075 [1] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102075 [1] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102075 [1] NCCL INFO init.cc:816 → 1
amd-5090:102006:102075 [1] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102075 [1] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102079 [5] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102079 [5] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102079 [5] NCCL INFO init.cc:816 → 1
amd-5090:102006:102079 [5] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102079 [5] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102080 [6] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102080 [6] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102080 [6] NCCL INFO init.cc:816 → 1
amd-5090:102006:102080 [6] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102080 [6] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102078 [4] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102078 [4] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102078 [4] NCCL INFO init.cc:816 → 1
amd-5090:102006:102078 [4] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102078 [4] NCCL INFO group.cc:73 → 1 [Async thread]

[2026-08-31 06:21:40] amd-5090:102006:102076 [2] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102076 [2] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102076 [2] NCCL INFO init.cc:816 → 1
amd-5090:102006:102076 [2] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102076 [2] NCCL INFO group.cc:73 → 1 [Async thread]
amd-5090:102006:102006 [7] NCCL INFO group.cc:475 → 1
amd-5090:102006:102006 [7] NCCL INFO group.cc:694 → 1
amd-5090:102006:102006 [7] NCCL INFO group.cc:104 → 1
amd-5090: Test NCCL failure common.cu:1279 'unhandled cuda error (run with NCCL_DEBUG=INFO for details) / ’
.. amd-5090 pid 102006: Test failure common.cu:1100

this is a really good example of the control plane saying everything is healthy while the actual workload says otherwise.

I’m working on a way to validate GPU setups from the workload side rather than just the fabric/status checks. if you want, I can give you a bounded command to run that treats successful P2P/NCCL execution as the pass condition and captures the failure evidence cleanly.

Sounds great, man! Please send it over.

We’re currently suspecting some physical layer issues (we recently touched the clock chip and PCIe retimers in this direct CPU-to-GPU setup), so having a clean test script to capture the actual workload failure would be super helpful. Thanks!