Hi everyone,
Need some help here with our NVIDIA B300 SXM6 setup.
The issue is: NVLink looks completely fine when idle, but fails as soon as we run any actual load.
What we checked so far:
-
Static / Control Plane Status: Everything looks super clean.
nvidia-smi nvlink -eshows healthy physical links (no major errors, link up). FabricManager and NVLSM (OpenSM) start without any issues. The logs literally say:“OpenSM Entering MASTER state” “Successfully configured all the available GPUs and NVSwitches to route NVLink traffic.”
-
The Problem: Despite FM and NVLSM claiming everything is successfully routed, the moment we trigger P2P traffic / workloads (like
p2pBandwidthLatencyTestor PyTorch/NCCL), NVLink communication breaks / hangs / falls back.
Environment Info:Intel(R) Xeon(R) 6730P
GPU: 8x NVIDIA B300 SXM6 AC
root@amd-5090:~/cuda-samples-master/build/bin/x86_64/linux/release# nvidia-smi
Mon Aug 31 06:12:13 2026
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 610.57.04 KMD Version: 610.57.04 CUDA UMD Version: 13.3 |
±----------------------------------------±-----------------------±---------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA B300 SXM6 AC On | 00000000:1A:00.0 Off | 0 |
| N/A 31C P0 183W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 1 NVIDIA B300 SXM6 AC On | 00000000:40:00.0 Off | 0 |
| N/A 35C P0 184W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 2 NVIDIA B300 SXM6 AC On | 00000000:62:00.0 Off | 0 |
| N/A 34C P0 184W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 3 NVIDIA B300 SXM6 AC On | 00000000:73:00.0 Off | 0 |
| N/A 32C P0 181W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 4 NVIDIA B300 SXM6 AC On | 00000000:9A:00.0 Off | 0 |
| N/A 32C P0 182W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 5 NVIDIA B300 SXM6 AC On | 00000000:BD:00.0 Off | 0 |
| N/A 35C P0 181W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 6 NVIDIA B300 SXM6 AC On | 00000000:DF:00.0 Off | 0 |
| N/A 35C P0 184W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 7 NVIDIA B300 SXM6 AC On | 00000000:F0:00.0 Off | 0 |
| N/A 32C P0 180W / 1100W | 0MiB / 275040MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
±----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
±----------------------------------------------------------------------------------------+
GPU 0: NVIDIA B300 SXM6 AC (UUID: GPU-8337aacd-dd11-a827-ae2c-9da20f8ea472)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 1: NVIDIA B300 SXM6 AC (UUID: GPU-bb5fe741-d156-7a95-8985-53829e581ac3)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 2: NVIDIA B300 SXM6 AC (UUID: GPU-9590ee95-e06d-59e8-15e4-5ba7fa9d47f2)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 3: NVIDIA B300 SXM6 AC (UUID: GPU-6aeb073e-15e8-ef5d-c2ca-023572e0bd61)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 4: NVIDIA B300 SXM6 AC (UUID: GPU-e822c480-886d-55e3-8f9a-4e39f8519279)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 5: NVIDIA B300 SXM6 AC (UUID: GPU-ffab3c6b-ff72-3301-6fdf-ebe8cede2f57)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 6: NVIDIA B300 SXM6 AC (UUID: GPU-6f11100c-e92f-830e-c248-de77bf05ea83)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
GPU 7: NVIDIA B300 SXM6 AC (UUID: GPU-71abcb2c-a244-4ca5-e077-2047cfdd5490)
Link 0: 53.125 GB/s
Link 1: 53.125 GB/s
Link 2: 53.125 GB/s
Link 3: 53.125 GB/s
Link 4: 53.125 GB/s
Link 5: 53.125 GB/s
Link 6: 53.125 GB/s
Link 7: 53.125 GB/s
Link 8: 53.125 GB/s
Link 9: 53.125 GB/s
Link 10: 53.125 GB/s
Link 11: 53.125 GB/s
Link 12: 53.125 GB/s
Link 13: 53.125 GB/s
Link 14: 53.125 GB/s
Link 15: 53.125 GB/s
Link 16: 53.125 GB/s
Link 17: 53.125 GB/s
root@amd-5090:~/cuda-samples-master/build/bin/x86_64/linux/release# systemctl status nvidia-fabricmanager.service
● nvidia-fabricmanager.service - NVIDIA fabric manager service
Loaded: loaded (/usr/lib/systemd/system/nvidia-fabricmanager.service; disabled; preset: enabled)
Active: active (running) since Mon 2026-08-31 05:54:35 UTC; 22min ago
Process: 89378 ExecStartPre=/usr/bin/nvidia-fabricmanager-start.sh --mode precheck (code=exited, status=0/SUCCESS)
Process: 89605 ExecStart=/usr/bin/nvidia-fabricmanager-start.sh --mode start (code=exited, status=0/SUCCESS)
Main PID: 89757 (nv-fabricmanage)
Tasks: 90 (limit: 629145)
Memory: 36.0M (peak: 42.4M)
CPU: 8.277s
CGroup: /system.slice/nvidia-fabricmanager.service
├─89735 /opt/nvidia/nvlsm/sbin/nvlsm -F /usr/share/nvidia/nvlsm/nvlsm.conf -B --pid_file /var/run/nvidia-fabricmanager/nvlsm.pid -g 0xac3ae20300b1773a
├─89738 osm_crashd
└─89757 /usr/bin/nv-fabricmanager -c /usr/share/nvidia/nvswitch/fabricmanager.cfg -g 0xac3ae20300b1773a
Aug 31 05:54:29 amd-5090 OpenSM[89732]: -I- Configuration loaded
Aug 31 05:54:29 amd-5090 nvidia-fabricmanager-start.sh[89605]: Started “Nvidia NVLink Subnet Manager”
Aug 31 05:54:29 amd-5090 OpenSM[89735]: /var/log/nvlsm.log log file opened
Aug 31 05:54:29 amd-5090 OpenSM[89735]: OpenSM 2025.10.14_42349a2_821fcee_ae214d2
Aug 31 05:54:29 amd-5090 OpenSM[89735]: Entering DISCOVERING state
Aug 31 05:54:32 amd-5090 OpenSM[89735]: Entering MASTER state
Aug 31 05:54:35 amd-5090 nv-fabricmanager[89757]: NodeId 0 partition id 57082 is activated.
Aug 31 05:54:35 amd-5090 nv-fabricmanager[89757]: Successfully configured all the available GPUs and NVSwitches to route NVLink traffic. NVLink Peer-to-Peer support will be enabled onc>
Aug 31 05:54:35 amd-5090 nvidia-fabricmanager-start.sh[89605]: Started “Nvidia Fabric Manager”
Aug 31 05:54:35 amd-5090 systemd[1]: Started nvidia-fabricmanager.service - NVIDIA fabric manager service.
root@amd-5090:~/cuda-samples-master/build/bin/x86_64/linux/release#
We are testing an 8x NVIDIA B300 SXM6 node and hit a critical roadblock: NVLink control plane looks 100% healthy at idle, but P2P traffic fails immediately under load.
-
Control Plane / Idle State (Looks Good):
-
All 8 GPUs are properly enumerated by
nvidia-smi. -
nvidia-smi nvlink -sshows all 12 links per GPU in Active state. -
FabricManager & NVLSM (OpenSM) start clean with no errors. Logs explicitly confirm:
“OpenSM Entering MASTER state”
“Successfully configured all the available GPUs and NVSwitches to route NVLink traffic.”
-
-
Data Plane / Workloads (Fails Immediately):
- The moment we trigger high-bandwidth P2P traffic (
p2pBandwidthLatencyTest,nccl-tests, or PyTorch training), NVLink communication hangs, drops, or falls back to PCIe / SYS.
- The moment we trigger high-bandwidth P2P traffic (
Below are our system outputs and runtime log details:
[P2P (Peer-to-Peer) GPU Bandwidth Latency Test]
Device: 0, NVIDIA B300 SXM6 AC, pciBusID: 1a, pciDeviceID: 0, pciDomainID:0
Device: 1, NVIDIA B300 SXM6 AC, pciBusID: 40, pciDeviceID: 0, pciDomainID:0
Device: 2, NVIDIA B300 SXM6 AC, pciBusID: 62, pciDeviceID: 0, pciDomainID:0
Device: 3, NVIDIA B300 SXM6 AC, pciBusID: 73, pciDeviceID: 0, pciDomainID:0
Device: 4, NVIDIA B300 SXM6 AC, pciBusID: 9a, pciDeviceID: 0, pciDomainID:0
Device: 5, NVIDIA B300 SXM6 AC, pciBusID: bd, pciDeviceID: 0, pciDomainID:0
Device: 6, NVIDIA B300 SXM6 AC, pciBusID: df, pciDeviceID: 0, pciDomainID:0
Device: 7, NVIDIA B300 SXM6 AC, pciBusID: f0, pciDeviceID: 0, pciDomainID:0
Device=0 CANNOT Access Peer Device=1
Device=0 CANNOT Access Peer Device=2
Device=0 CANNOT Access Peer Device=3
Device=0 CANNOT Access Peer Device=4
Device=0 CANNOT Access Peer Device=5
Device=0 CANNOT Access Peer Device=6
Device=0 CANNOT Access Peer Device=7
Device=1 CANNOT Access Peer Device=0
Device=1 CANNOT Access Peer Device=2
Device=1 CANNOT Access Peer Device=3
Device=1 CANNOT Access Peer Device=4
Device=1 CANNOT Access Peer Device=5
Device=1 CANNOT Access Peer Device=6
Device=1 CANNOT Access Peer Device=7
Device=2 CANNOT Access Peer Device=0
Device=2 CANNOT Access Peer Device=1
Device=2 CANNOT Access Peer Device=3
Device=2 CANNOT Access Peer Device=4
Device=2 CANNOT Access Peer Device=5
Device=2 CANNOT Access Peer Device=6
Device=2 CANNOT Access Peer Device=7
Device=3 CANNOT Access Peer Device=0
Device=3 CANNOT Access Peer Device=1
Device=3 CANNOT Access Peer Device=2
Device=3 CANNOT Access Peer Device=4
Device=3 CANNOT Access Peer Device=5
Device=3 CANNOT Access Peer Device=6
Device=3 CANNOT Access Peer Device=7
Device=4 CANNOT Access Peer Device=0
Device=4 CANNOT Access Peer Device=1
Device=4 CANNOT Access Peer Device=2
Device=4 CANNOT Access Peer Device=3
Device=4 CANNOT Access Peer Device=5
Device=4 CANNOT Access Peer Device=6
Device=4 CANNOT Access Peer Device=7
Device=5 CANNOT Access Peer Device=0
Device=5 CANNOT Access Peer Device=1
Device=5 CANNOT Access Peer Device=2
Device=5 CANNOT Access Peer Device=3
Device=5 CANNOT Access Peer Device=4
Device=5 CANNOT Access Peer Device=6
Device=5 CANNOT Access Peer Device=7
Device=6 CANNOT Access Peer Device=0
Device=6 CANNOT Access Peer Device=1
Device=6 CANNOT Access Peer Device=2
Device=6 CANNOT Access Peer Device=3
Device=6 CANNOT Access Peer Device=4
Device=6 CANNOT Access Peer Device=5
Device=6 CANNOT Access Peer Device=7
Device=7 CANNOT Access Peer Device=0
Device=7 CANNOT Access Peer Device=1
Device=7 CANNOT Access Peer Device=2
Device=7 CANNOT Access Peer Device=3
Device=7 CANNOT Access Peer Device=4
Device=7 CANNOT Access Peer Device=5
Device=7 CANNOT Access Peer Device=6
***NOTE: In case a device doesn’t have P2P access to other one, it falls back to normal memcopy procedure.
So you can see lesser Bandwidth (GB/s) and unstable Latency (us) in those cases.
P2P Connectivity Matrix
D\D 0 1 2 3 4 5 6 7
0 1 0 0 0 0 0 0 0
1 0 1 0 0 0 0 0 0
2 0 0 1 0 0 0 0 0
3 0 0 0 1 0 0 0 0
4 0 0 0 0 1 0 0 0
5 0 0 0 0 0 1 0 0
6 0 0 0 0 0 0 1 0
7 0 0 0 0 0 0 0 1
Unidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 5826.15 38.08 38.18 38.41 38.09 38.00 38.07 38.03
1 37.79 5896.23 38.21 38.14 37.66 37.81 38.11 38.12
2 37.78 37.65 5896.23 38.18 37.89 37.97 37.87 38.10
3 37.93 37.79 38.01 5896.23 38.06 38.13 38.06 37.91
4 37.85 37.95 37.78 38.02 5987.31 38.49 38.60 38.57
5 38.00 38.05 38.28 38.22 38.23 5919.26 38.38 38.61
6 37.71 38.08 38.22 37.99 38.39 38.55 5918.56 38.60
7 37.73 37.73 38.05 37.91 38.43 38.46 38.51 5941.06
Unidirectional P2P=Enabled Bandwidth (P2P Writes) Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 5896.23 38.10 37.93 38.22 37.98 37.93 38.20 37.92
1 37.99 5896.23 38.00 38.19 37.64 38.13 38.16 38.12
2 37.67 37.82 5896.23 38.01 37.86 38.07 38.02 38.07
3 37.67 37.81 38.00 5896.23 38.08 38.03 38.05 38.04
4 37.83 37.84 38.10 38.10 5940.36 38.44 38.60 38.55
5 37.89 38.19 38.12 38.29 38.27 5896.23 38.53 38.61
6 37.92 37.90 37.99 38.35 38.26 38.61 5941.77 38.54
7 37.82 37.79 37.95 38.28 38.29 38.22 38.40 5941.77
Bidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 6174.36 47.95 47.98 48.03 47.83 47.82 47.78 47.83
1 48.03 6150.06 48.18 48.21 47.76 47.85 47.87 47.81
2 48.20 48.09 6125.57 48.31 47.71 47.90 48.01 47.79
3 47.81 48.17 48.06 6150.06 47.62 47.83 47.98 47.89
4 47.62 47.76 47.91 47.93 6162.19 48.52 48.73 48.52
5 47.94 47.85 48.00 48.04 48.48 6150.06 48.55 48.67
6 47.92 48.00 48.06 47.99 48.51 48.60 6150.06 48.70
7 47.76 47.73 48.09 47.92 48.60 48.51 48.67 6150.06
Bidirectional P2P=Enabled Bandwidth Matrix (GB/s)
D\D 0 1 2 3 4 5 6 7
0 6125.57 48.00 48.23 47.98 47.58 47.67 47.88 47.74
1 47.71 6125.95 48.22 48.23 47.69 47.82 47.87 47.92
2 47.96 48.26 6150.06 48.17 47.76 48.02 47.86 47.97
3 48.16 48.33 48.08 6125.95 47.91 47.88 48.27 47.79
4 47.94 47.70 47.89 48.04 6150.06 48.53 48.61 48.80
5 47.95 47.90 48.04 48.08 48.58 6198.86 48.66 48.68
6 48.00 47.95 47.87 48.11 48.55 48.63 6125.95 48.76
7 47.79 47.76 47.74 48.06 48.68 48.52 48.76 6125.95
P2P=Disabled Latency Matrix (us)
GPU 0 1 2 3 4 5 6 7
0 1.77 18.51 18.66 18.67 19.55 19.48 19.60 18.87
1 18.53 1.40 18.50 18.51 19.34 19.12 19.44 19.40
2 18.55 18.53 1.38 18.51 18.86 18.87 18.92 19.15
3 18.46 18.50 18.43 1.40 18.55 18.54 18.64 18.58
4 18.55 18.52 18.54 18.53 1.77 18.49 18.44 18.52
5 18.54 18.55 18.53 18.46 18.39 1.40 18.50 18.53
6 18.55 18.46 18.54 18.46 18.52 18.52 1.37 18.45
7 18.45 18.45 18.55 18.46 18.47 18.43 18.45 1.37
CPU 0 1 2 3 4 5 6 7
0 2.56 9.05 8.94 8.82 8.03 8.04 8.07 8.00
1 9.01 2.45 8.88 8.83 8.05 7.96 8.01 8.02
2 9.15 8.84 2.44 8.85 8.12 8.06 8.12 8.01
3 8.94 8.78 8.81 2.46 8.08 7.96 7.83 8.06
4 8.38 8.24 8.28 8.17 2.16 7.35 7.49 7.54
5 8.22 8.20 8.20 8.06 7.51 2.18 7.56 7.52
6 8.32 8.27 8.24 8.10 7.42 7.63 2.18 7.59
7 8.29 8.18 8.23 8.18 7.42 7.55 7.65 2.20
P2P=Enabled Latency (P2P Writes) Matrix (us)
GPU 0 1 2 3 4 5 6 7
0 1.76 18.52 18.39 18.55 19.52 19.39 19.48 19.49
1 18.50 1.39 18.48 18.47 19.06 19.22 18.66 18.73
2 18.47 18.51 1.38 18.44 19.05 18.85 18.79 19.13
3 18.30 18.41 18.42 1.39 18.69 18.70 18.56 18.68
4 18.54 18.56 18.46 18.46 1.78 18.53 18.52 18.49
5 18.46 18.55 18.55 18.46 18.53 1.39 18.37 18.44
6 18.52 18.46 18.52 18.45 18.29 18.43 1.37 18.44
7 18.45 18.45 18.55 18.55 18.32 18.37 18.50 1.39
CPU 0 1 2 3 4 5 6 7
0 2.60 8.84 8.85 8.61 8.01 7.94 8.04 7.81
1 8.85 2.43 8.80 8.73 7.96 7.92 8.00 7.95
2 8.75 8.61 2.39 8.54 7.91 7.95 7.99 7.98
3 8.84 8.75 8.80 2.48 7.97 7.89 8.10 7.91
4 8.27 8.30 8.33 8.17 2.20 7.62 7.72 7.63
5 8.27 8.27 8.31 8.23 7.46 2.13 7.54 7.63
6 8.17 8.31 8.27 8.21 7.56 7.52 2.16 7.66
7 8.32 8.17 8.28 8.16 7.63 7.52 7.75 2.15
NOTE: The CUDA Samples are not meant for performance measurements. Results may vary when GPU Boost is enabled.
root@amd-5090:~/nccl-tests-master/build# ./all_gather_perf -g 8 -b 32m -e 4g -f 2
nccl-tests version 2.17.6 nccl-headers=22707 nccl-library=22707
Collective test starting: all_gather_perf
nThread 1 nGpus 8 minBytes 33554432 maxBytes 4294967296 step: 2(factor) warmup iters: 1 iters: 20 agg iters: 1 validation: 1 graph: 0
Using devices
Rank 0 Group 0 Pid 102006 on amd-5090 device 0 [0000:1a:00] NVIDIA B300 SXM6 AC
Rank 1 Group 0 Pid 102006 on amd-5090 device 1 [0000:40:00] NVIDIA B300 SXM6 AC
Rank 2 Group 0 Pid 102006 on amd-5090 device 2 [0000:62:00] NVIDIA B300 SXM6 AC
Rank 3 Group 0 Pid 102006 on amd-5090 device 3 [0000:73:00] NVIDIA B300 SXM6 AC
Rank 4 Group 0 Pid 102006 on amd-5090 device 4 [0000:9a:00] NVIDIA B300 SXM6 AC
Rank 5 Group 0 Pid 102006 on amd-5090 device 5 [0000:bd:00] NVIDIA B300 SXM6 AC
Rank 6 Group 0 Pid 102006 on amd-5090 device 6 [0000:df:00] NVIDIA B300 SXM6 AC
Rank 7 Group 0 Pid 102006 on amd-5090 device 7 [0000:f0:00] NVIDIA B300 SXM6 AC
amd-5090:102006:102006 [0] NCCL INFO Bootstrap: Using enx00e0bc492457:192.168.99.171<0>
amd-5090:102006:102006 [0] NCCL INFO cudaDriverVersion 13030
amd-5090:102006:102006 [0] NCCL INFO NCCL version 2.27.7+cuda13.0
amd-5090:102006:102075 [1] NCCL INFO NET/Plugin: Could not find: libnccl-net.so.
amd-5090:102006:102075 [1] NCCL INFO NET/IB : Using [0]mlx5_2:1/IB [1]mlx5_3:1/IB [2]mlx5_4:1/IB [3]mlx5_5:1/IB [RO]; OOB enx00e0bc492457:192.168.99.171<0>
amd-5090:102006:102075 [1] NCCL INFO Initialized NET plugin IB
amd-5090:102006:102075 [1] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102075 [1] NCCL INFO Using network IB
amd-5090:102006:102081 [7] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102081 [7] NCCL INFO Using network IB
amd-5090:102006:102078 [4] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102078 [4] NCCL INFO Using network IB
amd-5090:102006:102079 [5] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102079 [5] NCCL INFO Using network IB
amd-5090:102006:102074 [0] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102074 [0] NCCL INFO Using network IB
amd-5090:102006:102077 [3] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102077 [3] NCCL INFO Using network IB
amd-5090:102006:102076 [2] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102076 [2] NCCL INFO Using network IB
amd-5090:102006:102080 [6] NCCL INFO Assigned NET plugin IB to comm
amd-5090:102006:102080 [6] NCCL INFO Using network IB
amd-5090:102006:102075 [1] NCCL INFO DMA-BUF is available on GPU device 1
amd-5090:102006:102081 [7] NCCL INFO DMA-BUF is available on GPU device 7
amd-5090:102006:102078 [4] NCCL INFO DMA-BUF is available on GPU device 4
amd-5090:102006:102079 [5] NCCL INFO DMA-BUF is available on GPU device 5
amd-5090:102006:102074 [0] NCCL INFO DMA-BUF is available on GPU device 0
amd-5090:102006:102077 [3] NCCL INFO DMA-BUF is available on GPU device 3
amd-5090:102006:102076 [2] NCCL INFO DMA-BUF is available on GPU device 2
amd-5090:102006:102080 [6] NCCL INFO DMA-BUF is available on GPU device 6
amd-5090:102006:102075 [1] NCCL INFO ncclCommInitRankConfig comm 0x59ec45086370 rank 1 nranks 8 cudaDev 1 nvmlDev 1 busId 40000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102074 [0] NCCL INFO ncclCommInitRankConfig comm 0x59ec44f7e760 rank 0 nranks 8 cudaDev 0 nvmlDev 0 busId 1a000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102076 [2] NCCL INFO ncclCommInitRankConfig comm 0x59ec4518df80 rank 2 nranks 8 cudaDev 2 nvmlDev 2 busId 62000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102079 [5] NCCL INFO ncclCommInitRankConfig comm 0x59ec454a9cd0 rank 5 nranks 8 cudaDev 5 nvmlDev 5 busId bd000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102078 [4] NCCL INFO ncclCommInitRankConfig comm 0x59ec4539d7a0 rank 4 nranks 8 cudaDev 4 nvmlDev 4 busId 9a000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102075 [1] NCCL INFO RAS client listening socket at 127.0.0.1<28028>
amd-5090:102006:102077 [3] NCCL INFO ncclCommInitRankConfig comm 0x59ec45295b90 rank 3 nranks 8 cudaDev 3 nvmlDev 3 busId 73000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102080 [6] NCCL INFO ncclCommInitRankConfig comm 0x59ec455b8980 rank 6 nranks 8 cudaDev 6 nvmlDev 6 busId df000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102081 [7] NCCL INFO ncclCommInitRankConfig comm 0x59ec456c7630 rank 7 nranks 8 cudaDev 7 nvmlDev 7 busId f0000 commId 0x70b6d56682e5fb9e - Init START
amd-5090:102006:102076 [2] NCCL INFO Bootstrap timings total 0.001817 (create 0.000035, send 0.000096, recv 0.001232, ring 0.000324, delay 0.000000)
amd-5090:102006:102075 [1] NCCL INFO Bootstrap timings total 0.002059 (create 0.000044, send 0.000119, recv 0.000454, ring 0.001061, delay 0.000001)
amd-5090:102006:102080 [6] NCCL INFO Bootstrap timings total 0.000697 (create 0.000036, send 0.000085, recv 0.000277, ring 0.000210, delay 0.000000)
amd-5090:102006:102081 [7] NCCL INFO Bootstrap timings total 0.000591 (create 0.000052, send 0.000094, recv 0.000167, ring 0.000144, delay 0.000000)
amd-5090:102006:102078 [4] NCCL INFO Bootstrap timings total 0.001270 (create 0.000031, send 0.000090, recv 0.000223, ring 0.000321, delay 0.000000)
amd-5090:102006:102077 [3] NCCL INFO Bootstrap timings total 0.000966 (create 0.000040, send 0.000104, recv 0.000331, ring 0.000312, delay 0.000000)
amd-5090:102006:102079 [5] NCCL INFO Bootstrap timings total 0.001601 (create 0.000036, send 0.000131, recv 0.001078, ring 0.000252, delay 0.000000)
amd-5090:102006:102074 [0] NCCL INFO Bootstrap timings total 0.001927 (create 0.000045, send 0.000134, recv 0.000257, ring 0.000159, delay 0.000000)
amd-5090:102006:102080 [6] NCCL INFO MNNVL busId 0xdf000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102075 [1] NCCL INFO MNNVL busId 0x40000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102076 [2] NCCL INFO MNNVL busId 0x62000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102078 [4] NCCL INFO MNNVL busId 0x9a000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102079 [5] NCCL INFO MNNVL busId 0xbd000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102081 [7] NCCL INFO MNNVL busId 0xf0000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102074 [0] NCCL INFO MNNVL busId 0x1a000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
amd-5090:102006:102077 [3] NCCL INFO MNNVL busId 0x73000 fabric UUID 0.0 cliqueId 0x0 state 3 healthMask 0x80
[2026-08-31 06:21:40] amd-5090:102006:102074 [0] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102074 [0] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102074 [0] NCCL INFO init.cc:816 → 1
amd-5090:102006:102074 [0] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102074 [0] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102081 [7] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102081 [7] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102081 [7] NCCL INFO init.cc:816 → 1
amd-5090:102006:102081 [7] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102081 [7] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102077 [3] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102077 [3] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102077 [3] NCCL INFO init.cc:816 → 1
amd-5090:102006:102077 [3] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102077 [3] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102075 [1] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102075 [1] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102075 [1] NCCL INFO init.cc:816 → 1
amd-5090:102006:102075 [1] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102075 [1] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102079 [5] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102079 [5] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102079 [5] NCCL INFO init.cc:816 → 1
amd-5090:102006:102079 [5] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102079 [5] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102080 [6] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102080 [6] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102080 [6] NCCL INFO init.cc:816 → 1
amd-5090:102006:102080 [6] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102080 [6] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102078 [4] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102078 [4] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102078 [4] NCCL INFO init.cc:816 → 1
amd-5090:102006:102078 [4] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102078 [4] NCCL INFO group.cc:73 → 1 [Async thread]
[2026-08-31 06:21:40] amd-5090:102006:102076 [2] graph/paths.cc:350 NCCL WARN P2P is disabled between NVLINK connected GPUs 1 and 0. This should not be the case given their connectivity, and is probably due to a hardware issue. If you still want to proceed, you can set NCCL_IGNORE_DISABLED_P2P=1.
amd-5090:102006:102076 [2] NCCL INFO graph/paths.cc:648 → 1
amd-5090:102006:102076 [2] NCCL INFO init.cc:816 → 1
amd-5090:102006:102076 [2] NCCL INFO init.cc:1449 → 1
amd-5090:102006:102076 [2] NCCL INFO group.cc:73 → 1 [Async thread]
amd-5090:102006:102006 [7] NCCL INFO group.cc:475 → 1
amd-5090:102006:102006 [7] NCCL INFO group.cc:694 → 1
amd-5090:102006:102006 [7] NCCL INFO group.cc:104 → 1
amd-5090: Test NCCL failure common.cu:1279 'unhandled cuda error (run with NCCL_DEBUG=INFO for details) / ’
.. amd-5090 pid 102006: Test failure common.cu:1100