We have two NVIDIA DGX GB300 servers, each with 8× NVIDIA ConnectX-8 NICs connected through a NVIDIA Q3400 switch. When running NCCL tests across the two servers, the bandwidth is only around 180 GB/s. However, local NCCL tests on a single server can reach 800 GB/s, and point-to-point ib_read_bw tests can reach 480 GB/s.
elsthub
1
You can try adding a command to the NCCL parameter that bypasses the CPU.
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| DGX Spark NCCL Test: 10GB/s not 200 Gbps=25 GB/s | 3 | 1030 | November 5, 2025 | |
| DGX Spark NCCL Test: 15GB/s So Slow | 1 | 303 | March 4, 2026 | |
| NCCL single-cable test caps at 100Gbps | 15 | 487 | March 17, 2026 | |
| NCCL bandwidth capped at 3 GB/s, GPU PCIe topology reports Gen1 x1 on DGX Spark FE | 5 | 445 | April 14, 2026 | |
| DGX Spark ↔ EdgeXpert NCCL only ~17 GB/s over 200GbE | 4 | 420 | April 9, 2026 | |
| NCCL Test Bandwidth is only 3GB/s between 2 DGX Spark using QSFP cable | 9 | 580 | April 19, 2026 | |
| DGX A100, when 8 IB network cards use ib_write_bw to test the bandwidth at the same time, the rate decreases, which is not expected | 2 | 1340 | June 2, 2023 | |
| How can I improve the 'p2p enabled' bandwidth when testing NCCL performance with two A5000 GPU using PCIe 4.0 x16? | 2 | 1361 | September 15, 2023 | |
| ConnectX-7 RDMA write_bw does not meet performence expectation | 1 | 188 | August 20, 2025 | |
| 17GB bandwidth issue, 2x FE's + one newbie end user | 4 | 122 | June 30, 2026 |