and here’s one more from FS
So when I ordered my cable from NADDOD, I hadn’t discovered this post yet. Not knowing what to get I ordered the MCP1650-H00AE30 instead of the MCP1650-V00AE30. I assumed I wanted the Infiniband cable. I’ve been using it. It does seem to work. I think physically they are the same cable, but there is different identifier in the EEPROM. I don’t know if NADDOD maybe “bins” them differently. I think Infiniband assumes a lossless fabric, where ethernet can handle packet losses better. According to Gemini the DGX Spark is shipped and architected to use Ethernet for its clustering fabric, specifically RoCE v2 (RDMA over Converged Ethernet). The ConnectX-7 is a VPI (Virtual Protocol Interconnect) card, meaning it can be switched to InfiniBand mode. However, this is not a simple toggle for the Spark platform. InfiniBand requires a Subnet Manager (SM) to be running on one of the nodes to manage Local IDs (LIDs) and route traffic? It would be nice NVIDIA would document this fully. I guess they assume that DGX Spark owners are already familiar with the technology.
The only odd thing I see are these warn messages:
(EngineCore_DP0 pid=1896) (RayWorkerWrapper pid=2112) [2025-11-22 05:14:24] spark1:2112:4122 [0] p2p_plugin.c:421 NCCL WARN NET/IB : roceP2p1s0f0:1 unknown event type (18)
Otherwise I think it’s working:
(EngineCore_DP0 pid=1896) (RayWorkerWrapper pid=2112) spark1:2112:7697 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0
Granted, I’m not an expert in Infiniband stack, but Spark seems to use IB connection just fine:
$ ib_write_bw 192.168.177.12 -d rocep1s0f1 --report_gbits -q 4 -R --force-link IB
---------------------------------------------------------------------------------------
RDMA_Write BW Test
Dual-port : OFF Device : rocep1s0f1
Number of qps : 4 Transport type : IB
Connection type : RC Using SRQ : OFF
PCIe relax order: ON
ibv_wr* API : ON
TX depth : 128
CQ Moderation : 1
Mtu : 1024[B]
Link type : IB
Max inline data : 0[B]
rdma_cm QPs : ON
Data ex. method : rdma_cm
---------------------------------------------------------------------------------------
local address: LID 0000 QPN 0x03ec PSN 0xb680ae
local address: LID 0000 QPN 0x03ed PSN 0x808800
local address: LID 0000 QPN 0x03ee PSN 0x5b694a
local address: LID 0000 QPN 0x03ef PSN 0xe2efd1
remote address: LID 0000 QPN 0x03eb PSN 0x75f6ee
remote address: LID 0000 QPN 0x03ec PSN 0x436140
remote address: LID 0000 QPN 0x03ed PSN 0x81698a
remote address: LID 0000 QPN 0x03ee PSN 0x4a8b11
---------------------------------------------------------------------------------------
#bytes #iterations BW peak[Gb/sec] BW average[Gb/sec] MsgRate[Mpps]
65536 20000 111.72 111.71 0.213070
---------------------------------------------------------------------------------------
Where did you buy NJAAKKR-0006, @eugr ? Amphenol’s website?
$42 at Digikey, but looooooooong lead time.
ExxactCorp.
I get the same warning messages with the NADDOD MCP1650-V00AE30 cable, if it makes you feel any better.
Laser? Is that an active optical cable? Ours are copper.
A bit frustrating, to be fair. I’d appreciate a bit more clear guidance and participation from NVidia on this. For instance, whether we need to bond both interfaces for a physical port together to get 200G speeds, because I’m not able to get more than 100G with iperf3 or ib_write_bw.
Other frustrations are mostly related to the state of Blackwell support in VLLM and the worst mmap performance on any system I’ve used recently. When it takes 10 minutes to just load a 80GB model into vllm, it results in a lot of wasted time.
Haven’t done anything useful in terms of clustering so far…
We’re capped by the DGX Spark Physical PCle Link Hardware.
This is from Gemini 3 Pro:
sudo lspci -vvv | grep -A 20 "NVIDIA\|Mellanox" | grep -E "LnkCap|LnkSta"
LnkCap: Port #0, Speed 32GT/s, Width x4, ASPM not supported
LnkSta: Speed 32GT/s, Width x4
LnkCap: Port #0, Speed 32GT/s, Width x4, ASPM not supported
LnkSta: Speed 32GT/s, Width x4
LnkCap: Port #0, Speed 32GT/s, Width x4, ASPM not supported
LnkSta: Speed 32GT/s, Width x4
LnkCap: Port #0, Speed 32GT/s, Width x4, ASPM not supported
LnkSta: Speed 32GT/s, Width x4
pcilib: sysfs_read_vpd: read failed: No such device
LnkCap: Port #0, Speed 2.5GT/s, Width x16, ASPM L1, Exit Latency L1 <4us
PCIe Link Speed: Each Mellanox/NVIDIA NIC shows “Speed 32GT/s, Width x4” (PCIe Gen5 x4). The expected is x16 for full-bandwidth. Gen5 x4 maxes out at ~16 GB/s per lane, and only ~64 Gbps per cable—not sufficient for your 200G physical link, which explains the observed effective speed ceiling.
I wonder if NVIDIA will review BIOS/firmware options to enhance PCIe slot configuration bottleneck.
I can’t believe I’m spending this long in device setup before getting to work on anything meaningful.
When a cable connector is inserted in one of the physical ports two virtual links light-up, i.e. enp1s0f0np0 and enP2p1s0f0np0 Each virtual link is x4 wide capable of 32GT/s, so the combined output will be x8 and 64GT/s speed. That’s enough to send, theoretically, 200Gbps on the wire!
Even the Vital Product Data claims it:
Product Name: NVIDIA DGX Spark, P4242-0000, 2-port QSFP up to 200G, Ethernet, PCIe5
Use lspci -vvvs <port ID> for details.
This cable started selling at Microcenter:
https://www.microcenter.com/product/703095/pny-certified-stacking-dgx-spark-cable-40
Theoretically, the official certified cable for stacking DGX sparks by PNY.
That’s a longer version, NJAAKK-N911. The one they had before was NJAAKR one.
The only difference is cable length:
NJAAKK-N911 = 400mm (0.4 meters)
NJAAKK-0006 = 500mm (0.5 meters)
Both cables are made by Amphenol. The part after NJAAKK signifies the cable length. See details at https://www.amphenol-cs.com/catalogsearch/result?query=NJAAKK&page=1
Thanks for this suggestion. I struggled to find a suitable cable available here in the UK but the NADDOD one arrived quickly from China and is working for me: NVIDIA/Mellanox 0.5m 200G QSFP56 Passive Direct Attach Cable for DGX Spark (Powered by the GB10 Grace Blckawell Superchip) Dual-System Interconnect - NADDOD

