Suggested cable to link two Sparks?

and here’s one more from FS

So when I ordered my cable from NADDOD, I hadn’t discovered this post yet. Not knowing what to get I ordered the MCP1650-H00AE30 instead of the MCP1650-V00AE30. I assumed I wanted the Infiniband cable. I’ve been using it. It does seem to work. I think physically they are the same cable, but there is different identifier in the EEPROM. I don’t know if NADDOD maybe “bins” them differently. I think Infiniband assumes a lossless fabric, where ethernet can handle packet losses better. According to Gemini the DGX Spark is shipped and architected to use Ethernet for its clustering fabric, specifically RoCE v2 (RDMA over Converged Ethernet). The ConnectX-7 is a VPI (Virtual Protocol Interconnect) card, meaning it can be switched to InfiniBand mode. However, this is not a simple toggle for the Spark platform. InfiniBand requires a Subnet Manager (SM) to be running on one of the nodes to manage Local IDs (LIDs) and route traffic? It would be nice NVIDIA would document this fully. I guess they assume that DGX Spark owners are already familiar with the technology.

The only odd thing I see are these warn messages:

(EngineCore_DP0 pid=1896) (RayWorkerWrapper pid=2112) [2025-11-22 05:14:24] spark1:2112:4122 [0] p2p_plugin.c:421 NCCL WARN NET/IB : roceP2p1s0f0:1 unknown event type (18)

Otherwise I think it’s working:

(EngineCore_DP0 pid=1896) (RayWorkerWrapper pid=2112) spark1:2112:7697 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0

Granted, I’m not an expert in Infiniband stack, but Spark seems to use IB connection just fine:

$ ib_write_bw 192.168.177.12 -d rocep1s0f1 --report_gbits -q 4 -R --force-link IB
---------------------------------------------------------------------------------------
                    RDMA_Write BW Test
 Dual-port       : OFF          Device         : rocep1s0f1
 Number of qps   : 4            Transport type : IB
 Connection type : RC           Using SRQ      : OFF
 PCIe relax order: ON
 ibv_wr* API     : ON
 TX depth        : 128
 CQ Moderation   : 1
 Mtu             : 1024[B]
 Link type       : IB
 Max inline data : 0[B]
 rdma_cm QPs     : ON
 Data ex. method : rdma_cm
---------------------------------------------------------------------------------------
 local address: LID 0000 QPN 0x03ec PSN 0xb680ae
 local address: LID 0000 QPN 0x03ed PSN 0x808800
 local address: LID 0000 QPN 0x03ee PSN 0x5b694a
 local address: LID 0000 QPN 0x03ef PSN 0xe2efd1
 remote address: LID 0000 QPN 0x03eb PSN 0x75f6ee
 remote address: LID 0000 QPN 0x03ec PSN 0x436140
 remote address: LID 0000 QPN 0x03ed PSN 0x81698a
 remote address: LID 0000 QPN 0x03ee PSN 0x4a8b11
---------------------------------------------------------------------------------------
 #bytes     #iterations    BW peak[Gb/sec]    BW average[Gb/sec]   MsgRate[Mpps]
 65536      20000            111.72             111.71             0.213070
---------------------------------------------------------------------------------------

For the DGX Spark the QSFP56 cable is the best option. The QSFP112 is for the 400Gbps links and won’t help the Spark squeeze any more bits since it’s capped at 200Gbps system wide. I’m using this cable from FS and works very well:

Where did you buy NJAAKKR-0006, @eugr ? Amphenol’s website?
$42 at Digikey, but looooooooong lead time.

ExxactCorp.

I get the same warning messages with the NADDOD MCP1650-V00AE30 cable, if it makes you feel any better.

Laser? Is that an active optical cable? Ours are copper.

Hey @eugr and @maiia,

How’s the experience so far with the spark cluster, what type of workloads have you been experimenting? Any testing with distributed training yet?

A bit frustrating, to be fair. I’d appreciate a bit more clear guidance and participation from NVidia on this. For instance, whether we need to bond both interfaces for a physical port together to get 200G speeds, because I’m not able to get more than 100G with iperf3 or ib_write_bw.

Other frustrations are mostly related to the state of Blackwell support in VLLM and the worst mmap performance on any system I’ve used recently. When it takes 10 minutes to just load a 80GB model into vllm, it results in a lot of wasted time.

Haven’t done anything useful in terms of clustering so far…

We’re capped by the DGX Spark Physical PCle Link Hardware.

This is from Gemini 3 Pro:

sudo lspci -vvv | grep -A 20 "NVIDIA\|Mellanox" | grep -E "LnkCap|LnkSta"
		LnkCap:	Port #0, Speed 32GT/s, Width x4, ASPM not supported
		LnkSta:	Speed 32GT/s, Width x4
		LnkCap:	Port #0, Speed 32GT/s, Width x4, ASPM not supported
		LnkSta:	Speed 32GT/s, Width x4
		LnkCap:	Port #0, Speed 32GT/s, Width x4, ASPM not supported
		LnkSta:	Speed 32GT/s, Width x4
		LnkCap:	Port #0, Speed 32GT/s, Width x4, ASPM not supported
		LnkSta:	Speed 32GT/s, Width x4
pcilib: sysfs_read_vpd: read failed: No such device
		LnkCap:	Port #0, Speed 2.5GT/s, Width x16, ASPM L1, Exit Latency L1 <4us

PCIe Link Speed: Each Mellanox/NVIDIA NIC shows “Speed 32GT/s, Width x4” (PCIe Gen5 x4). The expected is x16 for full-bandwidth. Gen5 x4 maxes out at ~16 GB/s per lane, and only ~64 Gbps per cable—not sufficient for your 200G physical link, which explains the observed effective speed ceiling.

I wonder if NVIDIA will review BIOS/firmware options to enhance PCIe slot configuration bottleneck.

I can’t believe I’m spending this long in device setup before getting to work on anything meaningful.

When a cable connector is inserted in one of the physical ports two virtual links light-up, i.e. enp1s0f0np0 and enP2p1s0f0np0 Each virtual link is x4 wide capable of 32GT/s, so the combined output will be x8 and 64GT/s speed. That’s enough to send, theoretically, 200Gbps on the wire!

Even the Vital Product Data claims it:

Product Name: NVIDIA DGX Spark, P4242-0000, 2-port QSFP up to 200G, Ethernet, PCIe5

Use lspci -vvvs <port ID> for details.

This cable started selling at Microcenter:

https://www.microcenter.com/product/703095/pny-certified-stacking-dgx-spark-cable-40

Theoretically, the official certified cable for stacking DGX sparks by PNY.

That’s a longer version, NJAAKK-N911. The one they had before was NJAAKR one.

It says 1.64 ft. (0.50 m) for this one. The previous one says (NJAAKK-0006):

usage: Connector

the new one (NJAAKK-N911):

Usage: Stacking

On PNY’s website, they list both

:) Confusing

The only difference is cable length:

NJAAKK-N911 = 400mm (0.4 meters)

NJAAKK-0006 = 500mm (0.5 meters)

Both cables are made by Amphenol. The part after NJAAKK signifies the cable length. See details at https://www.amphenol-cs.com/catalogsearch/result?query=NJAAKK&page=1

Thanks for this suggestion. I struggled to find a suitable cable available here in the UK but the NADDOD one arrived quickly from China and is working for me: NVIDIA/Mellanox 0.5m 200G QSFP56 Passive Direct Attach Cable for DGX Spark (Powered by the GB10 Grace Blckawell Superchip) Dual-System Interconnect - NADDOD