Suggested cable to link two Sparks?

I’ve looked at the files in that repo, they seem to be generated by ChatGPT, so I wouldn’t put too much trust into it.

Not yet, but I ordered the “official” one from ExxactCorp, should arrive on Tuesday.

You seem to know what you’re doing (more than I do at least). Although my first computer use was in 1967–a teletype connected to a timesharing computer at Bolt, Beranek and Newman in Cambridge–I majored in Government and trained as a lawyer. This is my hobby (although I had a successful Internet-based litigation support business).

Let me know how it goes with 2 Sparken.

https://www.exxactcorp.com/Amphenol-NJAAKK-0006-E193723350?

I went with this one: Amphenol NJAAKR-0006. It’s the same one that Microcenter was selling as an official Spark stacking cable. The only difference with the one you linked is wire gauge - the NJAAKK one is 32AWG, NJAAKR is 30AWG.

This is a bleeding edge field, so we all have to learn as we go, really.

This week is going to be quite busy for me, so not sure if I am able to set up a second Spark, but I’ll definitely post my thoughts when I do!

Ha I used the BBN Butterfly at Purdue when it replaced the IMP as the DARPA Internet main (one of 3 at the time) Internet backbone). And I’m still learning. ; Right now I’m trying to figure out if I can connect 3 or 4 of them and if I should just get one cable or four (since I can’t afford a full switch at those prices!).

I managed to get the MCP1650-V00AE30 cable, purchased from fs.com to work, contrary to my previous report.

On the Nvidia performance tests, I manage to get about 75% of theoretical line-rate performance, without any kernel parameter tuning, kernel modules, etc. Reporting on my setup:

Looking at the Sparks from the rear, the cable is connected to the right-hand side port on each device. On both devices, when I run ibdev2netdev, I see this:

rocep1s0f0 port 1 ==> enp1s0f0np0 (Down)
rocep1s0f1 port 1 ==> enp1s0f1np1 (Up)
roceP2p1s0f0 port 1 ==> enP2p1s0f0np0 (Down)
roceP2p1s0f1 port 1 ==> enP2p1s0f1np1 (Up)

I had trouble getting to this point. If you don’t see *Up*, it’s a hardware connection issue and nothing below will work.

On the primary Spark, I create the file /etc/netplan/40-cx7.yaml with this:

network:
  version: 2
  ethernets:
    enp1s0f1np1:
      dhcp4: false
      link-local: []
      addresses:
        - 10.0.0.3/31

On the secondary Spark, I create the file /etc/netplan/40-cx7.yaml with this:

network:
  version: 2
  ethernets:
    enp1s0f1np1:
      dhcp4: false
      link-local: []
      addresses:
        - 10.0.0.2/31

Per ChatGPT, this is the standard way to do static IP addresses for peer-to-peer connections. If you do buy a second cable, you can add another entry for the other physical interface. Notice that I have not set jumbo frames, etc.

On both devices, run sudo netplan apply to apply the configuration. Ensure you can SSH between the two devices without a password prompt.

NCCL Performance Test

I built and run this test: GitHub - NVIDIA/nccl-tests: NCCL Tests .

After you clone the repository, you can build as follows:

sudo apt install libnccl-dev libopenmpi-dev
export MPI_HOME=/usr/lib/aarch64-linux-gnu/openmpi
make -j MPI=1

You need to do this on both devices in exactly the same path.

I ran the test on the primary Spark like this:

mpirun -np 2 -H 10.0.0.2:1,10.0.0.3:1 \
    --mca btl_tcp_if_include enp1s0f1np1 \
    ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1

The btl_tcp_if_include flag was essential to get MPI to launch successfully. Here is a clipped version of the output:

# Using devices

#  Rank  0 Group  0 Pid 324646 on spark-397a device  0 [000f:01:00] NVIDIA GB10

#  Rank  1 Group  0 Pid  29206 on spark-77da device  0 [000f:01:00] NVIDIA GB10

...




#

#                                                              out-of-place                       in-place          

#       size         count      type   redop    root     time   algbw   busbw  #wrong     time   algbw   busbw  #wrong 

#        (B)    (elements)                               (us)  (GB/s)  (GB/s)             (us)  (GB/s)  (GB/s)         

...

     8388608       2097152     float     sum      -1   666.67   12.58   12.58       0   668.90   12.54   12.54       0

    16777216       4194304     float     sum      -1  1297.15   12.93   12.93       0  1385.55   12.11   12.11       0

    33554432       8388608     float     sum      -1  2813.39   11.93   11.93       0  2043.23   16.42   16.42       0

    67108864      16777216     float     sum      -1  4270.90   15.71   15.71       0  4181.67   16.05   16.05       0

   134217728      33554432     float     sum      -1  8077.73   16.62   16.62       0  8071.75   16.63   16.63       0

   268435456      67108864     float     sum      -1  14933.3   17.98   17.98       0  15163.1   17.70   17.70       0

   536870912     134217728     float     sum      -1  28723.4   18.69   18.69       0  29383.5   18.27   18.27       0

Given large-enough data, the peak performance is 18 GB/s, which is 144 Gbps, or 75% of line rate.

I am getting much lower performance with PyTorch, but that’s a much larger software stack.

It seems like it is possible to get that extra 25% performance by following Connecting Two DGX Spark Systems via 200Gb/s RoCE Network for Multi-Node GPU Training | by Doran Gao | Oct, 2025 | Medium , though that involves tuning kernel parameters and installing more drivers. I’m happy to not mess with the system software right now.

I had my gpt-oss-120b write a go program that throws multiple chat completion requests at the stacked Sparks running Llama3.3-70B-Instruct. With 12 simultaneous requests the generated TPS was about 30 and with 15 requests it was very close to 40. With 20 I got 50 and with 40 I hit 88 - 92 tokens per second. Aggregate numbers, of course, so the per request speed was between 2 and 3 tps.

Still, you can get more work done per unit time with simultaneous requests, which is a nice aspect of vLLM. 40 requests took 45 seconds to process, which is a bit over 1 sec per request average. A single request alone (no simultaneity) took 35 seconds.

Try the above suggestion by @raphael.amorim I got 22 GB/s that way.

You could just use vllm bench serve to get all kind of stats :)

And yes, concurrent throughput is much higher, because the inference is batched together. So you are still limited by memory bandwidth for a single request speed, but you will be able to serve multiple requests at the same time thanks to batching and available GPU compute capacity.

It won’t help you in a single conversation setting when you have to wait for the request to finish before starting a new one, but will definitely help with running independent workflows, like processing documents in batches, etc.

That’s applicable to any machine, not just Spark.

I’m still waiting on my Naddod and FS cables to arrive but the account manager from Naddod shared this blog from this morning. where they’ve connected more than two Sparks to an Ethernet switch and thought I’d share.

I was in 6th grade in 1967. You must be old!

Those switches sure ain’t cheap!

The NADDOD rep also told me she had checked with their engineers and you can’t increase the connection speed much by bonding the two Spark CX7 ports together and using two cables to connect the stacked Sparks. I thought I had read otherwise in the forum, but the NADDOD engineers say:

Each DGX Spark is equipped with one ConnectX-7 NIC that has dual 200G ports. People might be considering combining these two ports to achieve 400G bandwidth, but that’s not technically possible. Due to hardware limitations, the NIC’s theoretical maximum bandwidth is around 2 × 128G (equivalent to two PCIe Gen5 x4 lanes), so it cannot reach the full 400G rate.

If anybody knows otherwise, please say so. I note that 2 x 128G is faster than 200G, but still.

Their generic brand is indeed pricey but it’s half the price of NVIDIA Mellanox Spectrum -3. An Ethernet switch is useful if you’re connecting more than 4 Sparks to avoid speed drops from a series connection.

Re connection speed: I got the same response from a few iterative answers from Perplexity Max’s deep research which said I can split 100G bidirectionally or orchestrate 200G one way per Spark.

My Spark is not recognizing the second manual IP for my second spark so either about to reset everything and re-try the connection with the original NIC or once the Naddod or the FS cable(s) arrive this week force the recognition of the 2nd port to build the NCCL and start using the actual combined dual-cluster.

After I upped the mtu to 9000 (sudo ip link set dev enp1s0f1np1 mtu 9000) I got:

max_mtu:		4096 (5)
active_mtu:		4096 (5)

And the test showed 22GB avg. bus bandwidth.

Finally, got to set up the second Spark last night.

The cable (NJAAKR-0006) seems to be working well, getting ~23GB/s bus bandwidth on nccl test. Tried to run vllm in cluster mode, but not using suggested Docker setup - the cluster started, but vllm failed.

I managed to run Qwen3-235B (the Q3_K_XL quant that I already had) using both Sparks with llama.cpp RPC backend, but since it uses a standard TCP/IP stack, looks like I’m losing some speed on latency - will try to play with coalescing parameters to see if I can improve that. Tested on GPT-OSS-120B, it goes from 57 t/s on a single Spark to 47 t/s on dual.

I need to focus on my actual work now, but will experiment again later.

Cool @eugr, let us know your progress on this topic.

For DGX Stacking cables, the supported cables are listed here: Spark Stacking — DGX Spark User Guide

  • Amphenol:

    • NJAAKK-N911 (QSFP to QSFP112, 32AWG, 400mm (0.4m), LSZH)
    • NJAAKK-0006 is the 500mm (0.5m) version of this cable
  • Luxshare:

    • LMTQF022-SD-R (QSFP112 400G DAC Cable, 400mm (0.4m), 30AWG)

Full list of suppliers at the bottom of NVIDIA DGX Spark

Here are a few links that you can use to order direct . All links for part#: NJAAKK-0006

You can add NJAAKKR-0006 - same as NJAAKK-0006, but with 30AWG gauge. This was the cable that Microcenter was selling.

I’m using this cable and it works fine. Is there some issue with QSFP56?