Triple stack

You can absolutely scale DGX Sparks out, but it’s “real and bounded” and I should be clear up front that I haven’t personally run these specific multi‑node configs yet—I’m basing this on my research of NVIDIA material and public writeups, not my own lab cluster.

Two-node scaling in practice
With two Sparks hooked up over the built‑in 200 Gb/s ConnectX‑7 links, you can effectively pool enough memory to run something in the 400B‑class range in FP4 and still see distributed training efficiency reported in the low‑90s percent range (around 93% or so). That efficiency figure is verifiable from current docs and community reports, but again, I haven’t reproduced it myself, so I treat it as “vendor/early‑adopter numbers,” not hard data from my own profiling.
Small clusters via Cluster Assistant
NVIDIA’s recent software updates lean into the “local cluster” story: there’s a Cluster Assistant in the DGX Spark stack that’s designed to walk you through wiring multiple Sparks together instead of leaving you to roll your own. The supported topologies today are up to three nodes in a switchless ring, and up to four nodes if you drop a suitable switch in front of them, which lines up with public examples of people doing ~700B‑scale fine‑tuning on a four‑Spark setup.
Why scaling is “real but bounded”
The catch is bandwidth and topology more than raw compute. Once you leave the box, the inter‑node fabric is far slower than local HBM, so cross‑node traffic is always the limiting factor. That’s why pipeline‑parallel and capacity‑driven workloads (big models, long contexts, lots of concurrent jobs) tend to scale reasonably well as you add Sparks, while tensor‑parallel, latency‑sensitive workloads are much more fragile and hit a wall sooner. In other words, you scale a Spark cluster to get more capacity and throughput across jobs, not to make one single, tightly coupled stream of tensors feel like it’s still living entirely on one box.

Hope this helps, from my side, this is a best‑effort synthesis of what’s publicly documented and discussed; it’d be genuinely useful to have an NVIDIA rep or a hands‑on technical contact sanity‑check that these statements hang together the way they intend. And looking forward, it’d be great to see the clustering roadmap open up even more—higher effective inter‑node bandwidth, better support for larger topologies, and more explicit guidance on which workload patterns will keep scaling cleanly as people push beyond the current three‑ and four‑node configurations. Cheers