The MikroTik CRS804 switch works very well for an 8x cluster of Sparks, but I’ve noticed that the Juniper QFX5130-48C-AFO can be bought new-in-box for a little over $6,000 and has 8x 400G ports, along with 48x 100G ports and 2x 10G ports that could be used for other things. It should work well for clustering 16x Sparks using the 8x 400G ports, for a total of 2TB of cluster memory. Of course, there are other enterprise switches that could even do 32x or 64x Sparks, but this switch seems relatively low power, can run off of 120v, is relatively compact, and is much cheaper than those other options. Has anyone been able to do anything with a cluster size above 8?
There are plenty of 8-bit and 16-bit models that can only be run on an 8x cluster, such as the recent DeepSeek-V4-Pro, but as an example of some of the models that would require a 16x cluster:
Then there are some models that would require a little more than a 16x cluster could offer, so would need a 32x cluster to run at 16-bit, or could be run at 8-bit versions on a 16x cluster:
This recent video implies that it’d be possible to network a 16x cluster by having 2x MikroTik CRS804 switches split the load, where each GB10 is connected to both of the switches. This is possible because the GB10 has 2x ConnectX-7 ports. It’d have the benefit of being cheaper and of using the MikroTik switches, which are quieter, lower power, and run cooler than any of the enterprise switch options, since the enterprise market doesn’t really care about those aspects as much.
I’m not sure if the networking set up would be much more complicated using two switches, or if there are any limits there. Has anyone done this? It seems like it could end up being more complicated than just using an enterprise switch with 400G to 2x 200G cables, but maybe not?
@Patrick-ServeTheHome, sorry for pinging you, but I think you might be uniquely able to help answer the question: would a cluster of 16x Sparks work with two CRS804 switches each handling 100Gbps using the 400G to 4x100G cables in my previous reply?
At 16x, I would probably just do a Dell Z9332F-ON or something like that and get a 32-port switch you can split out. That is your easy path. You can use DACs. It will be loud and use more power, but the setup is straightforward, and with everything on one switch, troubleshooting is easy.
400G to 4x 100G, we usually just convert to DR optics so we know each lane is running at 100Gbps, which makes life easy.
We do not have 16x GB10’s, but we have split out the MikroTik 400G ports with DR4. Then use MPO/MTP-12 to LC breakout OS2 cables. Four of the six LC pairs then go to 4x DR1 100G optics in QSFP ports. We have to do this to hook it up to our Keysight load generation cards, which are QSFP28 100G/ port.
This is one of those, you can do it, but I also think you would be wise to just get a bigger switch. Not everyone has drawers full of optics and other components at their disposal.
@Patrick-ServeTheHome, thanks a bunch. I’m not as knowledgeable as you, so I’ll continue to digest your reply. My takeaway so far is that you’re saying it’d be too difficult (but still possible?) to do it with 2x CRS804 switches? I think the Dell Z9332F-ON looks like a good option at less than $4,000 open box, thanks, but I am worried about the extra noise. As for using optical cables, I’m seeing prices that I’d rather not pay, but maybe I just don’t know what to look for. If it could be done with 2x CRS804s and DACs then all we’d need is ~$1,300 additional for the DACs. We wouldn’t mind putting some effort into it, assuming it’s achievable. Although having a switch capable of a 32x cluster could be worthwhile in the future, assuming larger open source models continue to be released…
IMO At 16x you are going to have complicated management and configuration of this setup, lots of different failure scenarios you need to manage, latency issues and it would make more sense to spend more and move to NVIDIA DGX Station and erase those problems from existence.
@raphael.amorim, thanks a lot for providing your perspective as a valuable community member, but we view it quite a bit differently. I also have to be skeptical that it’d be *so* much different from running 8x clusters, of which we’re running multiple.
We already have all of the hardware required other than the switch solution, for which there are two possibilities: either use one switch (new-to-us but used or openbox switch, more expensive) or two switches (new cables, less expensive). Using a two switch solution with MikroTik switches would be much better for home user and SMB usage and be more accessible using new hardware that’s still being sold through primary channels. We’re talking about it here in the interest of building / learning with and for the community.
As for the alternative of using a DGX Station, that has some downsides that are essentially dealbreakers for us. First, even using two in a cluster, which is the max that’s officially supported (same as the Spark used to be before this community), it can’t run as large of models as a 16x DGX Spark cluster could, despite costing way over 2x the cost, although who’s really to even say as there’s no transparent pricing. Second, it needs to be purchased through specialized (read annoying and non-dev focused) sales channels with much more upfront cost, whereas with our DGX Sparks we were able to gradually get into it, 1x, 2x, 4x, 8x and so on, which allowed us to do incremental validation of our use cases before making a larger commitment. Third, there’s not and never going to be as large of a community that grows around it compared with the GB10 community and we think that means a lot and will continue to buildup in interesting ways.
We have to say, we got the huge help of this community, a single amazing switch manufacturer in Europe (MikroTik), and @Patrick-ServeTheHome and others trailblazing to 4x and 8x clusters.
We like the idea of our GB10 clusters being reconfigurable, similar to what @Patrick-ServeTheHome has expressed, because it allows cluster size to be adjusted relatively granularly, which we also think will mean a lot over time.
Edit: I realized I didn’t include the cost of doing the DGX Station networking, which would be significantly more expensive than the Spark platform.
I do think you are right that the DGX Station GB300 makes sense for a lot of folks over these. Frankly, if you are running a model where everything fits in 250GB or so, you are going to be golden on that platform.
If you want to do a 750GB-ish model, then an 8x RTX Pro 6000 Blackwell Server Edition might be the right way to go, albeit at >10x the power and it is not going to sit in your office.
Both of these will also cost more than 2x what an 8x node GB10 cluster costs, and use more power, but there is an upside in speed and lowering complexity.
If you want to run a >1TB model, then that will be more challenging. May should be when we start to get on those, but I highly doubt the GB10’s will sit idle. I think of these more as adding token capacity rather than replacing.
If my goal were to try many different types of models, a large pool of GB10s that I could reconfigure would be a really neat capability.
First, I want this 16x node cluster numbers on https://spark-arena.com. Please contribute all the models you can test. We would be forever grateful.
In terms of costs, flexibility and larger memory pool, easy to acquire and cluster of course I would technically agree with both of you, because analyzing those aspects leads to a single mathematical, economical and logical conclusion. That’s why I said: it’s just an opinion and focused on configuration, cluster management, hardware failures/configurations (during training, for instance),data replication … etc
If you don’t look at the models you can exclusively on the 16x cluster, but you’ll be able to run most of them at FP8 (10 PFLOPS) or NVFP4 (15 PFLOPS without sparsity) on the workstation, you can expand it with RTX Pro 6000. you have 7.1 TB/s HBM3e memory bandwith, 396 GB/s LPDDR5X memory bandwidth with ECC … It’s just a more convenient and reliable system for 95% of the use cases, although more expensive as well).
Including 120B+ dense models and MoE’s with large number of active parameters that would never run that well and have reached scaling limits on 16X Spark Cluster. For training the Spark scales much better. I totally see 4x and 8x as very good clusters. 16x I think we’re going to see some diminishing returns, particularly on latency that will kinda defeat the purpose.
@ash.x.kingsley 16x cluster by having 2x MikroTik CRS804 switches split the load …that will be a NCCL configuration nightmare! NCCL is not plug-and-play.
@raphael.amorim and @elsaco, thank you both for your input. We’re only focused on inference with these, although maybe some fine-tuning training in the future, but really just inference. Maybe I’m way off base, but even though it’s not plug-and-play, my assumption would be that it’s conceptually similar to the set up for a 3x cluster without a switch, no? Each ConnectX-7 port is connected with a cable that takes a different route, with each providing 100G.
@eugr, sorry to ping you, but what do you think of this possibility with 16x nodes using two switches?
@eugr has also said that bandwidth is not so important, and that inference will work just fine with only 100G links. So another option could be 400G to 4x100G cables on a single CRS804 and just accept that the bandwidth will be halved. That’s actually the cheapest option for us to test.
I’m really grateful for everyone’s participation in this thread and worst case outcome here is we’ll decide to try a 16x cluster on a used / openbox enterprise switch, but I’d really like to fully explore the possibility of using 2x CRS804 switches, or just 1x using half bandwidth. We view this as a nice learning opportunity. If a 16x cluster ends up being way past diminishing returns then at least we’ll have learned that, but we’d like to give it a shot and maybe it ends up being useful for larger models.
Unless you intend to partition the cluster, to have all 16 nodes working together setup a trunk between the switches using one of the 400G ports and the remaining three ports split into 100G link, then connect eight Sparks on each switch. This is how I’d do it! Having a real switch (CSR804 it in the toy category) and connect all Sparks to it at 200G would make an awesome Spark cluster. For the record I only use two Sparks back-to-back, a.k.a the poor man’s cluster
Well, I mean, you could just use one switch in this case, as the bandwidth is not as important as latency, but two switches should work as long as you are using spark-vllm-docker that includes subnet-aware NCCL build. Assuming you want to use 4x splitter cables for each CRS804 400G port.
That’s great to hear @eugr! Thank you for your excellent work. Yes, with the CRS804 we want to use 400G to 4x100G breakouts. We have some first priorities while running our 8x clusters right now, but will definitely be testing this and report back within a month or two. If anyone else moves faster than us, please share your results too.
I don’t think it’s fair to call the MikroTik CRS804 a toy switch. It’s datacenter tech newly packaged with concern for power usage, heat, and noise, at a price point and with sales channels that make it easy to dip into. We’re very happy it exists and probably wouldn’t have embarked on this journey if the only option were buying traditional used enterprise switches. It’s just totally different markets. We see a higher-end of the “prosumer” segment opening for AI with all of these things. MikroTik, if you’re reading, pay attention to how this develops. Maybe it’d make sense for you to produce a new switch with 2x the capacity of the CRS804!
(CSR804 it in the toy category)
I mean, MikroTik is using the same Marvell Prestera lines that Cisco uses in their lower-end switches. An example, albeit much lower-end than this, is the Cisco Catalyst C1300 series. This is a higher-end segment than that. Also, Marvell makes hyper-scale 51.2T switches and what have you from the Innovium acquisition (we took apart the 51.2T and 12.8T generations in videos). The CPU in the CRS804 is an AWS/ Annapurna Labs one.
Granted, I would still say Tomahawk4, Spectrum, or newer would be better, but there are realistically only so many 1.6T PAM4 switch chips on the market. The challenge right now is that high-end switching is on fire. Low-end for WiFi-7 and other lower-end clients is relatively healthy. Between that, there is not a huge amount of investment. Just for a sense of scale, 1.6T is now a single switch port on higher-end switches that are available today.
@eugr Just to be clear, yes, but you would need a gearbox in there to do that, right?
If you are referring to splitting the load between two switches, you’d just need to make sure both CX7 ports are on different subnets (as they should be). Now, I haven’t tested that, but I don’t see why it wouldn’t work.
I doubt it will translate to any gains in vLLM though, compared to a single switch.
My brain was stuck on converting DR4 to DR1 since we are hooking the CRS504 to the Keysight cards today. Congrats again on the new gig!