Is it worth buying a second DGX Spark now?

I have a second DGX Spark reserved (ASUS, 1 TB SSD) for €5,100 including VAT. I bought my first one in March for €3,200.

I’m currently running the Qwen3.8-Flash-Next model. I’d also like to run faster models with around 6,000 tokens/s prefill, as well as models for speech recognition, speech generation, and image generation.

For those of you running two DGX Sparks: what kind of setup are you using?

Do you run a single model across both Sparks, or do you use them separately—for example, one Spark for LLM inference and the other for speech/image models?

And most importantly: does adding a second DGX Spark make a significant difference in practice?

I’m trying to decide whether it’s worth buying the second one at the current price

I makes sense if you want more quality, not necessarily more speed. Two models for two sparks that make sense: GLM 5.3 Flash and DeepSeek v4 Vision Exp.
DeepSeek is pretty fast but not as fast as you want - you may consider M5 Ultra for that (5090 speed). Image gen on Spark is not fast too (slow ram dictates slow decode which is almost entirely what image/video gen does - very little prefill). Sparks good for many concurrent and good prefill (very fast GPU but slow ram), but M5 Ultra is beating it at the prefill game too. So - if you have one, you may want second to run higher quality models at acceptable, not lightning speed across two. Otherwise either Apple or custom build with many 3090 etc.

In my opinion - if you can sell your spark for more than you bought it (should be possible) and buy M5 Ultra 256GB - you will be happier. Price would be similar or just slightly higher (depends mostly on what SSD you chose - very pricey at apple’s). Then you can run big model fast, small models super fast, image gen, video gen, few things at once, cache sessions to SSD with oMLX, change engines, come back to oMLX and restore context from SSD without re-prefilling.

Is the recipe/engine quality and breadth the same in the Apple space, though? The stability of most (good) Spark recipes has been astounding. I’m a little wary about Apple’s ecosystem.

Also I guess depends on your entry price, but I think a 256 Apple is just way more than you could possibly sell 2 sparks for right now. Probably quite region dependent. I got mine for $7k total, they would probably sell for maybe 8-10k if I’m lucky? But after Ebay fees, much less. Then the Apple would be like what, 12k for a 256? Decent chunk of money out, almost another Spark worth. Might be worth it depending on what you need though.

I’ll second what 0rand said, I’m using GLM 5.3F, EXL3 entrpi engine, and it’s been incredibly intelligent. I wasn’t originally going to keep my second Spark, but this model nudged me to do it. If you just need code gen (and you know exactly what you want) I guess you could just consider a 5090. Running some Qwen 3.8 variant, it should be incredibly fast. All of these suggestions vary depending on what you want, though I know “what you want” can kind of vary depending on what you have…

I see significant added value in a 2x cluster compared to a single Spark unit, as it enables you to run models in the ~300B (4-bit quantization) or even ~550B (2- to 3-bit quantization) range - models that operate in a completely different league than anything runnable on a single Spark. For me, DeepSeek v4.1 Flash currently represents the pinnacle of what is achievable with a 2x cluster.

While the 256 GB version of Apple’s M5 Ultra might currently be a strong contender, opting for it means sacrificing some of your informational self-determination and becoming part of an ecosystem that - to me - is questionable. Furthermore, comparing the price of such an M5 Ultra setup to a 2x Spark cluster isn’t really an apples-to-apples comparison.

Moreover, MLX currently lacks the maturity and functionality of CUDA; the expertise you gain with a DGX Spark is far more valuable in a professional context (scaling to B300 hardware etc.) than the knowledge acquired within the relatively limited Apple ecosystem.

I’ve run my two a couple different ways.

  • 2 x vLLM clustered with Qwen3 Coder Next, Qwen 3.6 27b + ollama on one system for small models, embedding, and testing + comfyui and automatic1111 on the other for occasional use. I would turn off qwen3 coder next vllm if I was doing any video gen, I had ok headroom for image gen with everything running.
  • 1 x vLLM clustered with MiniMax M2.7. Not a ton of headroom for other things.
  • 1 x vLLM clustered with DeepSeek v4 Flash. Very little headroom there to do anything else.
  • 1 x vLLM clustered with Qwen3.8 Flash Next and have a little headroom for ollama on one and comfyui / automatic1111 on the other. Haven’t had need for image/video gen recently but think I have plenty of headroom for image gen, though video gen is iffy with the amount of free RAM.

I have a Mac Studio M4 Max w/ 36GB RAM. The main difference I notice in inference is sparks are way faster due to prefill. Token generation is slower but I mostly do agentic / software dev stuff and the prefill is so much faster on sparks that it spanks my Mac.

I’m curious to see how the m5/m6s do on prefill as if the Mac caught up there it would be more tempting. The other reason I went Spark over bigger Mac is if you want to do more cutting edge image/video/audio stuff, it comes our for Nvidia hardware first and there were things I wanted to experiment with on my Mac that I couldn’t because the software was only available for Windows/Linux and cuda.

That’s pretty interesting to read from you, because I also have mac (m5 max 128) and 2 sparks and I only use sparks for inference and all other stuff including image gen (just for fun), tts/stt for my other systems - strictly on Mac.

As for performance, it was different 2 weeks ago and I would not have suggested it but the turntables have turned as MLX cracked open the in8 hw acceleration on m5 and everything went ballistic. And M5 Ultra is ~ 2x that performance.

Yes buy another on, more if you can

If you need it and have the money then buy it. One spark is not the best option. Should be two or more.

If you can wait you might see devices with 196gb ram and higher bandwidth at 5k range in the future.

If you like to sell many will buy from you, but then what are your options M5 ultra will start at 10k

How stable is it though? Is it still hard to find stable recipes? Also why is your 8x increase so disproportionate to your 2x and 4x?

Your dilemma is shared. The price inflation is about 2x from when I purchased it, and it’s been online like an absolute tank since it booted.

So, double the capacity or invest in new hardware that’s got more leg-room for improvements? M5 Ultra’s 256gb unit is a little less than the current cost of 2 sparks off-the-shelf. If you’re in the Apple ecosystem like me, that’s a wildly compelling idea.

However upgrading to a 2nd offers you an incredible infrastructure that is only getting faster by the week. By spring, a 2x cluster will be enterprise standard for small dev pool, or small/medium small brick and mortar shop.

If there’s another spark price increase, cutting loose and going M5 would be a lot more appealing.

Just follow omlx. Best you can get. Jun, the developer, was just hired by Hugging face to support mlx full time. Just like Eugr was by nvidia. What is funny nvidia is supposedly to buy hugging face.

Speed increase - just mtp hit optimized path. It’s not really increases much of mixed context, don’t bank on it. Single req numbers are correct. I kept working with it for days now, it was building my stacks for mimo. Zero cloud use. 390k context works, still 40+ t/s. Lightning mtp, best 8 bit tq, as much ssd cache as you can 4gb ram cache is optimal - it reserves activations on top and uses lot of ram for it, and ssd on Mac is insanely fast. You change model or reboot engine, come back and no reprefill - session loads from ssd in seconds

Now that’s very intriguing.

My next move is to swap/sell a 5090 to go for either a 2nd spark, or an M5. My brain is struggling to comprehend how fortunate I might be to make this decision, or short-sighted now that the fall updates to d/rtx are due.