At my disposal three devices 2pcs DGX Spark + 1 Asus DX10 and here it seems to buy one more and get into the club of owners of four devices and run large models! It seems reasonable and worth buying, but my mind stopped me!
why 4 devices to run large models? In terms of cost vs speed, it’s not reasonable. A large model will work, but with a low token generation rate, even if it runs 24/7, it won’t outweigh the cost of connecting the same model via API and using it for a couple of years.
The market for free large models is decreasing, for example, QWEN has taken away our joy, and I believe that in the future, others will also close access or release LLMs with free licenses less frequently.
A single DGX Sprk device is already impressive! You can run qwen 3.6-35b with a huge context and multimodal abilities - this is the golden mean (subjectively)
DGX Spark - x2 - also great, the model is average, for example, Minimax 2.7 - it’s not the smartest, but it already provides higher quality and doesn’t compete with large paid models or take away the bread from major corporations (which are likely to continue releasing new versions for free), who invest millions in training LLMs. That is, this segment also doesn’t bother anyone)
P.S. How much is it reasonable to go into the + 512gb VRAM segment, spending a large amount and running one slow LLM on 4 devices, or is it better to distribute tasks by creating an orchestra of small and medium LLMs + using paid APIs to solve complex point tasks.
What will your colleagues say?
#4 is the option I’ve went with Qwen-3.6-35B local ~ 55 t/s per concurrent use on average and using Gemini Pro for the brain power for OpenClaw, the combination is working out really well so far.
I completely agree with you, but when working on a project that is reasonable to finish as soon as possible + Unfortunately, there are big doubts that qwen3.5 397b will have an update, and this is a shame((
The frequency of releases for capable open-weight models is a function of the release cycle of API-gated frontier models. Open-weight players, and most notably Chinese companies, are playing smart by eroding with their models the market share of API-gated solutions. Ultimately, it is a matter of sustainability, or lack thereof: As long as proprietary models will retain the edge, China will play catch-up, pushing, in turn, more expenditures in this space. Of course, this goes to the detriment of those actors whose energy bills are higher (i.e., sure enough not China, for structural reasons). Notably, the capabilities of the released open-weight models can be deliberately chosen, so as to sustain this financial “blood bath” at will. Meanwhile, what is released is asymptotically more and more capable, and it will stay with us for the foreseeable future. We are living interesting times indeed.
Yes, strangely enough, but mostly we are happy only with China’s open models, and many thanks to them for this) European and American companies mostly play only towards the paid API, and sometimes crumbs in the form of small LLMs reach us. Of course, I don’t blame them for this, they have their own philosophy on this issue!
Open models are already starting to be smart enough for production work. The worst case is that they don’t release new models, and you stay at where you are now, and that is not that bad to be honest. They are smart enough for many works already.
More likely is that they just release the model slower, but they will still do so, let’s say release models two generations older, and that is still an improvement over the models we have now.
Given how much they invest on training these models, with all the hardware constraints they have due to politics, and they are releasing them for free, I do not see much to blame. I am still very optimistic for the future.
In China, competition is fierce, so if one company does not release free model, another company will, and so there is really nothing to worry about. We cannot always ask for everything they have. Perhaps the only concerning factor is when they are blocked completely from building or developing AI due to you know reasons, or when our hardware is blocked from running them, then we will have no choice but live with the paid closed models in the states. That is the only situation to worry about.
It’s important to note that outside of capacity, scaling is not linear. Added nodes increase complexity and overhead.
If you’re trying to go for 1T Parameters it’s important to know what model you’ll run and how it will scale. I don’t think it’s possible to get good speeds on any of the 1T sized models on spark meshes.