I bought a DGX Spark a few months ago I bought, I’m a developer and so far I haven’t found any model good enough, the only one roughly drinking is Qwen3.6 27b but is much too slow and makes 70% of the time bugs, moreover Qwen no longer seems to be interested in releasing OpenWeight models at least in the field of development.
Qwen3.6 35b all he knows to do is do dozens of multiturn sessions to write erase write correct erase recreate to finally have something that doesn’t work…
There seems to be no actor who wants to release models better than Qwen3.6 27b and faster.
All new models are now for at least 2 Dgx Spark now it’s frustrating…
I also wanted to try finetune models but there was never a single difference and 90% of the time is even worse (Ornith, Agent, Tess4 …) each time they lie there is no benefit in dev.
I also tried to fine-tune I also tried to fine-tune myself with my own claude sessions with GLM5.2 but with only 128G of ram the max that I can recalculate it’s 8k of context so it has no interest either and I’m 99% sure that finetuner does not bring anything at all for the dev.
Do you think a new actor will come in reinforcement one day or is it over for the Dgx Spark at least unless I have to spend €5,000 again until it’s obsolete again ?
It’s a good job that has been done unfortunately in my field I need at least 20-30 tok/s of decode at 64k context, it’s unfortunately a little bit too slow but Ill try it thank you :)
i think you mistaken the spark as inference mashine
wait a year and you need 16 sparks for the models .. or more kimi 3 is 2.6 t param .. obviously you wont run frontier stuff local .. - not even if you spend 500k (hgx box with 8 b300) .. we sizeing up .. yet that doesnt make them useless .. if it doesnt do it for you - you should have no problem to selling it .. plenty will take a discount on a unit
i intended to only buy 1 spark and ran a 2 or 3 small models simultaneously, but the fact that there is an expensive NIC in there that was not being used caused my dumb brain to rationalize getting a second. then i started seeing posts in here like “quad sparks is the new dual” (thanks mashie :P) lol. there will always be hardware upgrade lust
in the end, 2x sparks with dsv4 flash has actually stopped the fomo for me.
2x sparks has been amazing for me. I started with a strix halo. Halo is a good entry to local AI and can also be a very good workstation as well as a decent inference device but inference is much slower on a strix. Upgraded to 2x sparks and do not regret it at all. Dual sparks is fantastic for inference. I think if you only get a single spark you get a false impression that the spark is meh when in reality single spark can be meh and 2-4 sparks is fantastic you just have to get used to running a cluster.
In 1 year no more frontiere model will be released I think or else they’ll change the business model and switch to renting out access to an LLM, I don’t think it will really be open source anymore (at least for the biggest ones). That said, I think that even a year from now, an active10–20B model will still see improvements for developers. They just want to show off who has the biggest one by rolling out models capable of launching a rocket and doing a bit of everything, for a model focused exclusively on development, it is possible to keep it within smaller-scale models.
llms dont scale that way .. we are far over chinchilla optimum - also its not said that we can actually cramp way more into small models - i like your enthusiasm - but there are fundamentals in research
I think that if they create an active moe 12b model and total 120b without teaching him trivial things like what year was born Pierre Paul Jacques and mainly on topics in development (80% dev and 20% important general topics) I think there is a way to get something better.
I tested it, the decode is rather acceptable but there is a big delay each turn from the prefill and there is a big quality concern he made several errors calling tools forgetting parameters
A single Spark by itself is equivalent to a MacBook or Mac Studio with 128 GB of RAM, which can be difficult to find right now, except the Spark can cluster better and has Nvidia compatibility instead of the lower compatibility of the Mac ecosystem. You can run a number of perfectly useful LLMs on 128 GB, but of course not anything close to state of the art. A cluster of two Sparks is comparable to a 256 GB Mac Studio, which can’t be ordered currently.
Troll. You are just admitting you dont own one without saying. I am going 2 billion token a week . If those tokens were cloud services its already above 4k$ a month.
I could buy 1 spark every month. If you Own one , and learn to run good models , you wont be saying that.
Agree, best and pretty much only fast enough option for single spark is 122b. Best part of spark spec is that you can add a second one and then its a different ball game. If you ever plan to stay with one then m5 max could be a better choice. You can run 27b at 40 t/s in fp8, top quality. But the big downside of Mac influences never say on youtube: it’s terrible at concurrency and much slower on prefill than spark, but twice faster on decode. If you mostly use a single thread and stay under 150-200k you better off with m5 max 128gb. I use 2x spark with deepseek V4 DSpark for most tasks and use Mac to run subagents on small context on different models.
In one year the bubble will have well popped and no one will even want all the data center hardware because it uses way too much electricity to make any sense in a post-bubble world. The only buyers will be people who still want to mess around training large foundation models. That will quickly dry up too, because they’re slamming into the scaling wall anyone could have seen coming all along. The good news is the memory manufacturers will have upgraded all of their production lines to HBM and won’t want to spend money downgrading them to DDR. So that capacity will be available cheaply and vendors will work on developing edge compute products that use HBM. It’s likely the Spark 2 will use HBM. Its competition certainly will.
You confidently say bubble will pop. I wouldn’t be so sure that popping means everything goes down and hw cheaper. I think Google, Meta will eat openai and Anthropic as they own whole vertical. But this by no means indicates low demand in hw. Ai is not going anywhere. There is a bubble, but in certain names not in industry. Google, meta, ms, nvidia valuations are super healthy
1 year? Even with HBMs avalbile to Consumers it won’t burst. all those HBM will be taken over by giants .
AI Progress is just started. It will demand more and more compute not less.