Hey everyone, well I aped into all of this with 8 Sparks and a CRS804. I have all the hardware set up, but I’m at a loss for how to get DSV4 Pro running? Can anyone provide a recipe for that?
Or a recipe for any model large enough to require an 8x cluster? There’s obviously at least a fair bit of interest in this based on the posts and discussion here in the forum. Nvidia would be smart to provide some guidance on it since they’d be likely to have more people setting up large clusters so more buying.
GLM-5.2 would be an ideal model for 6-8x. NVIDIA never really intended to officially support more than two DGX Sparks being clustered, at GTC they expanded to four sparks because of market interest. The problem is if the DGX Spark scales out too well and is “too capable” then they can’t squeeze people for even more money to get their more expensive products.
if you have bought 8 sparks and are not going to make the effort to to figure out the recipes yourself, then you have wasted money on 7 sparks.
Not many people on these forums have 8, so anything they are able to share is fantastic, but really if you have just spent the money to buy 8, you also need to go spend your own time on getting 8 to work with your favourite recipe
@truxnor thanks for chiming in. I am making effort to figure it out myself and if and when I do I’ll share what I did. To me it doesn’t matter if that’s even in a month or two. Nothing has been “wasted” because 4 pairs of 2x running the models they can run is still better than 1 pair of 2x.
I’m just asking the people who have done it or Nvidia themselves because it’d be good if a recipe were posted here just like it’s good that recipes for dual Sparks are posted here.
Hopefully I’m wrong but I sense a bit of hostility toward people trying to discuss larger clusters here and I think that’s counterproductive if that’s the case. I find it quite odd because many people have cars, even multiple, that each cost more than a cluster of 8x Sparks. I’d personally prefer to have the cluster since I can do more with it than a second car, but I wouldn’t hold any ill will at people with multiple cars.
@sjug thanks. GLM-5.2 does look very nice.
There is no hostility.
I’m just suggesting you need to help yourself, very few people have 8 sparks. The ones that usually do, have usually spent the time and effort to get recipes working, and might be posting and sharing their recipes and asking for help to maybe see if others can improve upon them.
There is no sign you have done any of this and are asking for people to help you, I’m merely suggesting that you need to help yourself as I will re-iterate, very few people have 8 sparks, so you are not going to get a lot of help to run the recipes that you want to run.
@truxnor honestly your attitude right there is obviously hostile. The sign I’ve “done any of this” is that I literally just said I’ve been trying and will continue to do so. As for there not being many people, thanks, I already knew that based on the content here. You come off very weird. I never said people have to share any info or that I demand it or anything. My attitude has been perfectly pleasant, I’ll continue to work on it and I’ll share anything that’s successful. Thanks.
I have a cluster of 8. Look at this and other posts i shared my findings in. Welcom to 8x club. MiMo-V2.5-Pro-FP4-DFlash - #12 by ciprianveg
You might be a bit resource constrained with pro. Across the cluster you’ll probably need to reserve around10% for the os, and you’re looking at circa 850gb for weights. It’s not leaving a lot of room for context.
Also just wanted say best of luck - definitely interested to see how far clusters of 8 can be pushed and whether it’s worth the upgrade. It seems like a fair few people have 8 but not a lot of recipes floating around. Hopefully by asking some people will share.
Thank you for the welcome, I hadn’t seen your posts! I hadn’t been considering MiMo-V2.5-Pro because the non-quantized model is too large. I’ll download the quantized version 👍
@p33zy thank you. I agree and want to see the sub-community of 8x and above build up some resources. Hopefully I can help with that in due time. I had assumed DS4P would be OK because of the context memory usage reduction in DS4, but maybe you’re right. I’ll still try because I already have the weights, but I’m also going to be trying GLM-5.2 and MiMo-V2.5-Pro. There’s also the Kimi series. It’ll just take me some time to download the weights.
Have you tried Kimi 2.7?
@aidendle94 unfortunately not yet since I don’t enjoy the best internet speeds so right now I only have the DS4P weights finished, but it’s on my list with the others. Do you know of any success with it that you could point me to? Maybe I should prioritize downloading it?
Edit: Sorry, I just realized you weren’t replying to me. I’m new here 😂
Let me know how it goes. Especially with prefill.
Swap out the MTP one with the new Dflash nvidia dropped today
Thank you so much for that resource. That looks amazing. It’ll take me a bit to get the model downloaded but I’m intent on trying everything mentioned and will report back 😇
“Swap out the MTP one with the new Dflash nvidia dropped today”: I’m capable with code and containers and so on, but I’m new to local AI, so if you could briefly explain this that’d be a big help. I understand roughly what MTP and Dflash are but have no idea how I’d swap them.
Just tell Claude to pilot it. MTP allows models to improve output speeds by 2-5x. Dflash should get you 20-30 t/s output.
Nvidia released their own Dflash model today
But the Nvidia Dflash one seems to be 2.6 not 2.7?
Should work I don’t think the difference is big enough to where it’ll reduce acceptance.
OK, that’s very interesting. I didn’t realize it was possible to swap out a part of the model like that. Thanks for helping me understand.