Full Kimi K3 running on 16x GB10 cluster

I finally managed to start the full moonshotai/Kimi-K3 on 16x GB10 cluster, on a MikroTik Switch CRS804-4DDQ with 4x 400-to-4x100gbit breakout cables running with dspark runing on coherent corpus llama-benchy at 21t/s avg and 38t/s peak, 750tps prefill, as first try with dspark enabled using Inferact/Kimi-K3-DSpark. looking forward to improving this. I’ll be adding image and instructions to git soon:

model test t/s peak t/s ttfr (ms) est_ppt (ms) e2e_ttft (ms)
KIMI-K3 pp2048 @ d4000 654.89 ± 0.00 8288.20 ± 0.00 8286.85 ± 0.00 8288.20 ± 0.00
KIMI-K3 tg1500 @ d4000 21.71 ± 0.00 37.00 ± 0.00
KIMI-K3 pp2048 @ d16000 758.78 ± 0.00 21456.92 ± 0.00 21455.57 ± 0.00 21460.88 ± 0.00
KIMI-K3 tg1500 @ d16000 25.39 ± 0.00 38.00 ± 0.00

Tool-bench looking ok:

Model: /root/models/models115/Kimi-K3 ││ Score: 91 / 100 ││ Rating: ★★★★★ Excellent

Amazing, thank you!

16 DGX Spark, OMG…

Makes me wonder whether DGX Station plus a few DGX Spark would make more sense for these huge MoE models?

Very nice @ciprianveg ! On the the whole DGX Station vs DGX Spark - advantage of the Spark, I still think is flexibility. You can go anywhere from serving on 1 to 16 nodes, and have power usage advantages that way. Also, the amount of memory :) Great solution if you mostly do batch work.

Question: Do you think performance will be better with pp=2, tp=8?

Possible, it is on my test list.

well done! this is amazing! I’m speechless…

Also, it might be too early to ask, but how much context are you able to serve with the 16x DGX Spark?

External Image

How much bandwidth was used on the 100G during the benchmark?

What is the peak current/watt draw? I’m estimating ~3.8 kW including the switch.

More like 2.3kw

You could almost run that off one 10A circuit!

250k for the moment, but I think I can squeze more before going into dcp 2 or dcp 4..

How large is your KV cache ?

It’s hard not to be in awe. Well done and congratulations!

250k for now. Without dcp.

any concurrent stream testing?

Amazing work and looking very much forward to your git. We will try and run it on our 16x Spark cluster and report back our findings and potential improvements.

Are you planning on adding dcp support? How much wiggle room do you have left for ctx / kvcache?

Yes, dcp too, but dcp 1 with cca 250k will be the fastest

Deepseek just raised their price so I think writing is on the wall. Maybe it is time to mortgage the house, bite the bullet and buy DGX Station.

Good new for you, Gwen 3.8 drop next Wednesday. A95B will be interesting: ModelScope 魔搭社区