They should sell an decode companion to the spark to cover that one pain point. Preferably affordable.
A high performance decoder will work on disaggregated inference, and either have extremely high memory bandwidth, large memory pool and large and fast storage or be based on Language Processing Units (LPUs) which are more expensive than most GPUs. We’re also going to need thunderbolt + GPUDirect RDMA or equivalent for KV Cache Transfer and scale-out support for decode. Unified memory is primarily a proprietary feature today with Apple, AMD and NVIDIA, HBM is super overpriced due to high demand now. So, I think the minimal price you would see for such hardware would be around $2500-3000 per unit at least. So, around the same value you’re going to have in the future for a 128 GB M5 mac mini most likely. So, an approach might be using something like EXO with a high performance mac mini or mac studio