I have a 2x DGX Spark with Connect X7 + a couple of Intel Mini PC with 96GB ram. Is there a way to have a cluster them together to maximize both performance and the ability to run larger models ?
you could use llama.cpp but offloading on the mini pc’s will be super super slow like 0.5 t/s slow if not slower .. so theoretically yes - practically dont waste your time - network them together as kube cluster and pin non gpu workloads on the mini pc’s and gpu workload on the sparks