Feasibility of 2-3x Speedup via Speculative Decoding on High-Compute (1000 TFLOPS) / Low-BW Hardware

Lab to support speculative decoding on the spark

Thread with lots of performance numbers for inference in single and cluster mode: