Build that image with TF5 and adapt recipes for 35b-a3b or 122b-a10b which set up the correct templates, tool calling, and tweaks for Qwen3.5.
You will likely not see large gains in raw performance. From my personal quants of this model and Gemma3 27b dense models, the best single query rate is about 12 tok/s decode.
Assuming you have a Spark, you can make an int4 AutoRound quant of that model in around 4 hours. See a separate thread on how.
Note that I believe that quant dropped multimodal capabilities.
Performance-wise you’re probably better off doing a quant, etc.; however, to answer your original question, here is a recipe for use with sparkrun that builds on top of @eugr’s vllm docker repo.
sparkrun run @sparkrun-testing/jackrong-qwen3.5-27b-claude4.6-distill-vllm
The @sparkrun-testing/ prefix is required for “hidden” registries. I do that so that I can deploy recipes for particular use without them being part of the default tab completion, etc.
I tested that it ran and I was seeing ~4.5 tok/s, so not terribly impressive on performance with single node tensor parallel, but it’s interesting to see this new wave of opus distillation models! (Note that 27B dense model at BF16 would have a theoretical peak throughput of ~5.1 tok/s on a single spark Spark).