Run VLLM in Spark

Not sure if it would be helpful for you, but you could give this a shot for improving performance:

https://forums.developer.nvidia.com/t/vllm-on-gb10-gpt-oss-120b-mxfp4-slower-than-sglang-llama-cpp-what-s-missing/