# How do I run Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled on vllm community docker?

**URL:** <https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292>\
**Category:** DGX Spark / GB10\
**Tags:** llama\
**Created:** [March 12, 2026, 1:29pm UTC](https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292 "2026-03-12T13:29:58Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![tatamiso](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@tatamiso](https://forums.developer.nvidia.com/u/tatamiso)\
**Post date:** [March 12, 2026, 1:29pm UTC](https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292/1 "2026-03-12T13:29:58Z")

</div>

As the title,

how do I run this model using eugr’s community docker?

> **[Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled · Hugging Face](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled)**
>
> We’re on a journey to advance and democratize artificial intelligence through open source and open science.

I haven’t seen a recipe for this model version yet, only MoE ones.

I tried with llama.cpp + GGUF and the results are pretty good.  
Q4\_K\_M ~12tg/s  
Q8\_0 ~8tg/s

Would it be faster with vllm?

---

<div class="post-metadata">

**Author:** ![tim3233](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@tim3233](https://forums.developer.nvidia.com/u/tim3233)\
**Post date:** [March 12, 2026, 2:05pm UTC](https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292/2 "2026-03-12T14:05:14Z")

</div>

Hi,

would like to know this also. vLLM doenst work with Transformers 5.2. I think thats the point.

BG

---

<div class="post-metadata">

**Author:** ![jwarner](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@jwarner](https://forums.developer.nvidia.com/u/jwarner)\
**Post date:** [March 12, 2026, 5:37pm UTC](https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292/3 "2026-03-12T17:37:07Z")

</div>

Build that image with TF5 and adapt recipes for 35b-a3b or 122b-a10b which set up the correct templates, tool calling, and tweaks for Qwen3.5.

You will likely not see large gains in raw performance. From my personal quants of this model and Gemma3 27b dense models, the best single query rate is about 12 tok/s decode.

Assuming you have a Spark, you can make an int4 AutoRound quant of that model in around 4 hours. See a separate thread on how.

Note that I believe that quant dropped multimodal capabilities.

---

<div class="post-metadata">

**Author:** ![tatamiso](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@tatamiso](https://forums.developer.nvidia.com/u/tatamiso)\
**Post date:** [March 13, 2026, 3:55pm UTC](https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292/4 "2026-03-13T15:55:56Z")

</div>

Thank you, will try out this one and update here.

Also gotta look into that autoround conversion, sounds interesting!

---

<div class="post-metadata">

**Author:** ![dbsci](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/dbsci/32/467918_2.png) [@dbsci](https://forums.developer.nvidia.com/u/dbsci)\
**Post date:** [March 13, 2026, 10:17pm UTC](https://forums.developer.nvidia.com/t/how-do-i-run-jackrong-qwen3-5-27b-claude-4-6-opus-reasoning-distilled-on-vllm-community-docker/363292/5 "2026-03-13T22:17:18Z")

</div>

Performance-wise you’re probably better off doing a quant, etc.; however, to answer your original question, here is a recipe for use with sparkrun that builds on top of @eugr’s vllm docker repo.

`sparkrun run @sparkrun-testing/jackrong-qwen3.5-27b-claude4.6-distill-vllm`

The `@sparkrun-testing/` prefix is required for “hidden” registries. I do that so that I can deploy recipes for particular use without them being part of the default tab completion, etc.

You can check out the recipe file at: [sparkrun-recipe-registry/testing/recipes/qwen3.5/exotic/jackrong-qwen3.5-27b-claude4.6-distill-vllm.yaml at main · dbotwinick/sparkrun-recipe-registry · GitHub](https://github.com/dbotwinick/sparkrun-recipe-registry/blob/main/testing/recipes/qwen3.5/exotic/jackrong-qwen3.5-27b-claude4.6-distill-vllm.yaml)

I tested that it ran and I was seeing ~4.5 tok/s, so not terribly impressive on performance with single node tensor parallel, but it’s interesting to see this new wave of opus distillation models! (Note that 27B dense model at BF16 would have a theoretical peak throughput of ~5.1 tok/s on a single spark Spark).

You can learn more about how to install sparkrun in the forums at: [Sparkrun - central command with tab completion for launching inference on Spark Clusters](https://forums.developer.nvidia.com/t/sparkrun-central-command-with-tab-completion-for-launching-inference-on-spark-clusters/360832) or check out the docs at [https://sparkrun.dev](https://sparkrun.dev). sparkrun is designed to make it easier to run models and we’re working to make it easier to find recipes and understand baseline performance at [spark-arena.com](http://spark-arena.com).
