How to use eugr's docker?

I am a new user of dgx spark and currently not familiar with the performance of dgx spark. Through search engines, I learned about the project maintained by eugr in the community, thank you very much.

However, after reading the documentation, I still don’t know how to use it, or maybe I am using it incorrectly. I want to test the model Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF using the following startup command:
``

./launch-cluster.sh --solo exec \
  vllm serve \
    Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF \
    --port 8000 --host 0.0.0.0 \
    --gpu-memory-utilization 0.7 

``

But it returns an error and cannot start. The error is as follows:

``

Loading configuration from .env file...
Loaded .env variables: DOTENV_HF_TOKEN 
Solo mode enabled. Skipping node detection.
Head Node: 127.0.0.1
Worker Nodes: 
Container Name: vllm_node
Image Name: vllm-node
Action: exec
Starting Head Node on 127.0.0.1...
745701357e795a548674784d97a8d3dfa0af9d5934df0367c0cc503e3f068c42
Executing command: vllm serve Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF --port 8000 --host 0.0.0.0 --gpu-memory-utilization 0.7 --load-format fastsafetensors 
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:299] 
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:299]        █     █     █▄   ▄█
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:299]  ▄▄ ▄█ █     █     █ ▀▄▀ █  version 0.19.1rc1.dev36+g9a528260e.d20260405
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:299]   █▄█▀ █     █     █     █  model   Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:299]    ▀▀  ▀▀▀▀▀ ▀▀▀▀▀ ▀     ▀
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:299] 
(APIServer pid=29) INFO 04-07 07:54:22 [utils.py:233] non-default args: {'model_tag': 'Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF', 'host': '0.0.0.0', 'model': 'Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF', 'load_format': 'fastsafetensors', 'gpu_memory_utilization': 0.7}
(APIServer pid=29) WARNING 04-07 07:54:22 [envs.py:1783] Unknown vLLM environment variable detected: VLLM_BASE_DIR
(APIServer pid=29) INFO 04-07 07:54:28 [model.py:554] Resolved architecture: Qwen3_5ForConditionalGeneration
(APIServer pid=29) INFO 04-07 07:54:28 [model.py:1684] Using max model len 262144
(APIServer pid=29) INFO 04-07 07:54:29 [vllm.py:799] Asynchronous scheduling is enabled.
(APIServer pid=29) INFO 04-07 07:54:29 [kernel.py:199] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'])
(APIServer pid=29) Traceback (most recent call last):
(APIServer pid=29)   File "/usr/local/bin/vllm", line 10, in <module>
(APIServer pid=29)     sys.exit(main())
(APIServer pid=29)              ^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/cli/main.py", line 75, in main
(APIServer pid=29)     args.dispatch_function(args)
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/cli/serve.py", line 122, in cmd
(APIServer pid=29)     uvloop.run(run_server(args))
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/uvloop/__init__.py", line 96, in run
(APIServer pid=29)     return __asyncio.run(
(APIServer pid=29)            ^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/lib/python3.12/asyncio/runners.py", line 194, in run
(APIServer pid=29)     return runner.run(main)
(APIServer pid=29)            ^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/lib/python3.12/asyncio/runners.py", line 118, in run
(APIServer pid=29)     return self._loop.run_until_complete(task)
(APIServer pid=29)            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "uvloop/loop.pyx", line 1518, in uvloop.loop.Loop.run_until_complete
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/uvloop/__init__.py", line 48, in wrapper
(APIServer pid=29)     return await main
(APIServer pid=29)            ^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/api_server.py", line 684, in run_server
(APIServer pid=29)     await run_server_worker(listen_address, sock, args, **uvicorn_kwargs)
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/api_server.py", line 698, in run_server_worker
(APIServer pid=29)     async with build_async_engine_client(
(APIServer pid=29)   File "/usr/lib/python3.12/contextlib.py", line 210, in __aenter__
(APIServer pid=29)     return await anext(self.gen)
(APIServer pid=29)            ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/api_server.py", line 100, in build_async_engine_client
(APIServer pid=29)     async with build_async_engine_client_from_engine_args(
(APIServer pid=29)   File "/usr/lib/python3.12/contextlib.py", line 210, in __aenter__
(APIServer pid=29)     return await anext(self.gen)
(APIServer pid=29)            ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/api_server.py", line 136, in build_async_engine_client_from_engine_args
(APIServer pid=29)     async_llm = AsyncLLM.from_vllm_config(
(APIServer pid=29)                 ^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/async_llm.py", line 225, in from_vllm_config
(APIServer pid=29)     return cls(
(APIServer pid=29)            ^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/async_llm.py", line 135, in __init__
(APIServer pid=29)     self.renderer = renderer = renderer_from_config(self.vllm_config)
(APIServer pid=29)                                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/renderers/registry.py", line 83, in renderer_from_config
(APIServer pid=29)     tokenizer = cached_tokenizer_from_config(model_config, **kwargs)
(APIServer pid=29)                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/tokenizers/registry.py", line 227, in cached_tokenizer_from_config
(APIServer pid=29)     return cached_get_tokenizer(
(APIServer pid=29)            ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/tokenizers/registry.py", line 210, in get_tokenizer
(APIServer pid=29)     tokenizer = tokenizer_cls_.from_pretrained(tokenizer_name, *args, **kwargs)
(APIServer pid=29)                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/vllm/tokenizers/hf.py", line 85, in from_pretrained
(APIServer pid=29)     tokenizer = AutoTokenizer.from_pretrained(
(APIServer pid=29)                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/transformers/models/auto/tokenization_auto.py", line 1172, in from_pretrained
(APIServer pid=29)     tokenizer_class_py, tokenizer_class_fast = TOKENIZER_MAPPING[type(config)]
(APIServer pid=29)                                                ~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^
(APIServer pid=29)   File "/usr/local/lib/python3.12/dist-packages/transformers/models/auto/auto_factory.py", line 804, in __getitem__
(APIServer pid=29)     model_type = self._reverse_config_mapping[key.__name__]
(APIServer pid=29)                  ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^
(APIServer pid=29) KeyError: 'Qwen3_5Config'

Stopping cluster...
Stopping head node (127.0.0.1)...
Cluster stopped.

``

I have downloaded the model using ./hf-download.sh Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF and confirmed it exists locally. Where did I go wrong? This should not be a bug, so I did not ask on GitHub.

I wonder if anyone can help me solve this problem.thanks.

I am also new so might not be able to debug what you done wrong..
I use the eugr’s vllm with the sparkrun.

For eugr’s vllm I do the following for every 1~3days.
git pull
./build-and-copy.sh (and maybe with tag if you want to use models with tf5 or something)

For sparkrun
sparkrun update (everytime before I run the models)
I also create recipe repo for my own (it is for testing, if you want to use them, please copy or folk it)
I think you will also need to set ssh key and .cache folder’s permission
otherwise most of the steps are in both sparkrun and eugr’s git repo already (installation)

I’m currently not at home. If there is any problem, ask it and I’ll check it when I am back home.

i use clawcode right now for docker installs like that.

and before clawcode i used claude code with 20€ in API Tokens and that created my basic Qwen 3.5 35b a3b docker that builds me up.

I use an theclawbay subscription and an GLM 5 subscription for fallback situations in openclaw or when i need to close the docker for VRAM reasons and tests for bigger models :)

today they install me the docker from this thread dgx-spark-gb10-vllm-0-19-1-turboquant-kv-cache :) it should work fine with clawcode

vLLM isn’t really made for GGUF. You should switch to llama.cpp for best GGUF support.

Please note that GGUF support in vLLM is highly experimental and under-optimized at the moment, it might be incompatible with other features. Currently, you can use GGUF as a way to reduce memory footprint. If you encounter any issues, please report them to the vLLM team.

As that finetune is based on Qwen3.5 the instructions/recommendations for llama.cpp in here:

should apply, too.

I’m excited to see someone using their own registry!

Welcome to the community @codecrh !

I work with @eugr and am the maintainer for sparkrun.

If you’re just getting started, sparkrun is the easiest way to start.

# Install uv (you probably want that anyway...)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install sparkrun and set up your cluster in one step
# will guide you into setup wizard to do all the necessary configuration for your spark
uvx sparkrun setup

And then:
sparkrun run @sparkrun-testing/jackrong-qwen3.5-27b-claude4.6-distill-vllm

to run the model.

You could also install @kenny8379’s recipe registry to your sparkrun:

  1. sparkrun registry add https://github.com/NashiKanjou/nashi-sparkrun-registry.git
  2. sparkrun update

and then use his recipe:
sparkrun run @nashi-sparkrun/Jackrong-Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2


You can then run benchmarks and publish your benchmarks to Spark Arena. You can see what other people are running and what kind of performance they’re getting.

The Spark Arena Team (@eugr, @raphael.amorim and myself) are also working to publish more official recipes to help you be able to run more models, more easily!

More docs for sparkrun at: https://sparkrun.dev

Or just ask on the forums.

What’s the best way we can help with sparkrun? I’ve been getting a bunch of models running by creating custom spark-vllm-docker recipes/mods. Does it help to get these into sparkrun community recipes?

I think that might help

GitHub - spark-arena/community-recipe-registry · GitHub . Submit a PR to this repo (the README gives a rough guide). That repo is installed in sparkrun by default under @community, so if you provide a recipe e.g. in “recipes/qwen3-coder-next/zambonilli/qwen3-coder-next-bf16-vllm-zambonilli.yaml” in that repo, then people would be able to run it as
sparkrun run @community/qwen3-coder-next-bf16-vllm-zambonilli

It’s all still a bit rough but sparkrun flattens the recipe directories, so I think structure of:
recipes/{model}/{user}/{model}-{quant}-{runtime}-{user}.yaml

will balance making it that we can find things by tab completion and search at CLI plus making it sane when looking at the git repo – I think it’s overwhelming if there are 5000 files in one directory.

Nobody is using community recipe repo yet because honestly… haven’t been telling people about it – not sure what I was waiting for!


I think the other part that would help is more on us – to centralize our “standard” recipes so that they’re always available. I had been intending to be more proactive in recipe publishing and management, but I’ve spent a lot more time than intended on making sparkrun work for everyone across spectrum of different configurations and user experience levels – and that’s grown in complexity despite it being a relatively simple bit of code.


Another thing that I think would be good is if community/registry published recipes are then benchmarked on Spark Arena. And we’ll try to fix the cross-linking there – so that if you publish a community recipe and then benchmark it, the usage link should direct you to run it as:

sparkrun run @community/your-awesome-recipe-zambonilli

We’ve added direct integration such that running benchmarks can be done via sparkrun:
sparkrun arena benchmark <recipe>. See Benchmark on Spark Arena | sparkrun.

You’ll need to go to https://spark-arena.com/admin and request access – and then once you have access, you can do sparkrun arena login and link sparkrun. Then you can run benchmarks and send them to Spark Arena.


Sorry that was long winded, but I think it would be cool to get help and get more people involved. Even just getting bug reports of issues is great because I have been on a journey to make sure that it works for everyone.

Thank you very much, I will use the GitHub link you shared to test SparkRun.

Thank you for your patient answers. I will try using SparkRun to install vLLM and run it. Thank you very much for creating such a great tool.For me, this is a very good introductory tool, and I hope it can help me get started with Spark.

Thank you very much. Before this, I didn’t know about vLLM, and I didn’t know that it has poor support for GGUF. I will use other formats for testing. This is very helpful to me.