Has anyone managed to run this model on the Spark? nvidia/gpt-oss-puzzle-88B · Hugging Face
I did get it running with a little help from codex, but the output was pretty shoddy, so something definitely needs fixing.
paste of codex AI slop about what was done below:
nvidia/gpt-oss-puzzle-88B did not work:
- the generic nightly image I started with
- the repo’s existing mxfp4 GPT-OSS image
In both cases, vLLM fell back to the generic Transformers backend and failed with:
Transformers modeling backend does not support MXFP4 quantization yet.
What fixed it was following the model readme much more literally.
- Build a custom image from vllm/vllm-openai:v0.17.1.
- Inside that image, install the vLLM PR from the readme:
VLLM_USE_PRECOMPILED=1 pip install --no-build-isolation
‘git+https://github.com/vllm-project/vllm.git@refs/pull/38135/head’
- Install the matching FlashInfer packages:
pip install flashinfer-cubin==0.6.6 flashinfer-jit-cache==0.6.6
–extra-index-url Index of cu129
- Important extra fix: openai_harmony crashed during API init with:
error downloading or loading vocab file
The fix was to pre-download the tiktoken files and set:
mkdir -p /workspace/tiktoken_encodings
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken
-o /workspace/tiktoken_encodings/o200k_base.tiktoken
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken
-o /workspace/tiktoken_encodings/cl100k_base.tiktoken
export TIKTOKEN_ENCODINGS_BASE=/workspace/tiktoken_encodings
- I also had to lower --gpu-memory-utilization from 0.95 to 0.94, because 0.95 failed at startup on my machine due to slightly less free VRAM than requested.
Working launch command:
docker run --rm -d --name puzzle_manual
–gpus all --network host --ipc=host
-v ~/.cache/huggingface:/root/.cache/huggingface
-e PYTORCH_ALLOC_CONF=expandable_segments:True
vllm-node-gpt-oss-puzzle-readme
bash -lc ’
set -euo pipefail
mkdir -p /workspace/tiktoken_encodings
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken
-o /workspace/tiktoken_encodings/o200k_base.tiktoken
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken
-o /workspace/tiktoken_encodings/cl100k_base.tiktoken
export TIKTOKEN_ENCODINGS_BASE=/workspace/tiktoken_encodings
vllm serve nvidia/gpt-oss-puzzle-88B
-tp 1
–trust-remote-code
–kv-cache-dtype fp8
–max-num-batched-tokens 8192
–stream-interval 20
–gpu-memory-utilization 0.94
–max-num-seqs 8
–max-cudagraph-capture-size 8
–max-model-len 131072
–host 0.0.0.0
–port 8000
’
The key sign that it was finally on the right path was this log line:
Resolved architecture: GptOssPuzzleForCausalLM
After that, /v1/models responded correctly and inference requests succeeded.
Thank you so much, it worked for me
# Recipe: NVIDIA GPT-OSS Puzzle 88B
# Deployment-optimized MoE model derived from OpenAI gpt-oss-120b via NVIDIA Puzzle NAS
# Reference: https://forums.developer.nvidia.com/t/nvidia-gpt-oss-puzzle-88b/365006/2
# Requires vLLM PR #38135 for GptOssPuzzleForCausalLM architecture support
recipe_version: "1"
name: NVIDIA GPT-OSS Puzzle 88B
description: vLLM serving nvidia/gpt-oss-puzzle-88B with MXFP4 quantization and FP8 KV cache
# HuggingFace model to download (optional, for --download-model)
model: nvidia/gpt-oss-puzzle-88B
# Container image to use (dedicated build with PR #38135 applied)
# Build: ./build-and-copy.sh --apply-vllm-pr 38135 -t vllm-node-38135
container: vllm-node-38135
# Build arguments: apply PR #38135 for puzzle model support
build_args:
- --apply-vllm-pr
- "38135"
# No runtime mods needed - PR is applied at build time
mods: []
# Default settings (can be overridden via CLI)
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 2
gpu_memory_utilization: 0.78
max_model_len: 131072
max_num_batched_tokens: 8192
# Environment variables to set in the container
env:
PYTORCH_CUDA_ALLOC_CONF: "expandable_segments:True"
# The vLLM serve command template
command: |
vllm serve nvidia/gpt-oss-puzzle-88B \
--trust-remote-code \
--kv-cache-dtype fp8 \
--max-num-batched-tokens {max_num_batched_tokens} \
--max-model-len {max_model_len} \
--stream-interval 20 \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--enable-prefix-caching \
--load-format fastsafetensors \
--host {host} \
--port {port} \
--distributed-executor-backend ray
uvx llama-benchy --base-url http://127.0.0.1:8000/v1 --model nvidia/gpt-oss-puzzle-88B --tokenizer /home/csolutions_ai/.cache/llama-benchy/tokenizers/gpt-oss-puzzle-88B --pp 2048 --depth 4096 16000 32000
PyTorch was not found. Models won't be available and only tokenizers, configuration and file/data utilities can be used.
llama-benchy (0.3.5)
Date: 2026-03-29 23:42:49
Benchmarking model: nvidia/gpt-oss-puzzle-88B at http://127.0.0.1:8000/v1
Concurrency levels: [1]
Loading text from cache: /home/csolutions_ai/.cache/llama-benchy/cc6a0b5782734ee3b9069aa3b64cc62c.txt
Total tokens available in text corpus: 140865
Warming up...
Warmup (User only) complete. Delta: 67 tokens (Server: 88, Local: 21)
Warmup (System+Empty) complete. Delta: 74 tokens (Server: 95, Local: 21)
Running coherence test...
Coherence test PASSED.
Measuring latency using mode: api...
Average latency (api): 2.40 ms
Running test: pp=2048, tg=32, depth=4096, concurrency=1
Run 1/3 (batch size 1)...
Run 2/3 (batch size 1)...
Run 3/3 (batch size 1)...
Running test: pp=2048, tg=32, depth=16000, concurrency=1
Run 1/3 (batch size 1)...
Run 2/3 (batch size 1)...
Run 3/3 (batch size 1)...
Running test: pp=2048, tg=32, depth=32000, concurrency=1
Run 1/3 (batch size 1)...
Run 2/3 (batch size 1)...
Run 3/3 (batch size 1)...
Printing results in MD format:
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) |
|:--------------------------|----------------:|----------------:|-------------:|----------------:|----------------:|-----------------:|
| nvidia/gpt-oss-puzzle-88B | pp2048 @ d4096 | 4689.74 ± 19.61 | | 1312.73 ± 5.49 | 1310.33 ± 5.49 | 1400.45 ± 5.36 |
| nvidia/gpt-oss-puzzle-88B | tg32 @ d4096 | 33.96 ± 0.02 | 35.18 ± 0.02 | | | |
| nvidia/gpt-oss-puzzle-88B | pp2048 @ d16000 | 3602.26 ± 9.43 | | 5012.91 ± 13.14 | 5010.51 ± 13.14 | 5097.72 ± 13.35 |
| nvidia/gpt-oss-puzzle-88B | tg32 @ d16000 | 33.17 ± 0.06 | 34.36 ± 0.06 | | | |
| nvidia/gpt-oss-puzzle-88B | pp2048 @ d32000 | 2904.72 ± 2.47 | | 11724.26 ± 9.99 | 11721.86 ± 9.99 | 11806.65 ± 11.50 |
| nvidia/gpt-oss-puzzle-88B | tg32 @ d32000 | 32.22 ± 0.02 | 33.37 ± 0.02 | | | |
llama-benchy (0.3.5)
date: 2026-03-29 23:42:49 | latency mode: api
csolutions_ai@thinkmax3:~/spark-vllm-docker$
Try this with this in the env:
VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8=1
VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8=1 I tried it, it doesn’t work, problems with flashinter