nvidia/gpt-oss-puzzle-88B

Has anyone managed to run this model on the Spark? nvidia/gpt-oss-puzzle-88B · Hugging Face

I did get it running with a little help from codex, but the output was pretty shoddy, so something definitely needs fixing.
paste of codex AI slop about what was done below:
nvidia/gpt-oss-puzzle-88B did not work:

  • the generic nightly image I started with
  • the repo’s existing mxfp4 GPT-OSS image

In both cases, vLLM fell back to the generic Transformers backend and failed with:
Transformers modeling backend does not support MXFP4 quantization yet.

What fixed it was following the model readme much more literally.

  1. Build a custom image from vllm/vllm-openai:v0.17.1.
  2. Inside that image, install the vLLM PR from the readme:

VLLM_USE_PRECOMPILED=1 pip install --no-build-isolation
‘git+https://github.com/vllm-project/vllm.git@refs/pull/38135/head’

  1. Install the matching FlashInfer packages:

pip install flashinfer-cubin==0.6.6 flashinfer-jit-cache==0.6.6
–extra-index-url Index of cu129

  1. Important extra fix: openai_harmony crashed during API init with:
    error downloading or loading vocab file
    The fix was to pre-download the tiktoken files and set:

mkdir -p /workspace/tiktoken_encodings
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken
-o /workspace/tiktoken_encodings/o200k_base.tiktoken
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken
-o /workspace/tiktoken_encodings/cl100k_base.tiktoken
export TIKTOKEN_ENCODINGS_BASE=/workspace/tiktoken_encodings

  1. I also had to lower --gpu-memory-utilization from 0.95 to 0.94, because 0.95 failed at startup on my machine due to slightly less free VRAM than requested.

Working launch command:

docker run --rm -d --name puzzle_manual
–gpus all --network host --ipc=host
-v ~/.cache/huggingface:/root/.cache/huggingface
-e PYTORCH_ALLOC_CONF=expandable_segments:True
vllm-node-gpt-oss-puzzle-readme
bash -lc ’
set -euo pipefail
mkdir -p /workspace/tiktoken_encodings
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken
-o /workspace/tiktoken_encodings/o200k_base.tiktoken
curl -fsSL https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken
-o /workspace/tiktoken_encodings/cl100k_base.tiktoken
export TIKTOKEN_ENCODINGS_BASE=/workspace/tiktoken_encodings
vllm serve nvidia/gpt-oss-puzzle-88B
-tp 1
–trust-remote-code
–kv-cache-dtype fp8
–max-num-batched-tokens 8192
–stream-interval 20
–gpu-memory-utilization 0.94
–max-num-seqs 8
–max-cudagraph-capture-size 8
–max-model-len 131072
–host 0.0.0.0
–port 8000
’

The key sign that it was finally on the right path was this log line:
Resolved architecture: GptOssPuzzleForCausalLM

After that, /v1/models responded correctly and inference requests succeeded.

Thank you so much, it worked for me

# Recipe: NVIDIA GPT-OSS Puzzle 88B
# Deployment-optimized MoE model derived from OpenAI gpt-oss-120b via NVIDIA Puzzle NAS
# Reference: https://forums.developer.nvidia.com/t/nvidia-gpt-oss-puzzle-88b/365006/2
# Requires vLLM PR #38135 for GptOssPuzzleForCausalLM architecture support

recipe_version: "1"
name: NVIDIA GPT-OSS Puzzle 88B
description: vLLM serving nvidia/gpt-oss-puzzle-88B with MXFP4 quantization and FP8 KV cache

# HuggingFace model to download (optional, for --download-model)
model: nvidia/gpt-oss-puzzle-88B

# Container image to use (dedicated build with PR #38135 applied)
# Build: ./build-and-copy.sh --apply-vllm-pr 38135 -t vllm-node-38135
container: vllm-node-38135

# Build arguments: apply PR #38135 for puzzle model support
build_args:
  - --apply-vllm-pr
  - "38135"

# No runtime mods needed - PR is applied at build time
mods: []

# Default settings (can be overridden via CLI)
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.78
  max_model_len: 131072
  max_num_batched_tokens: 8192

# Environment variables to set in the container
env:
  PYTORCH_CUDA_ALLOC_CONF: "expandable_segments:True"

# The vLLM serve command template
command: |
  vllm serve nvidia/gpt-oss-puzzle-88B \
      --trust-remote-code \
      --kv-cache-dtype fp8 \
      --max-num-batched-tokens {max_num_batched_tokens} \
      --max-model-len {max_model_len} \
      --stream-interval 20 \
      --tensor-parallel-size {tensor_parallel} \
      --gpu-memory-utilization {gpu_memory_utilization} \
      --enable-prefix-caching \
      --load-format fastsafetensors \
      --host {host} \
      --port {port} \
      --distributed-executor-backend ray

uvx llama-benchy --base-url http://127.0.0.1:8000/v1 --model nvidia/gpt-oss-puzzle-88B --tokenizer /home/csolutions_ai/.cache/llama-benchy/tokenizers/gpt-oss-puzzle-88B --pp 2048 --depth 4096 16000 32000
PyTorch was not found. Models won't be available and only tokenizers, configuration and file/data utilities can be used.
llama-benchy (0.3.5)
Date: 2026-03-29 23:42:49
Benchmarking model: nvidia/gpt-oss-puzzle-88B at http://127.0.0.1:8000/v1
Concurrency levels: [1]
Loading text from cache: /home/csolutions_ai/.cache/llama-benchy/cc6a0b5782734ee3b9069aa3b64cc62c.txt
Total tokens available in text corpus: 140865
Warming up...
Warmup (User only) complete. Delta: 67 tokens (Server: 88, Local: 21)
Warmup (System+Empty) complete. Delta: 74 tokens (Server: 95, Local: 21)

Running coherence test...
Coherence test PASSED.
Measuring latency using mode: api...
Average latency (api): 2.40 ms
Running test: pp=2048, tg=32, depth=4096, concurrency=1
  Run 1/3 (batch size 1)...
  Run 2/3 (batch size 1)...
  Run 3/3 (batch size 1)...
Running test: pp=2048, tg=32, depth=16000, concurrency=1
  Run 1/3 (batch size 1)...
  Run 2/3 (batch size 1)...
  Run 3/3 (batch size 1)...
Running test: pp=2048, tg=32, depth=32000, concurrency=1
  Run 1/3 (batch size 1)...
  Run 2/3 (batch size 1)...
  Run 3/3 (batch size 1)...
Printing results in MD format:



| model                     |            test |             t/s |     peak t/s |       ttfr (ms) |    est_ppt (ms) |    e2e_ttft (ms) |
|:--------------------------|----------------:|----------------:|-------------:|----------------:|----------------:|-----------------:|
| nvidia/gpt-oss-puzzle-88B |  pp2048 @ d4096 | 4689.74 ± 19.61 |              |  1312.73 ± 5.49 |  1310.33 ± 5.49 |   1400.45 ± 5.36 |
| nvidia/gpt-oss-puzzle-88B |    tg32 @ d4096 |    33.96 ± 0.02 | 35.18 ± 0.02 |                 |                 |                  |
| nvidia/gpt-oss-puzzle-88B | pp2048 @ d16000 |  3602.26 ± 9.43 |              | 5012.91 ± 13.14 | 5010.51 ± 13.14 |  5097.72 ± 13.35 |
| nvidia/gpt-oss-puzzle-88B |   tg32 @ d16000 |    33.17 ± 0.06 | 34.36 ± 0.06 |                 |                 |                  |
| nvidia/gpt-oss-puzzle-88B | pp2048 @ d32000 |  2904.72 ± 2.47 |              | 11724.26 ± 9.99 | 11721.86 ± 9.99 | 11806.65 ± 11.50 |
| nvidia/gpt-oss-puzzle-88B |   tg32 @ d32000 |    32.22 ± 0.02 | 33.37 ± 0.02 |                 |                 |                  |

llama-benchy (0.3.5)
date: 2026-03-29 23:42:49 | latency mode: api
csolutions_ai@thinkmax3:~/spark-vllm-docker$ 

Try this with this in the env:

VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8=1

VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8=1 I tried it, it doesn’t work, problems with flashinter