New Nemotron 3.5 Lighting 30B-A3B

Tested Nemotron 3.5 Lightning 30B A3B NVFP4 on a single DGX Spark using the official ARM64 vLLM 0.27.1 path, both target-only and with the published DSpark draft model at speculative depth 3.

The deployment path worked successfully, including nemotron_v3 reasoning separation and native qwen3_coder tool calls. On the same deterministic prompt, target-only reached approximately 78.5 output tok/s, while DSpark reached 90.7 tok/s (+15.6%). vLLM reported 53.0% draft-token acceptance and 1.59 accepted tokens per draft in this small sample.

In tool-eval short, target-only scored 77/100 and DSpark 80/100. The recurring misses were multi-value extraction/tool-error recovery, an overly permissive follow-up after refusing an unsupported destructive request, unnecessary calculator use for simple arithmetic, and incomplete acknowledgement of a failed tool call. My current Qwen3.6 35B A3B FP8 reference scored 100/100 on the same short evaluation.

Will do some more testing…

It’s a model with no advanced skills, the goal of this model is to let people finetune on it without having too much hallucinations. It’s normal if it’s bad.

It’s fast AF. One spark, 120+ t/s decode. Recipe here:

recipe_version: "2"
name: nemotron-3.5-lightning-30b-a3b-nvfp4
description: "NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 — DSpark spec decode, fp8 KV, marlin MoE, mamba flashinfer, 1M ctx. From NVIDIA vLLM DGX Spark cookbook."
model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
runtime: vllm
container: vllm/vllm-openai:v0.27.1

metadata:
description: |
NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 (hybrid Mamba-2+MoE+Attention)
with DSpark speculative decoding. Adapted from the official NVIDIA vLLM
DGX Spark recipe (vLLM Nightly v0.27.1). 3B active MoE, ~21GiB weights incl
draft head. TP=1 on GB10.
maintainer: styles01
source: "
"
created: "2026-08-11"
updated: "2026-08-11"
tags:
- nemotron
- nemotron-3.5
- lightning
- 30b
- a3b
- nvfp4
- dspark
- mamba
- moe
- fp8-kv
- marlin
- dgx-spark

solo_only: true
cluster_only: false

min_nodes: 1
max_nodes: 1

defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.91
max_model_len: 1048576
kv_cache_dtype: fp8
speculative_config: '{"method":"dspark","model":"nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark","num_speculative_tokens":4}'
served_model_name: "nemotron-3.5-lightning nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4"

env:
HF_HOME: /cache/huggingface
PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True

executor_config:
entrypoint: ""
auto_remove: false
user: root

benchmark:
framework: llama-benchy
claimed_speed: "~124 tok/s single-stream (MiaAI-Lab DGX Spark SGLang+DSpark reference)"
notes: |
vLLM adaptation of the official NVIDIA DGX Spark cookbook recipe.
Verify tool-calling (qwen3_coder) and reasoning parser (nemotron_v3)
before agent use. sparkrun resolves {model} and the spec draft model
from the HF cache snapshot paths.

command: |
vllm serve {model} 
--served-model-name {served_model_name} 
--host {host} --port {port} 
--trust-remote-code 
--moe-backend marlin 
--kv-cache-dtype {kv_cache_dtype} 
--max-model-len {max_model_len} 
--enable-prefix-caching 
--gpu-memory-utilization {gpu_memory_utilization} 
--speculative-config '{speculative_config}' 
--mamba-backend flashinfer 
--mamba-cache-mode align 
--reasoning-parser nemotron_v3 
--tool-call-parser qwen3_coder 
--enable-auto-tool-choice

^---- this, if it fails on a prompt at least it fails instantly and I don’t have to wait 5+ minutes after looooong pages of blablah to get to the same fail (looking at you Qwen 35B and Gemma 31B)

That’s a good point, it’s very average model like all Nemotrons but has hybrid mamba attention and 1m context, meaning it’s very fast over very large context size. I doubt it can be fine tuned into something very good just because of space attention, hovewer, there probably are many tasks where good enough plus fast speeds make more sense than very good and slow. Not every model has to be a coding one or a literal sage and source of wisdom. It’s a corporate model. Every corporation employs a lot of barely useful idiots, only sensible to replace them with similar level of cognition model for pennies on the dollar

There’s a DSpark version too. Here are some numbers going past as I run benchmarks:

(APIServer pid=1) INFO 08-12 12:32:07 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 351.9 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:32:07 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.84, Accepted throughput: 227.96 tokens/s, Drafted throughput: 371.93 tokens/s, Accepted: 2280 tokens, Drafted: 3720 tokens, Per-position acceptance rate: 0.781, 0.600, 0.457, Avg Draft acceptance rate: 61.3%
(APIServer pid=1) INFO 08-12 12:32:17 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 369.2 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:32:17 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.95, Accepted throughput: 244.18 tokens/s, Drafted throughput: 374.97 tokens/s, Accepted: 2442 tokens, Drafted: 3750 tokens, Per-position acceptance rate: 0.812, 0.636, 0.506, Avg Draft acceptance rate: 65.1%
(APIServer pid=1) INFO 08-12 12:32:27 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 376.4 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:32:27 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.01, Accepted throughput: 251.47 tokens/s, Drafted throughput: 374.95 tokens/s, Accepted: 2515 tokens, Drafted: 3750 tokens, Per-position acceptance rate: 0.817, 0.668, 0.527, Avg Draft acceptance rate: 67.1%
(APIServer pid=1) INFO 08-12 12:32:37 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 362.9 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:32:37 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.93, Accepted throughput: 238.88 tokens/s, Drafted throughput: 371.96 tokens/s, Accepted: 2389 tokens, Drafted: 3720 tokens, Per-position acceptance rate: 0.815, 0.622, 0.490, Avg Draft acceptance rate: 64.2%
(APIServer pid=1) INFO 08-12 12:32:47 [loggers.py:310] Engine 000: Avg prompt throughput: 46.0 tokens/s, Avg generation throughput: 346.4 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:32:47 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.73, Accepted throughput: 219.67 tokens/s, Drafted throughput: 380.04 tokens/s, Accepted: 2197 tokens, Drafted: 3801 tokens, Per-position acceptance rate: 0.761, 0.560, 0.414, Avg Draft acceptance rate: 57.8%
(APIServer pid=1) INFO 08-12 12:32:57 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 350.2 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:32:57 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.76, Accepted throughput: 223.27 tokens/s, Drafted throughput: 380.94 tokens/s, Accepted: 2233 tokens, Drafted: 3810 tokens, Per-position acceptance rate: 0.771, 0.563, 0.424, Avg Draft acceptance rate: 58.6%
(APIServer pid=1) INFO 08-12 12:33:07 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 372.8 tokens/s, Running: 10 reqs, Waiting: 2 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=1) INFO 08-12 12:33:07 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 3.01, Accepted throughput: 248.86 tokens/s, Drafted throughput: 371.94 tokens/s, Accepted: 2489 tokens, Drafted: 3720 tokens, Per-position acceptance rate: 0.831, 0.662, 0.514, Avg Draft acceptance rate: 66.9%

However it’s only 51% through ifevalcode after 4 hours (whereas Gemma 4 did the entire thing in something like an hour), so while it might be fast, I don’t think it’s at all smart. I should’ve tried the bf16 version.