I’d like to share my success story with the community.
Model: LibertAIDAI/GLM-5.3-Flash-NVFP4
What is running
glm-5.3-flash via Bifrost :4000 → spark01:8045, TP=3 on 01/02/03, MTP-4, context 524.288, KV-pool 1.505.849 tok (2.87×), weights 63.79 GiB/rang, GPU clock 1500 MHz.
┌─────────────────────┬────────────────────────┐
│ Metric │ Value │
├─────────────────────┼────────────────────────┤
│ Decode │ 35.2 tok/s (32.1–39.0) │
├─────────────────────┼────────────────────────┤
│ TTFT │ 0.26 s │
├─────────────────────┼────────────────────────┤
│ Prefill │ ~1 800 tok/s │
├─────────────────────┼────────────────────────┤
│ Retrive 471.813 tok │ PASS, 274.6 s │
├─────────────────────┼────────────────────────┤
│ Xid │ 0 / 0 / 0 │
└─────────────────────┴────────────────────────┘
GPU cap test
┌──────────┬────────────┬─────────────┐
│ │ Decode │ Pfefill 64K │
├──────────┼────────────┼─────────────┤
│ 1500 MHz │ 35.2 tok/s │ 44.5 s │
├──────────┼────────────┼─────────────┤
│ 2100 MHz │ 36.8 tok/s │ 35.3 s │
├──────────┼────────────┼─────────────┤
│ │ +4.4% │ −26% │
└──────────┴────────────┴─────────────┘
Tools
tool-eval-bench --base-url http://spark01:8045 --hardmode --parallel 1 --seed 42
Note: default reasoning mode: max
│ Model: /models/glm-5.3-flash-nvfp4
│ Score: 90 / 100
│ Rating: ★★★★★ Excellent
│ Benchmark: tool-eval-bench v2.6.1.dev24+g845f15e6c
│ Engine: vLLM 0.1.dev20051+g487ecf187
│ Max context: 524,288 tokens
│
│ ✅ 74 passed ⚠️ 11 partial ❌ 3 failed
│ Points: 159/176
│
│ Quality: 90/100
│ Responsiveness: 31/100 (median turn: 5.2s)
│ Deployability: 72/100 (α=0.7)
│ Weakest: I Context & State (80%)
│
│ Completed in 1883.6s
│
│ 📊 Token Usage:
│ Total: 552,143 tokens │ Efficiency: 0.3 pts/1K tokens
│
│ 🛡️ SAFETY WARNINGS (1):
│ ⚠ TC-43 (Omitted Required Parameter): Called web_search with an empty query — violated required parameter constraint.
Partials
Most are earned (TC-63 never searched for a match, TC-57 answered without searching, TC-62 missing corrected revenue). But TC-28, TC-35, TC-47, TC-75 share one shape: “made an extra call” / “volunteered an unrequested conversion” / “asked for details but also guessed a date”. Those are over-helpfulness penalties, not wrong answers.
Bottom line
Of 17 lost points: 2 taken incorrectly (TC-68 should be 1/2), 2 taken pedantically (TC-43), the rest earned. True score is ~91–92, and the safety-critical warning should be ignored — it is category automation, not a risk assessment.
Performance
tool-eval-bench bench --perf-only \
--base-url http://spark01:8045/v1 --model glm-5.3-flash --backend vllm \
--tokenizer ~/.cache/glm53-tokenizer \
--pp 2048 --tg 512 \
--depth 6144,63488,161792 \
--concurrency 1,2,3 \
--benchy-runs 3 --benchy-latency-mode generation \
--timeout 1800 --no-live \
--label glm53-tp3-perf-1500mhz \
--output-dir runs/ --json-file runs/glm53-perf.json
╭──────────────────────────────────────────────── ⚡ llama-benchy Throughput Benchmark ────────────────────────────────────────────────╮
│ glm-5.3-flash │
│ pp=[2048] tg=[512] depth=[6144, 63488, 161792] concurrency=[1, 2, 3] runs=3 latency=generation │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
✓ Complete ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 27/27 0:58:29
llama-benchy 0.4.0
Estimated latency: 150.9 ms
llama-benchy Results
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┓
┃ Test ┃ c ┃ pp t/s ┃ tg t/s ┃ TTFT (ms) ┃ Total (ms) ┃ Tokens ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━┩
│ pp2048 tg512 @ d6144 │ c1 │ 1,591 │ 31.6 │ 5,309 │ 21,358 │ 2048+512 │
│ pp2048 tg512 @ d6144 │ c2 │ 1,522 │ 36.1 │ 8,068 │ 32,652 │ 2048+512 │
│ pp2048 tg512 @ d6144 │ c3 │ 1,516 │ 40.5 │ 10,796 │ 40,651 │ 2048+512 │
│ pp2048 tg512 @ d63488 │ c1 │ 1,509 │ 30.8 │ 43,578 │ 60,061 │ 2048+512 │
│ pp2048 tg512 @ d63488 │ c2 │ 1,488 │ 15.3 │ 65,824 │ 97,797 │ 2048+512 │
│ pp2048 tg512 @ d63488 │ c3 │ 1,482 │ 13.3 │ 88,085 │ 135,093 │ 2048+512 │
│ pp2048 tg512 @ d161792 │ c1 │ 1,470 │ 34.2 │ 111,631 │ 126,439 │ 2048+512 │
│ pp2048 tg512 @ d161792 │ c2 │ 1,447 │ 7.5 │ 169,173 │ 201,702 │ 2048+512 │
│ pp2048 tg512 @ d161792 │ c3 │ 1,444 │ 6.1 │ 225,916 │ 282,337 │ 2048+512 │
└───────────────────────────────────┴─────────┴────────────────┴────────────────┴──────────────────┴─────────────────┴─────────────────┘
Conslusion: this is the first NVFP4 quant I’m happy with :D
A couple of screenshot from the cluster whilePi is running.
And the Pelican of course:



