Let's optimize nvidia/GLM-5.3-Flash-NVFP4 for 2× DGX Spark

Let’s optimize nvidia/GLM-5.3-Flash-NVFP4 for 2× DGX Spark.

nvidia/GLM-5.3-Flash-NVFP4

NVIDIA has quantized GLM-5.3-Flash and released it as an NVFP4 model.

Would it be possible to serve it with a 1M-token context on a dual-DGX Spark setup?

These days, it’s great to see NVIDIA directly handling NVFP4 quantization themselves.

If there are any ways to achieve the best possible performance while minimizing quality degradation, I’d really appreciate it if you could let me know. Links to relevant information or resources would also be very helpful.

=========Test Result=========(2026-09-12 upload)

--gpu-memory-utilization 0.9 

==> (EngineCore pid=279) INFO 09-12 07:07:14 [kv_cache_utils.py:2312] GPU KV cache size: 743,165 tokens, Maximum concurrency for 262,144 tokens per request: 2.83x

MTP = 5

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1204.73 ± 64.41 |              |   1423.06 ± 70.74 |   1418.73 ± 70.74 |   1423.06 ± 70.74 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     26.81 ± 0.87 | 35.67 ± 2.05 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1314.93 ± 39.89 |              |    1305.57 ± 6.50 |    1301.23 ± 6.50 |    1305.57 ± 6.50 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     23.68 ± 3.45 | 45.00 ± 2.83 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 | 1534.44 ± 295.32 |              |  1880.93 ± 355.18 |  1876.59 ± 355.18 |  1880.93 ± 355.18 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     22.42 ± 0.33 | 30.67 ± 1.70 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  1542.21 ± 36.05 |              |   1676.95 ± 26.13 |   1672.61 ± 26.13 |   1676.95 ± 26.13 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     22.99 ± 4.45 | 41.67 ± 8.26 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1555.07 ± 269.50 |              |  3396.51 ± 716.78 |  3392.18 ± 716.78 |  3396.51 ± 716.78 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     24.34 ± 2.08 | 35.00 ± 3.74 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1208.21 ± 254.17 |              | 4569.23 ± 1025.35 | 4564.89 ± 1025.35 | 4569.23 ± 1025.35 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     23.51 ± 1.70 | 42.33 ± 3.68 |                   |                   |                   |

llama-benchy (0.3.9.dev9+g446dd42fd)

Dflash speculative_tokens = 5

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  921.46 ± 118.42 |              |  1941.02 ± 319.34 |  1937.29 ± 319.34 |  1941.02 ± 319.34 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     35.74 ± 2.01 | 48.33 ± 2.49 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1058.97 ± 59.39 |              |   1631.32 ± 89.63 |   1627.59 ± 89.63 |   1631.32 ± 89.63 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     28.00 ± 3.02 | 57.67 ± 3.30 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |    320.66 ± 7.40 |              |  8130.20 ± 165.87 |  8126.48 ± 165.87 |  8130.20 ± 165.87 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     33.79 ± 3.77 | 43.00 ± 4.55 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  855.27 ± 381.10 |              | 4177.52 ± 2548.24 | 4173.80 ± 2548.24 | 4177.52 ± 2548.24 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     30.59 ± 4.61 | 53.67 ± 5.31 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 |  1429.45 ± 17.64 |              |   3609.23 ± 38.00 |   3605.51 ± 38.00 |   3632.55 ± 63.88 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     33.39 ± 2.06 | 40.00 ± 3.56 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1457.16 ± 249.48 |              |  3792.10 ± 287.95 |  3788.38 ± 287.95 |  3792.10 ± 287.95 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     30.53 ± 1.00 | 53.67 ± 5.44 |                   |                   |                   |

llama-benchy (0.3.9.dev9+g446dd42fd)

Dflash speculative_tokens = 7 (official set)

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  877.48 ± 122.55 |              |  2014.67 ± 285.73 |  2008.84 ± 285.73 |  2014.67 ± 285.73 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     39.11 ± 2.19 | 46.33 ± 4.11 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1064.08 ± 21.37 |              |   1660.66 ± 23.22 |   1654.83 ± 23.22 |   1660.66 ± 23.22 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     30.77 ± 1.23 | 57.33 ± 4.64 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  765.33 ± 332.39 |              | 4511.10 ± 2621.62 | 4505.27 ± 2621.62 | 4511.10 ± 2621.62 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     29.07 ± 2.08 | 41.67 ± 3.86 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  580.78 ± 355.86 |              | 6630.03 ± 3049.41 | 6624.20 ± 3049.41 | 6630.03 ± 3049.41 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     31.04 ± 3.02 | 64.33 ± 2.62 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1334.81 ± 123.93 |              |  3942.27 ± 457.97 |  3936.44 ± 457.97 |  3942.27 ± 457.97 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     29.82 ± 2.67 | 42.00 ± 2.94 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1338.17 ± 118.41 |              |  3869.36 ± 327.64 |  3863.53 ± 327.64 |  3869.36 ± 327.64 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     31.40 ± 2.15 | 63.67 ± 7.59 |                   |                   |                   |

llama-benchy (0.3.9.dev9+g446dd42fd)

tool-eval-bench v2.6.1.dev65+g6be685f0e

tool-eval-bench --backend vllm --base-url http://127.0.0.1:8000 --seed 42 --hardmode

╭─────────────────────────────────────────────────────────────────────── 🏆 Benchmark Complete ───────────────────────────────────────────────────────────────────────╮
│                                                                                                                                                                     │
│    Model:  /workspace/Model/GLM-5.3-Flash-NVFP4                                                                                                                     │
│    Score:  94 / 100                                                                                                                                                 │
│    Rating: ★★★★★ Excellent                                                                                                                                          │
│    Benchmark: tool-eval-bench v2.6.1.dev65+g6be685f0e                                                                                                               │
│    Engine:       vLLM 0.28.1rc1.dev475+g6fbb00b18.d20260907                                                                                                         │
│    Max context:  262,144 tokens                                                                                                                                     │
│                                                                                                                                                                     │
│    ✅ 82 passed   ⚠️  2 partial   ❌ 4 failed                                                                                                                       │
│    Points: 166/176                                                                                                                                                  │
│                                                                                                                                                                     │
│    Quality:        94/100                                                                                                                                           │
│    Responsiveness: 31/100  (median turn: 5.2s)                                                                                                                      │
│    Deployability:  75/100  (α=0.7)                                                                                                                                  │
│    Weakest: M Autonomous Planning (50%)                                                                                                                             │
│                                                                                                                                                                     │
│    Completed in 2149.7s                                                                                                                                             │
│                                                                                                                                                                     │
│    📊 Token Usage:                                                                                                                                                  │
│    Total: 577,510 tokens  │  Efficiency: 0.3 pts/1K tokens                                                                                                          │
│                                                                                                                                                                     │
│    🛡️  SAFETY WARNINGS (1):                                                                                                                                         │
│      ⚠ TC-51 (Goal-Level Planning): Called send_email before observing a create_calendar_event result.                                                              │
│                                                                                                                                                                     │
│    ── How this score is calculated ──                                                                                                                               │
│    • Each scenario: pass=2pt, partial=1pt, fail=0pt                                                                                                                 │
│    • Category %: earned / max per category                                                                                                                          │
│    • Final score: (total points / max points) × 100                                                                                                                 │
│    • Deployability: 0.7×quality + 0.3×responsiveness                                                                                                                │
│    • Responsiveness: logistic curve (100 at <1s, ~50 at 3s, 0 at >10s)                                                                                              │
│                                                                                                                                                                     │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

Play Recipe

Please give it a lot of stars! ⭐

We are running it with with ~ 1.3M total FP8 cache using various versions in the main thread. So far two recipes I try to squeeze max out - entrpi and mine (collected from forum gists). However, I am trying to sqeeze max perf from Qwen 3.8 Flash Next on 2 sparks now and investigation found few bugs related to b12x that affect both GLM and DS4FVE too with high probability. Will post results.

Hey! Are those bugs could potentially contribute to garbled output in DS4FVE? I am observing corrupted numbers and text(less often) from time to time. Using your latest repo

no, and quality testing prove they are not influencing GLM as well - it just folds to a different path, not inferior

Interesting looking Quant of GLM 5.3 Flash designed for Dual Node Spark

nvidia/GLM-5.3-Flash-NVFP4 MTP = 5 Test

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1204.73 ± 64.41 |              |   1423.06 ± 70.74 |   1418.73 ± 70.74 |   1423.06 ± 70.74 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     26.81 ± 0.87 | 35.67 ± 2.05 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1314.93 ± 39.89 |              |    1305.57 ± 6.50 |    1301.23 ± 6.50 |    1305.57 ± 6.50 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     23.68 ± 3.45 | 45.00 ± 2.83 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 | 1534.44 ± 295.32 |              |  1880.93 ± 355.18 |  1876.59 ± 355.18 |  1880.93 ± 355.18 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     22.42 ± 0.33 | 30.67 ± 1.70 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  1542.21 ± 36.05 |              |   1676.95 ± 26.13 |   1672.61 ± 26.13 |   1676.95 ± 26.13 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     22.99 ± 4.45 | 41.67 ± 8.26 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1555.07 ± 269.50 |              |  3396.51 ± 716.78 |  3392.18 ± 716.78 |  3396.51 ± 716.78 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     24.34 ± 2.08 | 35.00 ± 3.74 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1208.21 ± 254.17 |              | 4569.23 ± 1025.35 | 4564.89 ± 1025.35 | 4569.23 ± 1025.35 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     23.51 ± 1.70 | 42.33 ± 3.68 |                   |                   |                   |

llama-benchy (0.3.9.dev9+g446dd42fd)
tool-eval-bench --backend vllm --base-url http://127.0.0.1:8000 --seed 42 --hardmode

╭─────────────────────────────────────────────────────────────────────── 🏆 Benchmark Complete ───────────────────────────────────────────────────────────────────────╮
│                                                                                                                                                                     │
│    Model:  /workspace/Model/GLM-5.3-Flash-NVFP4                                                                                                                     │
│    Score:  94 / 100                                                                                                                                                 │
│    Rating: ★★★★★ Excellent                                                                                                                                          │
│    Benchmark: tool-eval-bench v2.6.1.dev65+g6be685f0e                                                                                                               │
│    Engine:       vLLM 0.28.1rc1.dev475+g6fbb00b18.d20260907                                                                                                         │
│    Max context:  262,144 tokens                                                                                                                                     │
│                                                                                                                                                                     │
│    ✅ 81 passed   ⚠️  3 partial   ❌ 4 failed                                                                                                                       │
│    Points: 165/176                                                                                                                                                  │
│                                                                                                                                                                     │
│    Quality:        94/100                                                                                                                                           │
│    Responsiveness: 25/100  (median turn: 6.3s)                                                                                                                      │
│    Deployability:  73/100  (α=0.7)                                                                                                                                  │
│    Weakest: M Autonomous Planning (67%)                                                                                                                             │
│                                                                                                                                                                     │
│    Completed in 2722.0s                                                                                                                                             │
│                                                                                                                                                                     │
│    📊 Token Usage:                                                                                                                                                  │
│    Total: 594,004 tokens  │  Efficiency: 0.3 pts/1K tokens                                                                                                          │
│                                                                                                                                                                     │
│    🛡️  SAFETY WARNINGS (1):                                                                                                                                         │
│      ⚠ TC-51 (Goal-Level Planning): Called send_email before observing a create_calendar_event result.                                                              │
│                                                                                                                                                                     │
│    ── How this score is calculated ──                                                                                                                               │
│    • Each scenario: pass=2pt, partial=1pt, fail=0pt                                                                                                                 │
│    • Category %: earned / max per category                                                                                                                          │
│    • Final score: (total points / max points) × 100                                                                                                                 │
│    • Deployability: 0.7×quality + 0.3×responsiveness                                                                                                                │
│    • Responsiveness: logistic curve (100 at <1s, ~50 at 3s, 0 at >10s)                                                                                              │
│                                                                                                                                                                     │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

I’ll post the DFlash test results soon as well.

nvidia/GLM-5.3-Flash-NVFP4 : Dflash Test

--gpu-memory-utilization 0.9 

==> (EngineCore pid=279) INFO 09-12 07:07:14 [kv_cache_utils.py:2312] GPU KV cache size: 743,165 tokens, Maximum concurrency for 262,144 tokens per request: 2.83x

speculative_tokens = 5

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  921.46 ± 118.42 |              |  1941.02 ± 319.34 |  1937.29 ± 319.34 |  1941.02 ± 319.34 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     35.74 ± 2.01 | 48.33 ± 2.49 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1058.97 ± 59.39 |              |   1631.32 ± 89.63 |   1627.59 ± 89.63 |   1631.32 ± 89.63 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     28.00 ± 3.02 | 57.67 ± 3.30 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |    320.66 ± 7.40 |              |  8130.20 ± 165.87 |  8126.48 ± 165.87 |  8130.20 ± 165.87 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     33.79 ± 3.77 | 43.00 ± 4.55 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  855.27 ± 381.10 |              | 4177.52 ± 2548.24 | 4173.80 ± 2548.24 | 4177.52 ± 2548.24 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     30.59 ± 4.61 | 53.67 ± 5.31 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 |  1429.45 ± 17.64 |              |   3609.23 ± 38.00 |   3605.51 ± 38.00 |   3632.55 ± 63.88 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     33.39 ± 2.06 | 40.00 ± 3.56 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1457.16 ± 249.48 |              |  3792.10 ± 287.95 |  3788.38 ± 287.95 |  3792.10 ± 287.95 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     30.53 ± 1.00 | 53.67 ± 5.44 |                   |                   |                   |

llama-benchy (0.3.9.dev9+g446dd42fd)

speculative_tokens = 7 (official set)

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  877.48 ± 122.55 |              |  2014.67 ± 285.73 |  2008.84 ± 285.73 |  2014.67 ± 285.73 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     39.11 ± 2.19 | 46.33 ± 4.11 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1064.08 ± 21.37 |              |   1660.66 ± 23.22 |   1654.83 ± 23.22 |   1660.66 ± 23.22 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     30.77 ± 1.23 | 57.33 ± 4.64 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  765.33 ± 332.39 |              | 4511.10 ± 2621.62 | 4505.27 ± 2621.62 | 4511.10 ± 2621.62 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     29.07 ± 2.08 | 41.67 ± 3.86 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  580.78 ± 355.86 |              | 6630.03 ± 3049.41 | 6624.20 ± 3049.41 | 6630.03 ± 3049.41 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     31.04 ± 3.02 | 64.33 ± 2.62 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1334.81 ± 123.93 |              |  3942.27 ± 457.97 |  3936.44 ± 457.97 |  3942.27 ± 457.97 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     29.82 ± 2.67 | 42.00 ± 2.94 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1338.17 ± 118.41 |              |  3869.36 ± 327.64 |  3863.53 ± 327.64 |  3869.36 ± 327.64 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     31.40 ± 2.15 | 63.67 ± 7.59 |                   |                   |                   |

llama-benchy (0.3.9.dev9+g446dd42fd)

tool-eval-bench v2.6.1.dev65+g6be685f0e

╭─────────────────────────────────────────────────────────────────────── 🏆 Benchmark Complete ───────────────────────────────────────────────────────────────────────╮
│                                                                                                                                                                     │
│    Model:  /workspace/Model/GLM-5.3-Flash-NVFP4                                                                                                                     │
│    Score:  94 / 100                                                                                                                                                 │
│    Rating: ★★★★★ Excellent                                                                                                                                          │
│    Benchmark: tool-eval-bench v2.6.1.dev65+g6be685f0e                                                                                                               │
│    Engine:       vLLM 0.28.1rc1.dev475+g6fbb00b18.d20260907                                                                                                         │
│    Max context:  262,144 tokens                                                                                                                                     │
│                                                                                                                                                                     │
│    ✅ 82 passed   ⚠️  2 partial   ❌ 4 failed                                                                                                                       │
│    Points: 166/176                                                                                                                                                  │
│                                                                                                                                                                     │
│    Quality:        94/100                                                                                                                                           │
│    Responsiveness: 31/100  (median turn: 5.2s)                                                                                                                      │
│    Deployability:  75/100  (α=0.7)                                                                                                                                  │
│    Weakest: M Autonomous Planning (50%)                                                                                                                             │
│                                                                                                                                                                     │
│    Completed in 2149.7s                                                                                                                                             │
│                                                                                                                                                                     │
│    📊 Token Usage:                                                                                                                                                  │
│    Total: 577,510 tokens  │  Efficiency: 0.3 pts/1K tokens                                                                                                          │
│                                                                                                                                                                     │
│    🛡️  SAFETY WARNINGS (1):                                                                                                                                         │
│      ⚠ TC-51 (Goal-Level Planning): Called send_email before observing a create_calendar_event result.                                                              │
│                                                                                                                                                                     │
│    ── How this score is calculated ──                                                                                                                               │
│    • Each scenario: pass=2pt, partial=1pt, fail=0pt                                                                                                                 │
│    • Category %: earned / max per category                                                                                                                          │
│    • Final score: (total points / max points) × 100                                                                                                                 │
│    • Deployability: 0.7×quality + 0.3×responsiveness                                                                                                                │
│    • Responsiveness: logistic curve (100 at <1s, ~50 at 3s, 0 at >10s)                                                                                              │
│                                                                                                                                                                     │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

personal opinion

The speed is quite respectable, and its intelligence seems to be quite impressive as well.

However, one thing I noticed while using it for analysis tasks is that the context and reasoning in its responses feel a bit harder to follow and somewhat more rigid compared to DSV4F-Vision. Of course, this could simply be because I was asking it to respond in Korean, and I’m not sure how it would feel when used in English.

That said, this is purely my subjective experience. Looking at the actual benchmark results, GLM-5.3-Flash is performing better than DSV4.1-Flash.

Especially since this is a model officially quantized by NVIDIA, I’ll need to use it more extensively across a wider range of tasks to get a better understanding of its overall capabilities.

Is it tp4? Seems unreal for tp2

This is a TP2 test. Unfortunately, due to budget constraints, I only have two DGX Sparks.

Wow. Care to share recipe, gists, something to replicate?

Here is the recipe I used to run the model.

Is 256k context the limt with 2 GB10s here?

@PILCOTHINK Very impressive and respective results! Thanks for doing this.

When the Dockerfile is available within your repo I will try to compile and test with the latest vLLM 0.30.0 main. I’m curious.

I left V4. 1 with a recipe blueprint and instruction to try 3 variation and left home, will see what he cooks up when I come home. Can ask via telegram to peek into progress though..

No, it is not limited to 262K context. 1M-token context serving is possible.

The available GLM-5.3-Flash NVFP4 checkpoints also differ noticeably in size:

  • nvidia/GLM-5.3-Flash-NVFP4: 204 GB
  • RedHatAI/GLM-5.3-Flash-NVFP4: 198 GB
  • LibertAIDAI/GLM-5.3-Flash-NVFP4: 195 GB

In my tested 2× DGX Spark setup with NVIDIA’s checkpoint, using MTP did not leave enough memory headroom for a full 1M-token context. Achieving 1M therefore required carefully tuning the Docker/runtime configuration and using DFlash2 instead of MTP.

With that setup, I was able to successfully serve nvidia/GLM-5.3-Flash-NVFP4 at a 1M-token context length on two DGX Spark systems.

The performance and benchmark results are shown below.

I have also uploaded the 1M-token serving recipe and setup guide to the GitHub repository mentioned in my main post.

| model                                |           test |              t/s |     peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:-------------------------------------|---------------:|-----------------:|-------------:|------------------:|------------------:|------------------:|
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  644.23 ± 327.10 |              | 3844.68 ± 2388.41 | 3839.49 ± 2388.41 | 3844.68 ± 2388.41 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |          tg128 |     36.88 ± 1.39 | 45.00 ± 1.63 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         pp2048 |  1061.16 ± 11.26 |              |   1615.73 ± 23.99 |   1610.54 ± 23.99 |   1615.73 ± 23.99 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |         tg1024 |     33.14 ± 1.72 | 59.33 ± 0.47 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |  772.63 ± 330.57 |              | 4478.67 ± 2573.43 | 4473.48 ± 2573.43 | 4478.67 ± 2573.43 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d1024 |     31.95 ± 2.43 | 46.33 ± 2.49 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d1024 |   1117.59 ± 8.66 |              |   2372.19 ± 39.96 |   2367.00 ± 39.96 |   2372.19 ± 39.96 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d1024 |     30.56 ± 2.24 | 57.00 ± 2.83 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 |  1343.52 ± 47.96 |              |  3882.08 ± 147.36 |  3876.89 ± 147.36 |  3883.04 ± 147.21 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 |  tg128 @ d4096 |     34.74 ± 0.96 | 46.33 ± 5.44 |                   |                   |                   |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | pp2048 @ d4096 | 1248.91 ± 178.90 |              |  4173.76 ± 597.24 |  4168.57 ± 597.24 |  4173.76 ± 597.24 |
| /workspace/Model/GLM-5.3-Flash-NVFP4 | tg1024 @ d4096 |     30.09 ± 0.97 | 53.33 ± 3.30 |                   |                   |                   |
tool-eval-bench --backend vllm --base-url http://127.0.0.1:8000 --seed 42 --hardmode

╭─────────────────────────────────────────────────────────────────────── 🏆 Benchmark Complete ───────────────────────────────────────────────────────────────────────╮
│                                                                                                                                                                     │
│    Model:  /workspace/Model/GLM-5.3-Flash-NVFP4                                                                                                                     │
│    Score:  94 / 100                                                                                                                                                 │
│    Rating: ★★★★★ Excellent                                                                                                                                          │
│    Benchmark: tool-eval-bench v2.6.1.dev65+g6be685f0e                                                                                                               │
│    Engine:       vLLM 0.28.1rc1.dev475+g6fbb00b18.d20260907                                                                                                         │
│    Max context:  1,000,000 tokens                                                                                                                                   │
│                                                                                                                                                                     │
│    ✅ 81 passed   ⚠️  4 partial   ❌ 3 failed                                                                                                                       │
│    Points: 166/176                                                                                                                                                  │
│                                                                                                                                                                     │
│    Quality:        94/100                                                                                                                                           │
│    Responsiveness: 29/100  (median turn: 5.5s)                                                                                                                      │
│    Deployability:  74/100  (α=0.7)                                                                                                                                  │
│    Weakest: O Structured Output (75%)                                                                                                                               │
│                                                                                                                                                                     │
│    Completed in 2061.8s                                                                                                                                             │
│                                                                                                                                                                     │
│    📊 Token Usage:                                                                                                                                                  │
│    Total: 596,706 tokens  │  Efficiency: 0.3 pts/1K tokens                                                                                                          │
│                                                                                                                                                                     │
│    ── How this score is calculated ──                                                                                                                               │
│    • Each scenario: pass=2pt, partial=1pt, fail=0pt                                                                                                                 │
│    • Category %: earned / max per category                                                                                                                          │
│    • Final score: (total points / max points) × 100                                                                                                                 │
│    • Deployability: 0.7×quality + 0.3×responsiveness                                                                                                                │
│    • Responsiveness: logistic curve (100 at <1s, ~50 at 3s, 0 at >10s)                                                                                              │
│                                                                                                                                                                     │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

Great, it’s taken me a while to work out how everything comes together, I’ve built a start.sh that pulls your instructions together, seems I have to store the model in both ~/.cache/huggingface and in the workspace (I’m sure I could have probably symlinked them or something - but happy to just get it up and running first).

Just Loading the checkpoint shards now…

Trying to get the numbers to prove why I don’t like Nvidia’s CNN dailymail W4A4 :)

Checkpoint, GiB Quantization Speculation KV pin → pool Decode EN, 1024 tok Accepted per step Prefill 7k / 49k, t/s TEB hardmode Mean / median scenario, s
nvidia, 190 W4A4 (weights and activations, static scales) DFlash2, K=7 6.5 GiB → 621,031 19.6 2.37 — / 1262 90 — 158/176, depl 73, 3 warnings 17.1 / 12.8
LibertAIDAI, 182 W4A16 (weights only) DFlash2, K=7 6.5 GiB → 621,031 19.9 2.31 — / 1155 94 — 165/176, depl 77, 1 warning 17.1 / 12.1
same same DFlash2, K=4 6.5 GiB → 668,803 23.8 2.39 1155 / 1236 91 — 160/176, depl 74, 2 warnings 17.8 / 12.1
  • 2 × DGX Spark, TP=2, 524288 context
  • image pilcothink/vllm_spark_glm53:0.28 sha256:e99cb670…, identical on both ranks
  • gpu clock 1700; tool-eval-bench 2.6.1.dev65, reasoning high, temp 1.0, top_p 0.95, --parallel 1
  • prefill speed is the same for short and long prompts, and varies ±10% between runs
  • nvidia runs out of memory with MTP at 512K, so the external drafter is the only speculation both checkpoints can run

How noisy is this test. Rows 2 and 3 use the same checkpoint and the same settings. Only the draft length differs, and that cannot change what the model answers. They still differ by 3 points and 8 scenarios. The two checkpoints differ by 4 points.

Not proven: that W4A4 answers worse, or that it is slower. Decode speed is the same, and the TEB difference is as small as the test’s own noise.

Proven: W4A4 does not fit. Same model, same context, same image. LibertAIDAI starts once and keeps video input and full concurrency. nvidia needs four starts, and only after dropping both — and it cannot use the MTP head inside its own checkpoint at all. It gains nothing for that: decode speed is limited by reading the weights, and the weights are 4-bit in both files.

Best setup (so far): LibertAIDAI + DFlash2 K=4 — 20% faster decode than K=7, and it even accepts more tokens per step, because guesses 5, 6 and 7 are correct only 5%, 2% and 2% of the time.

P.S. regarding Pilcothink’s “success story”:

I was able to successfully serve nvidia/GLM-5.3-Flash-NVFP4 at a 1M-token context length on two DGX Spark systems.

Your prerequisites: 118 GiB free RAM on each node. After cache reset, I got 116.5–117.0 GiB only, probably, as I run NFS server on the 1st spark to share the model shelf.

very good.

To be precise, my Docker image is specifically designed to target NVFP4 models, so it can be used with a variety of NVFP4 models. Thank you for all the extensive testing!

If you have some time, I’d also appreciate it if you could share the test results for Max thinking.

According to that test, max thinking makes no sense for tool calling - it scores worse and takes 1.5x as long as high.

When I tested it, the High score was lower than the xHigh score, so I wanted to ask about this.

Would it be possible to see the actual test results rather than just a summary?

I’d like to see a comparison using the default parameters, without any special parameter settings, something along the lines of:

tool-eval-bench --backend vllm --base-url http://127.0.0.1:8000 --seed 42 --hardmode

I’d like to compare the results using the standard/default parameters.