NVIDIA Developer Forums
GLM 5.2 Hybrid-FP8+NVFP4+MXFP4 + Optimal runtime recipe
Accelerated Computing
DGX Spark / GB10 User Forum
DGX Spark / GB10
aidendle94
July 21, 2026, 8:15pm
4
Sorry, were you able to resolve this? What issues did you run into?
show post in topic
Related topics
Topic
Replies
Views
Activity
GLM-5.2 on a 3× GB10 cluster: ~16,13 tok/s decode, 215K ctx FP8 TP=3 + VISION!
DGX Spark / GB10 Projects
7
1274
July 26, 2026
GLM-5.2 on a 4× GB10 cluster: ~22 tok/s decode, 256K ctx, Recipe
DGX Spark / GB10
llama
,
deepseek
200
13279
August 9, 2026
GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode (generic - coding/structured)
DGX Spark / GB10
gaming
,
llama
29
2270
August 29, 2026
(Academic) GLM-5.2 on 2x DGX-Spark/GB10 nodes: Crazy 1-bit UD-IQ1_S + RPC llama.cpp + 256K context + 8 tok/s
DGX Spark / GB10
inception
,
llama
,
deepseek
8
3925
August 1, 2026
GLM-5.2 (unpruned) @ 200K context on 4× DGX Spark — 27 tok/s single / 52.5 tok/s @c4
DGX Spark / GB10 Projects
deepseek
3
748
August 4, 2026
GLM-5.3-Flash: 320B total parameters / 18B active
DGX Spark / GB10
286
11815
September 5, 2026
Fitting a high-quality REAP-less GLM-5.2 onto 4x DGX Spark
DGX Spark / GB10
deepseek
0
2084
June 29, 2026
Running GLM-4.7-FP8 (355B MoE) on 4x DGX Spark with SGLang + EAGLE Speculative Decoding
DGX Spark / GB10 Projects
38
2925
June 24, 2026
GLM-5.3-Flash weights released (Ox Alpha)
DGX Spark / GB10 Projects
10
4848
August 28, 2026
How to run GLM 4.7 on dual DGX Sparks with vLLM / mods support in spark-vllm-docker
DGX Spark / GB10
27
4813
January 2, 2026