@jeffccie How well does this work with tool calls and for long runs? Very interested to integrate this model into HOW-TO: setup-dgx-spark docker inference - A "Sane" Inference Stack for GB10 (Need Contributors!) - #3 by jd36
jd36
3
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Make GLM-4.7-Flash go BRRRRR | 18 | 2925 | March 25, 2026 | |
| GLM-4.7-Flash-NVFP4 was just released, but for Transformers 5.0 + vLLM 0.14...? | 89 | 5177 | February 13, 2026 | |
| [Request] GLM-4.7-Flash AWQ/NVFP4 Instructions | 7 | 1314 | January 26, 2026 | |
| Full GLM-4.7 (355B) NVFP4 at 64K context on 2× DGX Spark GB10 — working recipe (vLLM, TP=2) | 2 | 507 | August 6, 2026 | |
| GLM-5.3-Flash-NVFP4 on 2× DGX Spark — vLLM TP=2 -- docker compose | 3 | 1481 | September 1, 2026 | |
| GLM-4.7-NVFP4 (NOT Flash) served with TRT-LLM on 2x DGX Spark | 6 | 728 | January 26, 2026 | |
| How to run Gemma-4-NVFP4 in vLLM Docker? | 11 | 6407 | April 12, 2026 | |
| Running GLM-4.7-FP8 (355B MoE) on 4x DGX Spark with SGLang + EAGLE Speculative Decoding | 38 | 2952 | June 24, 2026 | |
| GLM 4.6V works on Spark! | 12 | 2529 | January 22, 2026 | |
| How to run GLM 4.7 on dual DGX Sparks with vLLM / mods support in spark-vllm-docker | 27 | 4830 | January 2, 2026 |