Summary
I run 1-bit quantized GLM-5.2 over 2x DGX Sparks just for testing and curiosity to observe whether a flagship class large MoE can even run across 2x nodes.
- This is a toy experiment that is not feasible to deploy.
- Very slow responsiveness. Token generation is tolerable, but prompt processing is making it unfeasible.
- Llama.cpp tensor split over RPC is unclean across nodes, hence such deployments suffer on DGX Spark.
- Interesting findings at this extreme quantization.
- Limited benchmarks show good capability, especially in safety. My initial expectation was to find a barely functioning LLM.
- Subjective novel riddle tests showed excellent resistance to making typical LLM mistakes.
- On another subjective assessment it can accurately continue a theoretical physics conversation.
Quantization
UD-IQ1_S is the smallest quantization Unsloth published, targeting log2(3)~1.58 BPW (bits per weight) that became a meme since its inception. They measured ~76.2% top-1 accuracy at ~14% of the size.
# 1-bit is more like 2.3 bits on average
0.00.670.132 I print_info: file type = IQ1_S - 1.5625 bpw
0.00.670.134 I print_info: file size = 201.82 GiB (2.30 BPW)
DGX-Spark/GB10 Recipie
As usual, DGX Sparkโs unified memory and the node distribution comes with a few caveats and gotchas that need sorting. This one is relatively straightforward to run with minor modifications on the vanilla configuration that Unsloth recommends.
Following is my installation of llama.cpp with custom flags. This needs to be compiled or cloned to the worker node as well.
# Clone the repository
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
# Create and enter the build directory
mkdir build && cd build
# Configure CMake for GB10 (sm_121) with RPC enabled
cmake .. -DGGML_CUDA=ON -DGGML_RPC=ON -DCMAKE_CUDA_ARCHITECTURES="121" -DCMAKE_BUILD_TYPE=Release
# Compile using all available CPU threads
make -j$(nproc)
Then run an RPC server on the worker node. I run the server on enp1s0f1np1 over ConnectX-7 NIC, which should pick up RDMA automatically.
CUDA_VISIBLE_DEVICES=0
./bin/rpc-server \
--host <worker-ip-of-enp1s0f1np1> \
--port 50052
Run llama.cpp on the main node.
CUDA_VISIBLE_DEVICES=0 \
./bin/llama-server \
--model ~/models/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00001-of-00006.gguf \
--alias "zai-org/GLM-5.2" \
--ctx-size 262144 \
--n-gpu-layers 999 \
--rpc <worker-ip-of-enp1s0f1np1>:50052 \
--host 0.0.0.0 \
--port 8019 \
--fit off \
--tensor-split 52,48 \
--ubatch-size 2048 \
--cache-ram 0 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--chat-template-kwargs '{"enable_thinking":false}'
# For reasioning-effort replace above line with the one below.
# Without thinking setting it defaults to reasoning_effort=max
# --chat-template-kwargs '{"reasoning_effort":"high"}'
--ctx-size 262144can be increased to 512K for empty DGX Spark nodes with no overhead.--fit offis necessary to prevent crashes. When fit is on llama.cpp tries to allocate whole memory, treating it as VRAM, starves the unified memory of DGX Spark and causes a hard out-of-memory event.--tensor-split 52,48hints llama.ccp to split the load manually 52% on the worker node and 48% on the main node. You may split it depending on your node utility.--ubatch-size 2048increases prompt-processing rate about 30%.--cache-ram 0saves ~8GB memory, helping 2x DGX Spark nodes starved of memory. Also unified memory does not benefit from spillover system memory behavior as much.--cache-type-k q8_0and--cache-type-v q8_0for FP8 KV cache quantization. Unsloth says Q4.1 might work but did not try my luck.--chat-template-kwargs '{"enable_thinking":false}'for disabling thinking.--chat-template-kwargs '{"reasoning_effort":"high"}'(highormax) for enabling reasoning. When unspecified the chat template defaults tomax.- I did not try
mtp, but I also do not think it is enabled for the model. - Do NOT use
--cache-reuse 256, it corrupts the KV cache in running conversations.
Benchmarks
I run some limited off-the-shelve benchmarks for comparison. The performance benchmarks are as abysmal as one would expect for the model size and llama.cpp partitioning over RPC .
llama-benchy Results
โโโโโโโโโโโโโโโโโโโโโโโโโณโโโโโโโณโโโโโโโโโณโโโโโโโโโณโโโโโโโโโโโโณโโโโโโโโโโโโโณโโโโโโโโโโโ
โ Test โ c โ pp t/s โ tg t/s โ TTFT (ms) โ Total (ms) โ Tokens โ
โกโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฉ
โ pp2048 tg128 @ d0 โ c1 โ 213 โ 7.7 โ 9,828 โ 26,240 โ 2048+128 โ
โ pp2048 tg128 @ d0 โ c2 โ 203 โ 8.0 โ 14,683 โ 39,810 โ 2048+128 โ
โ pp2048 tg128 @ d0 โ c4 โ 203 โ 8.9 โ 26,600 โ 66,018 โ 2048+128 โ
โ pp2048 tg128 @ d4096 โ c1 โ 189 โ 6.9 โ 32,977 โ 51,294 โ 2048+128 โ
โ pp2048 tg128 @ d4096 โ c2 โ 169 โ 4.6 โ 57,864 โ 90,371 โ 2048+128 โ
โ pp2048 tg128 @ d4096 โ c4 โ 158 โ 3.4 โ 98,634 โ 160,178 โ 2048+128 โ
โ pp2048 tg128 @ d8192 โ c1 โ 156 โ 6.1 โ 66,792 โ 87,418 โ 2048+128 โ
โ pp2048 tg128 @ d8192 โ c2 โ 131 โ 2.1 โ 113,779 โ 156,110 โ 2048+128 โ
โ pp2048 tg128 @ d8192 โ c4 โ 122 โ 1.5 โ 198,134 โ 286,851 โ 2048+128 โ
โ pp2048 tg128 @ d16384 โ c1 โ 112 โ 5.1 โ 164,316 โ 189,116 โ 2048+128 โ
โ pp2048 tg128 @ d16384 โ c2 โ 89 โ 0.8 โ 291,628 โ 348,321 โ 2048+128 โ
โ pp2048 tg128 @ d16384 โ c4 โ 82 โ 0.6 โ 504,770 โ 635,831 โ 2048+128 โ
โ pp2048 tg128 @ d32768 โ c1 โ 76 โ 4.1 โ 473,380 โ 504,629 โ 2048+128 โ
โ pp2048 tg128 @ d32768 โ c2 โ 55 โ 0.2 โ 832,804 โ 911,031 โ 2048+128 โ
โ pp2048 tg128 @ d32768 โ c4 โ 52 โ 0.2 โ 1,412,241 โ 1,596,258 โ 2048+128 โ
โโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโโดโโโโโโโโโโโโโดโโโโโโโโโโโ
โน Metrics sourced from llama-benchy โ see https://github.com/eugr/llama-benchy for methodology.
In tool-eval-bench I got much much better results than I expected.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโ
โ Category โ Score โ Bar โ Earned โ
โกโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฉ
โ Tool Selection โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Parameter Precision โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Multi-Step Chains โ 75% โ โโโโโโโโโโโโโโโโโโโโ โ 6/8 โ
โ Restraint & Refusal โ 83% โ โโโโโโโโโโโโโโโโโโโโ โ 5/6 โ
โ Error Recovery โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Localization โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Structured Reasoning โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Instruction Following โ 80% โ โโโโโโโโโโโโโโโโโโโโ โ 8/10 โ
โ Context & State โ 80% โ โโโโโโโโโโโโโโโโโโโโ โ 16/20 โ
โ Code Patterns โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Safety & Boundaries โ 92% โ โโโโโโโโโโโโโโโโโโโโ โ 24/26 โ
โ Toolset Scale โ 75% โ โโโโโโโโโโโโโโโโโโโโ โ 6/8 โ
โ Autonomous Planning โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Creative Composition โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 6/6 โ
โ Structured Output โ 100% โ โโโโโโโโโโโโโโโโโโโโ โ 12/12 โ
โ Hard Mode โ 73% โ โโโโโโโโโโโโโโโโโโโโ โ 22/30 โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโ
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ ๐ Benchmark Complete โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ โ
โ Model: zai-org/GLM-5.2 โ
โ Score: 88 / 100 โ
โ Rating: โ
โ
โ
โ
Good โ
โ โ
โ โ
69 passed โ ๏ธ 9 partial โ 6 failed โ
โ Points: 147/168 โ
โ โ
โ Quality: 88/100 โ
โ Responsiveness: 13/100 (median turn: 10.6s) โ
โ Deployability: 66/100 (ฮฑ=0.7) โ
โ Weakest: P Hard Mode (73%) โ
โ โ
โ Completed in 3380.1s โ tool-eval-bench v2.0.6 โ
โ โ
โ ๐ Token Usage: โ
โ Total: 262,797 tokens โ Efficiency: 0.6 pts/1K tokens โ
โ โ
โ โโ How this score is calculated โโ โ
โ โข Each scenario: pass=2pt, partial=1pt, fail=0pt โ
โ โข Category %: earned / max per category โ
โ โข Final score: (total points / max points) ร 100 โ
โ โข Deployability: 0.7รquality + 0.3รresponsiveness โ
โ โข Responsiveness: logistic curve (100 at <1s, ~50 at 3s, 0 at >10s) โ
โ โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Running trial 2/5โฆ 86/100
Running trial 3/5โฆ 89/100
Running trial 4/5โฆ 88/100
Running trial 5/5โฆ 89/100
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ ๐ Trial Statistics โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ โ
โ Trials: 5 โ
โ Score: 88.0 ยฑ 1.2 / 100 โ
โ Median: 88.0 โ
โ 95% CI: [87.0, 88.8] โ
โ Points: 148.0 ยฑ 2.1 โ
โ โ
โ Pass@5: 90.5% (capability ceiling) โ
โ Pass^5: 72.6% (reliability floor) โ
โ โ Gap: 17.9pp (high variance โ consistency issue) โ
โ โ
โ Categories with variance: โ
โ C Multi-Step Chains: 88% ยฑ 12.5% โ
โ H Instruction Following: 84% ยฑ 16.7% โ
โ I Context & State: 73% ยฑ 4.5% โ
โ K Safety & Boundaries: 90% ยฑ 2.2% โ
โ L Toolset Scale: 83% ยฑ 11.2% โ
โ N Creative Composition: 86% ยฑ 7.6% โ
โ P Hard Mode: 79% ยฑ 3.8% โ
โ โ
โ โก 16 unstable scenario(s): โ
โ TC-22: 1.2 ยฑ 1.1 (0,0,2,2,2) โ
โ TC-38: 1.4 ยฑ 0.6 (1,1,2,1,2) โ
โ TC-39: 1.2 ยฑ 0.5 (1,1,2,1,1) โ
โ TC-45: 1.2 ยฑ 1.1 (2,0,2,0,2) โ
โ TC-47: 1.6 ยฑ 0.6 (2,1,2,1,2) โ
โ TC-49: 1.8 ยฑ 0.5 (2,2,2,2,1) โ
โ TC-50: 1.2 ยฑ 0.5 (2,1,1,1,1) โ
โ TC-56: 1.2 ยฑ 0.5 (2,1,1,1,1) โ
โ TC-57: 1.4 ยฑ 0.6 (1,1,2,2,1) โ
โ TC-60: 1.2 ยฑ 0.5 (2,1,1,1,1) โ
โ TC-61: 1.0 ยฑ 1.0 (0,2,0,1,2) โ
โ TC-71: 1.6 ยฑ 0.9 (0,2,2,2,2) โ
โ TC-74: 1.8 ยฑ 0.5 (2,2,1,2,2) โ
โ TC-75: 0.4 ยฑ 0.9 (0,2,0,0,0) โ
โ TC-80: 1.6 ยฑ 0.9 (2,0,2,2,2) โ
โ TC-82: 0.2 ยฑ 0.5 (0,0,0,1,0) โ
โ โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Subjective Tests
Its thinking mode is completely intact even at this quantization. I tested it subjectively with my custom tests.
First test was solving nonsensical riddles and puzzles. These are similar to the โshould I drive or walk to car washโ paradox, but custom and unpublished.
- Did not make a single logical mistake:
- both thinking and non-thinking modes;
- across 4 fresh retries per riddle.
- The only issue was consistent lock-ups during one of the unsolvable paradoxes:
- it correctly identified paradoxes in its thinking traces;
- but it was unable to break out of the thinking loops;
- when thinking disabled, solved the paradox every single time with a valid logical correlation.
Second, I stressed it with a very long theoretical physics conversation I always use in testing LLMs that filled ~160K of its context.
- Wrote a consistent ~8000 word essay with proposals to extend it to ~20000 words.
- Correctly factored relativistic speeds in all scenarios.
- Correctly found a numerical error I injected and corrected one numerical typo in an assay I gave it to proof read blindly.
- Correctly argued counter-forces in a tricky question and proposed a safety net to my experiment.
- Correctly reasoned quite speculative hypothetical situations.
- It slightly beat DeepSeek-V4-Flash, which was the previous record holder.
- For reference, Qwen-3.6 MoE and Gemma-4 MoEs and Qwen-3.6-27B failed spectacularly, MiniMax-M3 (preliminary quantization for 2x nodes) hallucinated and refused to be steered out of the hallucination, Gemma-4 31B went quite well until it poisoned its context content with gibberish tokens and slowly broke down, Step-3.7 and Command-A+ were touch and go.
- Interesting that Gemini-3.5 also failed, not because its answers were bad, but it lost the thread.
- GPT-5.5 passed the test easily.
- Opus-4.8 also passed easily (while pissing me off with continuous backhand compliments and redundant arguments).
Conclusion
This results prove that there is a recipe that can tun GLM-5.2 over 2x DGX Sparks. The performance is abysmal as expected, but I was pleasantly surprised with the quality even at 1-bit quantization.
I hope you enjoyed the report. This is probably a useless recipe, but still wanted to share in case someone was wondering, or would like to experiment with it.