(Academic) GLM-5.2 on 2x DGX-Spark/GB10 nodes: Crazy 1-bit UD-IQ1_S + RPC llama.cpp + 256K context + 8 tok/s

Summary
I run 1-bit quantized GLM-5.2 over 2x DGX Sparks just for testing and curiosity to observe whether a flagship class large MoE can even run across 2x nodes.

  • This is a toy experiment that is not feasible to deploy.
    • Very slow responsiveness. Token generation is tolerable, but prompt processing is making it unfeasible.
    • Llama.cpp tensor split over RPC is unclean across nodes, hence such deployments suffer on DGX Spark.
  • Interesting findings at this extreme quantization.
    • Limited benchmarks show good capability, especially in safety. My initial expectation was to find a barely functioning LLM.
    • Subjective novel riddle tests showed excellent resistance to making typical LLM mistakes.
    • On another subjective assessment it can accurately continue a theoretical physics conversation.

Quantization
UD-IQ1_S is the smallest quantization Unsloth published, targeting log2(3)~1.58 BPW (bits per weight) that became a meme since its inception. They measured ~76.2% top-1 accuracy at ~14% of the size.

# 1-bit is more like 2.3 bits on average
0.00.670.132 I print_info: file type   = IQ1_S - 1.5625 bpw
0.00.670.134 I print_info: file size   = 201.82 GiB (2.30 BPW) 

DGX-Spark/GB10 Recipie
As usual, DGX Sparkโ€™s unified memory and the node distribution comes with a few caveats and gotchas that need sorting. This one is relatively straightforward to run with minor modifications on the vanilla configuration that Unsloth recommends.

Following is my installation of llama.cpp with custom flags. This needs to be compiled or cloned to the worker node as well.

# Clone the repository
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp

# Create and enter the build directory
mkdir build && cd build

# Configure CMake for GB10 (sm_121) with RPC enabled
cmake .. -DGGML_CUDA=ON -DGGML_RPC=ON -DCMAKE_CUDA_ARCHITECTURES="121" -DCMAKE_BUILD_TYPE=Release

# Compile using all available CPU threads
make -j$(nproc)

Then run an RPC server on the worker node. I run the server on enp1s0f1np1 over ConnectX-7 NIC, which should pick up RDMA automatically.

CUDA_VISIBLE_DEVICES=0
./bin/rpc-server \
	--host <worker-ip-of-enp1s0f1np1> \
	--port 50052

Run llama.cpp on the main node.

CUDA_VISIBLE_DEVICES=0 \
./bin/llama-server \
	--model ~/models/UD-IQ1_S/GLM-5.2-UD-IQ1_S-00001-of-00006.gguf \
	--alias "zai-org/GLM-5.2" \
	--ctx-size 262144 \
	--n-gpu-layers 999 \
	--rpc <worker-ip-of-enp1s0f1np1>:50052 \
	--host 0.0.0.0 \
	--port 8019 \
	--fit off \
	--tensor-split 52,48 \
	--ubatch-size 2048 \
	--cache-ram 0 \
	--cache-type-k q8_0 \
	--cache-type-v q8_0 \
	--chat-template-kwargs '{"enable_thinking":false}'

# For reasioning-effort replace above line with the one below.
# Without thinking setting it defaults to reasoning_effort=max
# --chat-template-kwargs '{"reasoning_effort":"high"}'
  • --ctx-size 262144 can be increased to 512K for empty DGX Spark nodes with no overhead.
  • --fit off is necessary to prevent crashes. When fit is on llama.cpp tries to allocate whole memory, treating it as VRAM, starves the unified memory of DGX Spark and causes a hard out-of-memory event.
  • --tensor-split 52,48 hints llama.ccp to split the load manually 52% on the worker node and 48% on the main node. You may split it depending on your node utility.
  • --ubatch-size 2048 increases prompt-processing rate about 30%.
  • --cache-ram 0 saves ~8GB memory, helping 2x DGX Spark nodes starved of memory. Also unified memory does not benefit from spillover system memory behavior as much.
  • --cache-type-k q8_0and --cache-type-v q8_0 for FP8 KV cache quantization. Unsloth says Q4.1 might work but did not try my luck.
  • --chat-template-kwargs '{"enable_thinking":false}' for disabling thinking.
  • --chat-template-kwargs '{"reasoning_effort":"high"}' (high or max) for enabling reasoning. When unspecified the chat template defaults to max.
  • I did not try mtp, but I also do not think it is enabled for the model.
  • Do NOT use--cache-reuse 256, it corrupts the KV cache in running conversations.

Benchmarks
I run some limited off-the-shelve benchmarks for comparison. The performance benchmarks are as abysmal as one would expect for the model size and llama.cpp partitioning over RPC .

                                llama-benchy Results
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”“
โ”ƒ Test                  โ”ƒ  c   โ”ƒ pp t/s โ”ƒ tg t/s โ”ƒ TTFT (ms) โ”ƒ Total (ms) โ”ƒ   Tokens โ”ƒ
โ”กโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ฉ
โ”‚ pp2048 tg128 @ d0     โ”‚  c1  โ”‚    213 โ”‚    7.7 โ”‚     9,828 โ”‚     26,240 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d0     โ”‚  c2  โ”‚    203 โ”‚    8.0 โ”‚    14,683 โ”‚     39,810 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d0     โ”‚  c4  โ”‚    203 โ”‚    8.9 โ”‚    26,600 โ”‚     66,018 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d4096  โ”‚  c1  โ”‚    189 โ”‚    6.9 โ”‚    32,977 โ”‚     51,294 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d4096  โ”‚  c2  โ”‚    169 โ”‚    4.6 โ”‚    57,864 โ”‚     90,371 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d4096  โ”‚  c4  โ”‚    158 โ”‚    3.4 โ”‚    98,634 โ”‚    160,178 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d8192  โ”‚  c1  โ”‚    156 โ”‚    6.1 โ”‚    66,792 โ”‚     87,418 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d8192  โ”‚  c2  โ”‚    131 โ”‚    2.1 โ”‚   113,779 โ”‚    156,110 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d8192  โ”‚  c4  โ”‚    122 โ”‚    1.5 โ”‚   198,134 โ”‚    286,851 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d16384 โ”‚  c1  โ”‚    112 โ”‚    5.1 โ”‚   164,316 โ”‚    189,116 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d16384 โ”‚  c2  โ”‚     89 โ”‚    0.8 โ”‚   291,628 โ”‚    348,321 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d16384 โ”‚  c4  โ”‚     82 โ”‚    0.6 โ”‚   504,770 โ”‚    635,831 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d32768 โ”‚  c1  โ”‚     76 โ”‚    4.1 โ”‚   473,380 โ”‚    504,629 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d32768 โ”‚  c2  โ”‚     55 โ”‚    0.2 โ”‚   832,804 โ”‚    911,031 โ”‚ 2048+128 โ”‚
โ”‚ pp2048 tg128 @ d32768 โ”‚  c4  โ”‚     52 โ”‚    0.2 โ”‚ 1,412,241 โ”‚  1,596,258 โ”‚ 2048+128 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

  โ„น Metrics sourced from llama-benchy โ€” see https://github.com/eugr/llama-benchy for methodology.

In tool-eval-bench I got much much better results than I expected.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”“
โ”ƒ Category                  โ”ƒ    Score     โ”ƒ Bar                      โ”ƒ   Earned    โ”ƒ
โ”กโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ฉ
โ”‚ Tool Selection            โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Parameter Precision       โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Multi-Step Chains         โ”‚     75%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘     โ”‚     6/8     โ”‚
โ”‚ Restraint & Refusal       โ”‚     83%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘     โ”‚     5/6     โ”‚
โ”‚ Error Recovery            โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Localization              โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Structured Reasoning      โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Instruction Following     โ”‚     80%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘     โ”‚    8/10     โ”‚
โ”‚ Context & State           โ”‚     80%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘     โ”‚    16/20    โ”‚
โ”‚ Code Patterns             โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Safety & Boundaries       โ”‚     92%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘     โ”‚    24/26    โ”‚
โ”‚ Toolset Scale             โ”‚     75%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘     โ”‚     6/8     โ”‚
โ”‚ Autonomous Planning       โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Creative Composition      โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚     6/6     โ”‚
โ”‚ Structured Output         โ”‚     100%     โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     โ”‚    12/12    โ”‚
โ”‚ Hard Mode                 โ”‚     73%      โ”‚ โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘     โ”‚    22/30    โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ ๐Ÿ† Benchmark Complete โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚                                                                                   โ”‚
โ”‚    Model:  zai-org/GLM-5.2                                                        โ”‚
โ”‚    Score:  88 / 100                                                               โ”‚
โ”‚    Rating: โ˜…โ˜…โ˜…โ˜… Good                                                              โ”‚
โ”‚                                                                                   โ”‚
โ”‚    โœ… 69 passed   โš ๏ธ  9 partial   โŒ 6 failed                                     โ”‚
โ”‚    Points: 147/168                                                                โ”‚
โ”‚                                                                                   โ”‚
โ”‚    Quality:        88/100                                                         โ”‚
โ”‚    Responsiveness: 13/100  (median turn: 10.6s)                                   โ”‚
โ”‚    Deployability:  66/100  (ฮฑ=0.7)                                                โ”‚
โ”‚    Weakest: P Hard Mode (73%)                                                     โ”‚
โ”‚                                                                                   โ”‚
โ”‚    Completed in 3380.1s  โ”‚  tool-eval-bench v2.0.6                                โ”‚
โ”‚                                                                                   โ”‚
โ”‚    ๐Ÿ“Š Token Usage:                                                                โ”‚
โ”‚    Total: 262,797 tokens  โ”‚  Efficiency: 0.6 pts/1K tokens                        โ”‚
โ”‚                                                                                   โ”‚
โ”‚    โ”€โ”€ How this score is calculated โ”€โ”€                                             โ”‚
โ”‚    โ€ข Each scenario: pass=2pt, partial=1pt, fail=0pt                               โ”‚
โ”‚    โ€ข Category %: earned / max per category                                        โ”‚
โ”‚    โ€ข Final score: (total points / max points) ร— 100                               โ”‚
โ”‚    โ€ข Deployability: 0.7ร—quality + 0.3ร—responsiveness                              โ”‚
โ”‚    โ€ข Responsiveness: logistic curve (100 at <1s, ~50 at 3s, 0 at >10s)            โ”‚
โ”‚                                                                                   โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

  Running trial 2/5โ€ฆ 86/100
  Running trial 3/5โ€ฆ 89/100
  Running trial 4/5โ€ฆ 88/100
  Running trial 5/5โ€ฆ 89/100
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ ๐Ÿ“Š Trial Statistics โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚                                                                                   โ”‚
โ”‚    Trials:  5                                                                     โ”‚
โ”‚    Score:   88.0 ยฑ 1.2 / 100                                                      โ”‚
โ”‚    Median:  88.0                                                                  โ”‚
โ”‚    95% CI:  [87.0, 88.8]                                                          โ”‚
โ”‚    Points:  148.0 ยฑ 2.1                                                           โ”‚
โ”‚                                                                                   โ”‚
โ”‚    Pass@5:  90.5%  (capability ceiling)                                           โ”‚
โ”‚    Pass^5:  72.6%  (reliability floor)                                            โ”‚
โ”‚    โš  Gap:    17.9pp  (high variance โ€” consistency issue)                          โ”‚
โ”‚                                                                                   โ”‚
โ”‚    Categories with variance:                                                      โ”‚
โ”‚      C Multi-Step Chains: 88% ยฑ 12.5%                                             โ”‚
โ”‚      H Instruction Following: 84% ยฑ 16.7%                                         โ”‚
โ”‚      I Context & State: 73% ยฑ 4.5%                                                โ”‚
โ”‚      K Safety & Boundaries: 90% ยฑ 2.2%                                            โ”‚
โ”‚      L Toolset Scale: 83% ยฑ 11.2%                                                 โ”‚
โ”‚      N Creative Composition: 86% ยฑ 7.6%                                           โ”‚
โ”‚      P Hard Mode: 79% ยฑ 3.8%                                                      โ”‚
โ”‚                                                                                   โ”‚
โ”‚    โšก 16 unstable scenario(s):                                                    โ”‚
โ”‚      TC-22: 1.2 ยฑ 1.1  (0,0,2,2,2)                                                โ”‚
โ”‚      TC-38: 1.4 ยฑ 0.6  (1,1,2,1,2)                                                โ”‚
โ”‚      TC-39: 1.2 ยฑ 0.5  (1,1,2,1,1)                                                โ”‚
โ”‚      TC-45: 1.2 ยฑ 1.1  (2,0,2,0,2)                                                โ”‚
โ”‚      TC-47: 1.6 ยฑ 0.6  (2,1,2,1,2)                                                โ”‚
โ”‚      TC-49: 1.8 ยฑ 0.5  (2,2,2,2,1)                                                โ”‚
โ”‚      TC-50: 1.2 ยฑ 0.5  (2,1,1,1,1)                                                โ”‚
โ”‚      TC-56: 1.2 ยฑ 0.5  (2,1,1,1,1)                                                โ”‚
โ”‚      TC-57: 1.4 ยฑ 0.6  (1,1,2,2,1)                                                โ”‚
โ”‚      TC-60: 1.2 ยฑ 0.5  (2,1,1,1,1)                                                โ”‚
โ”‚      TC-61: 1.0 ยฑ 1.0  (0,2,0,1,2)                                                โ”‚
โ”‚      TC-71: 1.6 ยฑ 0.9  (0,2,2,2,2)                                                โ”‚
โ”‚      TC-74: 1.8 ยฑ 0.5  (2,2,1,2,2)                                                โ”‚
โ”‚      TC-75: 0.4 ยฑ 0.9  (0,2,0,0,0)                                                โ”‚
โ”‚      TC-80: 1.6 ยฑ 0.9  (2,0,2,2,2)                                                โ”‚
โ”‚      TC-82: 0.2 ยฑ 0.5  (0,0,0,1,0)                                                โ”‚
โ”‚                                                                                   โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

Subjective Tests
Its thinking mode is completely intact even at this quantization. I tested it subjectively with my custom tests.

First test was solving nonsensical riddles and puzzles. These are similar to the โ€œshould I drive or walk to car washโ€ paradox, but custom and unpublished.

  • Did not make a single logical mistake:
    • both thinking and non-thinking modes;
    • across 4 fresh retries per riddle.
  • The only issue was consistent lock-ups during one of the unsolvable paradoxes:
    • it correctly identified paradoxes in its thinking traces;
    • but it was unable to break out of the thinking loops;
    • when thinking disabled, solved the paradox every single time with a valid logical correlation.

Second, I stressed it with a very long theoretical physics conversation I always use in testing LLMs that filled ~160K of its context.

  • Wrote a consistent ~8000 word essay with proposals to extend it to ~20000 words.
  • Correctly factored relativistic speeds in all scenarios.
  • Correctly found a numerical error I injected and corrected one numerical typo in an assay I gave it to proof read blindly.
  • Correctly argued counter-forces in a tricky question and proposed a safety net to my experiment.
  • Correctly reasoned quite speculative hypothetical situations.
  • It slightly beat DeepSeek-V4-Flash, which was the previous record holder.
    • For reference, Qwen-3.6 MoE and Gemma-4 MoEs and Qwen-3.6-27B failed spectacularly, MiniMax-M3 (preliminary quantization for 2x nodes) hallucinated and refused to be steered out of the hallucination, Gemma-4 31B went quite well until it poisoned its context content with gibberish tokens and slowly broke down, Step-3.7 and Command-A+ were touch and go.
    • Interesting that Gemini-3.5 also failed, not because its answers were bad, but it lost the thread.
    • GPT-5.5 passed the test easily.
    • Opus-4.8 also passed easily (while pissing me off with continuous backhand compliments and redundant arguments).

Conclusion
This results prove that there is a recipe that can tun GLM-5.2 over 2x DGX Sparks. The performance is abysmal as expected, but I was pleasantly surprised with the quality even at 1-bit quantization.

I hope you enjoyed the report. This is probably a useless recipe, but still wanted to share in case someone was wondering, or would like to experiment with it.

This is awesome, thanks for sharing!

Great work thank you for sharing!

I did another benchmark with an NVFP4 quantized version over a third-party subscription API, and the NVFP4 quality is noticeably worse than Unslothโ€™s 1-bit UD-IQ1_S quantization, just quicker. I think it makes sense if they quantized the wrong layers blindly to NVFP4, which will degrade the quality significantly. For reference, Unsloth does not quantize susceptible layers even when aggressively quantizing the rest to 1-bit.

For those of us who use LLMs over APIs, we should be careful choosing our providers.

Update: Nvidia just dropped their official NVFP4 version:

Hi @nerhun ,

Very nice work! Thanks for sharing!

If you want and have time then you can test and share the results of using vllm with gguf on 2 x DGX-Spark. This should give a better tps.

I tested successfully vllm with GGUF on 1 x DGX-Spark.

Many thanks!

Kind regards,

CM

Thank you for steering me to vLLM that I had at the back of my head.

I run the IQ1-S quantizattion of GLM-5.2 over vLLM 0.23, ~12 token/sec generation speed, tensor-parallel across 2 Sparks, but only ~40k context in total. I had to implement some patches to GGUF support, but it works.

I am now working on quantizing 0xSeroโ€™s REAP 504B with IQ2-XXS that will take a while to calculate the Imatrix. This is my first serious attempt at it and will take a week or two before I have results.

Cheers!

Hi @nerhun ,
You are welcome and thanks for sharing that vllm gguf is working with more than 30% faster token/sec (12 vs 8) speed than llama.cpp although with some implemented patches.

I am really interested about your upcoming results โ€œquantizing 0xSeroโ€™s REAP 504B with IQ2-XXSโ€ and running on two DGX Sparks. Please let us know and share the vLLM image and patches you have applied for your work once you are there.

Many thanks in advance!

Thanks for posting this โ€” I ran a close variant of your setup (2ร— Spark, CX7 direct link, GLM-5.2 UD-IQ1_M ~228.5 GB, llama.cpp RPC) and measured three things that line up with your recipe. Sharing in case they are useful.

1. โ€œshould pick up RDMA automaticallyโ€ โ€” it does, and that is worth pinning down

RDMA is not opt-in. If llama.cpp is built with libibverbs present, it switches to RDMA whenever a RoCE device matches the connection IP, even with GGML_RDMA_DEV unset. So โ€œunset means TCPโ€ is false, and a naive TCP-vs-RDMA comparison silently runs RDMA on both sides.

To confirm which transport you are actually on:

GGML_RPC_DEBUG=1 ./bin/rpc-server --host <worker-ip> --port 50052
# look for: RDMA probed: dev=...

That INFO line only appears on the rpc-server side, never on the client, so you cannot tell from the llama-server log. To force plain TCP, set a non-existent device (GGML_RDMA_DEV=disabled) so the probe fails, or build with -DGGML_RPC_RDMA=OFF.

What RDMA actually buys, A/B with everything else fixed (same hardware, same cmake flags, transport as the only variable):

workload TCP RDMA delta
GPT-OSS-120B, 2048-tok prompt eval 2174.5 tok/s 3163.3 tok/s +45%
GLM-5.2, 2048-tok prompt eval 171.85 tok/s 178.29 tok/s +3.8%
GLM-5.2, short-prompt QA soak (1,119 questions) 10.8 tok/s 15.9 tok/s +47%
GPT-OSS-120B, load time 64.0 s 80.0 s โˆ’25% (slower)

The pattern: the gain shrinks as GPU compute dominates. For GLM-5.2 on long prompts it is nearly nothing, which is consistent with your conclusion that prompt processing is what makes this unfeasible โ€” that cost is compute, not the link. Worth knowing before anyone spends time chasing RDMA as a fix for pp.

2. The c2/c4 rows may be measuring queuing rather than concurrency

If the server ran with the default --parallel 1, additional concurrent requests queue behind the first instead of running in parallel. On my side that looked like this:

config output tok/s mean TTFT
--parallel 1, c=1 7.32 3,175 ms
--parallel 1, c=2 8.21 23,909 ms

+12% throughput, 7.5ร— worse TTFT. The tell is ITL: it stayed at 119 ms in both cases, so generation speed never changed โ€” the second request simply waited. Your c1โ†’c4 rows have the same shape (tg 7.7โ†’8.9 while TTFT 9,828โ†’26,600 ms), which is why I suspect the slot count rather than the hardware.

Matching --parallel to the client concurrency, and raising --ctx-size since it is divided across slots, gave me:

config output tok/s mean TTFT mean ITL
p1, c=1 7.32 3,175 ms 119 ms
p2, c=2 10.86 3,603 ms 165 ms
p4, c=4 13.74 3,853 ms 266 ms

+87.7% throughput at 4-way with TTFT roughly flat. ITL degrades as the GPU time-shares, so it is a real trade, but it is a very different curve from the queuing case.

3. Stop order matters on the RDMA transport

If the workerโ€™s rpc-server is killed while the client is still tearing down (freeing remote buffers over RPC), the client aborts:

ggml-rpc.cpp: Remote RPC server crashed or returned malformed response
โ†’ ggml_abort() (from ggml_backend_rpc_buffer_free_buffer)

Deterministic here: kill client then immediately kill server = 11/11 crash; wait for the client PID to disappear first = 0/4. TCP handles the same drop gracefully via recv()==0 (6/6 clean). Do not use โ€œAPI port closedโ€ as the done signal โ€” the port closes before buffer teardown finishes, so the check has to be on the PID.

Caveats on my side

My prompt processing is lower than yours (171 vs 213 tok/s) โ€” I did not use --ubatch-size 2048, so your +30% tip is something I would have liked to test. Both machines have since gone back, so I cannot re-run. Also UD-IQ1_M vs your UD-IQ1_S, and my context was far smaller than 256K.

Full numbers, figures and scripts (including the PID-based stop script): GitHub - nabe2030/dgx-spark-2node-rpc: Running a 228.5GB GLM-5.2 GGUF across two DGX Sparks with llama.cpp RPC (RDMA/TCP A/B, concurrency sweep) ยท GitHub

Is it something you could upstream to VLLM?
Or can you share the patches with us?

FYI, idk if you heard of the Colibri project, but someone on their GH got 3.33 tok/sec running GLM 5.2 4-bit on a single Spark, with the remaining layers streaming off the SSD.
It lets you load different model shards from different drives, so if you use your 2nd Spark as a ramdisk over Connect-X7, thatโ€™s another ~120GB of shards read from something much faster than the SSD. Who knows, you might match the 8 tok/secs from the OP while being at Q4 instead of Q1.