New Quantized Models Drop: GLM-5 REAP 50% β€” How Many DGX Sparks Do You Need?

Exciting new quantized versions of GLM-5 REAP (50%) are now available on Hugging Face, courtesy of 0xSero:

Now, the big question for the community πŸ‘‡

I’m curious:

  • How many DGX Spark units would be needed to run the full GLM-5 REAP model efficiently?

  • Could any of these quantized variants (IQ2_M, IQ2_XXS, Q3_K_M) run comfortably on a single DGX Spark?

  • Has anyone already tested any of these GGUFs on a DGX Spark? What were your results β€” inference speed, memory usage, quality?

The IQ2_XXS and IQ2_M variants are the most aggressively compressed, so they might be the best candidates for single-node deployment.

Drop your benchmarks & configs in the comments.

  • How many DGX Spark units would be needed to run the full GLM-5 REAP model efficiently? 2
  • Could any of these quantized variants (IQ2_M, IQ2_XXS, Q3_K_M) run comfortably on a single DGX Spark? Only the first link works, no it won’t fit comfortably
  • Has anyone already tested any of these GGUFs on a DGX Spark? What were your results β€” inference speed, memory usage, quality? Not yet, we only have benchmarks for GLM 4.7 at Spark Arena cyankiwi/GLM-4.7-Flash-AWQ-4bit - Spark Arena Benchmark