Exciting new quantized versions of GLM-5 REAP (50%) are now available on Hugging Face, courtesy of 0xSero:
-
πΉ GLM-5-REAP-50pct-UD-IQ2_M-GGUF β Ultra-dynamic IQ2_M quantization
-
πΉ GLM-5-REAP-50pct-UD-IQ2_XXS-GGUF β Ultra-dynamic IQ2_XXS (extra compressed)
-
πΉ GLM-5-REAP-50pct-Q3_K_M-GGUF β Q3_K_M quantization for a balance of quality and size
Now, the big question for the community π
Iβm curious:
-
How many DGX Spark units would be needed to run the full GLM-5 REAP model efficiently?
-
Could any of these quantized variants (IQ2_M, IQ2_XXS, Q3_K_M) run comfortably on a single DGX Spark?
-
Has anyone already tested any of these GGUFs on a DGX Spark? What were your results β inference speed, memory usage, quality?
The IQ2_XXS and IQ2_M variants are the most aggressively compressed, so they might be the best candidates for single-node deployment.
Drop your benchmarks & configs in the comments.