Deepseek-v4-Flash 0731 GGUF (NEW model)

To run DeepSeek-V4-Flash-0731 in full precision lossless, run Q8 (UD-Q8_K_XL), which is 162GB and only 7GB bigger than Q4 (UD-Q4_K_XL).

DeepSeek-V4 and DeepSeek-V4-Flash-0731 are new open-weight models from DeepSeek. The Flash variant has 284B parameters (13B active) and V4-Pro has 1.6T parameters (49B active). DeepSeek-V4-Flash-0731 is the July 31 update to the Flash model, delivering the best performance for its size and outperforming V4-Pro (Preview). Designed for coding, agentic and chat workflows with a 1M context window.

Above is copied from Unsloth page. Disappointed I havo no time to test it now.
Excited.

Just downloaded the model, seems MTP is not yet compatible on the community vLLM. Anyone else see this as well?

Someone posted a NVFP4 variant: MJPansa/DeepSeek-V4-Flash-0731-NVFP4 · Hugging Face