# Introducing NVFP4 for Efficient and Accurate Low-Precision Inference

**URL:** <https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069>\
**Category:** Technical Blog\
**Created:** [June 24, 2025, 4:18pm UTC](https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069 "2025-06-24T16:18:52Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jwitsoe](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jwitsoe/32/16591_2.png) [@jwitsoe](https://forums.developer.nvidia.com/u/jwitsoe)\
**Post date:** [June 24, 2025, 4:18pm UTC](https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069/1 "2025-06-24T16:18:52Z")

</div>

Originally published at: [Introducing NVFP4 for Efficient and Accurate Low-Precision Inference | NVIDIA Technical Blog](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)  
   
To get the most out of AI, optimizations are critical. When developers think about optimizing AI models for inference, model compression techniques—such as quantization, distillation, and pruning—typically come to mind. The most common of the three, without a doubt, is quantization. This is typically due to its post-optimization task-specific accuracy performance and broad choice of…

---

<div class="post-metadata">

**Author:** ![TomNVIDIA](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/tomnvidia/32/14181_2.png) [@TomNVIDIA](https://forums.developer.nvidia.com/u/TomNVIDIA)\
**Post date:** [June 25, 2025, 5:35am UTC](https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069/2 "2025-06-25T05:35:45Z")

</div>



---

<div class="post-metadata">

**Author:** ![malkevich.alex](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@malkevich.alex](https://forums.developer.nvidia.com/u/malkevich.alex)\
**Post date:** [October 28, 2025, 12:08am UTC](https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069/4 "2025-10-28T00:08:14Z")

</div>

vLLM does not really support NVFP4 still. I’m unable to run NVFP4/Qwen3-Coder-30B-A3B-Instruct-FP4 on my DGX Spark using nightly vllm image.

---

<div class="post-metadata">

**Author:** ![nvidiasikkv](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@nvidiasikkv](https://forums.developer.nvidia.com/u/nvidiasikkv)\
**Post date:** [November 26, 2025, 11:59am UTC](https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069/5 "2025-11-26T11:59:58Z")

</div>

Why is there a sign bit on the scale factor? It seems redundant.

---

<div class="post-metadata">

**Author:** ![tgm1024](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@tgm1024](https://forums.developer.nvidia.com/u/tgm1024)\
**Post date:** [January 26, 2026, 6:08pm UTC](https://forums.developer.nvidia.com/t/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/337069/6 "2026-01-26T18:08:26Z")

</div>

Please don’t unnecessarily add animation to your images. There’s a place for such things, but not with those. You already have left-to-right flow as part of the image……causing the image to have parts disappear and reappear greatly hampers the reading of it.
