Some blackwell optimizations coming for llama.cpp
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4: DGX Spark vs. RTX PRO 6000 Blackwell Performance | 0 | 473 | August 12, 2026 | |
| Qwen3.6-27B AWQ INT4 on DGX Spark (GB10) — only 1.8-4.9 tok/s decode with 285k token prompt, how to improve? | 6 | 1558 | May 29, 2026 | |
| DGX Spark, Nemotron3, and NVFP4: Getting to 65+ tps | 14 | 2597 | December 22, 2025 | |
| NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 | 89 | 11913 | March 31, 2026 | |
| Does Qwen3.5-35B-A3B on GB10 leave a lot of performance on the table? | 40 | 7466 | March 16, 2026 | |
| We unlocked NVFP4 on the DGX Spark: 20% faster than AWQ! | 144 | 11604 | March 14, 2026 | |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 31 | 2976 | June 10, 2026 | |
| [Benchmark] nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 | 5 | 1726 | May 1, 2026 | |
| Nemotron-3-Super-120B at 20-22 tok/s Super Special Recipe | 4 | 1114 | May 30, 2026 | |
| NVIDIA folks -- where is this promised nvfp4 speedup? | 27 | 3348 | March 26, 2026 |