# What prompt processing speed can one expect above 500k ctx?

**URL:** <https://forums.developer.nvidia.com/t/what-prompt-processing-speed-can-one-expect-above-500k-ctx/355437>\
**Category:** DGX Spark / GB10\
**Tags:** llama, jetson, benchmarks, nemotron\
**Created:** [December 22, 2025, 11:59am UTC](https://forums.developer.nvidia.com/t/what-prompt-processing-speed-can-one-expect-above-500k-ctx/355437 "2025-12-22T11:59:05Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![raphael.amorim](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/raphael.amorim/32/446733_2.png) [@raphael.amorim](https://forums.developer.nvidia.com/u/raphael.amorim)\
**Post date:** [December 24, 2025, 3:28pm UTC](https://forums.developer.nvidia.com/t/what-prompt-processing-speed-can-one-expect-above-500k-ctx/355437/4 "2025-12-24T15:28:58Z")

</div>

Some blackwell optimizations coming for llama.cpp

> [@Llama.cpp experimental native mxfp4 support for blackwell PR](https://forums.developer.nvidia.com/t/llama-cpp-experimental-native-mxfp4-support-for-blackwell-pr/355639):
>
> Looks like we’re about to get some extra PP performance in llama.cpp via this PR [https://github.com/ggml-org/llama.cpp/pull/17906](https://github.com/ggml-org/llama.cpp/pull/17906).

---

_[View the full topic](https://forums.developer.nvidia.com/t/what-prompt-processing-speed-can-one-expect-above-500k-ctx/355437)._
