Hi all,
A question we keep getting from teams that run on a single workstation GPU: why use NIM instead of just running Ollama? So we measured it. Llama 3.1 8B Instruct, one RTX 5090, NIM 2.0.12, from one request at a time up to 128, with Ollama as the reference point. We also checked whether the faster configurations answer worse.
TL;DR
- One user: Ollama’s 4-bit build is faster (172–183 vs 87 tok/s), because it reads about 3.3× fewer bytes per token. NIM itself adds nothing on top of the vLLM 0.27.1 it ships with: same weights, same speed, the same 50 answers byte for byte.
- Once requests overlap, NIM pulls ahead. The crossover is between 2 and 4 concurrent requests. NIM is 3.5× ahead at 8 and 7.4× at 128 (5,458 vs 741 tok/s, against the fastest Ollama configuration we could build).
- NIM held the MLPerf Inference server latency target up to 128 concurrent requests. None of the Ollama configurations held it past one.
- NIM’s default on this card is FP8. It gives another 1.5× over bf16 (9,062 tok/s at 128). MMLU showed no measurable change; GSM8K was 1.5 points lower, which we could not separate from zero.
All throughput numbers: synthetic chat requests, 200 tokens in / 200 out, closed loop, one engine on the GPU at a time.
One RTX 5090 · Llama 3.1 8B Instruct · synthetic chat 200/200 · closed loop · NIM 2.0.12 (vLLM 0.27.1) / Ollama 0.34.4
Run it
This is the configuration we measured (FP8, NIM’s own choice on this card):
docker run -d --gpus all -p 8000:8000 \
-e NGC_API_KEY \
-e NIM_MODEL_PROFILE=c4789f7af56c770c1c88b73da666886365534d6980b6b922b41fd97036c77d73 \
-e NIM_MAX_MODEL_LEN=8192 \
-e VLLM_USE_V2_MODEL_RUNNER=0 \
-v ~/.cache/nim:/opt/nim/.cache \
nvcr.io/nim/meta/llama-3.1-8b-instruct:2.0.12
For bf16, use profile 092ed4213624e774d24cdaf84e3b6222839bab2008a21d3c214ab46626366f90. On our Docker Desktop / WSL2 host this image would not start without VLLM_USE_V2_MODEL_RUNNER=0 (RuntimeError: UVA is not available).
Results
Total throughput, tok/s (Llama 3.1 8B, synthetic chat 200/200)
| Concurrent requests | 1 | 8 | 32 | 128 |
|---|---|---|---|---|
| Ollama, 4-bit, 16 slots (tuned) | 172 | 175 | 755 | 741 |
| Ollama, 16-bit, 8 slots (tuned) | 79 | 317 | 390 | 402 |
| NIM, bf16 | 87 | 607 | 2,262 | 5,458 |
| NIM, FP8 (default) | 151 | 1,099 | 3,817 | 8,713 |
p99 time to first token at 128 concurrent requests: NIM bf16 1.84 s, NIM FP8 0.75 s, tuned 4-bit Ollama 31.9 s. The MLPerf server target is p99 TTFT ≤ 2 s and p99 TPOT ≤ 100 ms.
Measured with AIPerf. FP8 figures in this table are from the same run as the rest. A later recheck in a clean window measured 9,062 tok/s at 128, 1.50× bf16 NIM in the same window.
One RTX 5090 · Llama 3.1 8B Instruct · synthetic chat 200/200 · closed loop · NIM 2.0.12 (vLLM 0.27.1) / Ollama 0.34.4 · hollow points: p99 TTFT > 2 s
Why NIM pulls ahead
- For one user, precision decides the speed. Once requests overlap, the engine does. At matched 16-bit precision, Ollama and NIM are within 10% for one user (79 vs 87 tok/s). At 128 the gap is 13.6× (402 vs 5,458).
- The engine is the difference. NIM’s engine batches requests continuously and allocates the KV cache in pages. Ollama pre-allocates a full context per slot, so a 32 GB card caps the slot count, and extra requests only lengthen the queue.
- What NIM adds on top of vLLM is the configuration. It picks a validated profile for the card (FP8 here) and ships it pinned in one image. For this model, NIM 2.0.12 ships vLLM profiles only (support matrix). So for one user it is exactly as fast as the vLLM inside it, and it costs nothing extra.
Does faster mean worse answers?
We ran the four configurations on the same task sets (lm-evaluation-harness 0.4.13, temperature 0, identical requests) and compared each with bf16 NIM item by item.
| vs bf16 NIM (points, 95% CI) | FP8 NIM | 16-bit Ollama | 4-bit Ollama |
|---|---|---|---|
| MMLU, 2,850-question sample (bf16 NIM: 68.9%) | −0.04 [−0.81, +0.74] | +0.11 [−0.53, +0.74] | −0.25 [−1.23, +0.74] |
| GSM8K, all 1,319 (bf16 NIM: 85.6%) | −1.52 [−3.11, 0.00] | −1.29 [−2.88, +0.30] | −2.43 [−4.40, −0.53] |
General knowledge held everywhere. On multi-step math, the 4-bit build lost a measurable 2.4 points. Running 32 requests at once did not change FP8’s accuracy (84.4% vs 84.6% on GSM8K).
Things that will bite you
- Ollama picks 1 parallel slot by default on this card. Set
OLLAMA_NUM_PARALLELyourself, and check thatollama psshows 100% GPU. Slots × context has to fit in VRAM. - Pin
NIM_MODEL_PROFILEby its 64-character id. The display name with the suffix in parentheses is not accepted as an id. - Use
127.0.0.1, notlocalhost. On our Windows host,localhostcost about 2 s per request before the connection was made.
A note on our March post
We re-examined our first post with the arm-health check used throughout this one: a single-stream generation rate implies a memory bandwidth, and a healthy GPU decoder lands near the card’s peak. The 7.3× ratio that originally led the March 2026 Vol.1 compared two arms whose generation rates, converted to effective memory bandwidth, sit at 75% and 3.4% of this card’s peak — the denominator arm was not running at GPU speed. We do not know what the March 2026 machine was doing, and we do not claim to. The ratio therefore belongs to that configuration. Today’s 7.4× is close to it only by coincidence: it compares 128 concurrent requests with both arms healthy. The full re-examination is in the repo.
A few honest notes
- One RTX 5090 on Docker Desktop / WSL2. The card is not on NIM’s verified-GPU list. On data-center GPUs, and for models where NIM ships hardware-specific profiles (NVIDIA’s own NIM-off/NIM-on comparison), the picture can differ.
- The prompts are synthetic. This is a latency-target measurement, not a statement of how many users a system serves.
- “4-bit Ollama vs bf16 NIM” mixes precision and engine. That is why the matched 16-bit row is there.
Everything is in the repo
Pre-registrations, raw results, analysis scripts and the full boundary of every run, including the FP8 recheck, are in the Vol.1 README of NV-benchmark (tag vol1). Every number in this post has a row in the README’s claim-to-evidence table pointing to the file it comes from.
If you run this on another card (RTX PRO 6000, DGX Spark, H100), we’d love to compare notes.
QuanTuring Inc. · member of the NVIDIA Inception Program

