In this post, webenchmarked Moonshot AI’s Kimi K3 running on two NVIDIA B200 nodes (16 GPUs total) using vllm. The benchmarks cover three common inference scenarios:
- Prompt-heavy (long input, short output)
- Decode-heavy (short input, long output)
- Balanced (equal input and output lengths)
Prompt-heavy
vllm bench serve \
--model moonshotai/Kimi-K3 \
--dataset-name random \
--random-input-len 8000 \
--random-output-len 1000 \
--request-rate 10000 \
--num-prompts 16 \
--ignore-eos
============ Serving Benchmark Result ============
Successful requests: 16
Failed requests: 0
Request rate configured (RPS): 10000.00
Benchmark duration (s): 51.18
Total input tokens: 129408
Total generated tokens: 16000
Request throughput (req/s): 0.31
Output token throughput (tok/s): 312.63
Peak output token throughput (tok/s): 258.00
Peak concurrent requests: 16.00
Total token throughput (tok/s): 2841.22
---------------Time to First Token----------------
Mean TTFT (ms): 10619.92
Median TTFT (ms): 10662.66
P99 TTFT (ms): 13609.15
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 34.01
Median TPOT (ms): 33.51
P99 TPOT (ms): 43.45
---------------Inter-token Latency----------------
Mean ITL (ms): 235.88
Median ITL (ms): 64.08
P99 ITL (ms): 2254.20
---------------Speculative Decoding---------------
Acceptance rate (%): 14.81
Acceptance length: 2.04
Drafts: 7860
Draft tokens: 55020
Accepted tokens: 8147
Per-position acceptance (%):
Position 0: 48.64
Position 1: 25.64
Position 2: 13.68
Position 3: 7.68
Position 4: 4.13
Position 5: 2.35
Position 6: 1.53
==================================================
Decode-heavy
vllm bench serve \
--model moonshotai/Kimi-K3 \
--dataset-name random \
--random-input-len 1000 \
--random-output-len 8000 \
--request-rate 10000 \
--num-prompts 16 \
--ignore-eos
Output:
============ Serving Benchmark Result ============
Successful requests: 16
Failed requests: 0
Request rate configured (RPS): 10000.00
Benchmark duration (s): 338.50
Total input tokens: 17408
Total generated tokens: 128000
Request throughput (req/s): 0.05
Output token throughput (tok/s): 378.14
Peak output token throughput (tok/s): 257.00
Peak concurrent requests: 16.00
Total token throughput (tok/s): 429.57
---------------Time to First Token----------------
Mean TTFT (ms): 2056.04
Median TTFT (ms): 2156.18
P99 TTFT (ms): 2158.45
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 31.93
Median TPOT (ms): 34.85
P99 TPOT (ms): 41.66
---------------Inter-token Latency----------------
Mean ITL (ms): 1809.94
Median ITL (ms): 62.56
P99 ITL (ms): 65.80
---------------Speculative Decoding---------------
Acceptance rate (%): 13.66
Acceptance length: 1.96
Drafts: 65439
Draft tokens: 458073
Accepted tokens: 62560
Per-position acceptance (%):
Position 0: 42.35
Position 1: 21.77
Position 2: 12.61
Position 3: 7.83
Position 4: 5.00
Position 5: 3.51
Position 6: 2.53
==================================================
Balanced
vllm bench serve \
--model nvidia/Qwen3.6-27B-NVFP4 \
--dataset-name random \
--random-input-len 1000 \
--random-output-len 1000 \
--request-rate 10000 \
--num-prompts 16 \
--ignore-eos
Output:
============ Serving Benchmark Result ============
Successful requests: 16
Failed requests: 0
Request rate configured (RPS): 10000.00
Benchmark duration (s): 48.54
Total input tokens: 16000
Total generated tokens: 16000
Request throughput (req/s): 0.33
Output token throughput (tok/s): 329.64
Peak output token throughput (tok/s): 272.00
Peak concurrent requests: 16.00
Total token throughput (tok/s): 659.28
---------------Time to First Token----------------
Mean TTFT (ms): 1975.25
Median TTFT (ms): 2069.99
P99 TTFT (ms): 2072.92
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 44.87
Median TPOT (ms): 45.59
P99 TPOT (ms): 46.49
---------------Inter-token Latency----------------
Mean ITL (ms): 62.38
Median ITL (ms): 62.67
P99 ITL (ms): 64.32
---------------Speculative Decoding---------------
Acceptance rate (%): 5.61
Acceptance length: 1.39
Drafts: 11483
Draft tokens: 80381
Accepted tokens: 4508
Per-position acceptance (%):
Position 0: 24.98
Position 1: 8.61
Position 2: 3.23
Position 3: 1.41
Position 4: 0.66
Position 5: 0.24
Position 6: 0.12
==================================================
Kimi K3 demonstrates strong throughput across different inference workloads.