Ruuning Kimi K3 across two Nvidia 8xB200 nodes using vllm

In this post, webenchmarked Moonshot AI’s Kimi K3 running on two NVIDIA B200 nodes (16 GPUs total) using vllm. The benchmarks cover three common inference scenarios:

  • Prompt-heavy (long input, short output)
  • Decode-heavy (short input, long output)
  • Balanced (equal input and output lengths)

Prompt-heavy

vllm bench serve \
  --model moonshotai/Kimi-K3 \
  --dataset-name random \
  --random-input-len 8000 \
  --random-output-len 1000 \
  --request-rate 10000 \
  --num-prompts 16 \
  --ignore-eos
============ Serving Benchmark Result ============
Successful requests:                     16
Failed requests:                         0
Request rate configured (RPS):           10000.00
Benchmark duration (s):                  51.18
Total input tokens:                      129408
Total generated tokens:                  16000
Request throughput (req/s):              0.31
Output token throughput (tok/s):         312.63
Peak output token throughput (tok/s):    258.00
Peak concurrent requests:                16.00
Total token throughput (tok/s):          2841.22
---------------Time to First Token----------------
Mean TTFT (ms):                          10619.92
Median TTFT (ms):                        10662.66
P99 TTFT (ms):                           13609.15
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          34.01
Median TPOT (ms):                        33.51
P99 TPOT (ms):                           43.45
---------------Inter-token Latency----------------
Mean ITL (ms):                           235.88
Median ITL (ms):                         64.08
P99 ITL (ms):                            2254.20
---------------Speculative Decoding---------------
Acceptance rate (%):                     14.81
Acceptance length:                       2.04
Drafts:                                  7860
Draft tokens:                            55020
Accepted tokens:                         8147
Per-position acceptance (%):
  Position 0:                            48.64
  Position 1:                            25.64
  Position 2:                            13.68
  Position 3:                            7.68
  Position 4:                            4.13
  Position 5:                            2.35
  Position 6:                            1.53
==================================================

Decode-heavy

vllm bench serve \
  --model moonshotai/Kimi-K3 \
  --dataset-name random \
  --random-input-len 1000 \
  --random-output-len 8000 \
  --request-rate 10000 \
  --num-prompts 16 \
  --ignore-eos

Output:

============ Serving Benchmark Result ============
Successful requests:                     16
Failed requests:                         0
Request rate configured (RPS):           10000.00
Benchmark duration (s):                  338.50
Total input tokens:                      17408
Total generated tokens:                  128000
Request throughput (req/s):              0.05
Output token throughput (tok/s):         378.14
Peak output token throughput (tok/s):    257.00
Peak concurrent requests:                16.00
Total token throughput (tok/s):          429.57
---------------Time to First Token----------------
Mean TTFT (ms):                          2056.04
Median TTFT (ms):                        2156.18
P99 TTFT (ms):                           2158.45
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          31.93
Median TPOT (ms):                        34.85
P99 TPOT (ms):                           41.66
---------------Inter-token Latency----------------
Mean ITL (ms):                           1809.94
Median ITL (ms):                         62.56
P99 ITL (ms):                            65.80
---------------Speculative Decoding---------------
Acceptance rate (%):                     13.66
Acceptance length:                       1.96
Drafts:                                  65439
Draft tokens:                            458073
Accepted tokens:                         62560
Per-position acceptance (%):
  Position 0:                            42.35
  Position 1:                            21.77
  Position 2:                            12.61
  Position 3:                            7.83
  Position 4:                            5.00
  Position 5:                            3.51
  Position 6:                            2.53
==================================================


Balanced

vllm bench serve \
  --model nvidia/Qwen3.6-27B-NVFP4 \
  --dataset-name random \
  --random-input-len 1000 \
  --random-output-len 1000 \
  --request-rate 10000 \
  --num-prompts 16 \
  --ignore-eos

Output:

============ Serving Benchmark Result ============
Successful requests:                     16
Failed requests:                         0
Request rate configured (RPS):           10000.00
Benchmark duration (s):                  48.54
Total input tokens:                      16000
Total generated tokens:                  16000
Request throughput (req/s):              0.33
Output token throughput (tok/s):         329.64
Peak output token throughput (tok/s):    272.00
Peak concurrent requests:                16.00
Total token throughput (tok/s):          659.28
---------------Time to First Token----------------
Mean TTFT (ms):                          1975.25
Median TTFT (ms):                        2069.99
P99 TTFT (ms):                           2072.92
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          44.87
Median TPOT (ms):                        45.59
P99 TPOT (ms):                           46.49
---------------Inter-token Latency----------------
Mean ITL (ms):                           62.38
Median ITL (ms):                         62.67
P99 ITL (ms):                            64.32
---------------Speculative Decoding---------------
Acceptance rate (%):                     5.61
Acceptance length:                       1.39
Drafts:                                  11483
Draft tokens:                            80381
Accepted tokens:                         4508
Per-position acceptance (%):
  Position 0:                            24.98
  Position 1:                            8.61
  Position 2:                            3.23
  Position 3:                            1.41
  Position 4:                            0.66
  Position 5:                            0.24
  Position 6:                            0.12
==================================================

Kimi K3 demonstrates strong throughput across different inference workloads.