ok, I got a baseline to start with.
I compiled the latest pytorch, torchaudio, torchvision for TargetArch 8.7.
Got VLLM to compile, too.
It is compiling the CUDA graphs.
It seems that I am getting somewhere.
Prefill is abysmal, the system is running at 50W power setting. Question to the experts: are there additional options to further speed up the inference (I know, the context length is way too high, this was just a first test.)
(test) saskia@orin-agx:~/build/vllm/dist$ vllm serve /home/saskia/models/Qwen2.5-7B-Instruct-AWQ --gpu-memory-utilization 0.5
W0729 21:46:13.790000 92517 torch/_opaque_base.py:6] torch._opaque_base is deprecated, use torch._custom_class_base instead
W0729 21:46:13.792000 92517 torch/_library/opaque_object.py:288] register_opaque_type is deprecated, use register_custom_class instead
W0729 21:46:13.793000 92517 torch/_library/opaque_object.py:211] typ=‘value’ is deprecated, use typ=‘constant’ instead
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:345]
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:345] █ █ █▄ ▄█
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:345] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.26.1rc1.dev103+g381b69162.d20260729
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:345] █▄█▀ █ █ █ █ model /home/saskia/models/Qwen2.5-7B-Instruct-AWQ
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:345] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:345]
(APIServer pid=92517) INFO 07-29 21:46:32 [api_utils.py:273] non-default args: {‘model_tag’: ‘/home/saskia/models/Qwen2.5-7B-Instruct-AWQ’, ‘model’: ‘/home/saskia/models/Qwen2.5-7B-Instruct-AWQ’, ‘gpu_memory_utilization’: 0.5}
(APIServer pid=92517) INFO 07-29 21:46:53 [model.py:638] Resolved architecture: Qwen2ForCausalLM
(APIServer pid=92517) INFO 07-29 21:46:53 [model.py:1875] Using max model len 32768
(APIServer pid=92517) INFO 07-29 21:46:54 [vllm.py:1118] Asynchronous scheduling is enabled.
(APIServer pid=92517) INFO 07-29 21:46:54 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=[‘native’], fused_add_rms_norm=[‘native’])
W0729 21:47:03.768000 92634 torch/_opaque_base.py:6] torch._opaque_base is deprecated, use torch._custom_class_base instead
W0729 21:47:03.769000 92634 torch/_library/opaque_object.py:288] register_opaque_type is deprecated, use register_custom_class instead
W0729 21:47:03.771000 92634 torch/_library/opaque_object.py:211] typ=‘value’ is deprecated, use typ=‘constant’ instead
(EngineCore pid=92634) INFO 07-29 21:47:18 [core.py:121] Initializing a V1 LLM engine (v0.26.1rc1.dev103+g381b69162.d20260729) with config: model=‘/home/saskia/models/Qwen2.5-7B-Instruct-AWQ’, speculative_config=None, tokenizer=‘/home/saskia/models/Qwen2.5-7B-Instruct-AWQ’, skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.float16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=auto_awq, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend=‘auto’, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser=‘’, reasoning_parser_plugin=‘’, enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode=‘warn’, jit_monitor_verbose=False), seed=0, served_model_name=/home/saskia/models/Qwen2.5-7B-Instruct-AWQ, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={‘mode’: <CompilationMode.VLLM_COMPILE: 3>, ‘debug_dump_path’: None, ‘cache_dir’: ‘’, ‘compile_cache_save_format’: ‘binary’, ‘backend’: ‘inductor’, ‘custom_ops’: [‘none’], ‘ir_enable_torch_wrap’: True, ‘splitting_ops’: [‘vllm::unified_attention_with_output’, ‘vllm::unified_mla_attention_with_output’, ‘vllm::mamba_mixer2’, ‘vllm::mamba_mixer’, ‘vllm::short_conv’, ‘vllm::linear_attention’, ‘vllm::qwen_gdn_attention_core’, ‘vllm::gdn_attention_core_xpu’, ‘vllm::olmo_hybrid_gdn_full_forward’, ‘vllm::kda_attention’, ‘vllm::sparse_attn_indexer’, ‘vllm::rocm_aiter_sparse_attn_indexer’, ‘vllm::deepseek_v4_attention’, ‘vllm::hpc_rope_norm_forward’, ‘vllm::unified_kv_cache_update’, ‘vllm::unified_mla_kv_cache_update’], ‘compile_mm_encoder’: False, ‘cudagraph_mm_encoder’: False, ‘encoder_cudagraph_token_budgets’: , ‘encoder_cudagraph_max_vision_items_per_batch’: 0, ‘encoder_cudagraph_max_frames_per_batch’: None, ‘compile_sizes’: , ‘compile_ranges_endpoints’: [2048], ‘inductor_compile_config’: {‘enable_auto_functionalized_v2’: False, ‘combo_kernels’: True, ‘benchmark_combo_kernel’: True}, ‘inductor_passes’: {}, ‘cudagraph_mode’: <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, ‘cudagraph_num_of_warmups’: 1, ‘cudagraph_capture_sizes’: [1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 96, 104, 112, 120, 128, 136, 144, 152, 160, 168, 176, 184, 192, 200, 208, 216, 224, 232, 240, 248, 256, 272, 288, 304, 320, 336, 352, 368, 384, 400, 416, 432, 448, 464, 480, 496, 512], ‘cudagraph_copy_inputs’: False, ‘cudagraph_specialize_lora’: True, ‘use_inductor_graph_partition’: False, ‘pass_config’: {‘fuse_norm_quant’: False, ‘fuse_act_quant’: False, ‘fuse_attn_quant’: False, ‘enable_sp’: False, ‘fuse_gemm_comms’: False, ‘fuse_allreduce_rms’: False, ‘enable_qk_norm_rope_fusion’: False, ‘fuse_rope_kvcache_cat_mla’: False, ‘fuse_act_padding’: False, ‘fuse_qk_norm_rope_kvcache’: False}, ‘max_cudagraph_capture_size’: 512, ‘dynamic_shapes_config’: {‘type’: <DynamicShapesType.BACKED: ‘backed’>, ‘evaluate_guards’: False, ‘assume_32_bit_indexing’: False}, ‘local_cache_dir’: None, ‘fast_moe_cold_start’: False, ‘static_all_moe_layers’: }, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=[‘native’], fused_add_rms_norm=[‘native’]), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend=‘auto’, linear_backend=‘auto’)
(EngineCore pid=92634) INFO 07-29 21:47:18 [parallel_state.py:1640] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.178.41:45063 backend=nccl
(EngineCore pid=92634) INFO 07-29 21:47:18 [parallel_state.py:1977] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A
(EngineCore pid=92634) INFO 07-29 21:47:18 [gpu_worker.py:385] Using V2 Model Runner
(EngineCore pid=92634) INFO 07-29 21:47:19 [model_runner.py:295] Loading model from scratch…
(EngineCore pid=92634) INFO 07-29 21:47:19 [auto_awq.py:473] Using MarlinLinearKernel for AutoAWQMarlinLinearMethod
(EngineCore pid=92634) INFO 07-29 21:47:21 [cuda.py:482] Using FLASH_ATTN attention backend out of potential backends: [‘FLASH_ATTN’, ‘FLASHINFER’, ‘TRITON_ATTN’, ‘FLEX_ATTENTION’].
(EngineCore pid=92634) INFO 07-29 21:47:21 [flash_attn.py:789] Using FlashAttention version 2
(EngineCore pid=92634) INFO 07-29 21:47:22 [weight_utils.py:867] Filesystem type for checkpoints: EXT4. Checkpoint size: 5.19 GiB. Available RAM: 51.72 GiB.
(EngineCore pid=92634) INFO 07-29 21:47:22 [weight_utils.py:890] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
Loading safetensors checkpoint shards: 0% Completed | 0/2 [00:00<?, ?it/s]
Loading safetensors checkpoint shards: 50% Completed | 1/2 [00:01<00:01, 1.19s/it]
Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:01<00:00, 1.22it/s]
Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:01<00:00, 1.14it/s]
(EngineCore pid=92634)
(EngineCore pid=92634) INFO 07-29 21:47:24 [default_loader.py:430] Loading weights took 1.83 seconds
(EngineCore pid=92634) INFO 07-29 21:47:30 [model_runner.py:316] Model loading took 5.29 GiB and 10.699937 seconds
(EngineCore pid=92634) INFO 07-29 21:47:30 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
(EngineCore pid=92634) INFO 07-29 21:47:47 [backends.py:1094] Using cache directory: /home/saskia/.cache/vllm/torch_compile_cache/bb421a0b5e/rank_0_0/backbone for vLLM’s torch.compile
(EngineCore pid=92634) INFO 07-29 21:47:47 [backends.py:1155] Dynamo bytecode transform time: 15.79 s
(EngineCore pid=92634) INFO 07-29 21:48:02 [backends.py:393] Compiling a graph for compile range (1, 2048) takes 14.79 s
(EngineCore pid=92634) INFO 07-29 21:48:13 [backends.py:920] collected artifacts: 29 entries, 3 artifacts, 4049055 bytes total
(EngineCore pid=92634) INFO 07-29 21:48:13 [decorators.py:708] saved AOT compiled function to /home/saskia/.cache/vllm/torch_compile_cache/torch_aot_compile/060a295c7be4b828924994b87f9fcc9057b917b011652a16f16a31b22595e381/rank_0_0/model
(EngineCore pid=92634) INFO 07-29 21:48:13 [monitor.py:53] torch.compile took 41.91 s in total
(EngineCore pid=92634) INFO 07-29 21:48:13 [monitor.py:81] Initial profiling/warmup run took 0.18 s
(EngineCore pid=92634) INFO 07-29 21:48:17 [gpu_worker.py:563] Available KV cache memory: 22.32 GiB
(EngineCore pid=92634) INFO 07-29 21:48:17 [kv_cache_utils.py:2214] GPU KV cache size: 417,872 tokens
(EngineCore pid=92634) INFO 07-29 21:48:17 [kv_cache_utils.py:2215] Maximum concurrency for 32,768 tokens per request: 12.75x
(EngineCore pid=92634) INFO 07-29 21:48:23 [cutedsl_warmup.py:105] Skipping CuTeDSL warmup because no compile units were requested.
Capturing CUDA graphs (PIECEWISE): 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 51/51 [00:17<00:00, 2.88it/s]
Capturing CUDA graphs (FULL): 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 35/35 [00:06<00:00, 5.53it/s]
(EngineCore pid=92634) INFO 07-29 21:48:48 [model_runner.py:776] Graph capturing finished in 25 secs, took 1.17 GiB
(EngineCore pid=92634) INFO 07-29 21:48:48 [gpu_worker.py:789] Free memory on device (56.94/61.35 GiB) on startup. Desired GPU memory utilization is (0.5, 30.67 GiB). Actual usage is 7.32 GiB for consumed memory (weights + non-torch), 1.04 GiB for peak activation, and 1.17 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with --kv-cache-memory=22545779200 (21.0 GiB) to fit into requested memory, or --kv-cache-memory=50754564608 (47.27 GiB) to fully utilize gpu memory. Current kv cache memory in use is 22.32 GiB.
(EngineCore pid=92634) INFO 07-29 21:51:59 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
(EngineCore pid=92634) INFO 07-29 21:52:00 [core.py:348] init engine (profile, create kv cache, warmup model) took 270.15 s (compilation: 41.91 s)
(EngineCore pid=92634) INFO 07-29 21:52:01 [vllm.py:1118] Asynchronous scheduling is enabled.
(EngineCore pid=92634) INFO 07-29 21:52:01 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=[‘native’], fused_add_rms_norm=[‘native’])
(APIServer pid=92517) INFO 07-29 21:52:01 [api_server.py:676] Supported tasks: [‘generate’]
(APIServer pid=92517) WARNING 07-29 21:52:02 [model.py:1629] Default vLLM sampling parameters have been overridden by the model’s generation_config.json: {'repetition_penalty': 1.05, 'temperature': 0.7, 'top_k': 20, 'top_p': 0.8}. If this is not intended, please relaunch vLLM instance with --generation-config vllm.
(APIServer pid=92517) INFO 07-29 21:52:03 [hf.py:540] Detected the chat template content format to be ‘string’. You can set --chat-template-content-format to override this.
(APIServer pid=92517) INFO 07-29 21:52:03 [api_server.py:680] Starting vLLM server on http://0.0.0.0:8000
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:37] Available routes are:
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /docs, Methods: HEAD, GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /load, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /version, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /health, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /metrics, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /tokenize, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /detokenize, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/models, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /ping, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /ping, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /invocations, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/responses, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/completions, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/messages, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /generative_scoring, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/completions/render, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
(APIServer pid=92517) INFO 07-29 21:52:03 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
(APIServer pid=92517) INFO: Started server process [92517]
(APIServer pid=92517) INFO: Waiting for application startup.
(APIServer pid=92517) INFO: Application startup complete.
(APIServer pid=92517) INFO 07-29 21:55:24 [loggers.py:310] Engine 000: Avg prompt throughput: 6.0 tokens/s, Avg generation throughput: 2.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=92517) INFO 07-29 21:55:34 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 29.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.1%, Prefix cache hit rate: 0.0%
(APIServer pid=92517) INFO: 127.0.0.1:50056 - “POST /v1/chat/completions HTTP/1.1” 200 OK
(APIServer pid=92517) INFO 07-29 21:55:44 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.3 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=92517) INFO 07-29 21:55:54 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=92517) INFO 07-29 22:00:14 [loggers.py:310] Engine 000: Avg prompt throughput: 8.3 tokens/s, Avg generation throughput: 12.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.1%, Prefix cache hit rate: 10.1%
(APIServer pid=92517) INFO 07-29 22:00:24 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 29.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.1%, Prefix cache hit rate: 10.1%
(APIServer pid=92517) INFO: 127.0.0.1:33114 - “POST /v1/chat/completions HTTP/1.1” 200 OK
(APIServer pid=92517) INFO 07-29 22:00:34 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 9.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 10.1%
(APIServer pid=92517) INFO 07-29 22:00:44 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 10.1%
(APIServer pid=92517) INFO 07-29 22:02:44 [loggers.py:310] Engine 000: Avg prompt throughput: 0.3 tokens/s, Avg generation throughput: 3.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 43.4%
(APIServer pid=92517) INFO 07-29 22:02:54 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 29.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.1%, Prefix cache hit rate: 43.4%
(APIServer pid=92517) INFO 07-29 22:03:04 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 29.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.2%, Prefix cache hit rate: 43.4%
(APIServer pid=92517) INFO: 127.0.0.1:46764 - “POST /v1/chat/completions HTTP/1.1” 200 OK
(APIServer pid=92517) INFO 07-29 22:03:14 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 25.3 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 43.4%
(APIServer pid=92517) INFO 07-29 22:03:24 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 43.4%
Summary
This text will be hidden