Hi,
We are currently setting up Kubernetes (containerd) on an Ubuntu server and installing the NVIDIA GPU-Operator using Helm deployments as specified in the documentation at About the NVIDIA GPU Operator — NVIDIA GPU Operator. We are running video analysis on containers in Kubernetes, utilizing runtimeClasses set to nvidia.
Our models are converted from ONNX format to TensorRT format using the nvcr.io/nvidia/tensorrt:24.03-py3 image. However, when monitoring GPU Utilization through DCGM Exporter, nvitop, and nvidia-smi, we observed that we are not fully utilizing the NVIDIA A16 GPU. Despite experimenting with some Triton-specific parameters, we could not push utilization beyond a limit. We aim to make more effective use of the GPU.
In summary, we need information and support specifically related to Triton Inference Server in an environment managed by Kubernetes and GPU Operator.
Best regards,
Environment
TensorRT Version: tensorrt:24.03-py3
GPU Type: Nvidia A16
Nvidia Driver Version: 535.183.06
CUDA Version: 12.2
Operating System + Version: Ubuntu 20.04.6 LTS
Python Version (if applicable): 3.8
Baremetal or Container (if container which image + tag): nvcr.io/nvidia/tensorrt:24.03-py3, nvcr.io/nvidia/tritonserver:24.03-py3, KUbernetes (k3s) v1.29.4+k3s1