I use this (could be suboptimal, but it works quite fine for me):
Dockerfile:
ARG UBUNTU_VERSION=24.04
ARG CUDA_VERSION=13.3.1
ARG GCC_VERSION=14
ARG CUDA_ARCH=121a-real
# --------------------------------------
# Stage 1: build
# --------------------------------------
FROM nvidia/cuda:${CUDA_VERSION}-devel-ubuntu${UBUNTU_VERSION} AS builder
ARG GCC_VERSION
ARG CUDA_ARCH
RUN apt-get update && apt-get install -y --no-install-recommends \
openssh-client \
git \
wget \
gcc-${GCC_VERSION} \
g++-${GCC_VERSION} \
build-essential \
cmake \
&& rm -rf /var/lib/apt/lists/*
# Trust Github SSH key
RUN mkdir -p -m 0700 ~/.ssh && \
ssh-keyscan github.com >> ~/.ssh/known_hosts
WORKDIR /build
# Clone the repository via SSH, and drop its content in the working directory
RUN --mount=type=ssh \
git clone \
--branch master \
--depth 1 \
git@github.com:ggml-org/llama.cpp.git \
.
RUN cmake -B build \
-DGGML_CUDA=ON \
-DBUILD_SHARED_LIBS=OFF \
-DCMAKE_CUDA_ARCHITECTURES="${CUDA_ARCH}" \
-DCMAKE_BUILD_TYPE=Release
RUN cmake \
--build build \
--config Release \
--target llama-server \
--parallel $(nproc)
# --------------------------------------
# Stage 2: Runtime
# --------------------------------------
FROM nvidia/cuda:${CUDA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}
RUN apt-get update && apt-get install -y --no-install-recommends \
libgomp1 \
&& rm -rf /var/lib/apt/lists/*
COPY --from=builder /build/build/bin/llama-server /usr/local/bin/
WORKDIR /app
COPY ./entrypoint.sh .
RUN chmod +x ./entrypoint.sh
ENV LLAMA_ARG_HOST=0.0.0.0
ENV LLAMA_ARG_PORT=8000
ENV LLAMA_ARG_N_GPU_LAYERS=999
ENV LLAMA_ARG_N_GPU_LAYERS_DRAFT=999
ENTRYPOINT ["./entrypoint.sh"]
With a very simple entrypoint because not all properties are available as an ENV:
#!/bin/bash
args=()
if [ -n "${TEMP}" ]; then
args+=(--temperature "${TEMP}")
fi
if [ -n "${TOP_K}" ]; then
args+=(--top-k "${TOP_K}")
fi
if [ -n "${TOP_P}" ]; then
args+=(--top-p "${TOP_P}")
fi
if [ -n "${MIN_P}" ]; then
args+=(--min-p "${MIN_P}")
fi
exec llama-server "${args[@]}"
because I use an ssh key instead of via http, I need to enable the ssh agent beforehand eval "$(ssh-agent)" && ssh-add. I got random errors on having to authenticate myself via http git cloning, but like 50% of the time, so instead I switched to ssh based cloning (generate a key and add it to your github to make this work).
Locally cloning the repo and copying its content over to the builder step instead is also a good approach.
After that, I can build an image like this: docker buildx build --tag myName/llama-server --ssh default --no-cache . Compilation takes a few minutes.
and my docker compose looks like this:
services:
llama-server:
image: myName/llama-server
container_name: llama-server
restart: unless-stopped
ports:
- 8000:8000
volumes:
- ./models:/models:ro
- ./templates:/templates:ro
gpus: all
env_file: env/Qwen3.8-27B.env
which I can start by just doing docker compose up --detach.
I have in the root directory a directory models, for my models, templates for my templates, and env for all my env files (makes switching models easy).
This is an example env file for Qwen 3.8 27B:
LLAMA_ARG_MODEL=/models/Qwen3.8-27B-Q5_K_M.gguf
LLAMA_ARG_ALIAS="Qwen 3.8 27B - Thinkstation PGX"
LLAMA_ARG_SPEC_TYPE=draft-mtp
LLAMA_ARG_SPEC_DRAFT_N_MAX=3
LLAMA_ARG_CTX_SIZE=262144
LLAMA_ARG_MMPROJ=/models/qwen3.8-27b-mmproj-BF16.gguf
LLAMA_ARG_CHAT_TEMPLATE_FILE=/templates/qwen3.8-27b-chat_template.jinja
LLAMA_ARG_CHAT_TEMPLATE='{"reasoning_effort":"medium"}'
TEMP=1.0
TOP_P=0.95
TOP_K=20
MIN_P=0.0
Hope this helps. The big pro of using docker is so that you do not need to setup anything locally. I can switch the cuda version by simply editing one line of code, without touching anything else. Docker and all required dependencies for nvidia are by default already installed on your GB10 device.
If you are new to docker, it can be a bit of a steep learning curve.