Llama.cpp, how to install/setup?

I received my Nvidia GB10 last week and try my best to make it my AI workhorse.

I started with Ollama, combined with Claude and Claude Code Router, but I wasn’t quite impressed by the performance. The GPU load was 100% during LLM generation, but the CPU load went above 100% and tokens came in real slow. After a lot of tweaking, which didn’t really improve things, decided to switch to llama.cpp .

But that turned out to be less easy then expected.

I first followed this guidance: Build the GPU version of llama.cpp on GB10 | Arm Learning Paths

But that gave me llama.cpp that dind’t want to start. I discovered that after a couple of updates I now have 4 different versions of Cuda on my system: Cuda, Cuda13, Cuda-13.0 and Cuda-13.3. And which one you include in your path makes or breaks the build of llama.cpp.
At the end I got llama.cpp running, but it hardly used the GPU and performance was 2 token/sec or something.

Then i found this guide: https://build.nvidia.com/spark/llama-cpp/instructions

This instruction gives you a correctly build CPU llama.cpp, that fully neglects your beautiful GPU.

Does anyone have a link or other kind of help on how to build llama.cpp on a DGX Spark?

On the Sparks we mainly use vLLM and SGLang.

I highly recommend you look at sparkrun to get you up and running serving suitable models.

I use this (could be suboptimal, but it works quite fine for me):

Dockerfile:

ARG UBUNTU_VERSION=24.04
ARG CUDA_VERSION=13.3.1
ARG GCC_VERSION=14
ARG CUDA_ARCH=121a-real

# --------------------------------------
# Stage 1: build
# --------------------------------------
FROM nvidia/cuda:${CUDA_VERSION}-devel-ubuntu${UBUNTU_VERSION} AS builder

ARG GCC_VERSION
ARG CUDA_ARCH

RUN apt-get update && apt-get install -y --no-install-recommends \
        openssh-client \
        git \
        wget \
        gcc-${GCC_VERSION} \
        g++-${GCC_VERSION} \
        build-essential \
        cmake \
    && rm -rf /var/lib/apt/lists/*

# Trust Github SSH key
RUN mkdir -p -m 0700 ~/.ssh && \
    ssh-keyscan github.com >> ~/.ssh/known_hosts

WORKDIR /build

# Clone the repository via SSH, and drop its content in the working directory
RUN --mount=type=ssh \
    git clone \
        --branch master \
        --depth 1 \
        git@github.com:ggml-org/llama.cpp.git \
        .

RUN cmake -B build \
    -DGGML_CUDA=ON \
    -DBUILD_SHARED_LIBS=OFF \
    -DCMAKE_CUDA_ARCHITECTURES="${CUDA_ARCH}" \
    -DCMAKE_BUILD_TYPE=Release

RUN cmake \
    --build build \
    --config Release \
    --target llama-server \
    --parallel $(nproc)

# --------------------------------------
# Stage 2: Runtime
# --------------------------------------
FROM nvidia/cuda:${CUDA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}

RUN apt-get update && apt-get install -y --no-install-recommends \
        libgomp1 \
    && rm -rf /var/lib/apt/lists/*

COPY --from=builder /build/build/bin/llama-server /usr/local/bin/

WORKDIR /app
COPY ./entrypoint.sh .
RUN chmod +x ./entrypoint.sh

ENV LLAMA_ARG_HOST=0.0.0.0
ENV LLAMA_ARG_PORT=8000
ENV LLAMA_ARG_N_GPU_LAYERS=999
ENV LLAMA_ARG_N_GPU_LAYERS_DRAFT=999

ENTRYPOINT ["./entrypoint.sh"]

With a very simple entrypoint because not all properties are available as an ENV:

#!/bin/bash

args=()

if [ -n "${TEMP}" ]; then
    args+=(--temperature "${TEMP}")
fi

if [ -n "${TOP_K}" ]; then
    args+=(--top-k "${TOP_K}")
fi

if [ -n "${TOP_P}" ]; then
    args+=(--top-p "${TOP_P}")
fi

if [ -n "${MIN_P}" ]; then
    args+=(--min-p "${MIN_P}")
fi

exec llama-server "${args[@]}"

because I use an ssh key instead of via http, I need to enable the ssh agent beforehand eval "$(ssh-agent)" && ssh-add. I got random errors on having to authenticate myself via http git cloning, but like 50% of the time, so instead I switched to ssh based cloning (generate a key and add it to your github to make this work).

Locally cloning the repo and copying its content over to the builder step instead is also a good approach.

After that, I can build an image like this: docker buildx build --tag myName/llama-server --ssh default --no-cache . Compilation takes a few minutes.

and my docker compose looks like this:

services:
  llama-server:
    image: myName/llama-server
    container_name: llama-server
    restart: unless-stopped
    ports:
      - 8000:8000
    volumes:
      - ./models:/models:ro
      - ./templates:/templates:ro
    gpus: all
    env_file: env/Qwen3.8-27B.env

which I can start by just doing docker compose up --detach.

I have in the root directory a directory models, for my models, templates for my templates, and env for all my env files (makes switching models easy).

This is an example env file for Qwen 3.8 27B:

LLAMA_ARG_MODEL=/models/Qwen3.8-27B-Q5_K_M.gguf
LLAMA_ARG_ALIAS="Qwen 3.8 27B - Thinkstation PGX"
LLAMA_ARG_SPEC_TYPE=draft-mtp
LLAMA_ARG_SPEC_DRAFT_N_MAX=3
LLAMA_ARG_CTX_SIZE=262144
LLAMA_ARG_MMPROJ=/models/qwen3.8-27b-mmproj-BF16.gguf
LLAMA_ARG_CHAT_TEMPLATE_FILE=/templates/qwen3.8-27b-chat_template.jinja
LLAMA_ARG_CHAT_TEMPLATE='{"reasoning_effort":"medium"}'
TEMP=1.0
TOP_P=0.95
TOP_K=20
MIN_P=0.0

Hope this helps. The big pro of using docker is so that you do not need to setup anything locally. I can switch the cuda version by simply editing one line of code, without touching anything else. Docker and all required dependencies for nvidia are by default already installed on your GB10 device.

If you are new to docker, it can be a bit of a steep learning curve.

Thanks, but I am a bit reluctant to use docker on my GB10. Not only because it consumes additional clock cycles, but also as it needs additional effort to make it work with local sources.

The overhead of docker is completely negligible when the main ops you’re doing are forward passes. A lot of the value of the DGX Spark is that the base firmware, drivers, cuda version and OS “just work.” This includes the nvidia docker runtime.

Like someone mentioned above, most of us are using sglang or vllm via sparkrun, spark-vllm-docker or various handcrafted containers for a specific use case or model.

Someone is providing prebuilt binaries of the latest llama.cpp for arm64.

I also don’t like docker, but when using dgx spark, docker is the best solution.

I stumbled upon this link this morning: GitHub - Entrpi/ds4-on-spark: Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support · GitHub

And connected OpenCode to it. I am very happy with the performance, I think i’ll stick to this for a couple of days.

Not even an issue on these boxes and modern Linux in general with regards to resource overhead. All 3 of mine exclusively use Docker to run inferencing and diffusion tasks with no observable in the real world performance penalty. Docker is honestly the safest, repeatable way to go since you don’t have to worry about updating dependencies or packages on the OS of the host wrecking another installation you have. Every container you run should have exactly everything it needs to do the job built-in. Bind-mounting directories is super simple and if you get stuck, any free-tier LLM can help you work through it.

Good luck whichever way you choose to go and have fun learning!