Runbook: bu-30b-a3b-preview-AWQ-4bit Model on DGX Spark (Solo) with vLLM + Browser-Use

This post walks through a complete, working setup for deploying the cyankiwi/bu-30b-a3b-preview-AWQ-4bit model on a DGX Spark (GB10) system using vLLM and consuming it from a lightweight agent application running in Docker (ChatGPY style)

I tested this on a Single Spark…

bu-30b-a3b-preview-AWQ-4bit Model and Why It Matters

The bu-30b-a3b-preview-AWQ-4bit model is a quantized, browser-use–optimized variant of a larger Qwen-based architecture. It’s specifically tailored to work with the Browser-Use OSS library to provide strong web interaction and browsing capabilities for agent-style applications.

🔹 Browser-Use Purpose — This model is heavily trained to understand and reason about web content and DOM structures, making it ideal for tasks where an agent needs to interact with web pages, interpret page layouts, extract information, and perform task-oriented browsing.

🔹 Underlying Design — The full “BU-30B-A3B-Preview” base model comes from a mixture-of-experts (MoE) version of the Qwen-VL-30B-A3B architecture, with efficient instruction and vision-language reasoning capabilities suitable for complex browsing and web tasks.

The deployment is split into two parts:

  1. Model serving (vLLM on DGX Spark GPUs)

  2. Client application (Browser-Use agent calling the model via OpenAI-compatible API)


1. Environment Overview

Hardware

  • NVIDIA DGX Spark / GB10

  • NVIDIA GPUs with sufficient VRAM for AWQ 4-bit 30B model

Software Components

  • Docker + NVIDIA Container Toolkit

  • vLLM (OpenAI-compatible server)

  • Hugging Face model cache

  • Python 3.11 client container

  • Browser-Use + Playwright


2. Download the Model (Optional but Recommended)

Pre-downloading avoids repeated downloads inside containers.

curl -LsSf https://astral.sh/uv/install.sh | sh
uv --version

uvx hf download cyankiwi/bu-30b-a3b-preview-AWQ-4bit

This places the model in your local Hugging Face cache:

~/.cache/huggingface


3. Start the vLLM Model Server on DGX Spark

The following command launches a vLLM server optimized for DGX Spark, exposing an OpenAI-compatible API.

docker run -d --privileged --gpus all --rm --ipc=host --network host --name bu30b \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  scitrera/dgx-spark-vllm:0.13.0-t4 \
  vllm serve cyankiwi/bu-30b-a3b-preview-AWQ-4bit \
  --gpu-memory-utilization 0.70 \
  --served-model-name bu30b \
  --max-model-len 65536 \
  --host 0.0.0.0 --port 8000

Key Flags Explained

  • --served-model-name bu30b
    Used by clients when specifying the model.

  • --gpu-memory-utilization 0.70
    Leaves headroom for stability on GB10.

  • --max-model-len 65536
    Enables long-context workloads.

  • --network host
    Simplifies access from client containers.

Once running, the API is available at:

http://127.0.0.1:8000/v1


4. Client Application (Browser-Use Agent)

The client runs in a separate Docker container and connects to the vLLM server.

Dockerfile

FROM python:3.11-slim

ENV PYTHONDONTWRITEBYTECODE=1
ENV PYTHONUNBUFFERED=1

RUN apt-get update && apt-get install -y \
    curl \
    wget \
    gnupg \
    ca-certificates \
    libnss3 \
    libatk-bridge2.0-0 \
    libxkbcommon0 \
    libxcomposite1 \
    libxrandr2 \
    libxdamage1 \
    libgbm1 \
    libasound2 \
    libxshmfence1 \
    libgtk-3-0 \
    libdrm2 \
    libx11-xcb1 \
    libxcb1 \
    libxext6 \
    libxfixes3 \
    libxi6 \
    libglib2.0-0 \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

RUN playwright install chromium

COPY . .

CMD ["python", "main.py"]


requirements.txt

browser-use
python-dotenv
playwright


main.py

from dotenv import load_dotenv
from browser_use import Agent, ChatOpenAI

load_dotenv()

llm = ChatOpenAI(
    base_url='http://127.0.0.1:8000/v1',
    model='bu30b',
    temperature=0.6,
    top_p=0.95,
    dont_force_structured_output=True,  # improves latency
)

agent = Agent(
    task='Which is the top one topic of NVIDIA DGX Spark /GB10 User Forum - forums.developer.nvidia.com)',
    llm=llm,
)

agent.run_sync()


5. Build and Run the Client Container

docker build -t bu30b-client .
docker run --rm --network host bu30b-client

Because both containers use --network host, the client can directly reach the vLLM server on 127.0.0.1:8000.

Thanks for the overview! I have moved this to GB10 projects