This post walks through a complete, working setup for deploying the cyankiwi/bu-30b-a3b-preview-AWQ-4bit model on a DGX Spark (GB10) system using vLLM and consuming it from a lightweight agent application running in Docker (ChatGPY style)
I tested this on a Single Spark…
bu-30b-a3b-preview-AWQ-4bit Model and Why It Matters
The bu-30b-a3b-preview-AWQ-4bit model is a quantized, browser-use–optimized variant of a larger Qwen-based architecture. It’s specifically tailored to work with the Browser-Use OSS library to provide strong web interaction and browsing capabilities for agent-style applications.
🔹 Browser-Use Purpose — This model is heavily trained to understand and reason about web content and DOM structures, making it ideal for tasks where an agent needs to interact with web pages, interpret page layouts, extract information, and perform task-oriented browsing.
🔹 Underlying Design — The full “BU-30B-A3B-Preview” base model comes from a mixture-of-experts (MoE) version of the Qwen-VL-30B-A3B architecture, with efficient instruction and vision-language reasoning capabilities suitable for complex browsing and web tasks.
The deployment is split into two parts:
-
Model serving (vLLM on DGX Spark GPUs)
-
Client application (Browser-Use agent calling the model via OpenAI-compatible API)
1. Environment Overview
Hardware
-
NVIDIA DGX Spark / GB10
-
NVIDIA GPUs with sufficient VRAM for AWQ 4-bit 30B model
Software Components
-
Docker + NVIDIA Container Toolkit
-
vLLM (OpenAI-compatible server)
-
Hugging Face model cache
-
Python 3.11 client container
-
Browser-Use + Playwright
2. Download the Model (Optional but Recommended)
Pre-downloading avoids repeated downloads inside containers.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv --version
uvx hf download cyankiwi/bu-30b-a3b-preview-AWQ-4bit
This places the model in your local Hugging Face cache:
~/.cache/huggingface
3. Start the vLLM Model Server on DGX Spark
The following command launches a vLLM server optimized for DGX Spark, exposing an OpenAI-compatible API.
docker run -d --privileged --gpus all --rm --ipc=host --network host --name bu30b \
-v ~/.cache/huggingface:/root/.cache/huggingface \
scitrera/dgx-spark-vllm:0.13.0-t4 \
vllm serve cyankiwi/bu-30b-a3b-preview-AWQ-4bit \
--gpu-memory-utilization 0.70 \
--served-model-name bu30b \
--max-model-len 65536 \
--host 0.0.0.0 --port 8000
Key Flags Explained
-
--served-model-name bu30b
Used by clients when specifying the model. -
--gpu-memory-utilization 0.70
Leaves headroom for stability on GB10. -
--max-model-len 65536
Enables long-context workloads. -
--network host
Simplifies access from client containers.
Once running, the API is available at:
http://127.0.0.1:8000/v1
4. Client Application (Browser-Use Agent)
The client runs in a separate Docker container and connects to the vLLM server.
Dockerfile
FROM python:3.11-slim
ENV PYTHONDONTWRITEBYTECODE=1
ENV PYTHONUNBUFFERED=1
RUN apt-get update && apt-get install -y \
curl \
wget \
gnupg \
ca-certificates \
libnss3 \
libatk-bridge2.0-0 \
libxkbcommon0 \
libxcomposite1 \
libxrandr2 \
libxdamage1 \
libgbm1 \
libasound2 \
libxshmfence1 \
libgtk-3-0 \
libdrm2 \
libx11-xcb1 \
libxcb1 \
libxext6 \
libxfixes3 \
libxi6 \
libglib2.0-0 \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
RUN playwright install chromium
COPY . .
CMD ["python", "main.py"]
requirements.txt
browser-use
python-dotenv
playwright
main.py
from dotenv import load_dotenv
from browser_use import Agent, ChatOpenAI
load_dotenv()
llm = ChatOpenAI(
base_url='http://127.0.0.1:8000/v1',
model='bu30b',
temperature=0.6,
top_p=0.95,
dont_force_structured_output=True, # improves latency
)
agent = Agent(
task='Which is the top one topic of NVIDIA DGX Spark /GB10 User Forum - forums.developer.nvidia.com)',
llm=llm,
)
agent.run_sync()
5. Build and Run the Client Container
docker build -t bu30b-client .
docker run --rm --network host bu30b-client
Because both containers use --network host, the client can directly reach the vLLM server on 127.0.0.1:8000.
