Qwen-Image-2.1 on DGX Spark: 54 s per 1024² image (40 steps), 31.6 GB, and a fix for the pink vertical line

I’ve been running Qwen-Image-2.1 (bf16, diffusers) on a DGX Spark for the past few weeks, mostly to see whether it’s usable as a local text-to-image / editing box. I also ran the same setup on an RTX 4090 and a 16 GB M2 Mac for comparison, and ended up tracking down a rendering artifact along the way. Sharing the numbers in case anyone else is looking at this model on a Spark.

Full write-up with all the images (text rendering, reference images, retouching, transparent PNGs): GitHub - MindrLabs/image-generation-qwen-image-2.1: Qwen-Image-2.1 on NVIDIA DGX Spark and RTX 4090: quick start, speed, and a hands-on report of what works. · GitHub

This post is only the Spark-relevant part.

Setup

DGX Spark GB10, driver 580.126.09, CUDA 13.0, 120 GB unified. Shared machine, about 93 GB free during measurement
RTX 4090 24 GB, cpu_offload: model (needed to fit)
Software diffusers 0.41.0.dev0, transformers 5.17.0, bf16, torch 2.13.0 cu130 (aarch64) on the Spark, 2.14.0 cu126 on the 4090
Weights Qwen/Qwen-Image-2.1 rev 790c9263, 33 GB
Defaults 1024², 40 steps, true_cfg_scale 1.0 (CFG off), VAE tiled decoding on

No NGC container. torch on the Spark is the plain aarch64 wheel from the official PyTorch index (https://download.pytorch.org/whl/cu130), pip-installed into a venv on the host. model-compose picks the cu channel from the driver’s CUDA version reported by nvidia-smi, so it landed on 2.13.0+cu130 by itself.

Speed numbers on the Spark are the median of two runs per condition with the same seed (42), excluding the first request after server start. 4090 numbers are from a handful of manual requests, so they’re less precise.

Text-to-image speed

Resolution 20 steps 40 steps Peak memory
1024² 28.3 s 54.1 s 31.6 GB
2048² (officially recommended) 2 min 14 s 4 min 16 s 33.7 GB
  • Time is roughly linear in step count. 4× the pixels takes about 4.7× longer.
  • Run-to-run variance is tiny (four 1024²/20-step runs: 28.2–28.3 s), and the output files were byte-identical for the same seed.
  • Server start + model load is about 3 minutes.
  • Memory is the worker process’s used_memory from nvidia-smi --query-compute-apps, sampled every 0.5 s during the job, max taken (GiB). That includes the CUDA context and torch’s cache pool, but not the process’s CPU-side RSS, which on a Spark comes out of the same unified pool. Side note for GB10: total memory.used reports N/A on this driver, but the per-process value works. The 4090’s 19.8 GB was measured the same way.

With reference images (1024², 40 steps)

References DGX Spark RTX 4090 (offload)
0 54 s ~45 s
1 62.5 s 48.5 s
3 184 s not measured
10 598 s not measured

Ten references works (the model accepts up to 10) but takes 10 minutes per image and some items drift from their reference.

The 4090 was faster, even with CPU offload

This surprised me a bit. The 4090 can’t hold the 33 GB of weights, so it runs with cpu_offload: model (one whole sub-model on the GPU at a time, the rest in system RAM, ~45 GB RSS). It still came out ahead of the Spark running everything resident in unified memory.

My guess is memory bandwidth (273 GB/s vs 1,008 GB/s) plus raw compute, but note the torch/CUDA versions differ (cu130 vs cu126) and the 4090 sample is small, so treat this as a rough comparison rather than a controlled one.

Where the Spark wins is that offload isn’t needed at all. On the Spark, turning cpu_offload on for a 1-reference image took it from 62.5 s to about 4 min 20 s, so if you have the memory, leave it off.

The pink vertical line is VAE tiled decoding

Some outputs had a faint pink/purple vertical line at fixed x positions (≈438, 631, 824 in 1024² output, about 193 px apart), sometimes right across a face. Same column regardless of seed, steps, CFG, device (Spark and 4090), or API vs Gradio. The same symptom is reported upstream in QwenLM/Qwen-Image-2.1 #5.

It’s the seam of VAE tiled decoding. Tiles are 256 px placed every 192 px, and all three columns sit in the overlap zones. Decoding the same latent with tiling off removes it completely:

Image (DGX) Alpha seam depth, tiling on Tiling off Extra decode memory, tiling off
Coffee cup 1024², seed 42 10.4 0.0 not measured
Coffee cup 1024², seed 7 13.0 0.0 not measured
Lookbook 2048×1152, seed 42 8.6 0.0 15.9 GB
Lookbook 2048×1152, seed 7 4.1 0.0 15.9 GB
Coffee cup 2048², seed 42 7.0 0.0 28.2 GB

(Seam depth is how far alpha dips at that column, 0–255.)


This is the one place the Spark’s memory really pays off: decoding 2048² without tiling needs ~28 GB extra, which is fine on 120 GB unified but won’t fit on a 24 GB card. If you’re on a Spark, turn tiling off. In plain diffusers that means not calling enable_tiling() on the VAE (the default is already off). I had it on because the serving tool I used (model-compose) turns tiling on by default so the same config also runs on 24 GB cards. Every image in the write-up was decoded with it on.

What I’d recommend on a Spark

  • Draft at 1024², 20 steps (~28 s), then regenerate the picks at the official 2048², 40 steps (~4 min). 20 steps keeps the composition but can drop small prompt elements like text on a cup.
  • Leave cpu_offload off.
  • Turn VAE tiling off (see above). It costs up to ~28 GB extra at 2048², which you have.
  • If you feed the model an image it generated itself as a reference, use a different seed from the one that made it. At the same seed it copies the input back, oversaturated, instead of making a new scene.
  • English and Chinese text rendering is reliable; Korean gets individual characters wrong in a way that’s stable across seeds (rewording fixes it, reseeding doesn’t). Details and examples in the repo.

How to run it

pip install model-compose   # 0.4.111 or later
git clone https://github.com/MindrLabs/image-generation-qwen-image-2.1
cd image-generation-qwen-image-2.1
model-compose up

First run builds a venv and downloads the 33 GB checkpoint. Gradio on :8081, HTTP API on :8080. On a Spark, delete the cpu_offload: model line in model-compose.yml first.

curl -o out.png localhost:8080/api/workflows/runs \
  -F workflow_id=generate -F wait_for_completion=true -F output_only=true \
  -F "input.prompt=A poster that says HELLO in bold letters" -F input.seed=42 \
  -F input.width=1024 -F input.height=1024 -F input.inference_steps=40

Not measured

  • NVFP4 or any quantization on the Spark. Everything here is bf16.
  • vLLM-OMNI (see @cosinus’s thread). Would be interesting to compare against the diffusers path.
  • Whether the tiling artifact also shows up in the editing / reference paths, and tiling-off memory on the 4090.
  • Comparison with other image models. This was only about whether this one runs and what it’s good at.

Happy to run a specific prompt or setting if someone wants a data point.

1 Like