# Qwen-Image-2.1 on DGX Spark: 54 s per 1024² image (40 steps), 31.6 GB, and a fix for the pink vertical line

**URL:** <https://forums.developer.nvidia.com/t/qwen-image-2-1-on-dgx-spark-54-s-per-1024-image-40-steps-31-6-gb-and-a-fix-for-the-pink-vertical-line/384774>\
**Category:** DGX Spark / GB10 Projects\
**Tags:** pytorch, generative\_ai, spark\
**Created:** [September 30, 2026, 4:51pm UTC](https://forums.developer.nvidia.com/t/qwen-image-2-1-on-dgx-spark-54-s-per-1024-image-40-steps-31-6-gb-and-a-fix-for-the-pink-vertical-line/384774 "2026-09-30T16:51:14Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![mjookim](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mjookim](https://forums.developer.nvidia.com/u/mjookim)\
**Post date:** [September 30, 2026, 4:51pm UTC](https://forums.developer.nvidia.com/t/qwen-image-2-1-on-dgx-spark-54-s-per-1024-image-40-steps-31-6-gb-and-a-fix-for-the-pink-vertical-line/384774/1 "2026-09-30T16:51:15Z")

</div>

I’ve been running Qwen-Image-2.1 (bf16, diffusers) on a DGX Spark for the past few weeks, mostly to see whether it’s usable as a local text-to-image / editing box. I also ran the same setup on an RTX 4090 and a 16 GB M2 Mac for comparison, and ended up tracking down a rendering artifact along the way. Sharing the numbers in case anyone else is looking at this model on a Spark.

Full write-up with all the images (text rendering, reference images, retouching, transparent PNGs): [GitHub - MindrLabs/image-generation-qwen-image-2.1: Qwen-Image-2.1 on NVIDIA DGX Spark and RTX 4090: quick start, speed, and a hands-on report of what works. · GitHub](https://github.com/MindrLabs/image-generation-qwen-image-2.1)

This post is only the Spark-relevant part.

## Setup

| | |
| --- | --- |
| DGX Spark | GB10, driver 580.126.09, CUDA 13.0, 120 GB unified. Shared machine, about 93 GB free during measurement |
| RTX 4090 | 24 GB, `cpu_offload: model` (needed to fit) |
| Software | diffusers 0.41.0.dev0, transformers 5.17.0, bf16, torch 2.13.0 cu130 (aarch64) on the Spark, 2.14.0 cu126 on the 4090 |
| Weights | `Qwen/Qwen-Image-2.1` rev `790c9263`, 33 GB |
| Defaults | 1024², 40 steps, `true_cfg_scale` 1.0 (CFG off), VAE tiled decoding on |

No NGC container. torch on the Spark is the plain aarch64 wheel from the official PyTorch index (`https://download.pytorch.org/whl/cu130`), pip-installed into a venv on the host. model-compose picks the cu channel from the driver’s CUDA version reported by `nvidia-smi`, so it landed on 2.13.0+cu130 by itself.

Speed numbers on the Spark are the median of two runs per condition with the same seed (42), excluding the first request after server start. 4090 numbers are from a handful of manual requests, so they’re less precise.

## Text-to-image speed

| Resolution | 20 steps | 40 steps | Peak memory |
| --- | --- | --- | --- |
| 1024² | 28.3 s | 54.1 s | 31.6 GB |
| 2048² (officially recommended) | 2 min 14 s | 4 min 16 s | 33.7 GB |

- Time is roughly linear in step count. 4× the pixels takes about 4.7× longer.
- Run-to-run variance is tiny (four 1024²/20-step runs: 28.2–28.3 s), and the output files were byte-identical for the same seed.
- Server start + model load is about 3 minutes.
- Memory is the worker process’s `used_memory` from `nvidia-smi --query-compute-apps`, sampled every 0.5 s during the job, max taken (GiB). That includes the CUDA context and torch’s cache pool, but not the process’s CPU-side RSS, which on a Spark comes out of the same unified pool. Side note for GB10: total `memory.used` reports N/A on this driver, but the per-process value works. The 4090’s 19.8 GB was measured the same way.

## With reference images (1024², 40 steps)

| References | DGX Spark | RTX 4090 (offload) |
| --- | --- | --- |
| 0 | 54 s | ~45 s |
| 1 | 62.5 s | 48.5 s |
| 3 | 184 s | not measured |
| 10 | 598 s | not measured |

Ten references works (the model accepts up to 10) but takes 10 minutes per image and some items drift from their reference.

## The 4090 was faster, even with CPU offload

This surprised me a bit. The 4090 can’t hold the 33 GB of weights, so it runs with `cpu_offload: model` (one whole sub-model on the GPU at a time, the rest in system RAM, ~45 GB RSS). It still came out ahead of the Spark running everything resident in unified memory.

My guess is memory bandwidth (273 GB/s vs 1,008 GB/s) plus raw compute, but note the torch/CUDA versions differ (cu130 vs cu126) and the 4090 sample is small, so treat this as a rough comparison rather than a controlled one.

Where the Spark wins is that offload isn’t needed at all. On the Spark, turning `cpu_offload` on for a 1-reference image took it from 62.5 s to about 4 min 20 s, so if you have the memory, leave it off.

## The pink vertical line is VAE tiled decoding

Some outputs had a faint pink/purple vertical line at fixed x positions (≈438, 631, 824 in 1024² output, about 193 px apart), sometimes right across a face. Same column regardless of seed, steps, CFG, device (Spark and 4090), or API vs Gradio. The same symptom is reported upstream in [QwenLM/Qwen-Image-2.1 #5](https://github.com/QwenLM/Qwen-Image-2.1/issues/5).

It’s the seam of VAE tiled decoding. Tiles are 256 px placed every 192 px, and all three columns sit in the overlap zones. Decoding the same latent with tiling off removes it completely:

| Image (DGX) | Alpha seam depth, tiling on | Tiling off | Extra decode memory, tiling off |
| --- | --- | --- | --- |
| Coffee cup 1024², seed 42 | 10.4 | 0.0 | not measured |
| Coffee cup 1024², seed 7 | 13.0 | 0.0 | not measured |
| Lookbook 2048×1152, seed 42 | 8.6 | 0.0 | 15.9 GB |
| Lookbook 2048×1152, seed 7 | 4.1 | 0.0 | 15.9 GB |
| Coffee cup 2048², seed 42 | 7.0 | 0.0 | 28.2 GB |

(Seam depth is how far alpha dips at that column, 0–255.)

 ![tiling-on-off](https://global.discourse-cdn.com/nvidia/original/4X/3/8/5/3852d6ab6fd78b115e5a9db316b1da8ce528c20d.jpeg)  
 ![artifact-vertical-line-face](https://global.discourse-cdn.com/nvidia/original/4X/e/8/e/e8e0f09bfbd8bb9cc41234d45a2c01ddc3b65675.jpeg)

This is the one place the Spark’s memory really pays off: decoding 2048² without tiling needs ~28 GB extra, which is fine on 120 GB unified but won’t fit on a 24 GB card. If you’re on a Spark, turn tiling off. In plain diffusers that means not calling `enable_tiling()` on the VAE (the default is already off). I had it on because the serving tool I used (model-compose) turns tiling on by default so the same config also runs on 24 GB cards. Every image in the write-up was decoded with it on.

## What I’d recommend on a Spark

- Draft at 1024², 20 steps (~28 s), then regenerate the picks at the official 2048², 40 steps (~4 min). 20 steps keeps the composition but can drop small prompt elements like text on a cup.
- Leave `cpu_offload` off.
- Turn VAE tiling off (see above). It costs up to ~28 GB extra at 2048², which you have.
- If you feed the model an image it generated itself as a reference, use a different seed from the one that made it. At the same seed it copies the input back, oversaturated, instead of making a new scene.
- English and Chinese text rendering is reliable; Korean gets individual characters wrong in a way that’s stable across seeds (rewording fixes it, reseeding doesn’t). Details and examples in the repo.

## How to run it

```bash
pip install model-compose # 0.4.111 or later
git clone https://github.com/MindrLabs/image-generation-qwen-image-2.1
cd image-generation-qwen-image-2.1
model-compose up

```

First run builds a venv and downloads the 33 GB checkpoint. Gradio on :8081, HTTP API on :8080. On a Spark, delete the `cpu_offload: model` line in `model-compose.yml` first.

```bash
curl -o out.png localhost:8080/api/workflows/runs \
  -F workflow_id=generate -F wait_for_completion=true -F output_only=true \
  -F "input.prompt=A poster that says HELLO in bold letters" -F input.seed=42 \
  -F input.width=1024 -F input.height=1024 -F input.inference_steps=40

```

## Not measured

- NVFP4 or any quantization on the Spark. Everything here is bf16.
- vLLM-OMNI (see @cosinus’s [thread](https://forums.developer.nvidia.com/t/qwen-image-2-1-via-vllm-omni/383994)). Would be interesting to compare against the diffusers path.
- Whether the tiling artifact also shows up in the editing / reference paths, and tiling-off memory on the 4090.
- Comparison with other image models. This was only about whether this one runs and what it’s good at.

Happy to run a specific prompt or setting if someone wants a data point.
