GPU-PV (dxgkrnl): CUDA workload slows ~1.7× over 3–4 days of host uptime with accumulating dxgvmb_send_sync_msg: wait_for_completion failed in dmesg;

WSL2 GPU-PV: CUDA throughput degrades over days; wsl --shutdown does not recover, restarting the NVIDIA device does

(microsoft/WSL の issue と NVIDIA Developer Forums「CUDA on Windows Subsystem for Linux」に出す用の英文。数字は 2026-09-21〜25 の実測)


Title: GPU-PV (dxgkrnl): CUDA workload slows ~1.7× over 3–4 days of host uptime with accumulating dxgvmb_send_sync_msg: wait_for_completion failed in dmesg; wsl --shutdown does not recover, pnputil /restart-device on the GPU does

Environment

  • Windows 11 Home 10.0.26200.9457, host RAM 63.6 GB
  • WSL 2.7.11.0, kernel 6.18.33.2-2 (6.18.33.2-microsoft-standard-WSL2), WSLg 1.0.73.2, Direct3D 1.611.1-81528511, DXCore 10.0.26100.1
  • .wslconfig: memory=52GB, swap=16GB
  • NVIDIA GeForce RTX 5090 (32 GB), driver 610.88 (32.0.16.1088, 2026-07-22)
  • Workload: ComfyUI 0.33 (PyTorch 2.13.0+cu130) running a video diffusion model (MiniMax H3), 100% CUDA, no Vulkan/GL. ~20–30 GB VRAM, ~44 GB guest RSS.

Symptom

The same fixed job (576×1344×158 frames, 3 sampling steps, i2v) measured with a per-node timer:

host uptime total sampler node video VAE decode audio VAE decode
fresh (host reboot 09-21) 33 s 19.5 s ~12 s ~0.2 s
+4 days (09-25) 52–55 s 30–33 s 12–15 s 6–15 s
  • During the slow state the GPU is busy (98% util, 540 W) for only ~12 of the ~30 s the sampler node takes; the rest is the guest waiting on the host. Trivial kernels (audio VAE, <1 s of work) take 10+ s. nvidia-smi itself becomes slow to respond.
  • GPU clocks/temps/throttle reasons are normal (2.9 GHz, 70 °C, no throttling). Guest swap in/out ≈ 0 during the job, PSI memory 0, CPU load 1.5.
  • Windows System event log shows no nvlddmkm / DxgKrnl / Display errors during this period (only Hyper-V VmSwitch info events).

dmesg signature (guest)

Errors accumulate day by day (counts since VM boot on 09-22 03:28):

misc dxg: dxgk: dxgvmb_send_sync_msg: wait_for_completion failed: fffffe00   ×152

misc dxg: dxgk: process_completion_packet: did not find packet to complete ×152

misc dxg: dxgk: dxgkio_query_adapter_info: Ioctl failed: -512 / -2 ×63

09-22 (boot day) 09-23 09-24 09-25 (until noon)
31 0 185 175

In the healthy state the same job adds 0 errors; in the degraded state 4 runs add ~12. (~28 query_adapter_info errors appear at every VM boot and are not related.)

What does NOT recover it

  • Restarting the CUDA process (ComfyUI)
  • wsl --shutdown and starting the distro again — the fresh VM is still at 52–55 s, and starts logging the same dxg errors within minutes (28 in 21 min)
  • Win+Ctrl+Shift+B (WDDM graphics stack restart)
  • Disabling pinned host memory in the application

What DOES recover it (without a host reboot)

  1. wsl --shutdown
  2. Restart the GPU device: pnputil /restart-device "PCI\VEN_10DE&DEV_2B85&..." (or Device Manager disable → enable)
  3. Start the distro again

→ back to 34.0 / 34.1 / 34.2 s, audio VAE 0.2 s, 0 new dxg errors. A full host reboot also recovers it (09-21). Note: restarting the device while the VM is running pulls the GPU out from under the guest (nvidia-smi: GPU access blocked by the operating system, dxgvmb_send_create_process: create_process failed -75), so the VM must be shut down first.

Interpretation

The accumulated state lives on the host side of the GPU-PV path (nvlddmkm / dxgkrnl host), not in the guest kernel or the VM: a new VM inherits the slowness, and only re-initializing the WDDM device clears it. Looks related to microsoft/WSL#41682 (same dxgvmb_send_sync_msg timeouts, same WSL/kernel versions, Intel Arc) and to the RTX 5060 Ti + WSL2 nvlddmkm DPC-latency reports on the NVIDIA forums (369079, 370836).

Reproduction

Keep the host up and run CUDA video-generation jobs for 3–4 days (several hours of GPU time per day); measure the same job daily. I can provide: full dmesg, wsl --version, nvidia-smi -q, the per-node timing tool, and collect-wsl-logs output on request.

also posted: