DGX Spark GB10 shows only ~56GB VRAM inside AI Workbench (128GB expected)

This is invalid. I can run models using vLLM on docker with high-utilization. I suggest you start using llama.cpp over ollama, it’s much more performant.

Sample with vLLM: Running nvidia/Nemotron-Nano-VL-12B-V2-NVFP4-QAD on your spark

Performance numbers, using containers with high-utilization with 1 or 2 sparks: