We are planning to deploy an NVIDIA A10 GPU in an HPE DL380 Gen10 Plus server running Ubuntu 24.04 LTS on bare metal (no VMware, Citrix, Horizon, Hyper-V, or vGPU partitioning).
The server will be used for:
- OCR and document processing
- AI inference (LLMs)
- RAG (Retrieval-Augmented Generation)
- Embedding generation
- Docker containers (vLLM, PyTorch, CUDA workloads)
- PostgreSQL and vector databases
There will be no virtual desktops, virtual workstations, or GPU sharing between VMs.
Our understanding is that the standard NVIDIA datacenter driver is sufficient and that RTX vWS / vGPU licensing is only required for virtual workstation or vGPU scenarios.
Can anyone confirm:
- Is any NVIDIA RTX vWS, vPC, vApps, or vGPU license required for this bare-metal deployment?
- Are there any functional, performance, or support-related benefits of purchasing RTX vWS when the A10 is used solely for AI inference and CUDA workloads on Ubuntu?
- If RTX vWS licensing is purchased for a bare-metal deployment, how is Concurrent User (CCU) licensing interpreted when there are no virtual workstation users?
Any guidance from users running A10s in production AI environments would be appreciated.