# NVIDIA ACE: Model Archive

**URL:** <https://forums.developer.nvidia.com/t/nvidia-ace-model-archive/362682>\
**Category:** General Topics & Other SDKs\
**Tags:** llama, agentic-ai, gaming, nemotron\
**Created:** [March 10, 2026, 3:30pm UTC](https://forums.developer.nvidia.com/t/nvidia-ace-model-archive/362682 "2026-03-10T15:30:59Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![TomNVIDIA](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/tomnvidia/32/14181_2.png) [@TomNVIDIA](https://forums.developer.nvidia.com/u/TomNVIDIA)\
**Post date:** [March 10, 2026, 3:30pm UTC](https://forums.developer.nvidia.com/t/nvidia-ace-model-archive/362682/1 "2026-03-10T15:30:59Z")

</div>

Access previous NVIDIA ACE for Games models

#### Llama3.2-3B-Instruct

Agentic small language model that enables better role-play, retrieval-augmented generation (RAG) and function calling capabilities. This model is compatible with multi-vendor GPUs and CPUs.

[Access Model Card](https://developer.nvidia.com/downloads/assets/ace/model_card/llama-3.2-3b_for_Nv_IGI_SDK.pdf)

[Download On-Device Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/llama-3.2-3b_v1.0.1.7z)

[Access Cloud Model](https://build.nvidia.com/meta/llama-3.2-3b-instruct/deploy)

[Documentation](https://github.com/NVIDIA-RTX/NVIGI-Plugins/blob/main/docs/ProgrammingGuideGPT.md)

#### Mistral-Nemo-Minitron Family

Agentic small language models that enable better role-play, retrieval-augmented generation (RAG) and function calling capabilities. They come in 8B, 4B and 2B parameter models to fit your VRAM and performance requirements. The on-device models are compatible with multi-vendor GPUs and CPUs.

[Access Model Card](https://developer.nvidia.com/downloads/assets/ace/model_card/Mistral-NeMo-Minitron-8B-128K-Instruct.pdf)

[Download On-Device 2B Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/mistral-nemo-minitron-2b-128k-instruct_v1.0.0.7z)

[Download On-Device 4B Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/mistral-nemo-minitron-4b-128k-instruct_v1.0.0.7z)

[Download On-Device 8B Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/mistral-nemo-minitron-8b-128k-instruct_v1.0.0.7z)

[Documentation](https://github.com/NVIDIA-RTX/NVIGI-Plugins/blob/main/docs/ProgrammingGuideGPT.md)

#### Nemovision-4B-Instruct

Agentic vision-language model that combines visual understanding of on-screen elements and actions and reasons for better context aware responses. The on-device model is compatible with multi-vendor GPUs and CPUs.

[Access Model Card](https://developer.nvidia.com/downloads/assets/ace/model_card/Nemotron-Mini-4B-Instruct.pdf)

[Download On-Device Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/mistral-nemotron-vision-4b-instruct_vv1.7z)

[Documentation](https://github.com/NVIDIA-RTX/NVIGI-Plugins/blob/main/docs/ProgrammingGuideGPT.md#90-vlm-visual-lanuage-models)

### Qwen3

Open source dense models optimized for on-device inference. Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support. It’s compatible with multi-vendor GPUs and CPUs.

[Access Model Card](https://developer.nvidia.com/downloads/assets/ace/model_card/qwen3-8b-instruct.pdf)

[Download On-Device 8B Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/qwen3-8b-q4_k_m.gguf.zip)

[Documentation](https://huggingface.co/Qwen/Qwen3-8B)

#### Riva TTS

Takes a text output and converts it into natural and expressive voices in multiple languages in real time. Built for agentic workflows and compatible with multi-vendor GPUs and CPUs. FP16 quantization offers higher accuracy for higher VRAM usage.

[Access Model Card](https://developer.nvidia.com/downloads/assets/ace/model_card/Riva_TTS_A2-Flow_for_Nv_IGI_SDK.pdf)

[Download On-Device Model](https://developer.nvidia.com/downloads/rtx/In-Game-Inference-SDK/riva-magpie-tts-flow-ggml-1p5-fp16.zip)(FP16)

[Download On-Device Model](https://developer.nvidia.com/downloads/rtx/In-Game-Inference-SDK/riva-magpie-tts-flow-ggml-1p5-q4.zip)(Q4)

[Access Cloud Model](https://build.nvidia.com/nvidia/magpie-tts-flow)

[Documentation](https://github.com/NVIDIA-RTX/NVIGI-Plugins/blob/main/docs/ProgrammingGuideTTSASqFlow.md)

#### Whisper ASR

Takes an audio stream as input and returns a text transcript in real time. It’s compatible with multi-vendor GPUs and CPUs..

[Access Model Card](https://developer.nvidia.com/downloads/assets/ace/model_card/Whisper_ASR.pdf)

[Download On-Device Model](https://developer.nvidia.com/downloads/assets/ace/model_zip/whisper_asr_gguf_v1.0.7z)

[Access Cloud Model](https://build.nvidia.com/openai/whisper-large-v3)

[Documentation](https://github.com/NVIDIA-RTX/NVIGI-Plugins/blob/main/docs/ProgrammingGuideASRWhisper.md)

---

<div class="post-metadata">

**Author:** ![TomNVIDIA](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/tomnvidia/32/14181_2.png) [@TomNVIDIA](https://forums.developer.nvidia.com/u/TomNVIDIA)\
**Post date:** [March 6, 2026, 9:40pm UTC](https://forums.developer.nvidia.com/t/nvidia-ace-model-archive/362682/2 "2026-03-06T21:40:29Z")

</div>


