what is the reasoning behind keeping only 3 model profiles, each of them have max_num_tokens=5255, no Blackwell support. Qwen3-32b is not listed on the multi-LLM container (Supported Architectures for Multi-LLM NIM — NVIDIA NIM for Large Language Models (LLMs))
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Implementation Guide: DGX Spark with Qwen3.5-35B-A3B via llama.cpp for Claude Code | 3 | 2035 | April 2, 2026 | |
| Integrate and Deploy Tongyi Qwen3 Models into Production Applications with NVIDIA | 0 | 204 | May 2, 2025 | |
| OpenAI Compatible API does not work | 6 | 1065 | August 26, 2024 | |
| Give us qwen 3.6 | 1 | 369 | April 22, 2026 | |
| Serving Qwen3.5-397B-A17B at 1M Tokens on 2× DGX Spark — MiniMax M3 Is Next | 5 | 1127 | July 3, 2026 | |
| nvidia/Nemotron-Cascade-2-30B-A3B yet another model to test | 19 | 1685 | March 24, 2026 | |
| Request for NIM API rate limit increase (40 → 200 RPM) — getting started with multi-model exploration | 0 | 41 | May 25, 2026 | |
| Qwen3.5-397B-A17B + DGX Spark (duo) | 62 | 6599 | June 14, 2026 | |
| Request for NVIDIA NIM API Rate Limit Increase (40 → 200 RPM) | 0 | 70 | May 25, 2026 | |
| Token limit defaults to 4096 for all models | 1 | 247 | May 23, 2026 |