Hi,
Could you please add google/gemma-4-26b-a4b-it as a hosted API
on build.nvidia.com?
The NIM container already exists on NGC Catalog:
Problems with the current gemma-4-31b-it hosted API:
- Thinking mode cannot be fully disabled — even with thinkingBudget:0,
it still produces empty tags that break real-time
applications like subtitle translation
- Also, google/gemma-4-31b-it is unavailable most of the time on the Playground and via the hosted API (integrate.api.nvidia.com/v1/chat/completions)
- The 26B A4B has the same quality (-2/3 points) without
thinking issues on E2B/E4B models
The 26B A4B is MoE with only 3.8B active parameters —
very efficient to serve. The container is already ready!
Thank you so much,
Regards, Pedro
they allready added the model bro 😅🤷♂️
**Thank you for the response! However, DiffusionGemma 26B A4B
and Gemma 4 26B A4B are NOT equivalent models:
Architecture difference:
- Gemma 4 26B A4B → standard autoregressive model
- DiffusionGemma 26B → diffusion-based generation
(generates 256 tokens in parallel)
Benchmark comparison (official Google model card):
- Gemma 4 26B A4B: MMLU Pro 82.6% | AIME 88.3% | GPQA 82.3%
- DiffusionGemma 26B: MMLU Pro 77.6% | AIME 69.1% | GPQA 73.2%
DiffusionGemma is 4x faster but makes 6x more errors
in factual tasks (Reddit benchmark tests, June 2026).
For real-time subtitle translation, accuracy matters more
than speed. The Gemma 4 26B A4B (autoregressive) would
provide significantly better translation quality, and also other things , text answers, roleplay , etc.
The NIM container already exists on NGC:**
**
Could you please consider adding the standard
google/gemma-4-26b-a4b-it as a separate hosted API?**
and i think is not too much trouble for Nvidia to add.
The real benchmarks officials of Google — AIME 88.3% Gemma 4 26B A4B vs 69.1% **DiffusionGemma 26B
Thank you!**
So no one answers , no response at all? Are the responsible persons on Nvidia , even read these threads, on this forum???