Request for NVIDIA NIM API Rate Limit Increase (40 → 200 RPM)

Hello NVIDIA Support Team,

I am writing to request an increase to the rate limit on my NVIDIA NIM API free-tier account, from the default 40 RPM to 200 RPM (or the highest reasonable tier available for individual / small-team evaluation).

Account details

  • Account email: rickwu@msi.com
  • Team size: individual developer/Product PM of MSI GB10
  • Current limit: 40 RPM
  • Requested limit: 200 RPM

Project background
I am developing a self-hosted heterogeneous AI stack for agentic workloads. The orchestration layer is Hermes Agent (NousResearch), and I am using nvidia/nemotron-3-super-120b-a12b through the NIM hosted endpoint as the default agent-loop model, given its strong tool-calling and high throughput for multi-agent scenarios. My local inference hardware includes an NVIDIA DGX Spark (GB10), and I use the NIM hosted endpoints for rapid prototyping and evaluation before committing workloads to self-hosted NIM containers.

The agent relies on NIM-hosted models for:

  • multi-step reasoning and tool calling,
  • subagent spawning within a single task,
  • long-context conversations with persistent memory.

The problem
Hermes Agent’s learning loop and subagent features can issue many model calls per task. During internal testing I often run requests in parallel (agents, tools, retries, streaming), which causes me to hit the 40 RPM ceiling frequently. This results in repeated 429 errors and stalled sessions, making it hard to validate the end-to-end prototype and user experience.

Usage intent and budget
This usage is strictly for internal development and evaluation, not for any large-scale production deployment, and I fully understand and respect NVIDIA’s fair-use policies for the free tier. If I can successfully validate the stack on NVIDIA’s platform, my plan is to move to paid offerings later — higher-throughput NIM endpoints and eventually self-hosted NIM containers on NVIDIA GPUs.

A temporary or standing increase to ~200 RPM would allow me to complete this early-stage evaluation. I would also appreciate any guidance on the most practical path to scale beyond the free tier for a small-budget developer (e.g., self-hosting NIM for evaluation without large GPU investment).

Thank you very much for your time, and for providing such powerful tools for developers. Any help on the rate-limit increase would be greatly appreciated.

Best regards,

Rick Wu

Hey @rickwu, I don’t know if this helps, but a moderator previously replied with this to someone asking for the same thing as you:“Many of you are using free tier API access to NVIDIA NIMs. This usually involves a rate limit that is dependent on model, use-case and the amount of current overall traffic using the same access. There is no official way to circumvent this rate limit or to receive a rate limit increase on that same tier. And specifically here on the forums we do not have any influence on those rate limits. To make full use of a NIM blueprint you will need to deploy it. For more details on NVIDIA NIM refer to…”

Basically, this means that if you are using the free-tier API, you have absolutely no right to demand an RPM increase. The only way to legitimately request higher RPM limits is:

Not through the forum. Moderators have said it over and over again: the forum is not the place to request RPM increases, as there are other channels and processes for that.

You need to pay and deploy a model through NVIDIA NIM / NVIDIA Build if you require higher usage limits and production-level access.

First, go to the model you prefer, for example DeepSeek V4 Flash. In that section you’ll see three options: Experience / Model Card / Deploy.

Click on Deploy and you’ll see several options such as Partner Endpoints or Self-hosted Deployments.

From there, you choose the option you want, and it will show you the pricing and deployment costs.

You can start with DeepSeek Flash since it’s the cheapest one, so you can learn how the process works first. And if you need help, you can always contact a moderator privately and ask for guidance. There’s no problem with sending a direct message saying you want to pay and deploy properly.

I am not saying this to be toxic, rude, or disrespectful. My goal is simply to help you understand and follow NVIDIA’s rules and the guidance that has already been provided by moderators multiple times.

IF YOU ARE ALREADY PAYING FOR NVIDIA SERVICES AND DEPLOYED MODELS, THEN YOU SHOULD CONTACT NVIDIA DIRECTLY THROUGH THE APPROPRIATE SUPPORT CHANNELS, SUCH AS EMAIL OR PRIVATE COMMUNICATION WITH THE RELEVANT SUPPORT TEAM, RATHER THAN MAKING RPM INCREASE REQUESTS ON THE FORUM.