Hello NVIDIA Developer Support Team,
I am a developer currently working on a project using the NVIDIA API (glm5.2).
I am reaching out to formally request an increase in my API rate limit from the default 40 requests to the maximum of 200 requests, as the current limit is causing bottlenecks in my development process.
Reason for Request: (Choose or modify the one that fits you)
- Batch Data Processing: I am building a RAG pipeline and prompt engineering workflow that requires sending hundreds of data points at once to evaluate the results. The 40-request limit makes it impossible to complete a single batch test efficiently.
- Concurrency Testing: I am running multiple local pipelines in parallel to test response times and identify bottlenecks, which frequently exceeds the 40-request threshold.
- Prototype Service: I have deployed a small-scale internal test service, and concurrent accesses from team members are frequently hitting the rate limit, interrupting the service flow.
Request Details:
- Current Limit: 40 requests
- Requested Limit: 200 requests
Allowing this increase will enable me to utilize NVIDIA’s powerful models more effectively and stably to complete my project.
Account Information:
- NVIDIA Account Email: bangae2@gmail.com
Thank you for your time and assistance. Have a great day!