Proposal for Model Portfolio Optimization and Resource Reallocation within NVIDIA NIM

Proposal: NVIDIA NIM Model Portfolio Optimization and Infrastructure Efficiency Initiative

Executive Summary

This report proposes the deprecation of several NVIDIA NIM-hosted models and services that currently provide limited strategic value, low adoption, excessive hardware consumption, overlapping functionality, or outdated capabilities compared to newer alternatives.

The primary objectives are:

  • Reduce GPU resource consumption.

  • Improve inference availability and response latency.

  • Increase cluster utilization efficiency.

  • Simplify operational maintenance.

  • Introduce newer state-of-the-art models with superior performance-per-token and performance-per-watt metrics.

  • Mitigate increasing infrastructure abuse caused by free-tier AI orchestration frameworks.

This proposal is intended to improve overall service quality while preparing the platform for next-generation foundation models.


Part I — Models Recommended for Deprecation

1. PaliGemma

Reasoning

PaliGemma represented an important step in lightweight multimodal AI; however, the current multimodal ecosystem has significantly evolved.

Current limitations:

  • Lower vision-language accuracy compared to modern VLMs.

  • Inferior OCR capabilities.

  • Limited reasoning performance.

  • Relatively low community adoption.

  • Significant overlap with newer multimodal models.

Impact

Removing PaliGemma would:

  • Free GPU memory resources.

  • Reduce maintenance overhead.

  • Consolidate multimodal workloads onto newer architectures.

Recommendation

Deprecate and migrate workloads toward newer multimodal models.


2. Flux1-Kontext-dev

Reasoning

The development version primarily serves experimentation purposes.

Issues:

  • Low production usage.

  • High VRAM requirements.

  • Inferior efficiency compared to optimized successors.

  • Increased operational burden due to maintaining development-grade models.

Recommendation

Retain only stable production-ready image generation models.


3. GenMol

Reasoning

GenMol serves a highly specialized molecular generation niche.

Challenges:

  • Extremely low utilization compared to general AI workloads.

  • Significant compute reservation for a narrow audience.

  • Limited strategic value for most NIM users.

Recommendation

Move to an on-demand deployment model or discontinue public hosting.


4. VISTA-3D

Reasoning

Medical 3D segmentation is a highly specialized workload.

Issues:

  • Large GPU footprint.

  • Limited user base.

  • High maintenance complexity.

  • Specialized healthcare applications are not aligned with general-purpose AI infrastructure goals.

Recommendation

Transition to enterprise-only deployment.


5. Active Speaker Detection

Reasoning

Usage statistics indicate low demand relative to compute allocation.

Problems:

  • Narrow use cases.

  • Overlap with broader multimodal pipelines.

  • Infrastructure fragmentation.

Recommendation

Deprecate standalone deployment.


6. Background Noise Removal

Reasoning

Dedicated noise removal models are increasingly replaced by integrated speech-processing pipelines.

Issues:

  • Low standalone utilization.

  • Functionality available through newer speech systems.

  • Additional maintenance requirements.

Recommendation

Retire standalone endpoint.


7. AbacusAI Dracarys Llama 3.1 70B Instruct

Reasoning

Although capable, this model is increasingly outperformed by newer open-source alternatives.

Challenges:

  • Extremely high GPU cost.

  • Poor performance-per-dollar ratio.

  • Lower efficiency than modern Mixture-of-Experts architectures.

Alternatives:

  • Qwen 3.6

  • Kimi 2.7

  • GLM 5.2

Recommendation

Deprecate.


8. Solar 10.7B Instruct

Reasoning

Solar was highly competitive when introduced, but the landscape has changed dramatically.

Current shortcomings:

  • Inferior reasoning benchmarks.

  • Lower coding performance.

  • Smaller ecosystem adoption.

Recommendation

Retire and replace with modern small-to-mid-sized reasoning models.


9. NVIDIA Ising Calibration 1-35B-A3B

Reasoning

Highly specialized scientific model with limited audience reach.

Issues:

  • Low request volume.

  • Dedicated infrastructure costs.

  • Minimal impact on broader user experience.

Recommendation

Move to research-only hosting.


10. Flux.1-dev

Reasoning

Development-oriented model.

Problems:

  • Heavy resource consumption.

  • Inferior operational stability.

  • Duplication with production image generation offerings.

Recommendation

Deprecate.


11. Llama Guard 4 12B

Reasoning

Modern moderation systems increasingly favor:

  • Smaller classifiers.

  • Multi-stage moderation pipelines.

  • Ensemble filtering approaches.

Issues:

  • Excessive GPU allocation for moderation tasks.

  • Lower efficiency than lightweight safety classifiers.

Recommendation

Replace with smaller moderation systems.


12. GLM 5.1

Reasoning

GLM 5.2 provides measurable improvements across:

  • Reasoning

  • Coding

  • Tool usage

  • Multilingual capabilities

  • Inference efficiency

Maintaining both versions introduces unnecessary duplication.

Recommendation

Deprecate GLM 5.1 and migrate traffic to GLM 5.2.


Part II — Models Recommended for Addition

1. Qwen 3.6 Family

Recommended variants:

  • Qwen3.6-7B

  • Qwen3.6-14B

  • Qwen3.6-27B

  • Qwen3.6-72B

Benefits:

  • Strong coding performance.

  • Excellent multilingual support.

  • Competitive reasoning.

  • Efficient deployment characteristics.

  • Large community adoption.

Expected Impact:

  • Increased user satisfaction.

  • Reduced latency.

  • Better cost-performance ratio.


2. Kimi 2.7

Benefits:

  • Strong long-context reasoning.

  • Excellent coding capabilities.

  • High benchmark performance.

  • Growing developer adoption.

Strategic Value:

Provides a competitive alternative to current flagship open-source models.


3. GLM 5.2

Benefits:

  • Significant improvements over GLM 5.1.

  • Better reasoning quality.

  • Enhanced multilingual performance.

  • Improved tool-calling support.

Recommendation:

Adopt as primary GLM offering.


4. MiMo v2.5 and MiMo v2.5 Pro

Benefits:

  • Modern architecture.

  • Efficient inference.

  • Strong instruction following.

  • Competitive cost-performance ratio.

Expected Outcome:

Improved infrastructure efficiency while expanding model diversity.


Part III — Infrastructure Abuse and Free-Tier Framework Overload

Current Situation

An increasing number of AI orchestration frameworks and gateways are leveraging NVIDIA NIM free-tier endpoints as backend providers.

Examples include:

  • Multi-model gateways

  • Open-source orchestration systems

  • Agent frameworks

  • Browser-based AI platforms

  • Community-hosted chatbot services

Many deployments operate without:

  • Rate limiting

  • Caching

  • Traffic shaping

  • Request prioritization

As a result, a small number of automated systems may generate disproportionately large workloads.


Observed Effects

Queue Congestion

Users frequently encounter:

  • Long wait times

  • Increased queue depth

  • Delayed token generation

Higher Latency

System-wide latency increases due to:

  • GPU saturation

  • Scheduler contention

  • Excessive parallel requests

Reduced Availability

Consequences include:

  • Request failures

  • HTTP 429 responses

  • Temporary endpoint degradation

Resource Inefficiency

Many orchestration frameworks repeatedly query multiple models simultaneously, causing:

  • Duplicate inference workloads

  • Unnecessary GPU utilization

  • Increased operational costs


Recommendations

Enhanced Rate Limiting

Implement stricter:

  • Per-user limits

  • Per-IP limits

  • Per-application limits


Framework Identification

Require:

  • User-Agent validation

  • Application registration

  • API client attribution


Fair Scheduling

Introduce:

  • Priority queues

  • Dynamic resource allocation

  • Abuse detection mechanisms


Model Portfolio Optimization

Removing low-value, low-adoption, high-resource models will:

  • Release GPU capacity.

  • Reduce infrastructure complexity.

  • Improve responsiveness for high-demand models.


Conclusion

Deprecating outdated, redundant, specialized, or underutilized models while introducing modern high-efficiency alternatives such as Qwen3.6, Kimi 2.7, GLM 5.2, and MiMo v2.5 will significantly improve NVIDIA NIM infrastructure efficiency.

Combined with stronger anti-abuse measures against free-tier orchestration framework overuse, these changes are expected to:

  • Reduce queue times.

  • Lower latency.

  • Increase service stability.

  • Improve GPU utilization.

  • Enhance overall user experience.

This modernization effort represents a practical and scalable path toward maintaining sustainable growth of the NVIDIA NIM ecosystem.

I completely agree with everything, especially the suggestion to add the newer models that have been released, particularly Mimo 2.5 and Mimo 2.5 Pro.

We also have to be realistic: there are a lot of people heavily using NVIDIA’s free tier, and that includes many OpenClaw and similar framework users. There is a reason why there are so many RPM increase requests with the justification of “it’s not enough for my personal development,” as if that alone were a valid reason to demand significantly higher RPM limits than everyone else.

Furthermore, because of those same OpenClaw users, models such as Gemma 4 31B, DeepSeek V4 Pro, and DeepSeek V4 Flash have become practically unusable. It is time to put a stop to this abuse.

That is why whenever I see a post saying, “Increase my RPM to 200,” I write a detailed reply explaining the situation and why that request cannot be granted under the free tier.

I hope the NVIDIA team reads your post, @supervisorxscp. There is definitely a need for these newer models, and it is also important to improve resource management and, if possible, address the constant stream of RPM increase requests from free-tier users.

maybe we can flag this

I’d say yes, @supervisorxscp, but you know what I think works even better? What I’ve been doing. I basically reply with a detailed explanation and provide the relevant information. That approach is better because it leaves an idea behind, makes the situation clear, and helps educate users.

Just look around, the forum is full of people asking for RPM increases as if it were automatic. I think it’s more useful to respond to those posts and explain that it doesn’t work that way. That’s what I’ve been doing, and I intend to keep doing it as long as people continue making those requests.

I don’t know what you think, but my goal is to leave a clear message behind, so that users who keep asking for higher RPM limits understand how the system actually works and why those requests are often unrealistic under the free tier.

That’s great btw

bump, the main problem i believe are openclaw users overloading the models, but if Nvidia won’t even try to ban that platform, at least remove some older models out of the extensive list of models the platform has. just throwing 429 errors and reducing the quota isn’t helping.