An unsuable mess of a service

For the love of God, either KILL the service or repair it. It’s been unusable for almost two months now, way before GLM 5.2 was taken offline. None of the current AIs are good at doing anything except for maybe MiniMax M3, and even that takes 2+ minutes to get a response at around 32K tokens, which I might as well go take a ■■■■ and finish before that prompt even comes back with a response.

And if any fool here thinks “oh it’s free why are you complaining?”, no, it’s not free. NIM takes information we provide for AI training, obviously, so no, it’s not free.

Kimi K3 is here, yippee. It doesn’t work. The same for both DeepSeek Flash V4 and Pro models. It does NOT work, either of them. 5+ minutes of waiting time on one prompt just to receive an error. If the service is overloaded, at least be damn transparent about it. Y’all have flagship models on here and advertise them as 40 RPM, yet they literally do not function. And IF they do function, you fail to mention that SOME models have hidden limits on them, like GLM 5.2 did, where the user would request three times in one minute, and they would be on a 10-minute cooldown. A user is lucky to get one successful response with Kimi K3 even during off-peak hours like 11 PM BST time.

I know Z.AI has some new restrictions for companies that have 10+ billion in yearly revenue, so it’ll take some time for NVIDIA to deploy GLM 5.3, but holy ■■■■, at least have a backup like literally any other GLM models out there.

Other issues exist, like Kimi 2.6 being listed in the fetched models list, even though when requesting a prompt, it gives a “Gone” error. The service was decent back around GLM 5.1, but holy hell, it’s gone downright dogshit these days, and I can’t believe how little people are speaking up about this.

Do better, NVIDIA.

Oh yeah, totally. This isn’t free, because you’ve paid for it with your precious time and supposedly “valuable” data. And your projects must be extremely important too, which is why you’re relying on a free-tier service for them. Nvidia should obviously raise the RPM from 40 to 200 and pay us damages for the instability. What a good idea, right?

ikr. The sheer level of entitlement is just embarrassing to read. I won’t even be surprised if NVIDIA actually scale down the free endpoints in the future because of these users.

True. If a provider gives away free models, the service can’t be both stable and fast. What OP is basically saying is: ‘NVIDIA must give me free, fast, and stable models, and also pay me dollars to compensate for my time and data.’

OP acts like he’s a god and expects everyone to serve him, lol.

Tell me you haven’t read the terms of service without telling me. They explicitly said in many places the free endpoints are subjected to unannounced interruptions and downtimes. And to not use them for long-term production.

But yeah, keep being a whiny free-loader. I’m sure it’ll work out for you, bud. 🤡

Okay buddy. 😭 If you want to defend your billion-dollar company for feeding you dogshit, I won’t stop you.

Don’t get it twisted. We all want better and more reliable service.

The problem I have is that entitled and ignorant people like you keep coming in here just to complain and b*tch without actually contributing anything helpful. You haven’t even done the bare minimum of reading the ToS. So, maybe consider piping down.

(And in case you don’t know: Yes, we’re already getting scraps from NVIDIA. NIM is run on whatever available resources they can give us for free. Their priority is the paying customers, just like any other corp.)

True. For the vast majority of free-tier users, their data is virtually valueless because they can’t build meaningful corner cases or complex CoT prompts. It’s mostly just ‘hello world’ or ‘write a snake game’ test noise. The data cleaning cost alone would bankrupt the server team. So OP’s data isn’t secretly funding their trillion-dollar empire. It’s just the overhead of a free marketing campaign. If he wants stable and fast, deploy on NIM Cloud.

If you think even one person is simply using this service for ‘hello world’ or ‘write a snake game’ than you are that person, no one else. Don’t assume your mediocrity is the mean; it’s not.

First: it’s ‘if… THEN’, not ‘if… THAN’. Second: ‘hello world’ was a metaphor for low-entropy noise, not a census. Third: your prompts’ training value is zero whether you’re writing a snake game or curing cancer, because NVIDIA trains on labeled data, not on your intentions. Don’t assume your project’s seriousness is the data’s value; it’s not.

not to defend nvidia here since i have worse experience right now on free endpoints but beggars shouldn’t be entitled, no one is forcing you to use their service, either pay or leave