The model deepseek-ai/deepseek-v4.1-flash (featured on build.nvidia.com ’s homepage
under “Free inference with leading models”) is not responding reliably.
1. Via API (curl, chat completions endpoint): requests either time out completely
with 0 bytes received (tested up to 90s) or, rarely, succeed with 200 OK.
2. Via the official Playground UI (build.nvidia.com/deepseek-ai/deepseek-v4.1-flash/playground ):
same result — “Error: Retries exhausted: 3/3” after a long wait, with no
NVIDIA code involved on my end at all.
This rules out any issue with my account, API key, or client code — even NVIDIA’s
own first-party Playground can’t get a reliable response from this model.
For comparison: GET /v1/models always responds instantly and correctly, and
models not assigned to my account correctly return a fast 404. This looks like
a capacity/serving issue specific to this model’s backend.
Could you check the health/capacity of this model’s serving infrastructure?
Screenshot attached: Playground showing “Retries exhausted: 3/3”.
b_casco
September 25, 2026, 4:59pm
2
I am getting “Gateway error” everytime in that model…
The deepseek-v4.1-flash, glm-5.3-flash, and Nemotron models have been throwing a “READ TIMEOUT” error (HTTPSConnectionPool(host=‘integrate.api.nvidia.com ’, port=443): Read timed out. (read timeout=90)) since this morning, and there is no solution for it.
dapolch
September 26, 2026, 2:21am
4
deepseek models currently seem to be fully down
rootsystemtechnology:
The model deepseek-ai/deepseek-v4.1-flash (featured on build.nvidia.com ’s homepage
under “Free inference with leading models”) is not responding reliably.
1. Via API (curl, chat completions endpoint): requests either time out completely
with 0 bytes received (tested up to 90s) or, rarely, succeed with 200 OK.
2. Via the official page Playground UI (build.nvidia.com/deepseek-ai/deepseek-v4.1-flash/playground ):
same result — “Error: Retries exhausted: 3/3” after a long wait, with no
NVIDIA code involved on my end at all.
This rules out any issue with my account, API key, or client code — even NVIDIA’s
own first-party Playground can’t get a reliable response from this model.
For comparison: GET /v1/models always responds instantly and correctly, and
models not assigned to my account correctly return a fast 404. This looks like
a capacity/serving issue specific to this model’s backend.
Could you check the health/capacity of this model’s serving infrastructure?
Screenshot attached: Playground showing “Retries exhausted: 3/3”.
I’m seeing the same kind of issue: if the official Playground also returns “Retries exhausted: 3/3,” it points more toward a model-serving/capacity problem than your client or API setup.
Hopefully NVIDIA can check the backend health and capacity for deepseek-ai/deepseek-v4.1-flash; intermittent 200s alongside long timeouts definitely suggests something unstable on the serving side.