Deepseek-ai/deepseek-v4.1-flash fails even in official Playground — "Retries exhausted: 3/3"

The model deepseek-ai/deepseek-v4.1-flash (featured on build.nvidia.com’s homepage

under “Free inference with leading models”) is not responding reliably.

1. Via API (curl, chat completions endpoint): requests either time out completely

with 0 bytes received (tested up to 90s) or, rarely, succeed with 200 OK.

2. Via the official Playground UI (build.nvidia.com/deepseek-ai/deepseek-v4.1-flash/playground):

same result — “Error: Retries exhausted: 3/3” after a long wait, with no

NVIDIA code involved on my end at all.

This rules out any issue with my account, API key, or client code — even NVIDIA’s

own first-party Playground can’t get a reliable response from this model.

For comparison: GET /v1/models always responds instantly and correctly, and

models not assigned to my account correctly return a fast 404. This looks like

a capacity/serving issue specific to this model’s backend.

Could you check the health/capacity of this model’s serving infrastructure?

Screenshot attached: Playground showing “Retries exhausted: 3/3”.

I am getting “Gateway error” everytime in that model…

The deepseek-v4.1-flash, glm-5.3-flash, and Nemotron models have been throwing a “READ TIMEOUT” error (HTTPSConnectionPool(host=‘integrate.api.nvidia.com’, port=443): Read timed out. (read timeout=90)) since this morning, and there is no solution for it.

deepseek models currently seem to be fully down

I’m seeing the same kind of issue: if the official Playground also returns “Retries exhausted: 3/3,” it points more toward a model-serving/capacity problem than your client or API setup.
Hopefully NVIDIA can check the backend health and capacity for deepseek-ai/deepseek-v4.1-flash; intermittent 200s alongside long timeouts definitely suggests something unstable on the serving side.