LLM API: overloaded_error (529) / The server is overloaded / Service unavailable — model temporarily overloaded
The provider is temporarily at capacity; the request was rejected even though your key and quota are fine.
Seen on:
REST API
Meaning
Overload errors (HTTP 529 at Anthropic, 503/“The engine is currently overloaded” elsewhere) are transient. Retry with exponential backoff and jitter, and consider fallback models or batch/off-peak processing.
Common causes
- Provider capacity spike
- Very large requests during peak hours
- No retry/backoff in the client
⚡ Quick fix
- Retry with exponential backoff and jitter (SDKs do this; raise max_retries)
- Use a fallback model or region
- Use batch processing for non-urgent work
Detailed fix by platform
Python
client = anthropic.Anthropic(max_retries=5) # SDK retries 429/5xx/529 with backoff
How to diagnose
- Status — 529/503 vs 429
- Timing — Peak periods?
- Retries — Backoff configured?
🔧 Still not fixed?
Many errors look alike. If the steps above didn’t solve it, one of these is probably what you’re facing:
Similar errors
Most viewed in API
Other ways to find it
🧠 Still stuck? Analyze your error
Paste the full message, response headers or stack trace — we'll detect the platform and point to the most likely cause.
Was this page helpful?
Report a correction or suggest an improvement
Last updated 7 Oct 2026