LLM API: This model's maximum context length is x tokens. However, your messages resulted in y tokens (context_length_exceeded / prompt is too long)
The prompt plus requested output exceed the model’s context window.
Seen on:
REST API
Meaning
Long chat histories, big documents or tool outputs accumulate tokens. Reserve room for the response (max_tokens) and trim, summarise or retrieve only relevant chunks.
Common causes
- Conversation history growing unbounded
- Whole documents pasted into the prompt
- max_tokens too high leaving no room
- Large tool/function outputs appended
⚡ Quick fix
- Truncate or summarise old messages
- Use retrieval to include only relevant chunks
- Count tokens before sending and lower max_tokens
Detailed fix by platform
Python
- python
count = client.messages.count_tokens(model=MODEL, messages=messages).input_tokens while count > LIMIT - MAX_OUTPUT: messages.pop(1) # drop oldest turns after the first count = client.messages.count_tokens(model=MODEL, messages=messages).input_tokens
How to diagnose
- Tokens — Prompt size vs limit
- History — Growing per turn?
- Output — max_tokens setting
🔧 Still not fixed?
Many errors look alike. If the steps above didn’t solve it, one of these is probably what you’re facing:
Similar errors
Most viewed in API
Other ways to find it
🧠 Still stuck? Analyze your error
Paste the full message, response headers or stack trace — we'll detect the platform and point to the most likely cause.
Was this page helpful?
Report a correction or suggest an improvement
Last updated 7 Oct 2026