context_length_exceeded 🔌 API

LLM API: This model's maximum context length is x tokens. However, your messages resulted in y tokens (context_length_exceeded / prompt is too long)

The prompt plus requested output exceed the model’s context window.

Seen on: REST API

Meaning

Long chat histories, big documents or tool outputs accumulate tokens. Reserve room for the response (max_tokens) and trim, summarise or retrieve only relevant chunks.

Common causes

  • Conversation history growing unbounded
  • Whole documents pasted into the prompt
  • max_tokens too high leaving no room
  • Large tool/function outputs appended

⚡ Quick fix

  1. Truncate or summarise old messages
  2. Use retrieval to include only relevant chunks
  3. Count tokens before sending and lower max_tokens

Detailed fix by platform

Python

  1. python
    count = client.messages.count_tokens(model=MODEL, messages=messages).input_tokens
    while count > LIMIT - MAX_OUTPUT:
        messages.pop(1)   # drop oldest turns after the first
        count = client.messages.count_tokens(model=MODEL, messages=messages).input_tokens

How to diagnose

  1. Tokens — Prompt size vs limit
  2. History — Growing per turn?
  3. Output — max_tokens setting

🔧 Still not fixed?

Many errors look alike. If the steps above didn’t solve it, one of these is probably what you’re facing:

🧠 Still stuck? Analyze your error

Paste the full message, response headers or stack trace — we'll detect the platform and point to the most likely cause.