Claude 4 "Maximum Token" Error in Long Documents
What You'll Learn
- How `max_tokens` interacts with thinking tokens and the context window
- Why changing a thinking budget can cause a 400 error or cache miss
- How to process long documents and PDFs without guessing at a line limit
- Which API checks and request reductions provide evidence before a retry
What the maximum token error actually means
The phrase “maximum token” can describe more than one failure. A request may exceed the model's context window, leave too little room for a requested response, use a thinking configuration that the selected model does not support, or send a document payload that is too large for the request. The visible error may mention tokens, but the corrective action depends on which boundary was crossed.
Anthropic's API separates the context window from the output setting. The context window includes the system prompt, message history, tool results, images, documents, and the output Claude generates. The `max_tokens` parameter controls the output ceiling for the turn. On models with thinking, the output includes thinking and response text, so the final answer has less room when internal reasoning consumes part of the ceiling.
| Boundary | What counts | Typical response |
| Request size | Messages, tools, images, and documents | Reduce or split the input |
| Context window | History plus current input and generated output | Compact or start a focused thread |
| Output ceiling | Thinking plus visible response tokens | Raise `max_tokens` or reduce thinking demand |
| Model feature support | Allowed thinking and parameter configuration | Use the model's documented request shape |
How `max_tokens` interacts with thinking
In manual extended thinking, Anthropic uses a `thinking` object with `type` set to `enabled` and a `budget_tokens` value. The documentation says the budget must be at least 1,024 tokens and normally less than `max_tokens` because thinking tokens count toward the total output limit. A request that leaves no practical room for the final answer can stop early or fail validation even when the input document fits.
The budget is a target rather than a promise that Claude will consume every token. Actual usage varies by task. Anthropic recommends starting near the minimum for simple tasks and using a larger starting point for complex work, then measuring quality, latency, and token usage on the actual workload. Do not copy a thinking budget from a different model without checking its supported mode.
The official extended thinking documentation also warns that manual extended thinking is deprecated on Claude 4.6 models and rejected on Claude 4.7 and later models. Where adaptive thinking is supported, remove the manual budget and use the documented adaptive configuration instead.
Why a long document can trigger the error
Long-document work consumes more than the visible document text. Every prior user and assistant message, system instruction, tool result, image, PDF page, and tool definition counts toward the request context. Anthropic describes context as working memory and warns that more context is not automatically better because accuracy and recall can degrade as the prompt grows.
A conversation that started with a short question can become a large request after many revisions. If you then attach a PDF and request a long answer with thinking enabled, the new input and output ceiling may cross the model's available boundary. The correct fix is not always “increase max tokens.” Sometimes the history must be compacted, the document split, or the answer narrowed.
Keep a separate document-processing thread for a large report. Save the source and extracted notes outside the chat, then ask focused questions against the smallest relevant section. This is similar to the context discipline used in the site's reasoning model guide, where task cost and successful output are measured rather than assumed.
Fix 1: Read the complete error and request payload
Start with the HTTP status, error type, message, model ID, `max_tokens`, thinking object, output configuration, message count, and document attachments. A 400 response caused by an unsupported parameter needs a different repair from a context overflow. Do not remove fields at random because that makes the next result harder to interpret.
Log the request shape without logging private document contents or API keys. Record whether the request used manual extended thinking, adaptive thinking, tools, prompt caching, PDFs, images, or a long conversation. This gives you a reproducible test and allows a safe comparison after one setting changes.
If the request is sent through a framework, inspect the final JSON after defaults are applied. A client library or helper can add a thinking field, an old model name, or a large output setting even when the application code appears correct. The Claude agents guide provides useful context for checking model and tool configuration separately.
Fix 2: Make the output budget leave room for the answer
When manual thinking is supported, choose a thinking budget that leaves enough of `max_tokens` for the visible response. For a short extraction, start with a modest budget and request a concise answer. For a difficult analysis, increase the ceiling and measure the result. A large input does not automatically require a large visible answer, so specify the desired format and length.
| Workload | Safer first request | What to measure |
| Short extraction | Small thinking budget and concise output | Accuracy and completion |
| Long summary | Focused section with explicit format | Coverage and truncation |
| Code review | One module and a bounded review list | Actionable findings |
| Tool workflow | Minimal tools and a clear stop condition | Tool calls and final response room |
Do not claim that one number solves every request. The correct value depends on the model, task, document, thinking mode, tool loop, and desired response. Record the actual `usage` fields returned by the API and compare successful requests with failed ones.
Fix 3: Remove obsolete manual thinking settings during migration
Anthropic's migration documentation is explicit about model-specific behavior. Some current models use adaptive thinking, while manual `thinking: {type: "enabled", budget_tokens: N}` is deprecated or rejected depending on the model generation. Disabling thinking can also return a 400 error on models where adaptive thinking is always on.
For a migration, update the model name, remove an unsupported manual budget, use the documented adaptive configuration, and move depth control to the model's supported effort setting when available. Then re-run a small request before restoring the full document. This isolates a parameter compatibility issue from a context issue.
Thinking configuration also affects conversations and caching. Anthropic notes that changing a manual budget can invalidate cache breakpoints. A cache miss is not the same as a token-limit failure, but it can change latency and cost during a troubleshooting test. Keep the setting stable while comparing requests.
Fix 4: Split long documents by task and preserve citations
For a long report, split by meaningful structure such as chapter, contract section, table group, or date range. Give each section an identifier and ask for a bounded result. Then combine the section notes in a separate synthesis request. This reduces the active context and makes it easier to locate a missing or conflicting passage.
Do not split blindly at arbitrary character counts when a table or legal clause depends on the next page. Keep headings, footnotes, table labels, and references with the relevant section. If the final answer needs citations, preserve page or section identifiers in every intermediate note.
Use a focused question for each pass. “Extract obligations and cite their section numbers” is more testable than “understand this entire PDF.” Once the extraction is complete, ask a separate synthesis question that uses the notes rather than the original full conversation.
PDF upload limits and visual document handling
Anthropic's PDF support documentation says Claude can analyze text, pictures, charts, and tables in standard PDFs. It lists a maximum request size of 32 MB and up to 600 PDF pages for requests with a 1M-token context window, or 100 pages when the context window is 200k tokens. Those are request limits, not a guarantee that a dense document will fit comfortably.
| PDF condition | Risk | Practical response |
| Dense small-font pages | Context fills before page limit | Split or downsample embedded images |
| Large repeated file | Payload and encoding overhead | Use the Files API when appropriate |
| Charts and visual tables | Text-only extraction misses layout | Use visual PDF processing and verify citations |
| Encrypted PDF | Document cannot be processed normally | Provide an accessible standard PDF |
The official PDF support guide says dense PDFs can fill the context window before reaching the page limit. Treat page count as one input signal, not as a safe promise that the request will succeed.
Context window overflow versus output truncation
These outcomes look similar but need different fixes. If the input alone exceeds the context window, Anthropic documents a 400 `invalid_request_error` with a prompt-too-long message. On newer model families, a request can be accepted when input plus `max_tokens` exceeds the window, then stop with a context-window stop reason if generation reaches the limit. The response metadata is essential.
Output truncation means the request began but the generated content ended before the intended answer was complete. Narrow the format, lower unnecessary thinking demand, or continue from a structured checkpoint. Context overflow means the request could not safely fit or reached the context boundary, so reduce history, documents, tools, or requested output.
| Evidence | Likely problem | First repair |
| 400 prompt-too-long error | Input exceeds context | Reduce history or split the document |
| Unsupported thinking error | Model configuration mismatch | Use documented thinking mode |
| Context-window stop reason | Generation reached the boundary | Lower input or output scope |
| Short complete response | Task may be narrower than expected | Check the requested format and evidence |
Token counting, usage logs, and cache behavior
Use Anthropic's token counting API before sending a large request when you need an estimate. After the request, inspect the usage object. Thinking tokens are billed as output tokens and count toward rate limits. In a cached conversation, cache creation and cache read tokens still count toward the context window even though their billing treatment differs.
Store only the measurements needed for diagnosis. A useful record includes model, input token estimate, output tokens, thinking token details when exposed, cache status, document size, page count, stop reason, and latency. Avoid storing the document body in ordinary application logs unless your privacy controls allow it.
Changing several settings at once prevents a useful comparison. Change one of history length, document split, thinking mode, output ceiling, or tool set, then rerun the same small test. The site's AI coding cost guide uses the same principle of measuring successful work instead of reading one token number in isolation.
When to use compaction or a fresh conversation
If the thread has accumulated many tool results and revisions, a fresh conversation may be the fastest safe test. Carry forward a compact brief with the document identifier, verified findings, unresolved questions, and citation map. Do not paste every failed attempt into the new context.
Anthropic documents server-side compaction for long-running conversations on supported models. Compaction can summarize earlier context so work continues past the point where a raw transcript becomes inefficient. It should preserve the facts and references needed for the next task, but you should verify critical details after compaction.
For a one-off document, splitting and summarizing may be simpler than redesigning the application. For a recurring agent workflow, context editing, compaction, token counting, and prompt caching should be part of the architecture. Choose the smallest solution that matches the recurring failure. The same staged approach is useful in long-running AI agent workflows, where state and context need explicit control.
Safe request patterns for long-document analysis
Use a staged pattern. First, count or estimate the input. Second, extract facts from bounded sections. Third, validate section coverage and citations. Fourth, synthesize the final answer with a clear length and evidence requirement. This creates checkpoints that survive a failed final request.
For PDFs, preserve page references. For code, preserve file names and line ranges. For contracts, preserve clause identifiers. For research reports, preserve source URLs and dates. A short answer with traceable evidence is more useful than a long answer that silently dropped half the document.
If the central problem is a parameter error, test the same model with a tiny text-only message. If that succeeds, add the thinking configuration, then the document, then the tool set one layer at a time. This order identifies the first failing layer without exposing private content unnecessarily.
Conclusion: fix the boundary, not the symptom
Claude 4 Maximum Token Error in Long Documents is not one universal failure. Check whether the request hit an unsupported thinking mode, an output ceiling, a context window, a PDF payload limit, or a long-running conversation boundary. Preserve the original document, reduce one variable at a time, inspect usage and stop reasons, and use the model's current documentation before restoring the full workload.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles