Skip to Content

Claude 4 "Maximum Token" Error in Long Documents

5 Verified Fixes for PDF Uploads and Context Window Failures
2026-08-20 21:37:18 Updated 2026-08-20 21:37:18.513454 — min read 279 views
Claude 4 "Maximum Token" Error in Long Documents
Claude 4 Maximum Token Error in Long Documents usually means the request has conflicting output, thinking, or context settings. Anthropic documents `max_tokens` as a hard output ceiling and says thinking tokens count toward it. This guide shows how to diagnose the request, protect long documents, and choose a safe fix without hiding the real limit.

What You'll Learn

  • How `max_tokens` interacts with thinking tokens and the context window
  • Why changing a thinking budget can cause a 400 error or cache miss
  • How to process long documents and PDFs without guessing at a line limit
  • Which API checks and request reductions provide evidence before a retry

What the maximum token error actually means

The phrase “maximum token” can describe more than one failure. A request may exceed the model's context window, leave too little room for a requested response, use a thinking configuration that the selected model does not support, or send a document payload that is too large for the request. The visible error may mention tokens, but the corrective action depends on which boundary was crossed.

Anthropic's API separates the context window from the output setting. The context window includes the system prompt, message history, tool results, images, documents, and the output Claude generates. The `max_tokens` parameter controls the output ceiling for the turn. On models with thinking, the output includes thinking and response text, so the final answer has less room when internal reasoning consumes part of the ceiling.

BoundaryWhat countsTypical response
Request sizeMessages, tools, images, and documentsReduce or split the input
Context windowHistory plus current input and generated outputCompact or start a focused thread
Output ceilingThinking plus visible response tokensRaise `max_tokens` or reduce thinking demand
Model feature supportAllowed thinking and parameter configurationUse the model's documented request shape

How `max_tokens` interacts with thinking

In manual extended thinking, Anthropic uses a `thinking` object with `type` set to `enabled` and a `budget_tokens` value. The documentation says the budget must be at least 1,024 tokens and normally less than `max_tokens` because thinking tokens count toward the total output limit. A request that leaves no practical room for the final answer can stop early or fail validation even when the input document fits.

The budget is a target rather than a promise that Claude will consume every token. Actual usage varies by task. Anthropic recommends starting near the minimum for simple tasks and using a larger starting point for complex work, then measuring quality, latency, and token usage on the actual workload. Do not copy a thinking budget from a different model without checking its supported mode.

The official extended thinking documentation also warns that manual extended thinking is deprecated on Claude 4.6 models and rejected on Claude 4.7 and later models. Where adaptive thinking is supported, remove the manual budget and use the documented adaptive configuration instead.

Why a long document can trigger the error

Long-document work consumes more than the visible document text. Every prior user and assistant message, system instruction, tool result, image, PDF page, and tool definition counts toward the request context. Anthropic describes context as working memory and warns that more context is not automatically better because accuracy and recall can degrade as the prompt grows.

A conversation that started with a short question can become a large request after many revisions. If you then attach a PDF and request a long answer with thinking enabled, the new input and output ceiling may cross the model's available boundary. The correct fix is not always “increase max tokens.” Sometimes the history must be compacted, the document split, or the answer narrowed.

Keep a separate document-processing thread for a large report. Save the source and extracted notes outside the chat, then ask focused questions against the smallest relevant section. This is similar to the context discipline used in the site's reasoning model guide, where task cost and successful output are measured rather than assumed.

Fix 1: Read the complete error and request payload

Start with the HTTP status, error type, message, model ID, `max_tokens`, thinking object, output configuration, message count, and document attachments. A 400 response caused by an unsupported parameter needs a different repair from a context overflow. Do not remove fields at random because that makes the next result harder to interpret.

Log the request shape without logging private document contents or API keys. Record whether the request used manual extended thinking, adaptive thinking, tools, prompt caching, PDFs, images, or a long conversation. This gives you a reproducible test and allows a safe comparison after one setting changes.

If the request is sent through a framework, inspect the final JSON after defaults are applied. A client library or helper can add a thinking field, an old model name, or a large output setting even when the application code appears correct. The Claude agents guide provides useful context for checking model and tool configuration separately.

Fix 2: Make the output budget leave room for the answer

When manual thinking is supported, choose a thinking budget that leaves enough of `max_tokens` for the visible response. For a short extraction, start with a modest budget and request a concise answer. For a difficult analysis, increase the ceiling and measure the result. A large input does not automatically require a large visible answer, so specify the desired format and length.

WorkloadSafer first requestWhat to measure
Short extractionSmall thinking budget and concise outputAccuracy and completion
Long summaryFocused section with explicit formatCoverage and truncation
Code reviewOne module and a bounded review listActionable findings
Tool workflowMinimal tools and a clear stop conditionTool calls and final response room

Do not claim that one number solves every request. The correct value depends on the model, task, document, thinking mode, tool loop, and desired response. Record the actual `usage` fields returned by the API and compare successful requests with failed ones.

Fix 3: Remove obsolete manual thinking settings during migration

Anthropic's migration documentation is explicit about model-specific behavior. Some current models use adaptive thinking, while manual `thinking: {type: "enabled", budget_tokens: N}` is deprecated or rejected depending on the model generation. Disabling thinking can also return a 400 error on models where adaptive thinking is always on.

For a migration, update the model name, remove an unsupported manual budget, use the documented adaptive configuration, and move depth control to the model's supported effort setting when available. Then re-run a small request before restoring the full document. This isolates a parameter compatibility issue from a context issue.

Thinking configuration also affects conversations and caching. Anthropic notes that changing a manual budget can invalidate cache breakpoints. A cache miss is not the same as a token-limit failure, but it can change latency and cost during a troubleshooting test. Keep the setting stable while comparing requests.

Fix 4: Split long documents by task and preserve citations

For a long report, split by meaningful structure such as chapter, contract section, table group, or date range. Give each section an identifier and ask for a bounded result. Then combine the section notes in a separate synthesis request. This reduces the active context and makes it easier to locate a missing or conflicting passage.

Do not split blindly at arbitrary character counts when a table or legal clause depends on the next page. Keep headings, footnotes, table labels, and references with the relevant section. If the final answer needs citations, preserve page or section identifiers in every intermediate note.

Use a focused question for each pass. “Extract obligations and cite their section numbers” is more testable than “understand this entire PDF.” Once the extraction is complete, ask a separate synthesis question that uses the notes rather than the original full conversation.

PDF upload limits and visual document handling

Anthropic's PDF support documentation says Claude can analyze text, pictures, charts, and tables in standard PDFs. It lists a maximum request size of 32 MB and up to 600 PDF pages for requests with a 1M-token context window, or 100 pages when the context window is 200k tokens. Those are request limits, not a guarantee that a dense document will fit comfortably.

PDF conditionRiskPractical response
Dense small-font pagesContext fills before page limitSplit or downsample embedded images
Large repeated filePayload and encoding overheadUse the Files API when appropriate
Charts and visual tablesText-only extraction misses layoutUse visual PDF processing and verify citations
Encrypted PDFDocument cannot be processed normallyProvide an accessible standard PDF

The official PDF support guide says dense PDFs can fill the context window before reaching the page limit. Treat page count as one input signal, not as a safe promise that the request will succeed.

Context window overflow versus output truncation

These outcomes look similar but need different fixes. If the input alone exceeds the context window, Anthropic documents a 400 `invalid_request_error` with a prompt-too-long message. On newer model families, a request can be accepted when input plus `max_tokens` exceeds the window, then stop with a context-window stop reason if generation reaches the limit. The response metadata is essential.

Output truncation means the request began but the generated content ended before the intended answer was complete. Narrow the format, lower unnecessary thinking demand, or continue from a structured checkpoint. Context overflow means the request could not safely fit or reached the context boundary, so reduce history, documents, tools, or requested output.

EvidenceLikely problemFirst repair
400 prompt-too-long errorInput exceeds contextReduce history or split the document
Unsupported thinking errorModel configuration mismatchUse documented thinking mode
Context-window stop reasonGeneration reached the boundaryLower input or output scope
Short complete responseTask may be narrower than expectedCheck the requested format and evidence

Token counting, usage logs, and cache behavior

Use Anthropic's token counting API before sending a large request when you need an estimate. After the request, inspect the usage object. Thinking tokens are billed as output tokens and count toward rate limits. In a cached conversation, cache creation and cache read tokens still count toward the context window even though their billing treatment differs.

Store only the measurements needed for diagnosis. A useful record includes model, input token estimate, output tokens, thinking token details when exposed, cache status, document size, page count, stop reason, and latency. Avoid storing the document body in ordinary application logs unless your privacy controls allow it.

Changing several settings at once prevents a useful comparison. Change one of history length, document split, thinking mode, output ceiling, or tool set, then rerun the same small test. The site's AI coding cost guide uses the same principle of measuring successful work instead of reading one token number in isolation.

When to use compaction or a fresh conversation

If the thread has accumulated many tool results and revisions, a fresh conversation may be the fastest safe test. Carry forward a compact brief with the document identifier, verified findings, unresolved questions, and citation map. Do not paste every failed attempt into the new context.

Anthropic documents server-side compaction for long-running conversations on supported models. Compaction can summarize earlier context so work continues past the point where a raw transcript becomes inefficient. It should preserve the facts and references needed for the next task, but you should verify critical details after compaction.

For a one-off document, splitting and summarizing may be simpler than redesigning the application. For a recurring agent workflow, context editing, compaction, token counting, and prompt caching should be part of the architecture. Choose the smallest solution that matches the recurring failure. The same staged approach is useful in long-running AI agent workflows, where state and context need explicit control.

Safe request patterns for long-document analysis

Use a staged pattern. First, count or estimate the input. Second, extract facts from bounded sections. Third, validate section coverage and citations. Fourth, synthesize the final answer with a clear length and evidence requirement. This creates checkpoints that survive a failed final request.

For PDFs, preserve page references. For code, preserve file names and line ranges. For contracts, preserve clause identifiers. For research reports, preserve source URLs and dates. A short answer with traceable evidence is more useful than a long answer that silently dropped half the document.

If the central problem is a parameter error, test the same model with a tiny text-only message. If that succeeds, add the thinking configuration, then the document, then the tool set one layer at a time. This order identifies the first failing layer without exposing private content unnecessarily.

Conclusion: fix the boundary, not the symptom

Claude 4 Maximum Token Error in Long Documents is not one universal failure. Check whether the request hit an unsupported thinking mode, an output ceiling, a context window, a PDF payload limit, or a long-running conversation boundary. Preserve the original document, reduce one variable at a time, inspect usage and stop reasons, and use the model's current documentation before restoring the full workload.

Frequently Asked Questions

It can mean an output ceiling, context window overflow, unsupported thinking configuration, or an oversized document request. Check the status, error type, model, max_tokens, thinking settings, and document payload before changing the request.
Yes. Anthropic documents thinking tokens as part of the total output governed by max_tokens. A manual budget must normally leave room for the final visible response.
Copy the source and split the document by meaningful sections. Extract bounded facts with page or section identifiers, then synthesize the notes in a separate focused request.
A fixed limit should not be assumed across Claude models or platforms. Output limits, thinking support, context windows, and model generations differ, so check the current model documentation and the request response metadata.
Yes. PDF pages, images, charts, tables, message history, tools, and requested output all consume request capacity. Anthropic documents a 32 MB request limit and page limits that depend on context window size.
Changing a manual thinking budget can invalidate prompt-cache breakpoints. It can also be unsupported on newer models where adaptive thinking replaces manual extended thinking, so confirm the model's supported configuration.
Use a focused new conversation when the old thread contains many revisions, tool results, or oversized document context. Carry forward a compact evidence brief, citation map, and unresolved questions instead of the entire transcript.
SK Jabedul Haque
Written by

SK Jabedul Haque

Founder & Chief Editor

Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.

Read full bio

Never miss an update

Get our clearest explainers on schemes, markets and money — read what matters, without the noise.

Explore more articles
In this article