GPT-5.3-Codex Pricing Explained
What You'll Learn
- The difference between GPT-5.3-Codex API prices and ChatGPT Codex credits
- The current input, cached-input, and output rates for both paths
- How context size, output length, fast mode, and task complexity affect usage
- How to estimate and control the cost of a real Codex workflow
GPT-5.3-Codex pricing is easy to misread because OpenAI publishes a dollar rate card for the API and a credit rate card for Codex inside ChatGPT. They describe related usage, but they are not the same invoice. A developer using an API key sees dollars per million tokens. A user signed in with ChatGPT sees plan access, usage limits, and credits when flexible billing applies.
The original article treated a set of April 2026 claims as permanent rules. It stated that OpenAI had changed every plan from message limits to one universal credit system, that Plus and Pro had fixed credit behaviour, and that GPT-5.3-Codex was a cheaper alternative to a newer flagship. The current OpenAI documentation is more conditional. Codex is included across ChatGPT plans, usage limits vary by plan, credits may be available after included limits, and API pricing is separate.
This is also a dated model guide. The current OpenAI model catalogue and pricing page now list newer model families. GPT-5.3-Codex remains documented as an agentic coding model with a 400,000-token context window and a 128,000-token maximum output, but a new project should compare it with the current supported model rather than assuming it is the default choice in August 2026.
GPT-5.3-Codex Pricing at a Glance
There are four numbers people usually need first. The direct API price is $1.75 per million input tokens, $0.175 per million cached input tokens, and $14 per million output tokens under standard processing. Fast mode uses a higher rate of $3.50 per million input tokens, $0.35 per million cached input tokens, and $28 per million output tokens.
Inside ChatGPT Work and Codex, OpenAI's rate card expresses the same model in credits. GPT-5.3-Codex uses 43.75 credits per million input tokens, 4.375 credits per million cached input tokens, and 350 credits per million output tokens. These credits are not dollars. The relationship between a credit balance and a user's payment depends on the plan, workspace, and credit arrangement.
| Usage path | Input | Cached input | Output |
|---|---|---|---|
| OpenAI API standard | $1.75 per 1M tokens | $0.175 per 1M tokens | $14.00 per 1M tokens |
| OpenAI API fast mode | $3.50 per 1M tokens | $0.35 per 1M tokens | $28.00 per 1M tokens |
| ChatGPT Work and Codex credits | 43.75 credits per 1M | 4.375 credits per 1M | 350 credits per 1M |
| Model page context | 400,000-token context and 128,000-token maximum output | ||
The same model name can therefore appear in two different cost conversations. Do not multiply ChatGPT credits by the API dollar rate or describe 43.75 credits as $43.75. Use the API table for an API-key deployment and the credit table for supported ChatGPT Work or Codex usage.
Two Billing Systems You Must Separate
The API is a metered developer service. The application sends requests to an OpenAI endpoint and is billed for the tokens used by the selected model. The current model page lists GPT-5.3-Codex on Chat Completions, Responses, and other supported endpoints, and shows standard input, cached input, and output prices on the model page.
Codex inside ChatGPT is a product feature included across Free, Go, Plus, Pro, Business, and Enterprise plans. The official help page says usage limits vary by plan. The current Codex pricing page describes limited trial access for Free and Go, expanded usage for Plus, and maximum Codex tasks for Pro. The exact allowance is not a single permanent number in the article because OpenAI can vary limits by plan, task type, model, and product surface.
Business and Enterprise workspaces can also have a flexible credit arrangement. OpenAI's rate card says standard Business seats use included plan limits first and can continue from a shared credit pool after an included limit when credits have been purchased. Codex-only seats include Codex usage and require workspace credits from first use. Enterprise and Edu workspaces on flexible pricing scale usage with credits instead of relying on one fixed rate limit.
The site's Codex versus Claude Code comparison is a useful related read for separating a model price from the full cost of the surrounding coding product.
GPT-5.3-Codex API Price per Million Tokens
For a direct API integration, the official OpenAI pricing page lists GPT-5.3-Codex under specialized Codex models. Standard processing is $1.75 per million input tokens, $0.175 for cached input, and $14 for output. Fast mode doubles those visible rates to $3.50, $0.35, and $28.
The model page adds important technical context. GPT-5.3-Codex accepts text and image input, returns text, supports reasoning effort settings of low, medium, high, and xhigh, and has a 400,000-token context window with a 128,000-token maximum output. The page describes it as optimized for agentic coding tasks in Codex or similar environments.
API pricing is not a monthly subscription. A request with a small prompt and a long generated patch has a different cost from a request that sends a large repository context and receives a short answer. Tool calls can add their own fees where applicable. OpenAI's pricing page also says that tokens used by built-in tools are billed at the chosen model's rates, while some tools have a separate per-call charge.
For teams using the Responses API, Chat Completions, or Batch API, record the endpoint and processing mode in the cost log. The model page lists Batch as a supported endpoint. The price table on the API page is the source of truth for the selected service tier on the day of the request, not an old calculator or a third-party tracker.
| API variable | Standard | Fast mode | Planning effect |
|---|---|---|---|
| Input tokens | $1.75 per 1M | $3.50 per 1M | Repository context and prompt size matter |
| Cached input | $0.175 per 1M | $0.35 per 1M | Repeated context can cost less |
| Output tokens | $14 per 1M | $28 per 1M | Long patches and explanations cost more |
| Context limit | 400,000 tokens | 400,000 tokens | Capacity is not the same as successful retrieval |
| Maximum output | 128,000 tokens | 128,000 tokens | Set an application-level output budget |
ChatGPT Codex Credits and Plan Access
ChatGPT Codex usage is not described by one fixed monthly token allowance that applies to every plan. The official pricing page says Codex is included in Free, Go, Plus, Pro, Business, and Enterprise. Free and Go have limited trial access. Plus provides expanded Codex usage. Pro provides maximum Codex tasks with the allowance shown for the selected plan and product setting.
The help documentation says Codex usage limits vary by plan and tells users to check the pricing page, the usage dashboard, or the limit banner. A user may be able to add credits, use a reset, upgrade, or wait for the displayed reset time. Eligible Plus and Pro users can buy credits without changing plans. This is different from claiming that every user automatically receives a fixed number of credits per month.
For Business and Enterprise, the shared pool matters. Codex, ChatGPT Work, Workspace Agents, ChatGPT for Excel, and ChatGPT for PowerPoint can draw from a common agentic usage and credit pool when those features are available on the plan. The actual debit depends on the model, task size, cached input, output, automations, fast mode, and concurrent instances.
| Plan or route | What the official pages say | Cost interpretation |
|---|---|---|
| Free | Limited trial access to Codex | Use the displayed allowance and reset information |
| Go | Limited trial access with more access than Free | Limits are product-specific and can change |
| Plus | Expanded Codex usage | Included usage first, with credits or other options where eligible |
| Pro | Maximum Codex tasks | Check the current plan page for the selected Pro tier |
| Business and Enterprise | Included limits or flexible credit pricing depending on workspace setup | Admin and workspace settings determine the billing route |
Checking Limits Before a Long Run
Before starting a large repository task, open Codex settings or the usage dashboard and record the allowance, remaining balance, reset time, model, and client. OpenAI says the displayed limit notice can show whether the next step is to add credits, use a reset, upgrade, or wait. This check prevents a long cloud task from stopping halfway through a migration.
Limits also depend on where the work runs. Local messages, cloud tasks, code reviews, automations, and delegated workers can use different counters or shared pools depending on the plan. A plan comparison should therefore name the surface being measured rather than treating one message count as the total capacity of the account.
Cached Input and Prompt Reuse
Cached input is the lowest token rate in both price systems. For the API, GPT-5.3-Codex cached input is $0.175 per million tokens under standard processing and $0.35 under fast mode. For supported ChatGPT Work and Codex usage, the rate card lists 4.375 credits per million cached input tokens compared with 43.75 credits for ordinary input.
A cache discount is not the same as a guarantee that the entire workspace will be cached. The request must contain reusable context in a form the service can recognize. Cache writes are not charged in the Business and Enterprise rate card, but the input and output tokens used by the task still count. The effective saving depends on how much context is repeated and how often the cached path is reused.
For a coding agent, the best cache candidate is stable repository context, instructions, or a repeated file set. The least useful candidate is a constantly changing prompt that invalidates most of the prior context. Keep the system instructions stable, separate dynamic task data from reusable project context, and record cache-read tokens in the request usage object.
Do not promise that a warm terminal session always gives a 90 percent saving. The official rate card gives the token rates, but the actual mix depends on the model, task, context, and product surface. A workflow can still be expensive when output is long, tools are called repeatedly, or multiple Codex instances run at once.
Context Window, Output Limit, and Reasoning
GPT-5.3-Codex has a 400,000-token context window and a 128,000-token maximum output according to its OpenAI model page. Those limits are separate. A request can fit inside the context window and still produce a shorter answer because the application, endpoint, or model configuration sets a lower output budget.
The model supports low, medium, high, and xhigh reasoning effort settings. A higher reasoning setting can change latency, token consumption, and completion quality. It should be part of the test record rather than treated as a free switch. If a task is a small code edit, the extra reasoning may not pay for itself. If the task involves a multi-file dependency change, a stronger setting may reduce later correction turns.
The context window also does not mean that every token is equally useful. Large generated repository listings, lock files, compiled assets, and repeated tool output can crowd out the information needed for a decision. The practical cost is therefore a function of selection, compression, cache reuse, output length, and failure recovery, not just the maximum context size.
For a broader look at agent tool boundaries, the site's agentic coding guide covers the surrounding workflow rather than one token price in isolation.
Fast Mode and Other Cost Modifiers
OpenAI's API pricing page lists a fast mode for GPT-5.3-Codex at $3.50 per million input tokens, $0.35 per million cached input tokens, and $28 per million output tokens. The page also says that priority processing was renamed Fast mode on July 30, 2026, while the request parameter can use either the legacy priority name or the fast name.
Fast mode is a speed choice, not a different model. Use it for latency-sensitive work where the faster response has measurable value. Do not enable it across every development request without measuring the effect on user wait time and total spend. A queue that is already fast enough may not justify the higher token rate.
Regional processing can also alter the price. OpenAI says eligible models released on or after March 5, 2026 can receive a 10 percent uplift when a data-residency endpoint is selected. OpenAI models hosted through Amazon Bedrock are billed through AWS and may differ from direct OpenAI pricing. These are deployment choices, so they belong in the cost record beside the model name.
Tool use is another modifier. The API pricing page lists web search at $10 per 1,000 calls plus search content tokens at model rates and lists a separate tool-call price. A Codex task that reads a repository, calls tools, retries after a failed command, and generates a patch can cost more than a plain text completion even when the final answer is short.
How to Estimate a Codex Task
Use separate rows for ordinary input, cached input, output, fast mode, tool calls, and retries. For the API, multiply each token category by its dollar rate and add any applicable tool or regional charge. For ChatGPT Work and Codex, use credits per million tokens and then compare the result with the workspace credit balance or included allowance.
As a simple API example, suppose a task sends 200,000 ordinary input tokens, 800,000 cached input tokens, and generates 20,000 output tokens under standard processing. The token charge is $0.35 for ordinary input, $0.14 for cached input, and $0.28 for output, for a total of $0.77 before any tool fee. The same task under fast mode is $0.70 plus $0.28 plus $0.56, for $1.54 before tools.
The equivalent ChatGPT Work or Codex credit estimate is 8.75 credits for ordinary input, 3.5 credits for cached input, and 7 credits for output, for a total of 19.25 credits. This is a usage estimate, not a dollar conversion. The amount paid for a credit depends on the plan or workspace arrangement.
| Example component | Quantity | API standard calculation | Codex credit calculation |
|---|---|---|---|
| Ordinary input | 200,000 tokens | 0.2 × $1.75 = $0.35 | 0.2 × 43.75 = 8.75 credits |
| Cached input | 800,000 tokens | 0.8 × $0.175 = $0.14 | 0.8 × 4.375 = 3.5 credits |
| Output | 20,000 tokens | 0.02 × $14 = $0.28 | 0.02 × 350 = 7 credits |
| Total before tools | 1,020,000 tokens | $0.77 | 19.25 credits |
The example shows why output control matters. A coding agent that returns a short patch and a concise test report may cost less than one that rewrites an entire file in the answer. At the same time, artificially limiting output can create another tool call or correction turn. The target is the lowest total cost for an accepted result, not the smallest first response.
What the Old Article Got Wrong
The old article said GPT-5.3-Codex used 43.75 credits for input, 350 credits for output, and a 90 percent cached-input discount. Those values are consistent with the current ChatGPT Work and Codex rate card, but they are not API dollar prices. The rewrite keeps the credit figures and adds the direct API rates of $1.75, $0.175, and $14.
The old article also claimed that OpenAI had moved every Plus, Pro, Business, and Enterprise user from message limits to a single credit system on April 2, 2026. The current help page says Codex is included across plans and that limits vary by plan. The rate card explains that flexible pricing applies to particular Business and Enterprise arrangements. The rewrite therefore avoids presenting one billing path as universal.
The old article used $20 per month for Plus, $100 per month for Pro, and a five-times Pro limit as fixed facts. The current Codex pricing page presents plan access and relative usage statements, including 5x or 20x more usage than Plus for Pro, while the help page points users to the current plan page and usage dashboard. Static monthly prices and one fixed multiplier are not used here as universal numbers.
The old article called GPT-5.3-Codex a faster, cheaper alternative to GPT-5.4 and recommended it for 90 percent of daily programming tasks. The official model page supports its agentic coding use and current specifications, but it does not establish that universal task share or a complete model-to-model cost verdict. The decision should be measured against the current model catalogue and the team's own workload.
How to Reduce Codex Credit and API Spend
Start with context selection. Exclude generated files, dependency caches, build output, and unrelated documentation from the agent's working set. A smaller relevant context reduces input tokens and can improve the chance that the model sees the right instruction. The site's 30-day coding assistant review provides a related task-testing perspective, but the cost calculation must use the current rate card.
Next, control output. Ask for a patch, a test result, and a short explanation rather than a full copy of every changed file. Set a practical maximum output for each endpoint. For a repository task, ask the agent to inspect first, propose a plan, make the smallest change, and run only the relevant checks.
Use caching deliberately. Keep stable instructions and repeated repository context in the reusable part of the request. Track the ratio of cached input to ordinary input. If cache reuse is low, a prompt restructure may deliver more saving than a plan upgrade.
Finally, monitor failures. A failed command, repeated tool call, or unnecessary retry can cost more than the successful generation. Store model ID, service tier, token counts, cache counts, reasoning setting, tool calls, latency, and accepted-result status for each representative task. This creates a real cost baseline rather than a rate-card guess.
Teams building at the edge can compare these hosted API economics with the site's Cloudflare Workers AI guide. The deployment model changes the cost categories, so a token price alone cannot decide the architecture.
Final Verdict on GPT-5.3-Codex Pricing
GPT-5.3-Codex has a clear direct API price: $1.75 per million input tokens, $0.175 per million cached input tokens, and $14 per million output tokens under standard processing. Fast mode is priced at $3.50, $0.35, and $28. In ChatGPT Work and Codex, the same model is represented by 43.75, 4.375, and 350 credits per million tokens.
Codex is included across the current ChatGPT plan range, but included access and usage limits vary. Business and Enterprise workspaces may use flexible credit pricing, shared pools, or plan limits depending on the workspace arrangement. Plus and Pro users can have additional options when a limit is reached, including credits or a reset, but the exact choice is account-specific.
The model page lists a 400,000-token context window, a 128,000-token maximum output, image input, and several reasoning effort settings. Those capabilities can make GPT-5.3-Codex suitable for agentic coding, but they do not make it the automatic cheapest or best model for every task. Compare the current model catalogue, then measure accepted patches, correction turns, token mix, cache reuse, tool calls, and latency.
The correct answer to “Does ChatGPT Codex cost money?” is conditional. Codex is included in ChatGPT plans, while usage is governed by plan limits and may draw on credits. The correct answer to “How much does a GPT-5.3-Codex token cost?” depends on whether the user means API dollars or ChatGPT credits. Keep those units separate and the estimate becomes auditable.
For a related view of autonomous development systems, see the site's AI agent swarms guide and Computer Use and MCP guide.
Published: April 22, 2026 | Last Updated: August 20, 2026 | Author: SK Jabedul Haque
For more updates on AI and technology, join our community on WhatsApp.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles