Skip to Content

xAI Model Retirement May 15, 2026

xAI Model Retirement: Redirects, Pricing, and Migration Steps
2026-05-19 07:52:07 Updated 2026-08-20 23:15:37.778448 — min read 870 views
xAI Model Retirement May 15, 2026
“The xAI model retirement on 15 May 2026 moved eight retired API slugs onto documented replacement paths. This guide explains each redirect, the Grok 4.3 pricing effect, reasoning settings, context limits, and a safe migration checklist for developers who need predictable behavior and billing.

What You'll Learn

  • Which eight xAI model slugs were retired and where each one is routed.
  • How automatic routing changes reasoning behavior and effective model pricing.
  • What Grok 4.3 provides for context length, reasoning control, and long prompts.
  • How to migrate, test, and monitor an xAI API integration without relying on silent compatibility behavior.

The xAI model retirement was an API compatibility change, not a simple deletion event. xAI's official migration notice says the eight listed slugs were retired on 15 May 2026 at 12:00 PM PT, while the slugs continued to resolve through automatic routing. That distinction matters because an application can remain available while its model, reasoning setting, and billable rate change underneath the request. Developers should read the official xAI migration notice as the source of truth, then update production code explicitly. For background on how model choice affects multi-step systems, see our guide to agentic AI.

How the xAI Model Retirement Changed API Behavior

The practical change is a layer of documented routing. A request that still names a retired slug may not fail with an unknown-model error. Instead, xAI can serve the request through a replacement path. This helps prevent an immediate outage, but it also means the old model name no longer describes the complete behavior of the request. A compatibility path is useful during a migration window. It is not a substitute for a controlled model configuration.

The migration notice separates reasoning, non-reasoning, coding, and image workloads. That separation corrects a common misunderstanding in older coverage, which treated every retired slug as a synonym for Grok 4.3. The code and image paths have different targets, so teams should inventory model names by workload rather than applying a single replacement across an entire repository.

This matters most in systems that record model names for cost reporting, evaluation, or compliance. If the configured slug is retired but the effective model is different, dashboards that group only by the requested string can understate the change. A safe migration records the requested model, the effective replacement when available, the reasoning effort, the token volume, and the date of the configuration change. Readers comparing broader API economics can also review our AI token pricing comparison, while keeping its figures separate from the xAI documentation used here.

The Eight Retired Model Slugs

The xAI documentation lists eight retired slugs. The table below keeps the retirement list and the documented post-retirement path together. The target is the behavior described by xAI, not a recommendation to leave production traffic on a compatibility redirect.

Retired slugDocumented replacement pathEffective setting
grok-4-1-fast-reasoninggrok-4.3low reasoning effort
grok-4-1-fast-non-reasoninggrok-4.3none reasoning effort
grok-4-fast-reasoninggrok-4.3low reasoning effort
grok-4-fast-non-reasoninggrok-4.3none reasoning effort
grok-4-0709grok-4.3low reasoning effort
grok-code-fast-1grok-build-0.1coding replacement
grok-3grok-4.3none reasoning effort
grok-imagine-image-progrok-imagine-image-qualityimage replacement

Two rows deserve special attention. The code model does not route to Grok 4.3 in the official recommendation. It routes to Grok Build 0.1. The image model also has its own destination, Grok Imagine Image Quality. Treating those two workloads as ordinary text requests can produce a migration that looks complete in configuration review but still points a coding or image pipeline at the wrong capability family.

How Automatic Redirects Work

For the reasoning entries, xAI documents a redirect to Grok 4.3 with low reasoning effort. For the non-reasoning entries, the redirect uses Grok 4.3 with none reasoning effort. These defaults preserve a broad distinction between deeper reasoning and lower-latency behavior, but they do not preserve every characteristic of the retired model. The effective response can differ in latency, output style, tool behavior, and cost.

Workload classAutomatic targetWhat to verify
Reasoning textgrok-4.3low effort is applied by the redirect
Non-reasoning textgrok-4.3none effort is applied by the redirect
Codinggrok-build-0.1coding tools and evaluation results
Image generationgrok-imagine-image-qualityimage endpoint and output behavior

Automatic routing is therefore best understood as a safety net. It gives teams time to change configuration, but it can hide a model transition from application code that checks only for HTTP success. Test logs should capture the returned model information if the API exposes it, and application-level telemetry should keep enough context to compare requests before and after migration.

Pricing Impact for API Users

The xAI documentation states that requests sent through a deprecated slug are billed at the effective replacement pricing. For Grok 4.3, the migration notice gives a standard rate of $1.25 per 1M input tokens and $2.50 per 1M output tokens. This is the key financial consequence of leaving an old slug in place. The request can continue to work while the pricing basis changes.

The current Grok 4.3 documentation also identifies cached input tokens at $0.20 per 1M tokens. The pricing page adds a long-context rate of $2.50 per 1M input tokens and $5.00 per 1M output tokens when the prompt reaches the documented long-context threshold of 200K tokens or more. Teams that send long documents should therefore test both normal and long-context paths instead of estimating cost from a short prompt alone.

Grok 4.3 pricing caseInputOutput
Standard context$1.25 per 1M tokens$2.50 per 1M tokens
Cached input$0.20 per 1M tokensNot applicable
Long context at 200K tokens or more$2.50 per 1M tokens$5.00 per 1M tokens

These figures describe xAI's published rates and not a forecast of a team's monthly bill. Actual spend depends on request volume, prompt length, output length, caching, long-context usage, and the replacement model used for code or image traffic. For a financial control, compare invoices and token telemetry with the model configuration rather than applying a percentage copied from an earlier article.

Why a Redirect Can Change Production Costs

A redirect changes the cost model because the bill follows the replacement path. The original request string can still say `grok-4-fast-non-reasoning`, while the service bills the request according to Grok 4.3. A code request can follow Grok Build 0.1 instead. The right question is not whether the old slug still responds. The right question is which model served the request and which pricing table applies.

Monitoring signalRisk if ignoredPractical control
Requested model slugOld names hide a retirementAlert on every retired slug
Effective modelCost and behavior are misclassifiedRecord the served model when available
Reasoning effortLatency and output depth driftSet the intended value explicitly
Context sizeLong prompts enter a different rate bandTrack prompt tokens and threshold crossings

This monitoring pattern is also useful when an application uses an agentic workflow. A planner can create more calls than a simple chat request, and each call may use a different model setting. Our embeddings API comparison explains why token and retrieval measurements need to remain separate from the text model decision.

The Recommended Replacement for Each Workload

Migration should follow the workload, not the old brand label. Reasoning text workloads should move to Grok 4.3 with an explicitly selected reasoning effort. Non-reasoning text workloads can use Grok 4.3 with none effort when the priority is lower latency. Coding workloads should evaluate Grok Build 0.1. Image workloads should use Grok Imagine Image Quality.

  • For reasoning text, set the model to `grok-4.3` and choose the effort that matches the task.
  • For non-reasoning text, set `grok-4.3` with none effort when that behavior is intentional.
  • For coding, assess `grok-build-0.1` with repository permissions, tests, and review controls.
  • For image generation, update the endpoint to `grok-imagine-image-quality` and recheck output requirements.

Do not treat a replacement name as proof of equivalence. Re-run the evaluations that matter for the product, such as structured output validity, tool-call success, latency, refusal behavior, code tests, image quality, or long-document extraction. If the application depends on a feature that was not part of the retired model, make that dependency explicit in the migration record.

Reasoning Effort and Context Window

Grok 4.3 provides a 1,000,000-token context window and supports none, low, medium, and high reasoning effort. The automatic retirement path uses low for the documented reasoning models and none for the documented non-reasoning models. Teams that need deeper reasoning should not assume the compatibility default is the correct production setting. Set the parameter deliberately and measure the resulting quality and latency.

Context capacity is not the same as context efficiency. A large window can hold a long document, but a long prompt still affects cost and can make retrieval, caching, and evaluation harder. The pricing documentation says long-context rates apply when the prompt reaches 200K tokens or more. This makes prompt segmentation, retrieval filtering, and cache reuse relevant engineering decisions, not just prompt-writing preferences.

For developers building retrieval systems, compare the model context policy with the data pipeline. A million-token limit can support broad analysis, while a smaller retrieved set may be easier to audit. Readers who want a broader explanation of autonomous workflows can use our agentic AI guide as conceptual background, but the xAI documentation remains the authority for Grok 4.3 settings.

What Developers Should Change in Code

Start with configuration, then inspect every code path that can override it. Search repositories, environment variables, deployment manifests, test fixtures, and prompt routers for all eight retired strings. Include image and coding integrations because their destinations differ from the ordinary text paths. A single hard-coded model in a background worker can keep the retirement problem alive after the main application has been updated.

  1. Replace retired model strings with workload-specific model identifiers.
  2. Set reasoning effort explicitly wherever a text workflow depends on response depth or latency.
  3. Record model, reasoning effort, prompt tokens, output tokens, and request outcome in telemetry.
  4. Run regression tests for structured outputs, function calling, retries, timeouts, and application error handling.
  5. Deploy the change behind a controlled release and compare cost and quality against the previous baseline.

Do not update only the user-facing application. Scheduled jobs, evaluation scripts, administrative tools, and emergency runbooks often retain old model names. A repository-wide search followed by a configuration diff is a simple way to reduce that risk. For developers comparing related AI workflows, our computer-use guide provides a separate example of why permissions and evaluation belong in the deployment plan.

A Safe Migration Plan

A safe migration can be staged without treating the redirect as an emergency outage. First, create an inventory of requested slugs and workload classes. Second, map each path to the official replacement. Third, select reasoning effort and context policies. Fourth, test representative traffic with production-like prompts. Fifth, release gradually with billing and quality monitoring enabled.

Use a small test set that includes short prompts, long prompts, tool calls, structured outputs, retries, and any image or coding flow the application supports. Keep the prompts fixed while comparing the old baseline with the explicit replacement where that comparison is meaningful. Record pass criteria before the test begins so that a successful HTTP response does not become the only definition of success.

After rollout, keep a temporary alert for retired slugs. The alert should fire if any service, worker, or third-party integration continues to request a retired identifier. Once the alert remains quiet and the billing data matches the expected model mix, remove the compatibility exception from the runbook. Keep the retirement date and source link in the change record.

Billing and Testing Checklist

Before closing the migration ticket, review the request path and the invoice path together. Confirm that the configured model matches the intended workload, that long-context calls are visible, and that cached input is measured separately if the application relies on it. A cost review should explain changes in tokens and model mix rather than merely comparing two monthly totals.

  • Confirm that no service still requests a retired slug.
  • Confirm the explicit target for text, coding, and image workloads.
  • Confirm the chosen reasoning effort for each text route.
  • Confirm prompt and output token telemetry.
  • Confirm tests for structured outputs and tool calls.
  • Confirm long-context behavior at the documented threshold.
  • Confirm billing alerts and a rollback plan.

Keep a copy of the official migration page with the engineering record. xAI can change model availability and pricing over time, so future maintenance should use the current documentation rather than this article as a permanent configuration reference. The same principle applies to adjacent coverage such as AI search optimization, where implementation details can change as platform behavior evolves.

Common Misreadings to Avoid

The first misreading is that every retired model becomes Grok 4.3. The official table does not say that. Coding and image requests have separate targets. The second misreading is that a successful response proves nothing changed. Compatibility routing can preserve availability while changing the model and bill. The third misreading is that the standard Grok 4.3 rate applies to every prompt. The pricing page documents a long-context band at 200K tokens or more.

The fourth misreading is that low reasoning effort is a universal best setting. It is the documented redirect default for reasoning models, not a quality guarantee for every application. The fifth misreading is that a large context window removes the need for retrieval discipline. More capacity can be useful, but it also increases the importance of token monitoring, data selection, and evaluation.

Finally, avoid presenting the retirement as a claim about xAI's entire product strategy. The official sources support a precise API migration story. They do not require speculation about corporate motives, future model launches, or guaranteed savings. A narrow article with traceable facts is more useful to an engineer than a dramatic article built from unsupported predictions. For a separate comparison of research workflows, see our Gemini and Perplexity guide.

Conclusion: Migrate Explicitly Before Costs Surprise You

The xAI model retirement left a compatibility path in place, but it also changed the effective model path for eight retired slugs. Reasoning and non-reasoning text requests route to Grok 4.3 with different defaults, coding requests route to Grok Build 0.1, and image requests route to Grok Imagine Image Quality. Explicit model selection, intentional reasoning settings, workload-specific tests, and billing telemetry are the practical response.

Update the configuration, test the replacement, monitor the served model and token mix, and keep the official migration notice in the change record. That approach protects availability without confusing availability with equivalence. It also gives the engineering team a repeatable method for the next model retirement or pricing change.

Frequently Asked Questions

xAI retired grok-4-1-fast-reasoning, grok-4-1-fast-non-reasoning, grok-4-fast-reasoning, grok-4-fast-non-reasoning, grok-4-0709, grok-code-fast-1, grok-3, and grok-imagine-image-pro.
The retired slugs continue to resolve through automatic routing after the retirement. That can prevent an immediate break, but developers should still select an explicit replacement so the effective model, settings, and pricing are intentional.
Requests to the retired reasoning models are served by grok-4.3 with low reasoning effort. Teams that need deeper reasoning should select the effort level explicitly instead of relying on the compatibility default.
Requests to the retired non-reasoning models are served by grok-4.3 with none reasoning effort. This preserves a lower-reasoning route, but it does not prove that every behavior of the old model is identical.
xAI recommends grok-build-0.1 for coding workloads. A migration should test repository access, coding tools, tests, review controls, latency, and output quality before the new model is used in production.
xAI documents grok-imagine-image-quality as the replacement path for grok-imagine-image-pro. Image applications should update the endpoint and recheck output requirements rather than treating the change as an ordinary text-model migration.
The standard Grok 4.3 rate is $1.25 per 1M input tokens and $2.50 per 1M output tokens. The model page lists a 1,000,000-token context window, and the pricing page documents a long-context rate when a prompt reaches 200K tokens or more.
SK Jabedul Haque
Written by

SK Jabedul Haque

Founder & Chief Editor

Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.

Read full bio

Never miss an update

Get our clearest explainers on schemes, markets and money — read what matters, without the noise.

Explore more articles
In this article