Cohere Aya Expanse and Tiny Aya 2026
What You'll Learn
- The difference between Aya research, Aya Expanse and Tiny Aya
- Current language, platform and API facts from Cohere sources
- How Aya Expanse token pricing compares with a simple workload
- How to choose between a hosted API and a compact local model
Aya Expanse and Tiny Aya are different products
Cohere's Aya work covers more than one model. The Cohere Labs research page describes Aya as an open multilingual research model covering 101 languages, including more than 50 underserved languages. Cohere's model documentation separately describes Aya Expanse as an API family covering 23 languages. Tiny Aya is a compact multilingual model listed by Cohere as supporting 70 languages.
These numbers should not be merged into one headline. The 101-language figure refers to the wider Aya research work, 23 refers to Aya Expanse in the current model documentation and 70 refers to Tiny Aya. For a broader view of current model pricing, see this AI model pricing comparison.
What Aya Expanse is designed to do
Cohere positions the Aya family around multilingual generative AI. The model documentation says Aya Expanse is available through the Chat endpoint and is intended to expand the number of languages covered by generative AI. This makes it relevant to translation, multilingual support, regional content and cross-language retrieval tasks.
Aya Expanse should not be described as a universal translation engine without testing. Language quality can differ by language pair, domain, prompt format and evaluation set. A support team should test terminology, names, numbers, formatting and safety behavior for each target language before production use.
What Tiny Aya adds
Tiny Aya is the compact option in this comparison. Cohere's model search result describes it as a 3.35 billion parameter multilingual model supporting 70 languages. A smaller model can be useful when memory, latency or local deployment matters more than maximum general capability.
Compact does not mean cost-free. A local deployment still needs suitable hardware, inference software, storage, monitoring and model updates. Compare the total operating cost with a hosted API rather than looking only at the absence of a per-token invoice. The AI deployment comparison provides a useful way to think about hosted versus self-managed tools.
Verified Aya Expanse API pricing
| Model family | Input per 1M tokens | Output per 1M tokens | Source status |
|---|---|---|---|
| Aya Expanse 8B and 32B | $0.50 | $1.50 | Cohere pricing page result |
| Tiny Aya | Not stated in the fetched pricing result | Not stated in the fetched pricing result | Compact model, verify access route |
The $0.50 input and $1.50 output rates are the figures returned for Aya Expanse 8B and 32B on Cohere's pricing page. They are US dollar list figures, not a complete invoice estimate. Taxes, account terms, platform charges, quotas and any marketplace billing can change the final amount.
Aya Expanse cost example
Suppose a multilingual support workflow sends 10 million input tokens and produces 2 million output tokens in one month. At the listed rates, input costs 10 multiplied by $0.50, or $5. Output costs 2 multiplied by $1.50, or $3. The token subtotal is therefore $8 before tax, retries, tool calls or other charges.
This example is only a calculation from the published rates. A real project should record language, prompt length, output length, retries, successful answers and human corrections. The relevant measure is cost per accepted result, not only cost per million tokens.
Languages, evaluation and translation quality
Language count is a useful starting point but not a quality score. A model integration may support a language at the API level while producing weaker results for legal terms, local names, mixed scripts or technical instructions. Build a test set for each language and include both direct translation and task completion prompts.
For multilingual support, test whether the model preserves dates, currencies, product names and HTML or Markdown structure. Also test code switching. A customer may ask in one language, quote a product name in English and include a local address. A production evaluation must reflect that real input. See the cost testing guide for a comparable workflow.
Hosted API or compact deployment
| Requirement | Hosted Aya Expanse | Compact Tiny Aya path |
|---|---|---|
| Fast start | Usually simpler through an API endpoint | Requires model and runtime setup |
| Per-request billing | Yes, based on usage and account terms | May shift cost into hardware and operations |
| Data control | Review provider and platform terms | More control when run inside approved infrastructure |
| Scaling | Depends on quotas and platform capacity | Depends on available compute and concurrency |
Choose the hosted path when the team needs a quick integration, managed capacity and a clear API workflow. Consider a compact deployment when data must stay within controlled infrastructure and the team can operate the model stack. Cohere's documentation lists its proprietary platform, Amazon SageMaker, Amazon Bedrock, Microsoft Azure and Oracle GenAI Service as available platforms for its model portfolio.
Use cases for regional and multilingual teams
Aya Expanse can be tested for multilingual customer support, cross-language search, translation assistance, regional content drafts and classification. The exact task should determine the evaluation. A model that translates a short FAQ well may still need more testing for long documents or domain-specific answers.
Use retrieval when the answer depends on current company information. The model should not be treated as a live source of product policy, prices or availability. Keep the source documents separate, record their date and test whether citations or extracted fields remain accurate.
API cost controls
Set input and output limits, remove repeated instructions, cache stable reference text where the platform supports it and log usage by feature. Add usage alerts before a multilingual campaign starts. Long prompts multiplied across many languages can increase the bill quickly even when each individual response looks small.
Review failed calls and retries. A low list price can become expensive when the application repeats a request after a timeout or asks several models to produce the same answer. Measure the full completed workflow and compare it with a compact local test when that option is practical.
Open research model versus commercial API
The Aya research work and the Aya Expanse API serve different needs. A research release can help laboratories study multilingual modeling and open evaluation. An API product is intended for an integration with access, billing and platform terms. Tiny Aya adds a compact option, but local operation still requires technical ownership.
Do not copy the language count, parameter count or price from one Aya product into another. Keep the model name and source URL beside each claim. Cohere's Aya research page and model documentation should be read separately.
How to evaluate Aya for a real project
Start with a fixed set of prompts across the languages that matter to the business. Score factual accuracy, translation fidelity, formatting, refusal behavior, latency and human correction time. Then compare Aya Expanse with the current production model and, if relevant, Tiny Aya on the same test set.
Use a cost sheet that records input tokens, output tokens, API retries, hardware time and review time. This prevents a language count or benchmark headline from deciding a deployment without evidence. The Claude Fable cost guide uses the same cost-per-accepted-result approach.
Cohere Aya Expanse and Tiny Aya 2026 decision guide
Use Aya Expanse when a hosted multilingual API fits the project's language, privacy and billing requirements. Test Tiny Aya when a compact model and controlled infrastructure are more important than managed API convenience. Treat Aya's 101-language research figure, Aya Expanse's documented 23 languages and Tiny Aya's 70 languages as separate facts.
The verified pricing snapshot lists Aya Expanse 8B and 32B at $0.50 per million input tokens and $1.50 per million output tokens. Check Cohere's pricing and model documentation before deployment because model availability, language coverage, quotas and rates may change.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles