Crowdin
/Blog

List of the Best LLMs for Translation: Choosing the Right Model

•Last updated: •31 min read
Best LLMs for translation

There’s a big variety of Large Language Models, but which one to consider for your tasks? This guide provides a comparison of the top large language models for translation, information about how top LLMs differ in speed, cost, and the list of supported languages.

You will also explore how to build an effective localization strategy and get insights about a hybrid approach of using different LLMs for different tasks, depending on your needs.

This guide covers general-purpose and translation-specialized large language models. If you’re comparing dedicated machine translation engines instead – DeepL, Google Translate, Microsoft Translator and similar – see our separate guide to the best machine translation software.

So, stay with us to discover how Crowdin offers flexibility in selecting various LLMs within a single platform.

The list of best llms for translation that Crowdin platform utilizes

What are large language models?

LLMs are general-purpose text models trained on large volumes of human language, which is why they can translate well despite none of them being built specifically for translation. That’s exactly why a comparison like this one is needed: unlike a dedicated MT engine, every LLM strikes its own balance between quality, speed and cost, and you can route different content types to different models.

LLMs for translation at a glance

If you only want the numbers, this is the whole article in one table. Everything below it is the reasoning behind these picks.

ModelContext windowInput / Output (per 1M tokens)Approx. cost per 10,000 wordsUse it for
Claude Fable 5.11M$10.00 / $50.00~$1.04 *The hardest multi-stage reasoning and technical localization
GPT-6 Astra1.05M$10.00 / $50.00~$0.80Legal, medical, high-value marketing; complex agentic workflows
Claude Opus 5.51M$4.00 / $20.00~$0.42 *Creative and brand-sensitive content, long-form consistency
Claude Sonnet 51M$2.00 / $10.00~$0.21 *General-purpose business localization
Gemini 3.1 Pro1.05M$2.00 / $12.00~$0.19Nuanced translation plus multimodal (screenshots, video, audio)
GPT-6 Sol1.05M$2.00 / $10.00~$0.16The everyday default: websites, UI strings, documentation
Qwen-MT-Plus16K≈$2.54 / $7.62~$0.14Chinese, Japanese, Korean, Vietnamese; strict terminology control
Claude Haiku 4.5200K$1.00 / $5.00~$0.08High-frequency strings, support tickets, user reviews
Muse Spark 1.31.05M$1.25 / $4.25~$0.073Meta’s flagship: long-horizon agentic and multimodal work
Gemini 3.8 Flash1.05M$0.75 / $3.75~$0.06High-volume continuous localization (promotional rate until 2027)
Mistral Large 3256K$0.50 / $1.50~$0.027Quality work that must stay in the EU
Gemini 3.1 Flash-Lite1.05M$0.25 / $1.50~$0.023Bulk, low-risk streams; live chat
Llama 4 Maverick1M~$0.25 / ~$0.70 (hosted)~$0.013Open-weight, multimodal, self-hostable
Mistral Small 4256K$0.15 / $0.60~$0.010Cheap multilingual bulk processing, Apache 2.0
Qwen-MT-Turbo16K$0.16 / $0.49~$0.0087High-throughput Asian-language pipelines
GPT-6 Luna1.05M$0.10 / $0.50~$0.008Social feeds, chat logs, first-pass drafts
Llama 4 Scout10M native~$0.09 / ~$0.28 (hosted)~$0.005Very long documents
Meta NLLB-200512Free (self-hosted)Compute only200 languages, including low-resource ones

* Claude figures include a tokenizer adjustment

Context windows are the vendors’ published input limits. Two of them are outliers worth noticing before you plan a workflow: Qwen-MT tops out at 16,384 tokens and NLLB-200 at 512, because both are purpose-built translation models that work sentence by sentence rather than holding a whole document in memory. That is fine for string-based localization and a real constraint for long-form consistency.

Costs assume 13,333 tokens × (input rate + output rate) at standard tier, and they are a floor, not an estimate: translating into Ukrainian, Japanese, Korean, Arabic or other token-hungry languages costs more than the table shows, because that extra lands on the output side. Full method and caveats.

Let’s look at what each of the current families offers.

GPT-6 family (Astra, Sol, Luna)

The latest models from OpenAI reached general availability with GPT-6, a three-tier lineup (Astra, Sol, and Luna) that replaced the GPT-5.5/GPT-5.4 generation as OpenAI’s current flagship line. Pricing dropped meaningfully compared to the previous generation, especially at the entry tier, and every model in the family carries the same 1,050,000-token context window (922K input plus 128K output). GPT-5.6 remains available for select cybersecurity (“Daybreak”) workloads, but for general localization work GPT-6 is now the current recommendation.

GPT-6 models comparison for translation needs

GPT-6 AstraGPT-6 SolGPT-6 Luna
Ideal forDeep reasoning, complex tasks, and content like legal or medical documents. Also the strongest choice for agentic, multi-step localization workflows.General-purpose translation and high-volume, moderately complex tasks like website content and product documentation.High-volume, low-risk content, such as social media feeds and real-time chat. Perfect for initial drafts and simple tasks like text classification.
Cost (standard, per 1M tokens)Short context: $10.00 input / $50.00 output. Long-context requests: $20.00 input / $75.00 output.Short context: $2.00 input / $10.00 output. Long-context requests: $4.00 input / $15.00 output.Short context: $0.10 input / $0.50 output. Long-context requests: $0.20 input / $0.75 output.
SpeedBuilt for deep, thorough reasoning; the slowest of the three tiers.Balanced speed and quality; OpenAI’s recommended default tier.Fastest model in the family, optimized for ultra-low latency.
Context Window1,050,000 tokens (922K input + 128K output). Long-context rates apply above 272K input tokens.1,050,000 tokens (922K input + 128K output), same threshold.1,050,000 tokens (922K input + 128K output), same threshold.
Key StrengthOpenAI’s deepest-reasoning tier, aimed at nuanced and multi-step work.The best price-to-performance tier for the bulk of everyday localization work.Lowest per-token cost in the current OpenAI lineup.

Source: OpenAI’s official pricing page – always double-check current rates there.

GPT-6 Astra: Flagship model

This is the most capable model in the GPT-6 family, designed for deep reasoning and complex tasks.

When to Use GPT-6 Astra for Translation:

  • High-Stakes and Nuanced Content: Use Astra for translating legal documents, medical reports, or high-value marketing campaigns where a single misinterpretation can have serious consequences.
  • Complex Workflows: If your localization process involves multiple steps – such as analyzing a technical diagram, extracting the text, translating it, and then placing it back into a new image – Astra is the best choice.
  • Long-Form Content: Its 1,050,000-token window holds entire books or large technical manuals at once, keeping terminology and style consistent. Note that requests above 272K input tokens switch to the higher long-context rate.

Cost: $10.00 per million input tokens and $50.00 per million output tokens for short-context requests (approx. $0.80 per 10,000 words); long-context requests run $20.00 / $75.00 (approx. $1.27 per 10,000 words).

GPT-6 Sol: Workhorse model

This model offers an excellent balance of performance, speed, and cost, and is OpenAI’s suggested default for most production traffic.

When to Use GPT-6 Sol for Translation:

  • General-Purpose Translation: Sol is a good choice for the bulk of your website content, UI strings, and product documentation.
  • High-Volume, Moderately Complex Tasks: For projects that need both speed and good quality, such as translating a large batch of knowledge base articles or support tickets.

Cost: $2.00 per million input tokens and $10.00 per million output tokens, costing only approx. $0.16 per 10,000 words for short-context work.

GPT-6 Luna: Speed and cost champion

This is the smallest and fastest model in the family, optimized for ultra-low latency and minimal cost.

When to Use GPT-6 Luna for Translation:

  • High-Volume, Low-Risk Content: Use Luna for content where speed and cost are the primary concerns – social media feeds, user-generated content, or real-time chat.
  • Initial Drafting: A very quick, machine-generated first pass on a large document, which a human translator can then refine.
  • Simple Tasks: Text classification (e.g., categorizing customer reviews by language) or short-form summarization.

Cost: $0.10 per million input tokens and $0.50 per million output tokens, making massive localization projects extremely cheap at approx. $0.008 per 10,000 words.

  • Critical Content: Use GPT-6 Astra for your most important, high-value translation work, such as legal or creative content.
  • High-Volume Content: Use GPT-6 Sol for the bulk of your general translation needs.
  • Low-Risk Content: Use GPT-6 Luna for basic, high-volume translations where minimal cost is the main concern.

Choose right GPT-6 model for translation

Do you need to move off GPT-5.x?

Not necessarily. OpenAI keeps the previous generations on its price list rather than retiring them the moment a new family ships, and for translation specifically the jump from GPT-5.x to GPT-6 is less dramatic than the version numbers suggest. If your existing pipeline is tuned around a GPT-5.x model and your quality checks pass, there’s no urgency to migrate – the practical reasons to move are cost and context window, not a cliff in translation quality.

GPT-5.6 Sol in particular is still worth a look at $4.00 input / $20.00 output per million tokens (approx. $0.32 per 10,000 words), and OpenAI notes this is promotional pricing available at least through November 21, 2026. The older GPT-5.5 and GPT-5.4 families remain listed too.

One caveat before you budget around any of them: third-party pricing trackers disagree noticeably on the older GPT-5.x rates, and we found figures for the same model that differed by several times. Open the All models tab on OpenAI’s pricing page and read the rate there rather than trusting a comparison table – including ours.

Gemini 3 family (3.1 Pro, 3.8 Flash, 3.1 Flash-Lite)

Gemini 3 family from Google continues to expand quickly – Gemini 3.8 Flash, released September 2, 2026, is now the newest and most capable Flash-tier model, sitting alongside the still-current Gemini 3.1 Pro and the budget-focused Gemini 3.1 Flash-Lite. Built natively as multimodal engines, these models process and reason across text, images, video, and audio simultaneously.

Choosing the right Gemini model for your translation tasks

Gemini 3.1 ProGemini 3.8 FlashGemini 3.1 Flash-Lite
Ideal ForHigh-stakes and nuanced translations, complex structural reasoning, massive long-form documents, and rich multimodal datasets.High-volume, low-latency enterprise localization, general-purpose translation (websites, UI), and continuous integration pipelines.High-throughput, cost-sensitive automated streams, real-time customer support chat, and lightning-fast initial drafts.
Cost (per 1M tokens)$2.00 input / $12.00 output for requests ≤200K tokens; $4.00 / $18.00 above 200K.$0.75 input / $3.75 output (promotional rate through Dec 31, 2026; doubles to $1.50 / $7.50 on Jan 1, 2027).$0.25 input / $1.50 output (text).
SpeedTailored for deep analytical thinking and multi-step tasks, resulting in a more measured response time.Fast, agent-oriented tier, built for long-horizon software and localization workflows.The fastest model in the lineup, engineered for ultra-low latency and high-frequency queries.
Context Window1,048,576 input tokens, 64K output.1,048,576 input tokens, 64K output.1,048,576 input tokens, 65,535 output.
Key StrengthAdvanced reasoning and near-perfect structural alignment.Newest, most intelligent Flash tier – built for agentic, long-horizon work at a fraction of Pro’s cost.Maximum throughput and cost efficiency for continuous data streams and repetitive string updates.

Source: Google’s official Gemini API pricing page

Gemini 3.1 Pro: Most capable

Gemini 3.1 Pro remains Google’s flagship reasoning model for translation-adjacent work.

When to Use Gemini 3.1 Pro for Translation:

  • Complex Reasoning and Nuance: Best choice for translating legal documentation, compliance frameworks, medical manuals, and creative marketing copy.
  • Massive Context: With a 1,048,576-token input window, it can ingest entire product repositories or massive books at once. The whole Gemini 3 line shares this window, so the choice between Pro, Flash and Flash-Lite comes down to reasoning quality, speed and cost – not how much text you can fit.
  • Multimodal Tasks: Natively multimodal, so it can read text, graphics, and video simultaneously – useful for localizing UI screenshots or subtitled video.

Cost: $2.00 per million input tokens and $12.00 per million output tokens for requests up to 200K tokens. For a 10,000-word localization task, the cost is approx. $0.19.

Gemini 3.8 Flash: New workhorse

Gemini 3.8 Flash is Google’s newest Flash-tier model – “our most intelligent Flash model,” in Google’s own description – built for long-horizon software engineering, autonomous agents, and complex enterprise workflows, at the same promotional price as the 3.7 and 3.6 Flash models that came before it.

When to Use Gemini 3.8 Flash for Translation:

  • High-Volume, Low-Latency Tasks: Ideal for high-frequency agile pipelines, continuous software localization, and dynamic web updates.
  • Cost-Efficiency: A strong candidate for moving the bulk of a standard localization workload off the Pro tier while the current promotional pricing lasts.
  • Automation: Runs complex automated localization tasks inside workflows without friction.

Cost: $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026 (approx. $0.06 per 10,000 words); rising to $1.50 / $7.50 from January 1, 2027.

Gemini 3.1 Flash-Lite: Cheapest option

Gemini 3.1 Flash-Lite is a cost-efficient model designed for handling massive datasets and high-frequency translation queues on a limited budget. (Google also lists a slightly pricier Gemini 3.5 Flash-Lite at $0.30 / $2.50 – the 3.1 tier remains the cheaper of the two.)

When to Use Gemini 3.1 Flash-Lite for Translation:

  • Bulk Translation: Highly repetitive, low-risk translation needs where cost is the overriding factor, such as internal wikis or large legacy logs.
  • Real-Time Data Streams: Perfect for user-generated product reviews, forum moderation, and live chat support.

Cost: $0.25 per million input tokens and $1.50 per million output tokens, reducing massive datasets to approx. $0.023 per 10,000 words.

Choose right Gemini model for translation

  • Critical Content: Route your core brand books, legal agreements, and complex master-files to Gemini 3.1 Pro.
  • High-Volume Content: Deploy Gemini 3.8 Flash for the vast majority of your active product documentation, applications, and web UI.
  • Low-Risk Content: Offload real-time user chat logs and customer reviews to Gemini 3.1 Flash-Lite.

Anthropic Claude family

Claude models from Anthropic remain highly regarded in AI localization for contextual fluency and stylistic nuance. The lineup you’d reference today is meaningfully different from a year ago: Claude Sonnet 5, Claude Opus 5.5, and Claude Haiku 4.5 are the current self-serve models, alongside the higher-cost Claude Fable 5.1 tier. All prices below are Anthropic’s own official rates – and worth noting, Sonnet 5’s launch price is now the permanent standard price rather than an introductory one.

Claude model comparison for translation

Claude Fable 5.1Claude Opus 5.5Claude Sonnet 5Claude Haiku 4.5
Ideal ForThe most demanding automated and multi-stage reasoning tasks, including complex software/technical localization.High-stakes corporate documentation, literary or creative marketing translation, and multi-document synthesis.General-purpose business localization, high-volume documentation, UI strings, and continuous integration pipelines.High-frequency string processing, real-time customer support chat logs, user reviews, and low-risk bulk datasets.
Cost (per 1M tokens)Input: $10.00. Output: $50.00.Input: $4.00. Output: $20.00.Input: $2.00. Output: $10.00.Input: $1.00. Output: $5.00.
Cached input$0.25 (2.5% of base input – the cheapest cache rate Anthropic offers).$0.20 (5% of base input).$0.20 (10% of base input).$0.10 (10% of base input).
Context Window1M tokens at standard pricing.1M tokens at standard pricing.1M tokens at standard pricing.200,000 tokens.
Key StrengthAnthropic’s highest-cost tier, positioned above Opus for the hardest reasoning and localization-pipeline tasks.Anthropic’s newest Opus-tier model, positioned for creative and brand-sensitive work – and cheaper than the Opus tier used to be.The industry-standard price-to-performance tier, now with its launch price made permanent.Fastest, cheapest current Claude tier – built for high-throughput, low-risk content.

Source: Anthropic’s official pricing page – this is the single most reliable source for Claude pricing since Anthropic updates this page directly.

Claude Fable 5.1: Premium tier

Fable sits above Opus in Anthropic’s lineup and is the model to reach for when quality has to be as close to certain as possible.

When to Choose Claude Fable 5.1:

  • Complex Software and Technical Localization: Built for engineering contexts with intricate syntax, embedded code, or highly structured schemas.
  • Advanced, Multi-Stage Pipelines: Ideal where the AI needs to assess extensive regulatory or brand criteria before executing localization changes.

Cost: $10.00 per million input tokens and $50.00 per million output tokens. Translating a 10,000-word dataset costs approx. $1.04 once the tokenizer adjustment is applied ($0.80 on the raw word-to-token rule of thumb).

Claude Opus 5.5: Standard for creative nuance

Opus 5.5 is Anthropic’s current Opus-tier model, priced notably lower than the Opus 5/4.8 generation it replaces.

When to Choose Claude Opus 5.5:

  • High-Stakes and Creative Content: The preferred option for marketing collateral, financial audits, and legal agreements, minimizing the need for human post-editing.
  • Long-form Consistency: Effective for continuous documentation streams that need stylistic alignment across separate files.

Cost: $4.00 per million input tokens and $20.00 per million output tokens – approx. $0.42 per 10,000-word translation task with the tokenizer adjustment applied.

Claude Sonnet 5: Enterprise workhorse

Sonnet 5 remains the standard choice for high-volume enterprise translation, and its once-introductory $2/$10 rate is now Anthropic’s permanent price (the previously scheduled increase to $3/$15 did not happen).

When to Choose Claude Sonnet 5:

  • General-Purpose Production: Suitable for website translation, application UIs, customer knowledge bases, and general corporate communications.
  • Agile Integration: Low latency fits smoothly into continuous localization setups needing immediate, automated translation.

Cost: $2.00 per million input tokens and $10.00 per million output tokens, keeping high-volume queues efficient at approx. $0.21 per 10,000 words with the tokenizer adjustment applied.

Claude Haiku 4.5: High-throughput alternative

For workflows where rapid processing and low operational cost matter more than advanced stylistic prose, Anthropic’s current Haiku is Claude Haiku 4.5.

When to Use Claude Haiku 4.5:

  • Real-Time Communications: Efficient for live customer support tickets, platform notifications, or chat transcripts.
  • High-Volume Low-Risk Assets: Useful for processing large text queues like user-generated reviews or e-commerce feedback streams.

Cost: $1.00 per million input tokens and $5.00 per million output tokens, clearing a 10,000-word queue for approx. $0.08. Haiku 4.5 uses the earlier tokenizer, so no adjustment applies here.

Choose right Claude model for translation

An effective localization approach relies on dynamic model routing rather than a single engine:

  • Premium Layer: Direct high-stakes legal contracts, corporate announcements, and core creative messaging to Claude Fable 5.1 or Claude Opus 5.5.
  • Core Layer: Route the bulk of your application interfaces, service documentation, and standard web content through Claude Sonnet 5.
  • Scale Layer: Offload real-time user-generated data, chat utilities, and raw data dumps to Claude Haiku 4.5.

Meta: Muse Spark, Llama 4, and NLLB-200

Meta’s line-up has changed more than any other vendor’s since this guide was last updated. Meta spent 2026 moving away from a purely open-weight strategy: its new proprietary flagship, released under the Meta Superintelligence Labs banner (initially code-named “Avocado”), shipped publicly as Muse Spark, now at version 1.3, available only through Meta’s own API. The open-weight Llama 4 family (Scout and Maverick) is still available and unchanged, and NLLB-200 remains Meta’s free, open specialist for very broad language coverage.

A quick comparison: Muse Spark 1.3, Llama 4 Maverick, Llama 4 Scout, and NLLB-200

Muse Spark 1.3Llama 4 MaverickLlama 4 ScoutNLLB-200
Ideal ForMeta’s new proprietary flagship – long-horizon agentic and coding-adjacent localization workflows, multimodal (image/video/PDF) content.Creative and nuanced content, and handling multimodal inputs like images and video.Deep analysis of extremely long documents, where consistency is critical.High-volume, cost-effective translation across a massive number of languages.
AccessProprietary – Meta’s own API only (public preview).Open-weight – hosted by third-party providers (Together, Fireworks, DeepInfra, etc.).Open-weight – hosted by third-party providers.Open-source, self-hosted or via a compatible cloud provider.
Cost (per 1M tokens)$1.25 input / $4.25 output (Meta Model API, official).Roughly $0.20–$0.30 input / $0.60–$0.85 output, depending on host (hosted rates vary).Roughly $0.08–$0.10 input / $0.25–$0.30 output, depending on host.Free to use; cost is hardware only.
Context Window1,048,576 tokens.1 million tokens.10 million tokens natively (most hosts cap lower, e.g. 128K–1M).Up to 512 tokens.
LanguagesMultilingual support across 100+ languages.Excels in a dozen core languages, foundational understanding of 200+.Same multilingual capability as Maverick, focused on long-form analysis.Specialist model for 200 distinct languages.
Key StrengthMeta’s strongest model to date; native multimodal understanding plus built-in search.Balance of intelligence, speed, and cost for general-purpose use.Ability to “remember” and reason over entire books or codebases.The single model providing quality translation across 200 languages.

Source: Meta’s official Muse Spark launch post for Muse Spark specs and pricing; hosted Llama 4 rates vary by provider (e.g. DeepInfra, Together AI, Fireworks) so treat those two ranges as approximate and check your provider directly.

Muse Spark 1.3: Meta’s new proprietary flagship

This is the model that didn’t exist the last time this article was updated, and it changes the shape of Meta’s offering considerably. Muse Spark is proprietary (no downloadable weights), reached the public via Meta’s new Model API in mid-2026, and is now on its third point release.

When to Use Muse Spark for Translation:

  • Complex, Multi-Step Localization Pipelines: Built for long-horizon agentic work – planning, delegating, and self-correcting across a multi-file project.
  • Multimodal Localization: Native understanding of images, video, and PDFs, useful for localizing UI screenshots, instructional video, or scanned documents.
  • Large Context Consistency: A 1M-token context window keeps terminology and tone consistent across a big batch of strings.

Cost: $1.25 per million input tokens and $4.25 per million output tokens – approx. $0.073 per 10,000 words.

Llama 4 Maverick & Llama 4 Scout: Still here, unchanged

Meta has not announced end-of-life for Llama 4, and the Scout/Maverick pair remains a solid open-weight option, particularly if you want to self-host or need a provider outside the big three.

  • Llama 4 Maverick – the creative, multimodal all-rounder. Hosted rates run roughly $0.20–$0.30 input / $0.60–$0.85 output per million tokens depending on provider (approx. $0.011–$0.015 per 10,000 words).
  • Llama 4 Scout – the long-document specialist, with a native 10-million-token context window (most hosts cap this lower). Hosted rates run roughly $0.08–$0.10 input / $0.25–$0.30 output per million tokens (approx. $0.005 per 10,000 words).

NLLB-200: Champion of linguistic diversity

NLLB-200 (“No Language Left Behind”) is Meta’s open model specifically designed for machine translation, unaffected by the Muse Spark launch. It still provides high-quality results for 200 different languages, including many low-resource languages.

  • Wide language coverage: Unparalleled breadth – a single model translating across 200 languages, including low-resource ones like Høgnorsk or Icelandic.
  • Cost-Effective by Design: Open-source, so you only pay for the compute to run it.
  • A Translation Specialist, Not a Generalist: Because it’s trained for translation alone, it stays closer to the source than a general-purpose LLM asked to translate. Quality is strongest for high-resource language pairs and drops off for the rarest ones, so sample your specific pairs before committing.

Making the right choice for Meta’s lineup

Choose right Meta model for translation

  • If you need Meta’s strongest available model and don’t mind a proprietary API, Muse Spark 1.3 is now the answer – it wasn’t an option a year ago.
  • If your focus is on premium quality, creative content, and multimodal tasks with an open-weight model you can self-host, Llama 4 Maverick is still your partner.
  • If your project demands absolute consistency across massive documents, Llama 4 Scout remains the specialized solution.
  • If your priority is cost-effective, high-volume translation across a diverse set of languages, NLLB-200 is still the unmatched, free champion.

Mistral AI

Mistral is the European entry in this list, and for a lot of localization teams that is the whole point: it’s a Paris-based lab with an open-weight heritage, and it offers EU data residency through regional inference endpoints (billed at a 10% uplift). If your legal team has opinions about where translation data is processed, this is usually the first name on the shortlist. Mistral is available natively in Crowdin, with your own credentials.

The naming can be confusing, because the model IDs you’ll see in Crowdin’s provider dropdown are dated snapshots rather than marketing names. Here’s the mapping:

Model ID in CrowdinMarketing nameContext windowInput / Output (per 1M tokens)Cost per 10,000 words
mistral-large-2512Mistral Large 3256K$0.50 / $1.50approx. $0.027
mistral-medium-2604Mistral Medium 3.5256K$1.50 / $7.50approx. $0.12
mistral-small-2603Mistral Small 4256K$0.15 / $0.60approx. $0.010
ministral-14b-2512Ministral 3 (14B)256K$0.20 / $0.20approx. $0.005
ministral-8b-2512Ministral 3 (8B)256K$0.15 / $0.15approx. $0.004
ministral-3b-2512Ministral 3 (3B)131K$0.10 / $0.10approx. $0.003

Sources: Mistral’s official API pricing page for rates; context windows from Mistral’s own model cards and release changelog. The -latest variants in the dropdown point to whichever dated snapshot is current, so they can change under you – pin the dated ID if you need reproducible results.

Which Mistral model for which content

Choose right Mistral model for translation

  • Mistral Large 3 is the flagship: open-weight, multimodal, and explicitly positioned as multilingual. At $0.50 / $1.50 it is startlingly cheap for a flagship tier – roughly a sixth of what Claude Sonnet 5 costs per translated word. Use it for the content where quality matters and EU processing is a requirement.
  • Mistral Medium 3.5 is, confusingly, more expensive than Large 3 on output. It’s tuned for long-horizon agentic work and tool calling rather than for raw translation quality, so for straightforward localization Large 3 is usually the better buy.
  • Mistral Small 4 is the volume workhorse – multimodal, multilingual, Apache 2.0 licensed, and cheap enough for bulk string processing.
  • Ministral 3 (3B, 8B, 14B) is a separate family of small edge models built to run on a phone, a laptop or a single modest server. For a cloud localization workflow they’re rarely the right pick – Mistral Small 4 costs little more and is meaningfully more capable – but they matter when the constraint is where the translation happens rather than what it costs.

Every Mistral model in the list above carries a 256K context window, with one exception: ministral-3b-2512 is served at 131K on Mistral’s own API, even though the Ministral 3 architecture supports 256K when you self-host it.

Alibaba Qwen

Qwen is Alibaba’s model family, and it earns a place in a roundup of LLMs for a reason that isn’t obvious from the name. Alongside the general-purpose Qwen3.7 line, Alibaba maintains Qwen-MT – a translation-specialized branch built by fine-tuning Qwen3 on translation data, not a separate machine translation architecture. So it is still an LLM, just one that has given up general reasoning in exchange for being good at one job.

That trade-off is the whole point of including it, because it shows what a specialist buys you and what it costs. What you get is terminology control, format preservation and domain adaptation built into the model rather than bolted on through instructions. What you give up is everything the general models in this article do around translation: the context window drops to 16,384 tokens, so Qwen-MT works string by string rather than holding a document in memory, and it can’t reason about a screenshot or restructure a file.

The other reason to look at it is language coverage skewed differently from the rest of the list. Alibaba trains and positions Qwen-MT for Asian language pairs in particular – Chinese, Japanese, Korean, Vietnamese – with 92 supported languages and localized dialect mapping. If those markets are a large share of your volume, it is worth benchmarking against your current engine rather than assuming a Western-built flagship wins by default.

Qwen-MT family: Specialized machine translation

Qwen-MT-PlusQwen-MT-TurboQwen-MT-Lite
Ideal ForHigh-stakes localization in Asian language pairs – legal, technical, and terminology-heavy content.Everyday translation tasks needing a balance of speed and quality – website content and knowledge bases.High-volume, low-risk content such as real-time chat, user reviews, or initial drafts for MTPE.
Supported Languages92 languages, upgraded to Qwen3 architecture.Broad language coverage optimized for throughput.32 languages, optimized for speed and cost.
Context Window16,384 tokens.16,384 tokens.16,384 tokens.
Cost (per 1M tokens, official Alibaba international endpoint)≈$2.54 input / $7.62 output on Alibaba’s own Model Studio/PAI-EAS pricing pages. (Some resold routes, e.g. Novita AI, list Qwen-MT-Plus far lower at $0.25 / $0.75 – worth checking both before you commit.)$0.16 input / $0.49 output.$0.12 input / $0.36 output.
Key StrengthAlibaba’s flagship translation specialist – terminology control, format preservation, domain-specific adaptation.High-speed, balanced translation across a wide range of language pairs.Lowest-cost tier for rapid translation across its 32 supported languages.

Source: Alibaba Cloud Model Studio’s official pricing pages for Qwen-MT-Plus, Qwen-MT-Turbo and Qwen-MT-Lite – check the current rates there.

Qwen-MT-Plus: Specialized machine translation

Qwen-MT-Plus is Alibaba’s flagship specialized machine translation model, upgraded to Qwen3 architecture, fine-tuned on trillions of translation-specific tokens with reinforcement training that emphasizes terminology control and domain-specific vocabulary mapping.

Qwen-MT-Plus is Best For:

  • High-stakes enterprise documentation, including regulatory compliance data and legal contracts, where domain alignment is critical.
  • Deep localization updates for major Asian markets (Chinese, Japanese, Korean, Vietnamese) requiring localized cultural awareness.
  • Enforcing strict glossary matching via customized domain instructions to prevent stylistic drift.

Cost: On Alibaba’s own international pricing pages, ≈$2.54 input / $7.62 output per million tokens – for a 10,000-word file, that’s approx. $0.14. Some third-party resellers list the same model at $0.25 / $0.75 (≈$0.013 per 10,000 words); the gap is large enough that it’s worth confirming which route you’d actually be billed through.

Qwen-MT-Turbo: Balanced workhorse

Qwen-MT-Turbo is the mid-tier option in Alibaba’s translation-specialist line-up, built for high-speed, high-concurrency translation across a wide range of language pairs.

Cost: $0.16 per million input tokens and $0.49 per million output tokens – approx. $0.0087 per 10,000 words.

Qwen-MT-Lite: Fastest and cheapest

Engineered for ultra-low latency, Qwen-MT-Lite clears high-volume translation queues at minimal cost across its 32 supported languages.

Cost: $0.12 per million input tokens and $0.36 per million output tokens – translating a 10,000-word dataset for a fraction of a cent (approx. $0.0064).

General-purpose alternative: Qwen 3.7

If you need a more general-purpose model rather than a translation specialist, Alibaba’s current flagship general model is Qwen 3.7-Max, priced at $2.50 input / $7.50 output per million tokens with a 1M-token context window (list price on the official Model Studio pricing page; promotional routes at roughly half that price have also been reported, so check the console for current offers).

Hybrid approach for Alibaba Qwen models

Choose right Alibaba Qwen model for translation

Maximizing your return on localization deployment still calls for a tiered setup:

  • Premium Tier: Route your core legal document pools, marketing masterfiles, and technical glossaries through Qwen-MT-Plus – but confirm which pricing route applies to your account first.
  • Production Tier: Allocate standard help desks and localized web structures to Qwen-MT-Turbo for stable, well-balanced processing.
  • Throughput Tier: Stream real-time user-generated comments and system logs into Qwen-MT-Lite to keep costs near zero.

Build a multi-model strategy with Crowdin

Everything above points the same way. The spread between the cheapest and most expensive model in this article is more than a hundredfold. No vendor leads on every language pair. Published benchmarks lag model releases by months. So the useful question isn’t “which model is best” – it is “which model for which job, and how do I check”.

System AI Providers in Crowdin

One model per step, not one model per project

Translation isn’t a single operation. Analyzing context, mapping terminology, producing the translation, and running QA are different tasks with different difficulty and different volume – and Crowdin’s AI Pipelines let you assign a different model to each stage rather than routing everything through one engine.

AI models in Crowdin AI Pipeline

This is where the cost table stops being an abstraction. A reasoning-heavy step that runs once per file can afford GPT-6 Astra. The translation step that runs on every string can’t – that’s where Gemini 3.8 Flash or Mistral Large 3 belong. And a QA pass over finished output is a different calculation again.

What teams actually run

The case studies in this article are more instructive than any benchmark, because they show real production choices:

  • Bitrefill runs Claude Sonnet 4.6 through Crowdin’s Bring Your Own API Key, with custom instructions that translate each new string as it lands – and keeps testing other models rather than settling.
  • Holafly also uses Claude 4.6 Sonnet, paired with a glossary and a style guide inside Crowdin, with human linguists proofreading for brand tone.
  • MyHeritage has processed nearly 60 million words across 50 languages, with 25 million of them handled by Google Gemini, OpenAI and others – and 5 million reused from translation memory rather than translated at all.

Two things stand out. First, nobody is running the newest model on the list: Sonnet 4.6 is a generation behind Sonnet 5, and these workflows work fine (these were the models in use when each case study was published – any of these teams may well have moved on since). Second, MyHeritage’s cheapest 5 million words came from translation memory, not from any model in this article.

Test it on your own content

Crowdin connects GPT, Gemini, Claude, Mistral, Qwen and others from one place, and Bring Your Own API Key means you use your own account, your own rates and your own agreement with the provider. You can also connect a custom AI model, including one you host yourself.

What that buys you is the only answer to model choice that doesn’t go stale: run the same file through two or three engines, in your language pairs, with your glossary, and compare the post-editing effort. That number is specific to your content, and it will tell you more than this article can.

Localize your product with Crowdin

Automate content updates, boost team collaboration, and reach new markets faster.
Free 14-day Trial

How we calculated the cost per 10,000 words

Every “per 10,000 words” figure in this article uses the same simple model, so the numbers stay comparable across vendors:

  • 10,000 words ≈ 13,333 tokens. The usual rule of thumb for English is about 0.75 words per token.
  • We bill that volume twice – once as input (the source text you send) and once as output (the translation you get back). So the formula is 13,333 × (input rate + output rate) ÷ 1,000,000.
  • Standard tier, short context, no caching. Batch APIs typically halve these rates, and cached input can cut the input side by 90%, so a well-configured pipeline will pay less than the table shows.

Two things will move your real bill away from these numbers, and both matter more in localization than in most other AI use cases:

Your target language changes the arithmetic. The 0.75 words-per-token ratio is an English figure. Languages with non-Latin scripts or rich morphology – Ukrainian, Japanese, Korean, Arabic, Thai, Finnish – consistently need more tokens to express the same content, and that cost lands on the output side, which is the expensive one. If you translate from English into these languages, treat the comparison table as a floor, not an estimate. The only reliable way to know your multiplier is to run a sample file through the vendor’s tokenizer and compare.

Tokenizers differ between vendors, and even between generations. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text than earlier Claude models. We’ve applied that adjustment to Claude Fable 5.1, Opus 5.5 and Sonnet 5; Haiku 4.5 predates the change, so its figure is unadjusted. Without this correction, Claude would look about a third cheaper than it really is next to OpenAI and Google.

FAQ

What is the best LLM for translation?

There is no single “best” LLM for all translation tasks. The ideal approach is to use a hybrid strategy that uses different models for different needs. For example, a high-reasoning model like GPT-6 Astra or Claude Opus 5.5 can be used for critical content, while a faster, more cost-effective model like GPT-6 Luna or Gemini 3.1 Flash-Lite can handle high-volume, low-risk content. If you need a dedicated translation specialist with glossary and terminology control built in, a purpose-built model like Qwen-MT is worth testing alongside the general-purpose options.

Why use a multi-model approach for translation?

A multi-model strategy allows you to optimize your workflow and budget by using the right model for the right task. This approach ensures you get the necessary quality for each piece of content while managing costs and maintaining efficiency. Platforms like Crowdin allow you to connect and switch between various AI engines to suit different content types and project needs.

Which LLMs are the most cost-effective for translation?

For high-volume, low-cost translation, you should choose models optimized for minimal expense. Examples include GPT-6 Luna, Gemini 3.1 Flash-Lite, and the Qwen-MT-Lite and Qwen-MT-Turbo models. Open-source models like Meta NLLB-200 are also a very cheap option, as the only cost is the computational resources to run them.

What is a “context window” and why is it important for translation?

The context window is the number of tokens an LLM can process at once. A larger context window allows the model to maintain consistency across long documents, such as books or technical manuals. This is a key challenge in translation, as it helps ensure that terminology and style remain consistent throughout the entire text. Meta Llama 4 Scout has an industry-leading native context window of 10 million tokens.

Yuliia Makarenko

Yuliia Makarenko

Yuliia Makarenko is a marketing specialist with over a decade of experience, and she’s all about creating content that readers will love. She’s a pro at using her skills in SEO, research, and data analysis to write useful content. When she’s not diving into content creation, you can find her reading a good thriller, practicing some yoga, or simply enjoying playtime with her little one.

Share this post: