How well do AI models really master different languages?

New linguist-designed benchmark, M-GATE, gives enterprises an independent way to compare how leading AI models perform across languages.

Inhaltsübersicht

Reading time: 02:45 minutes

Global AI solutions company RWS (AIM: RWS.L) has launched M-GATE (Multilingual Grammar, Accuracy in Translation & Efficiency), an independent benchmark that measures whether frontier AI models genuinely understand the languages they claim to support.

M-GATE evaluates 70 models on grammar, translation accuracy, and tokenizer efficiency across 30 languages – from widely spoken ones to under-served ones such as Kinyarwanda, Basque, and Fijian.
  

What the benchmark reveals

M-GATE tests grammar using a binary test – a sentence is either correct or incorrect – so random guessing scores around 50%. Yet results diverge sharply even within the same language. For example, on Fijian grammar, xAI's Grok 4.20 tops the entire field, while OpenAI's GPT-5.5, the benchmark's best translator, falls below random chance, and Meta's Muse Spark scores just 23%.

Google’s Gemini 3.1 Pro Preview currently tops the grammar leaderboard overall across 70 models and 30 languages, while OpenAI’s GPT-5.5 leads on roundtrip translation, which takes a sentence from English into the target language and back – but strength in one language rarely carries to the next.
  

Four patterns stood out

  1. Flagship status doesn’t predict grammar performance. M-GATE’s "stumper" sentences, designed to probe tricky language-specific grammatical features, humble nearly every model, with many – including some top-tier models – scoring at or below random guessing in some languages. One example stumper sentence, “Everything I told you is what I thought I had said I would,” sounds awkward enough to be wrong, but it isn’t. That’s exactly the kind of distinction the models are being asked to make on M-GATE.
  2. Tokenizer costs vary sharply by language. The heaviest tokenizers use more than ten times as many tokens to process some languages, such as Khmer, as they do English. Claude's tokenizer usage looks balanced overall. But in absolute terms, it packs the fewest characters per token of any major lab in the benchmark – which drives costs up.
  3. Speed varies enormously. The slowest models average roughly 100 times the latency of the fastest, with worst-case responses measured in minutes.
  4. The gap is closing where it matters most. Translation quality on low-resource languages has roughly doubled across the models benchmarked – progress that reasoning benchmarks don’t measure – but meaningful gaps remain in the hardest languages.
      

How M-GATE works

M-GATE is an automated, linguist-designed benchmark that evaluates models on three dimensions:

  • Grammar proficiency: 100 linguist-crafted "stumper" sentences per language – half containing deliberate errors, half correct but tricky – designed to probe tricky grammar rules specific to each language.
  • Roundtrip translation accuracy: 100 source sentences spanning seven challenge categories – discourse pragmatics, non-compositional, implicit content, structural complexity, lexical disambiguation, referential precision (and control) – to test whether meaning survives a translate-and-back journey.
  • Tokenizer efficiency: a measure of how many tokens each model consumes to process similar text and how quickly it responds – both tied directly to cost and speed.

"The uncomfortable truth is that some of the most capable models on the market perform worse than random guessing in certain languages – and the benchmarks everyone quotes can never tell you that. We built M-GATE to measure the thing every buyer assumes and no one checks – whether a model actually holds up in the language you’re about to ship it in,” said Tomáš Burkert, Head of Innovation at TrainAI by RWS.
  

Why this matters for enterprises

The findings come as businesses rapidly expand their use of AI-generated and AI-translated content across global markets. RWS’s Content Unlocked 2026 study of 200 senior enterprise content leaders found that while 86% said AI had accelerated content creation, 65% said it had simultaneously slowed localization through additional rework. And most multilingual benchmarks test reasoning, math, or knowledge expressed in a language, not command of the language itself.

M-GATE gives organizations another way to interrogate model performance before deployment – comparing model performance at the individual language level rather than having to assume a strong overall benchmark score will translate into equally strong multilingual performance.

Visit M-GATE to explore the benchmark and see how today's top models really perform across languages.