AI Models Fail Grammar Tests in Key Languages, RWS Benchmark Reveals

  • RWS launched M-GATE, a benchmark evaluating 70 AI models across 30 languages on grammar, translation, and tokenizer efficiency.
  • Google's Gemini 3.1 Pro Preview topped the grammar leaderboard, while OpenAI's GPT-5.5 led in roundtrip translation.
  • Some models, including Meta's Muse Spark, scored below random chance in certain languages.
  • Tokenizer costs vary sharply by language, with some models using more than ten times as many tokens for languages like Khmer compared to English.
  • Translation quality on low-resource languages has roughly doubled, but significant gaps remain.

RWS's M-GATE benchmark highlights significant variability in AI model performance across languages, challenging the assumption that flagship models excel uniformly. As businesses increasingly rely on AI for multilingual content, the need for independent benchmarks like M-GATE becomes critical for ensuring accurate and efficient model selection. The findings underscore the importance of evaluating models on language-specific capabilities rather than relying on broad benchmark scores.

Model Selection
How enterprises will use M-GATE to make more informed decisions about AI model selection for specific languages.
Performance Gaps
Whether the identified performance gaps in low-resource languages will narrow as models continue to evolve.
Cost Efficiency
The impact of tokenizer efficiency on the cost and speed of AI model deployment across different languages.