AI Models Fail Grammar Tests in Key Languages, RWS Benchmark Reveals
Event summary
- RWS launched M-GATE, a benchmark evaluating 70 AI models across 30 languages on grammar, translation, and tokenizer efficiency.
- Google's Gemini 3.1 Pro Preview topped the grammar leaderboard, while OpenAI's GPT-5.5 led in roundtrip translation.
- Some models, including Meta's Muse Spark, scored below random chance in certain languages.
- Tokenizer costs vary sharply by language, with some models using more than ten times as many tokens for languages like Khmer compared to English.
- Translation quality on low-resource languages has roughly doubled, but significant gaps remain.
The big picture
RWS's M-GATE benchmark highlights significant variability in AI model performance across languages, challenging the assumption that flagship models excel uniformly. As businesses increasingly rely on AI for multilingual content, the need for independent benchmarks like M-GATE becomes critical for ensuring accurate and efficient model selection. The findings underscore the importance of evaluating models on language-specific capabilities rather than relying on broad benchmark scores.
What we're watching
- Model Selection
- How enterprises will use M-GATE to make more informed decisions about AI model selection for specific languages.
- Performance Gaps
- Whether the identified performance gaps in low-resource languages will narrow as models continue to evolve.
- Cost Efficiency
- The impact of tokenizer efficiency on the cost and speed of AI model deployment across different languages.
