Most industry benchmarks rely on English-centric data, creating a blind spot for companies deploying AI agents globally. LILT CEO Spence Green warns that while poor translation was once a minor inconvenience, agentic errors in customer-facing roles now carry significant commercial risk. AURORA shifts the focus to native-language performance, utilizing tasks verified by domain experts across sectors like software development, retail, and banking.
The platform evaluates models through four core benchmarks: Multilingual Terminal-bench for localized coding, Multilingual τ³-bench for customer support, Multilingual MultiChallenge for instruction-following, and Multilingual GAIA-v2-LILT for agentic reasoning. Initial analysis from LILT’s PhD-led research team highlights the volatility of model quality, noting that top-performing models shift depending on the language—such as GPT 5.5 excelling in Spanish, Claude Opus 5.5 leading in Japanese, and Muse Spark 1.3 performing best in Serbian.

Comments (0)
No comments yet. Be the first!