—
00:00
Growing Money
Growing Money
USD/RUB—
EUR/RUB—
Releases

LILT Unveils AURORA to Benchmark AI Performance in Non-English Languages

San Francisco-based LILT has launched AURORA, a benchmarking platform designed to test frontier AI models on agentic, multimodal, and socio-cultural tasks beyond English. By moving past traditional translation-heavy metrics, the leaderboard aims to provide enterprises with data on how AI agents function in specific regional and cultural contexts.

LILT Unveils AURORA to Benchmark AI Performance in Non-English Languages

Most industry benchmarks rely on English-centric data, creating a blind spot for companies deploying AI agents globally. LILT CEO Spence Green warns that while poor translation was once a minor inconvenience, agentic errors in customer-facing roles now carry significant commercial risk. AURORA shifts the focus to native-language performance, utilizing tasks verified by domain experts across sectors like software development, retail, and banking.

The platform evaluates models through four core benchmarks: Multilingual Terminal-bench for localized coding, Multilingual τ³-bench for customer support, Multilingual MultiChallenge for instruction-following, and Multilingual GAIA-v2-LILT for agentic reasoning. Initial analysis from LILT’s PhD-led research team highlights the volatility of model quality, noting that top-performing models shift depending on the language—such as GPT 5.5 excelling in Spanish, Claude Opus 5.5 leading in Japanese, and Muse Spark 1.3 performing best in Serbian.

Share

Comments (0)

Leave a comment

No comments yet. Be the first!