What if I told you your AI is not as smart as it could be simply because you're
speaking another language?
This is currently a possibility for anyone using AI in a language other than English.
When you ask your AI to do work - answering customers in Tokyo or
processing documents in Berlin - you might notice that the output quality
measures lower against industry benchmarks.
You might assume the model is simply "dumber" in that language… but what if we told you it's
the benchmark that's broken, not the model?
This was the throughline of our recent AI Circle roundtable dinner in New York, where LILT
CEO Spence Green walked a room of frontier researchers and builders through a finding that
changes how we measure multilingual AI.
LILT decided to re-audit one of the most-cited agent benchmarks (GAIA) for languages, and the
results were astounding: measured performance jumped by an average
of 20.7%.
The "fluent yet broken" Paradox
The problem is relatively simple, most multilingual AI benchmarks aren't built in the languages
they claim to test. The benchmark datasets (the "tests") - are written in English, then run
through auto-translation and this is where everything starts to go wrong.
A translated benchmark question can be flawless German or Korean and
still be classified as "wrong." The reason lies in the answer key - which
grades the responses against an answer written for the English original. A minor nuance gets
completely misinterpreted. Here are two examples:
So when an AI agent "fails" one of these questions, it's incorrectly recorded as a capability
gap when in reality the benchmark just measures the translation failure, not the model's
intelligence itself.
Here's a visual walk-through of the paradox I'm describing:
To solve this, LILT rebuilt the GAIA multilingual benchmark with proper auditing. The tasks
were written in the native language, locale-correct conventions were applied, and - most
importantly - culturally grounded context was added. They've called it
GAIA-v2-LILT, and the measured improvement was a lot bigger than expected.
+20.7pp
Average recovery in measured agent performance across languages
So why does this matter for you?
The internal benchmarks you use to measure performance are likely underselling you.
If you're deploying AI anywhere outside English - customer support in Tokyo, coding agents in
São Paulo - you may be unnecessarily holding back product launches, over-engineering fixes, or
picking the wrong models, all based on numbers that aren't entirely accurate.
But the language issue is just one of many verticals that could be affected.
Multilingual benchmarking is just the most visible example because the failure mode is obvious:
translate a task badly and humans can usually tell. But the deeper concern is that any
benchmark built on assumptions that no longer hold - outdated tool APIs, contrived task
framings, cultural defaults baked in - is likely mismeasuring the thing it claims to evaluate.
This matters more as AI shifts from chatbots to agents. LLM benchmarks mostly test whether an
answer is good. An agent benchmark tests whether a multi-step task is
completed - which means it's testing the environment, the tools, and the cultural and locale
assumptions baked into "success," not just the model. The surface area for measurement error
gets pretty large.
What practitioners can do
Before trusting a score, ask: who built the test? In what language? Against which
assumptions? Treat every capability gap as something to triple-check against the
benchmark answer key.
The most valuable role in the next phase of AI isn't just one focused on building smarter
models - it's building better instruments to measure them - and having the
discipline to question the numbers everyone else takes on good faith.