Your AI agent isn't wrong. Your benchmark is.

What a Park Avenue dinner with LILT's CEO revealed about a benchmark error that's distorting how we evaluate AI in every language but English.

What if I told you your AI is not as smart as it could be simply because you're speaking another language?

This is currently a possibility for anyone using AI in a language other than English.

When you ask your AI to do work - answering customers in Tokyo or processing documents in Berlin - you might notice that the output quality measures lower against industry benchmarks.

You might assume the model is simply "dumber" in that language… but what if we told you it's the benchmark that's broken, not the model?

This was the throughline of our recent AI Circle roundtable dinner in New York, where LILT CEO Spence Green walked a room of frontier researchers and builders through a finding that changes how we measure multilingual AI.

LILT decided to re-audit one of the most-cited agent benchmarks (GAIA) for languages, and the results were astounding: measured performance jumped by an average of 20.7%.

The "fluent yet broken" Paradox

The problem is relatively simple, most multilingual AI benchmarks aren't built in the languages they claim to test. The benchmark datasets (the "tests") - are written in English, then run through auto-translation and this is where everything starts to go wrong.

A translated benchmark question can be flawless German or Korean and still be classified as "wrong." The reason lies in the answer key - which grades the responses against an answer written for the English original. A minor nuance gets completely misinterpreted. Here are two examples:

✓ Correct in German

27.04.2026

✕ English answer key

April 27, 2026

Same date but scored as wrong

✓ Correct in German

€1.299,00

✕ English answer key

$1,299.00

Exact same price but scored as wrong

So when an AI agent "fails" one of these questions, it's incorrectly recorded as a capability gap when in reality the benchmark just measures the translation failure, not the model's intelligence itself.

Here's a visual walk-through of the paradox I'm describing:

To solve this, LILT rebuilt the GAIA multilingual benchmark with proper auditing. The tasks were written in the native language, locale-correct conventions were applied, and - most importantly - culturally grounded context was added. They've called it GAIA-v2-LILT, and the measured improvement was a lot bigger than expected.

+20.7pp

Average recovery in measured agent performance across languages

Measured recovery after re-auditing

Percentage-point gain once the benchmark was rebuilt in-language.

A fifth of the apparent "capability gap" between English and non-English AI was an error.

So why does this matter for you?

The internal benchmarks you use to measure performance are likely underselling you.

If you're deploying AI anywhere outside English - customer support in Tokyo, coding agents in São Paulo - you may be unnecessarily holding back product launches, over-engineering fixes, or picking the wrong models, all based on numbers that aren't entirely accurate.

But the language issue is just one of many verticals that could be affected.

Multilingual benchmarking is just the most visible example because the failure mode is obvious: translate a task badly and humans can usually tell. But the deeper concern is that any benchmark built on assumptions that no longer hold - outdated tool APIs, contrived task framings, cultural defaults baked in - is likely mismeasuring the thing it claims to evaluate.

This matters more as AI shifts from chatbots to agents. LLM benchmarks mostly test whether an answer is good. An agent benchmark tests whether a multi-step task is completed - which means it's testing the environment, the tools, and the cultural and locale assumptions baked into "success," not just the model. The surface area for measurement error gets pretty large.

Chatbot benchmark 1 thing tested

The answer

Agent benchmark 5 things tested

The answer
The environment
The tools
Locale conventions
Cultural assumptions

Same model on both sides. The agent score reflects everything in the stack - four new ways to mismeasure.

What practitioners can do

Before trusting a score, ask: who built the test? In what language? Against which assumptions? Treat every capability gap as something to triple-check against the benchmark answer key.

The most valuable role in the next phase of AI isn't just one focused on building smarter models - it's building better instruments to measure them - and having the discipline to question the numbers everyone else takes on good faith.

Going deeper

LILT's work on GAIA-v2-LILT and language-grounded evaluation The benchmarks under discussion: Terminal-Bench & tau³-bench AI Circle's next roundtable - request an invite