Skip to content

Benchmarks

Language performance

How well does each AI model perform when prompted in different languages? Within each language, models are ranked by head-to-head comparison on the same native prompts — 1000 is the average of the models measured in that language.

Tests run weekly · Models with image/audio/TTS capabilities excluded

🇬🇧

English

prompts: 2 · 430 scored runs · English

🇳🇱

Nederlands

prompts: 2 · 328 scored runs · Dutch

🇩🇪

Deutsch

prompts: 1 · 169 scored runs · German

Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.

🇫🇷

Français

prompts: 1 · 365 scored runs · French

Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.

🇪🇸

Español

prompts: 1 · 363 scored runs · Spanish

Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.

🇹🇷

Türkçe

prompts: 1 · 120 scored runs · Turkish

Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.