Benchmarks
Language performance
How well does each AI model perform when prompted in different languages? Within each language, models are ranked by head-to-head comparison on the same native prompts — 1000 is the average of the models measured in that language.
Tests run weekly · Models with image/audio/TTS capabilities excluded
English
prompts: 2 · 430 scored runs · English
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Claude Opus 4.6Anthropic | 93%±3 | 8 |
| 1 | Claude Opus 4.5Anthropic | 92%±4 | 8 |
| 1 | gpt-5-search-api-2025-10-14OpenAI | 91%±4 | 8 |
| 1 | Gemini 2.5 Flash-LiteGoogle Gemini | 90%±4 | 8 |
| 1 | Gemini Flash-Lite LatestGoogle Gemini | 89%±5 | 8 |
Nederlands
prompts: 2 · 328 scored runs · Dutch
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | gpt-4oOpenAI | 80%±5 | 6 |
| 1 | gpt-4.1-mini-2025-04-14OpenAI | 78%±6 | 6 |
| 1 | Claude Fable 5Anthropic | 77%±6 | 3 |
| 1 | Claude Sonnet 5Anthropic | 77%±6 | 2 |
| 1 | gpt-5OpenAI | 77%±6 | 2 |
Deutsch
prompts: 1 · 169 scored runs · German
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 54%±11 | 1 |
| 1 | Claude Haiku 4.5Anthropic | 54%±11 | 3 |
| 1 | Claude Opus 4.5Anthropic | 54%±11 | 3 |
| 1 | Claude Opus 4.6Anthropic | 54%±11 | 3 |
| 1 | Claude Opus 4.7Anthropic | 54%±11 | 3 |
Français
prompts: 1 · 365 scored runs · French
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Gemini Flash-Lite LatestGoogle Gemini | 86%±7 | 6 |
| 1 | gpt-5.1OpenAI | 86%±7 | 3 |
| 1 | o3OpenAI | 86%±7 | 3 |
| 1 | o4-mini-2025-04-16OpenAI | 86%±7 | 3 |
| 1 | Gemini 3.1 Flash LiteGoogle Gemini | 85%±7 | 6 |
Español
prompts: 1 · 363 scored runs · Spanish
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Gemini Flash-Lite LatestGoogle Gemini | 80%±6 | 5 |
| 1 | gpt-3.5-turboOpenAI | 80%±6 | 6 |
| 1 | gpt-3.5-turbo-0125OpenAI | 80%±6 | 6 |
| 1 | gpt-3.5-turbo-1106OpenAI | 80%±6 | 6 |
| 1 | gpt-3.5-turbo-16kOpenAI | 80%±6 | 5 |
Türkçe
prompts: 1 · 120 scored runs · Turkish
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 57%±11 | 2 |
| 1 | Claude Haiku 4.5Anthropic | 57%±11 | 2 |
| 1 | Claude Opus 4.5Anthropic | 57%±11 | 2 |
| 1 | Claude Opus 4.6Anthropic | 57%±11 | 2 |
| 1 | Claude Opus 4.7Anthropic | 57%±11 | 2 |