Benchmarks
Language performance
How well does each AI model perform when prompted in different languages? Within each language, models are ranked by head-to-head comparison on the same native prompts — 1000 is the average of the models measured in that language.
Tests run weekly · Models with image/audio/TTS capabilities excluded
English
prompts: 2 · 493 scored runs · English
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 97%±3 | 1 |
| 1 | gpt-5.4-nano-2026-03-17OpenAI | 91%±4 | 3 |
| 1 | gpt-oss-120bOVH AI Endpoints (GRA) | 91%±4 | 7 |
| 1 | Claude Opus 4.8Anthropic | 91%±4 | 6 |
| 2 | gpt-5-search-apiOpenAI | 88%±4 | 9 |
Nederlands
prompts: 2 · 325 scored runs · Dutch
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | gpt-4oOpenAI | 80%±5 | 6 |
| 1 | gpt-4.1-mini-2025-04-14OpenAI | 78%±6 | 6 |
| 1 | Claude Fable 5Anthropic | 77%±6 | 3 |
| 1 | Claude Sonnet 5Anthropic | 77%±6 | 2 |
| 1 | gpt-5OpenAI | 77%±6 | 2 |
Deutsch
prompts: 1 · 244 scored runs · German
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 63%±11 | 1 |
| 1 | Claude Sonnet 5Anthropic | 63%±11 | 1 |
| 1 | Claude Fable 5Anthropic | 59%±8 | 2 |
| 1 | Claude Haiku 4.5Anthropic | 59%±8 | 4 |
| 1 | Claude Opus 4.5Anthropic | 59%±8 | 4 |
Français
prompts: 1 · 359 scored runs · French
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Gemini Flash-Lite LatestGoogle Gemini | 85%±7 | 6 |
| 1 | gpt-5.1OpenAI | 85%±7 | 3 |
| 1 | o3OpenAI | 85%±7 | 3 |
| 1 | o4-mini-2025-04-16OpenAI | 85%±7 | 3 |
| 1 | Gemini 3.1 Flash LiteGoogle Gemini | 85%±7 | 6 |
Español
prompts: 1 · 358 scored runs · Spanish
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Gemini Flash-Lite LatestGoogle Gemini | 79%±6 | 5 |
| 1 | gpt-3.5-turboOpenAI | 79%±6 | 6 |
| 1 | gpt-3.5-turbo-0125OpenAI | 79%±6 | 6 |
| 1 | gpt-3.5-turbo-1106OpenAI | 79%±6 | 6 |
| 1 | gpt-3.5-turbo-16kOpenAI | 79%±6 | 5 |
Türkçe
prompts: 1 · 119 scored runs · Turkish
Only 1 prompt in this language so far — this shows the evidence collected, not a ranking. More prompts are needed before the order means anything.
| # | Model | Win rate | Runs |
|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 56%±11 | 2 |
| 1 | Claude Haiku 4.5Anthropic | 56%±11 | 2 |
| 1 | Claude Opus 4.5Anthropic | 56%±11 | 2 |
| 1 | Claude Opus 4.6Anthropic | 56%±11 | 2 |
| 1 | Claude Opus 4.7Anthropic | 56%±11 | 2 |