Skip to content
Tier B — Production
Runs in:FranceMade in:France
OVH AI Endpoints (GRA)

Mistral-Small-3.2-24B-Instruct-2506

Tier B — Production

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
7036467222107981437408-2109-14ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

66%
Coding
judge mean 98
41%
Creative
judge mean 86
46%
Factual
judge mean 70
50%
Multilingual
judge mean 98
61%
Reasoning
judge mean 96

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing

What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.

💰
API rates — Mistral-Small-3.2-24B-Instruct-2506
$0.2400 per 1M input tokens
$0.7300 per 1M output tokens
≈ $0.0003 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.2400
per 1M output tokens$0.7300
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)1149 / avg 1317
282184

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: mistralai
Section 06

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

100.0%

n=819

Last 30 days

100.0%

n=2,307

Median response time

1,905ms

n=2,307

Based on 2,692 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

2,307

OK responses (30d)

2,307

Total calls (7d)

819

OK responses (7d)

819

Image quality control pilot (2026-06-10)

Recall

9.4%

n=300

False-alarm rate

12.1%

n=300

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-591/100 · 83 runs
73 correct7 partial3 wrong88% accuracy
2026-09-13

Quality plummets 17.1 points to 70.7 amid doubled latency and coverage shift

Mistral-Small-3.2-24B-Instruct-2506 experiences a significant performance decline in this benchmark window, with overall quality dropping from 87.8 to 70.7 points. This 17.1-point decrease represents the continuation of a downward trend observed in the previous period. Latency has deteriorated substantially, with the median response time nearly doubling from 4102ms to 8162ms, a 99% increase that will impact real-world application responsiveness. Category performance reveals an unusual pattern. Coding remains exceptional at 96, maintaining near-perfect scores across both windows. Reasoning achieved a perfect 100 score in current testing. However, factual accuracy collapsed to just 16 points, representing a critical weakness. The benchmark window also shows a shift in category coverage, with multilingual and creative categories from the previous period replaced by factual and reasoning assessments in the current window. The dramatic decline in factual performance combined with significantly slower response times raises concerns about model reliability for knowledge-based tasks. While coding and reasoning capabilities remain strong, the overall trajectory suggests potential infrastructure or configuration issues affecting both speed and accuracy.

Quality

70.7

Latency p50

8,162 ms

Test runs

5

Quality dropped 17.1 points Latency doubled to 8162ms Factual accuracy critically low Perfect reasoning score achieved
Last automated test
Sep 14, 2026 · 20:02 UTC · Speed benchmark
P50 latency
174 ms
P95 latency
691 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 14, 2026