Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing
What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.
Last 7 days
100.0%
n=819
Last 30 days
100.0%
n=2,307
Median response time
1,905ms
n=2,307
Based on 2,692 measurements over the last 30 days.
Technical details
Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.
Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.
Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.
Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.
Total calls (30d)
2,307
OK responses (30d)
2,307
Total calls (7d)
819
OK responses (7d)
819
Tokonomix benchmark verdicts
Quality plummets 17.1 points to 70.7 amid doubled latency and coverage shift
Mistral-Small-3.2-24B-Instruct-2506 experiences a significant performance decline in this benchmark window, with overall quality dropping from 87.8 to 70.7 points. This 17.1-point decrease represents the continuation of a downward trend observed in the previous period. Latency has deteriorated substantially, with the median response time nearly doubling from 4102ms to 8162ms, a 99% increase that will impact real-world application responsiveness. Category performance reveals an unusual pattern. Coding remains exceptional at 96, maintaining near-perfect scores across both windows. Reasoning achieved a perfect 100 score in current testing. However, factual accuracy collapsed to just 16 points, representing a critical weakness. The benchmark window also shows a shift in category coverage, with multilingual and creative categories from the previous period replaced by factual and reasoning assessments in the current window. The dramatic decline in factual performance combined with significantly slower response times raises concerns about model reliability for knowledge-based tasks. While coding and reasoning capabilities remain strong, the overall trajectory suggests potential infrastructure or configuration issues affecting both speed and accuracy.
Quality
70.7
Latency p50
8,162 ms
Test runs
5
Mistral-Small-3.2-24B-Instruct-2506
by OVH AI Endpoints (GRA)
- Released
- June 20, 2025
- Context window
- — tokens
- Input price
- $0.2400 / 1M
- Output price
- $0.7300 / 1M
- Tier
- Tier B — Production
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 522
More from OVH AI Endpoints (GRA)