Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.0900
input / 1M
— stable
$0.2800
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.
Last 7 days
100.0%
n=620
Last 30 days
100.0%
n=1,696
Median response time
1,617ms
n=1,696
Based on 2,076 measurements over the last 30 days.
Technical details
Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.
Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.
Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.
Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.
Total calls (30d)
1,696
OK responses (30d)
1,696
Total calls (7d)
620
OK responses (7d)
620
Tokonomix benchmark verdicts
Quality surges 11.5 points to 91.6 with balanced performance across categories
Mistral-Small-3.2-24B-Instruct-2506 demonstrates a significant recovery in this benchmark window, achieving an overall quality score of 91.6, up 11.5 points from the previous period's 80.2. This represents a strong rebound after the prior window's substantial decline. The model now shows remarkably consistent performance across all tested categories, with scores tightly clustered between 91 and 92 for coding, creative writing, factual accuracy, and reasoning tasks. This balanced profile marks a dramatic improvement from the previous window where factual performance had plummeted to 53, creating significant inconsistency. Latency has also improved by 18 percent, with the median response time dropping from 7337ms to 6053ms. The addition of creative and reasoning categories in this window prevents direct comparison for those dimensions, but the factual category's recovery from 53 to 91 points is particularly noteworthy. Coding performance dipped slightly from 96 to 92, though this remains a strong absolute score. Users can expect reliable, well-rounded performance across diverse task types with improved response times compared to the previous evaluation period.
Quality
91.6
Latency p50
6,053 ms
Test runs
5
Mistral-Small-3.2-24B-Instruct-2506
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.0900 / 1M
- Output price
- $0.2800 / 1M
- Tier
- Tier B — Production
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 474
More from OVH AI Endpoints (GRA)