Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.1000
input / 1M
— stable
$0.1000
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Quality falls to 63.8 as creative performance collapses to 6
Mistral-7B-Instruct-v0.3 on OVH AI Endpoints continues a troubling downward trajectory, dropping another 5.3 points to an overall quality score of 63.8. This marks the second consecutive window of quality decline, bringing the total drop to over 10 points. The most alarming change is the creative category, which has plummeted to just 6 points, indicating severe degradation in tasks requiring imaginative or open-ended responses. While coding performance improved from 83 to 92 and factual capabilities surged from 32 to 94, these gains cannot offset the creative collapse. Reasoning stabilized at 63 points, a moderate level that suggests inconsistent logical processing. Multilingual testing was not included in the current window, making it impossible to assess whether that previously strong category remains stable. Latency improved slightly from 6.7 to 5.9 seconds at the median, offering a minor silver lining. Users requiring creative text generation, content writing, or storytelling should look elsewhere, as this model appears fundamentally compromised in that domain. Those focused purely on coding assistance or factual retrieval may still find value, but the overall instability raises concerns about reliability across deployment scenarios.
Quality
63.8
Latency p50
5,879 ms
Test runs
5
Mistral-7B-Instruct-v0.3
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.1000 / 1M
- Output price
- $0.1000 / 1M
- Tier
- Tier C — Specialist
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 474
More from OVH AI Endpoints (GRA)