Skip to content
Tier C — Specialist
Runs in:FranceMade in:France
OVH AI Endpoints (GRA)

Mistral-7B-Instruct-v0.3

Tier C — Specialist

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency104 runs
8955110131475193708-1009-05ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

29%
Coding
judge mean 91
13%
Creative
judge mean 59
25%
Factual
judge mean 59
30%
Multilingual
judge mean 87
30%
Reasoning
judge mean 54

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Mistral-7B-Instruct-v0.3
$0.1000 per 1M input tokens
$0.1000 per 1M output tokens
≈ <$0.0001 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.1000
per 1M output tokens$0.1000

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.1000

input / 1M

— stable

$0.1000

output / 1M

— stable

2026-06-142026-07-192026-08-30
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)1087 / avg 1212
2223196

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: mistralai
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-573/100 · 72 runs
37 correct17 partial18 wrong51% accuracy
2026-08-30

Quality falls to 63.8 as creative performance collapses to 6

Mistral-7B-Instruct-v0.3 on OVH AI Endpoints continues a troubling downward trajectory, dropping another 5.3 points to an overall quality score of 63.8. This marks the second consecutive window of quality decline, bringing the total drop to over 10 points. The most alarming change is the creative category, which has plummeted to just 6 points, indicating severe degradation in tasks requiring imaginative or open-ended responses. While coding performance improved from 83 to 92 and factual capabilities surged from 32 to 94, these gains cannot offset the creative collapse. Reasoning stabilized at 63 points, a moderate level that suggests inconsistent logical processing. Multilingual testing was not included in the current window, making it impossible to assess whether that previously strong category remains stable. Latency improved slightly from 6.7 to 5.9 seconds at the median, offering a minor silver lining. Users requiring creative text generation, content writing, or storytelling should look elsewhere, as this model appears fundamentally compromised in that domain. Those focused purely on coding assistance or factual retrieval may still find value, but the overall instability raises concerns about reliability across deployment scenarios.

Quality

63.8

Latency p50

5,879 ms

Test runs

5

Creative performance collapsed to 6 Quality dropped 5.3 points Coding improved to 92 Factual accuracy up to 94
Last automated test
Sep 5, 2026 · 08:01 UTC · Speed benchmark
P50 latency
184 ms
P95 latency
214 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 5, 2026