Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.6700
input / 1M
— stable
$0.6700
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.
Last 7 days
100.0%
n=10
Last 30 days
100.0%
n=10
Median response time
3,552ms
n=10
Based on 390 measurements over the last 30 days.
Technical details
Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.
Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.
Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.
Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.
Total calls (30d)
10
OK responses (30d)
10
Total calls (7d)
10
OK responses (7d)
10
Tokonomix benchmark verdicts
Quality climbs 6.8 points to 92.0 with consistent coding performance
Meta-Llama-3.3-70B-Instruct continues its upward trajectory, posting a 92.0 overall quality score, up 6.8 points from the previous window's 85.3. This marks the second consecutive improvement period for this model. Coding performance remains rock-solid at 92 across both windows, demonstrating reliable capability in programming tasks. The current window shows multilingual and creative work both scoring 92, a dramatic improvement in creative output from the previous 53. However, the previous window's exceptional factual score of 100 and reasoning score of 96 are not represented in current category results, making direct comparison incomplete. Latency has increased from 7416ms to 8405ms at the median, a 13% slowdown that users should account for in time-sensitive applications. The model appears to have achieved more balanced performance across categories, trading some excellence in specific domains for broader competency. With five test runs in each window, the results provide reasonable confidence in these trends. Users requiring strong creative capabilities will benefit from recent improvements, while those prioritizing speed may need to evaluate whether the quality gains justify the latency increase.
Quality
92.0
Latency p50
8,405 ms
Test runs
5
Meta-Llama-3_3-70B-Instruct
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.6700 / 1M
- Output price
- $0.6700 / 1M
- Tier
- Tier B — Production
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 499
More from OVH AI Endpoints (GRA)