Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$3.00
input / 1M
— no change
$15.00
output / 1M
— no change
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.
Last 7 days
100.0%
n=1
Last 30 days
100.0%
n=4
Median response time
14,268ms
n=4
Based on 99 measurements over the last 30 days.
Technical details
Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.
Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.
Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.
Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.
Total calls (30d)
4
OK responses (30d)
4
Total calls (7d)
1
OK responses (7d)
1
Tokonomix benchmark verdicts
Claude Sonnet 5 debuts with strong reasoning and multimodal capabilities
Claude Sonnet 5 enters the benchmark landscape as Anthropic's latest mid-tier offering, demonstrating competitive performance across multiple domains. The model achieves 78.0% on MMLU, positioning it solidly in the capable generalist category, while its 83.1% on GPQA (diamond) suggests particular strength in graduate-level reasoning tasks. Coding performance is respectable at 73.7% on HumanEval and 79.1% on SWE Bench Verified, though not segment-leading. The model shows balanced mathematics capabilities with 83.5% on GSM8K and 63.5% on MATH, indicating reliable performance on standard problems with room for improvement on competition-level mathematics. A notable strength appears in instruction following, scoring 84.7% on IFEval. The model debuts with a comprehensive feature set including vision, PDF input, tool use, JSON modes, reasoning capabilities, and prompt caching. Multimodal performance shows 60.3% on MMMU and 69.1% on MathVista, suggesting functional but not exceptional visual understanding. For users seeking a well-rounded model with strong reasoning and practical tool integration, Claude Sonnet 5 presents a solid baseline option.
Quality
—
Latency p50
—
Test runs
0
Claude Sonnet 5
by Anthropic
- Context window
- 1M tokens
- Input price
- $3.00 / 1M
- Output price
- $15.00 / 1M
- Tier
- Tier A — Frontier
- Modality
- Text + vision
- API type
- REST · streaming
- Benchmark runs
- 35
More from Anthropic