Skip to content
Tier A — Frontier
Runs in:USMade in:United States
Anthropic

Claude Sonnet 5

Tier A — Frontier · 1M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency30 runs
1159215931604160516008-0408-11ms
Section 02

Quality scores

Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.

97
Coding
100
Creative
100
Multilingual
100
Reasoning
Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Claude Sonnet 5
$3.00 per 1M input tokens
$15.00 per 1M output tokens
≈ $0.0048 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$3.00
per 1M output tokens$15.00

Pricing over time

Input & output per 1M tokens · step-line = price changes

$3.00

input / 1M

— no change

$15.00

output / 1M

— no change

2026-08-092026-08-092026-08-09
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)63 / avg 89
17150

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

toolssource: manualvisionjson modepdf inputreasoningjson schemaprompt cachingmax output tokens: 128000
Section 06

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

100.0%

n=1

Last 30 days

100.0%

n=4

Median response time

14,268ms

n=4

Based on 99 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

4

OK responses (30d)

4

Total calls (7d)

1

OK responses (7d)

1

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-599/100 · 5 runs
5 correct0 partial0 wrong100% accuracy
2026-08-09

Claude Sonnet 5 debuts with strong reasoning and multimodal capabilities

Claude Sonnet 5 enters the benchmark landscape as Anthropic's latest mid-tier offering, demonstrating competitive performance across multiple domains. The model achieves 78.0% on MMLU, positioning it solidly in the capable generalist category, while its 83.1% on GPQA (diamond) suggests particular strength in graduate-level reasoning tasks. Coding performance is respectable at 73.7% on HumanEval and 79.1% on SWE Bench Verified, though not segment-leading. The model shows balanced mathematics capabilities with 83.5% on GSM8K and 63.5% on MATH, indicating reliable performance on standard problems with room for improvement on competition-level mathematics. A notable strength appears in instruction following, scoring 84.7% on IFEval. The model debuts with a comprehensive feature set including vision, PDF input, tool use, JSON modes, reasoning capabilities, and prompt caching. Multimodal performance shows 60.3% on MMMU and 69.1% on MathVista, suggesting functional but not exceptional visual understanding. For users seeking a well-rounded model with strong reasoning and practical tool integration, Claude Sonnet 5 presents a solid baseline option.

Quality

Latency p50

Test runs

0

Strong GPQA reasoning performance Comprehensive multimodal feature set Solid instruction following capability Moderate competition math scores
Last automated test
Aug 11, 2026 · 08:05 UTC · Speed benchmark
P50 latency
3189 ms
P95 latency
3370 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·August 11, 2026