Skip to content
Tier A — Frontier
Runs in:USMade in:United States

Archived

This model has been discontinued by the provider. Historical data is preserved.

No longer available since June 30, 2027.

Anthropic

Claude Sonnet 5

Tier A — Frontier · 1M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency104 runs
8483324580082751075108-1609-11ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

62%
Coding
judge mean 96
89%
Creative
judge mean 92
61%
Factual
judge mean 81
65%
Multilingual
judge mean 98
53%
Reasoning
judge mean 67

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Claude Sonnet 5
$3.00 per 1M input tokens
$15.00 per 1M output tokens
≈ $0.0048 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$3.00
per 1M output tokens$15.00

Pricing over time

Input & output per 1M tokens · step-line = price changes

$3.00

input / 1M

— stable

$15.00

output / 1M

— stable

2026-08-092026-08-232026-09-06
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)212 / avg 149
23427

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

toolssource: manualvisionjson modepdf inputreasoningjson schemaprompt cachingmax output tokens: 128000
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-590/100 · 25 runs
23 correct0 partial2 wrong92% accuracy
2026-09-06

Claude Sonnet 5 shows no benchmark data despite seven new capabilities

Claude Sonnet 5 continues to report no benchmark performance data across any standard evaluation metrics for the second consecutive window. The model has maintained its seven recently added capabilities: tools, vision, json_mode, pdf_input, reasoning, json_schema, and prompt_caching. However, without quantitative performance measurements, users cannot assess how this model compares to alternatives or previous versions on key dimensions like accuracy, reasoning quality, or task completion rates. The absence of benchmark data makes it impossible to verify whether the added capabilities translate into measurable improvements in real-world performance. For organizations evaluating Claude Sonnet 5, the lack of transparent metrics presents a challenge in making informed deployment decisions. The model's actual capabilities in areas like visual understanding, structured output generation, or tool use remain unquantified through independent testing. Users considering this model should seek empirical validation through their own testing protocols before committing to production use, as public benchmark performance remains unavailable to guide selection decisions.

Quality

Latency p50

Test runs

0

Seven capabilities now available No benchmark data available Performance remains unverified No comparative metrics provided
Last automated test
Sep 11, 2026 · 02:04 UTC · Speed benchmark
P50 latency
943 ms
P95 latency
1342 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 11, 2026