Skip to content
Tier A — Frontier
Runs in:USMade in:United States
Anthropic

Claude Sonnet 5

Tier A — Frontier · 1M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan·
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency105 runs
805229337815268675608-3009-25ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

53%
Coding
judge mean 91
88%
Creative
judge mean 92
52%
Factual
judge mean 73
62%
Multilingual
judge mean 98
60%
Reasoning
judge mean 70

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing

What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.

💰
API rates — Claude Sonnet 5
$2.60 per 1M input tokens
$13.00 per 1M output tokens
≈ $0.0042 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$2.60
per 1M output tokens$13.00
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)216 / avg 163
24627

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

toolssource: manualvisionjson modepdf inputreasoningjson schemaprompt cachingmax output tokens: 128000
Section 06

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

100.0%

n=4

Last 30 days

100.0%

n=19

Median response time

31,148ms

n=19

Based on 399 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

19

OK responses (30d)

19

Total calls (7d)

4

OK responses (7d)

4

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-586/100 · 35 runs
29 correct3 partial3 wrong83% accuracy
● 2026-09-20

Claude Sonnet 5 adds multimodal capabilities, no performance data available

Claude Sonnet 5 has expanded its feature set significantly since the previous benchmark window, adding seven new capabilities: tool use, vision, JSON mode, PDF input, reasoning, JSON schema support, and prompt caching. These additions represent a substantial enhancement to the model's functionality, particularly in multimodal processing and structured output generation. However, no benchmark performance data is available for either the current or previous window, making it impossible to assess the model's actual performance across standard evaluation metrics. Without quantitative results on tasks like reasoning, coding, or knowledge retrieval, users cannot compare Claude Sonnet 5's capabilities against competing models or evaluate whether the new features come with performance tradeoffs. The addition of prompt caching and structured output modes suggests optimization for production use cases, while vision and PDF input expand the model's applicability to document-heavy workflows. Users interested in these specific features may find value in the expanded capabilities, but the absence of benchmark data means adoption decisions must rely on qualitative testing rather than objective performance comparisons.

Quality

—

Latency p50

—

Test runs

0

✓ Added vision and PDF support✓ New tool use capabilities✓ JSON schema and caching added✗ No benchmark data available
Section 08

Full model profile

Claude Sonnet 5: the mid-tier Claude built to act, not just answer

Claude Sonnet 5 is Anthropic's mid-tier model, released June 30, 2026 as the successor to Sonnet 4.6. Anthropic frames it as the most agentic Sonnet the company has shipped: a model meant to plan a task, use tools like a browser or a terminal, and carry a multi-step job forward with less step-by-step supervision than earlier Sonnet generations needed. Anthropic's own positioning is that Sonnet 5 closes much of the gap to the flagship Opus line while staying in the Sonnet tier.

Where it shines

The model is built around sustained tool use rather than single-turn chat. It supports function calling, structured JSON output against a schema, PDF input, and vision, and it exposes a reasoning mode so it can work through a problem before answering. That combination fits agentic workflows well: a task that requires reading a document, calling a tool, checking the result, and deciding what to do next is exactly the shape Anthropic says this release targets. Anthropic reports 78.5% on OSWorld-Verified, its computer-use benchmark, and 46.8% on Humanity's Last Exam when the model is allowed to use tools — both are Anthropic's own published figures, not independently reproduced numbers, so they describe the vendor's account of the model's progress rather than a neutral third-party result.

Under the hood

Sonnet 5 carries a 1,000,000-token context window with a maximum synchronous output of 128,000 tokens per response. Output is text-only — this generation of Sonnet does not generate images or audio. Prompt caching is supported, which matters for agentic loops that resend a large, mostly-unchanged context (a codebase, a long document, a tool schema) across many turns of the same session.

Where it falls short

Anthropic's own announcement of this release does not state a knowledge cutoff date, so anyone relying on the model for very recent events or newly released libraries should verify rather than assume currency. The model's strengths are concentrated in tool-using, multi-step work; for short single-turn questions or simple classification, the agentic machinery is overhead rather than benefit, and a lighter model in the same lineup will usually answer just as well. Anthropic's own benchmark claims, since they come from the vendor testing its own model, are also worth treating as directional rather than settled — the scale of an improvement over Sonnet 4.6 is Anthropic's characterization until it is checked against independent workloads.

When to pick it

Sonnet 5 fits coding agents, research agents, and any workflow where the model needs to keep working across several tool calls without a human re-prompting it at every step — debugging a failing test suite, working through a multi-file refactor, or running a browse-and-summarize task end to end. It also fits workloads that need a genuinely large context window alongside tool use, such as reviewing a large codebase or a long document set in one session.

For work that does not need that autonomy — short-form generation, simple extraction, single-call classification — the agentic overhead is not doing anything useful, and a smaller, faster model will typically be the more sensible default.

Alternatives worth comparing

Within Anthropic's own lineup, Claude Opus 5 sits above Sonnet 5 as the flagship for the most complex agentic and enterprise work, and Claude Fable 5 sits above that as Anthropic's most capable widely released model — both are options if a workload turns out to need more headroom than Sonnet 5 provides. Anthropic's own positioning frames Sonnet 5 as approaching Opus-class performance on many tasks, which makes it worth starting there before reaching for a larger model, and only stepping up if the workload specifically demonstrates it needs the extra capability.

Deployment notes

Because Sonnet 5 is built around long, tool-heavy sessions, integrations should expect and handle multi-turn tool-call loops rather than a single request/response exchange — timeouts, retry logic, and tool-result formatting all need to account for a task that may take several rounds to complete. The 128,000-token output ceiling is generous but finite; a workflow that expects one call to return an entire large generated artifact should confirm it fits under that cap before relying on it in production.

Last automated test
Sep 25, 2026 · 14:00 UTC · Speed benchmark
P50 latency
924 ms
P95 latency
1400 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 14, 2026