Skip to content
Tier A — Frontier
Runs in:USMade in:United States

Retirement date

The provider lists June 9, 2027 as a tentative retirement date. The model is available as usual today.

Anthropic

Claude Fable 5

Tier A — Frontier · 1M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan·

Claude Fable 5 is Anthropic's vision-and-reasoning model with a one-million-token context variant. In our own image-QC pilot it was the steadiest pair of eyes we tested — so we made it the default vision proposer in our image-consensus panel. Here is what we measured, and where we did not.

The property we wanted from a vision proposer was not cleverness but steadiness — and on our pilot set, Fable 5 was the steadiest.

Tokonomix vision-QC pilot, 9 June 2026
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
2251388555207154878808-2109-14ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

64%
Coding
judge mean 97
67%
Creative
judge mean 89
59%
Factual
judge mean 67
65%
Multilingual
judge mean 99
69%
Reasoning
judge mean 98

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing

What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.

💰
API rates — Claude Fable 5
$13.00 per 1M input tokens
$65.00 per 1M output tokens
≈ $0.0208 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$13.00
per 1M output tokens$65.00
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)69 / avg 60
8832

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Strengths & weaknesses

Drawn from benchmark results and aggregated community feedback on real use-cases.

Strengths

Most stable vision proposer in our pilot (88% run-identical)Low false-alarm rate (7.1% on the 300-image baseline)Caught blind spots other panel models missed1M-token context variant for long, image-heavy material

Weaknesses

Solo recall 66.9% — a panel is needed to reach 87.5%Reasoning/coding numbers are early (single run, n=1)
Section 06

Capabilities

toolssource: manualvisionjson modepdf inputreasoningjson schemamodel class: mythosprompt cachingmax output tokens: 128000
Section 07

Frequently asked questions

On our 9 June 2026 pilot it was run-identical 88% of the time, flipped its answer only 3.9% of the time, and produced zero false positives — the steadiness you want at the front of a QC pipeline. It was a deliberate product decision based on that pilot.

We would rather hand you the texture — small pilot, one-day baseline, early reasoning numbers — than a clean story that does not survive the next run.

Tokonomix editorial
Section 08

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

62.5%

n=48

Last 30 days

70.5%

n=61

Median response time

16,774ms

n=43

Based on 448 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

61

OK responses (30d)

43

Total calls (7d)

48

OK responses (7d)

30

Image quality control pilot (2026-06-10)

Recall

66.9%

n=300

False-alarm rate

7.1%

n=300

Section 09

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 2 runs
2 correct0 partial0 wrong100% accuracy
claude-sonnet-4-589/100 · 65 runs
55 correct4 partial6 wrong85% accuracy
2026-09-13

Claude Fable 5 shows reasoning gains but latency degrades 20%

Claude Fable 5 demonstrates a modest overall quality improvement, rising from 86.9 to 88.0 across the benchmark window. The most significant development is perfect reasoning performance at 100, showing strong capabilities in logical tasks and problem-solving scenarios. Coding performance remains exceptionally stable, improving only marginally from 95 to 96, indicating consistent strength in programming tasks. However, the model faces a notable performance tradeoff. Latency has increased substantially from 7361ms to 8825ms at the median, representing a 20% slowdown in response times. This degradation may impact user experience in time-sensitive applications despite the quality improvements. Category coverage has shifted between windows, with factual question performance now measured at 68, suggesting room for improvement in knowledge retrieval tasks. Previous strengths in creative writing (66) and multilingual capabilities (100) are not reflected in current testing, making direct comparison difficult. The overall quality gain of 1.1 points suggests incremental refinement rather than transformative changes. Users should weigh the reasoning improvements against increased wait times when evaluating this model for production deployment.

Quality

88.0

Latency p50

8,825 ms

Test runs

5

Perfect reasoning score achieved Coding remains exceptionally strong Latency increased 20% Factual performance needs improvement
Section 10

Full model profile

Claude Fable 5: the steadiest pair of eyes we tested

Anthropic's Claude Fable 5 is a vision-and-reasoning model, and the variant we lean on most — claude-fable-5[1m] — carries a one-million-token context window. We are not going to recite a spec sheet. What we can offer instead is something most write-ups cannot: numbers from our own measurements, with the dates and sample sizes attached, and an honest note about where the evidence runs thin.

Why we made it our default vision proposer

On 9 June 2026 we ran a vision-QC pilot — a set of images, each with a known media-quality flaw to catch, fed to several vision models so we could compare how they behaved. Fable 5 stood out for one quality that matters more than raw cleverness: it was stable.

Across repeated runs on the pilot set it was run-identical 88% of the time, with a flip-rate of 3.9% — meaning it rarely changed its mind about the same image from one pass to the next. On that pilot set it produced zero false positives: it did not invent defects that were not there. It also caught blind spots that other models in the panel missed.

Stability is the property you want from a proposer. A model that flags a real defect today and shrugs at the identical image tomorrow is hard to build a process on. One that answers the same way every time — and does not cry wolf — is one you can put at the front of a pipeline. So we did: Fable 5 is the default vision proposer in our image-consensus panel. (That was a deliberate product decision, made on the strength of the pilot.) It is not in our text-judge pools — its job here is looking at images, not scoring text.

Today's baseline: 300 images, measured live

A pilot is a small thing. So we put the model through a larger run and we publish the result, live, on our vision-QC benchmark.

On 10 June 2026, against the mediaqc-v3-2026-06-10 dataset of 300 images, Fable 5 solo scored:

  • 66.9% recall — the share of real defects it caught. That tied for the best single-model result in the run.
  • 7.1% false-alarm rate — how often it flagged a clean image. For comparison, another strong vision model on the same run sat more than twice as high on false alarms, which is exactly the kind of difference that decides whether a QC step saves time or wastes it.
  • 60.3% class-matched — cases where it not only caught the defect but named the right category of problem.

Recall of two-thirds, solo, is useful but not a finished story — which is the point of running a panel rather than a single model. With Fable 5 sitting in the consensus panel, recall rose to 87.5%. The proposer's steadiness gives the council something reliable to build on; the council closes the gap that one model alone leaves open.

Reasoning and coding: early, honest signal

Vision is where we have the most evidence. On general capability we have less, and we will say so plainly.

In an internal evaluation on 9 June 2026 — a 682-test harness we run across models — Fable 5's results separated sharply from the field: it scored 91% and 73% on the two halves of that harness where the models we compared against scored 3% and 0% on the identical tasks. That is a large gap, and we treat it as a strong signal rather than a settled verdict, because a single harness can flatter or punish a model in ways a second one would correct.

Our first intelligence run on the platform, today, returned a reasoning score of 100 and a coding score of 97. Those are preliminary — a single run, n=1. We are reporting them because hiding early numbers is its own kind of dishonesty, but one run is one run. Watch the leaderboard as the sample grows; that is where these figures will firm up or move.

What the 1M-context variant is for

The claude-fable-5[1m] variant exists for the cases where context is the bottleneck rather than reasoning. A window this wide lets you hold a long document set, a large body of reference material, or an extended interaction in front of the model at once instead of chunking it and hoping nothing important falls out of view. Paired with vision, it suits work where you need the model to reason over a lot of material and look at the images embedded in it — long reports, document sets with figures, screenshots in context.

How to read these numbers

Everything above is measured, dated, and bounded by a sample size we have named. The vision-QC pilot was small; the 300-image baseline is larger but is one dataset on one day; the reasoning and coding figures are early. We would rather hand you that texture than a clean story, because the clean story is usually the one that does not survive the next run.

Claude Fable 5 is available through the Tokonomix gateway and catalogue, and you can put it in front of your own images on our vision-QC benchmark or watch it across tasks on the leaderboard.

Last automated test
Sep 14, 2026 · 20:02 UTC · Speed benchmark
P50 latency
2913 ms
P95 latency
3228 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·June 10, 2026