Skip to content
Tier A — Frontier
Runs in:USMade in:United States

Archived

This model has been discontinued by the provider. Historical data is preserved.

No longer available since June 9, 2027.

Anthropic

Claude Fable 5

Tier A — Frontier · 1M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan·

Claude Fable 5 is Anthropic's vision-and-reasoning model with a one-million-token context variant. In our own image-QC pilot it was the steadiest pair of eyes we tested — so we made it the default vision proposer in our image-consensus panel. Here is what we measured, and where we did not.

The property we wanted from a vision proposer was not cleverness but steadiness — and on our pilot set, Fable 5 was the steadiest.

Tokonomix vision-QC pilot, 9 June 2026
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency105 runs
2251956316876241883150008-1009-05ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

67%
Coding
judge mean 97
71%
Creative
judge mean 93
49%
Factual
judge mean 67
62%
Multilingual
judge mean 99
68%
Reasoning
judge mean 98

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Claude Fable 5
$10.00 per 1M input tokens
$50.00 per 1M output tokens
≈ $0.0160 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$10.00
per 1M output tokens$50.00

Pricing over time

Input & output per 1M tokens · step-line = price changes

$10.00

input / 1M

— stable

$50.00

output / 1M

— stable

2026-06-102026-08-092026-08-30
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)61 / avg 56
8829

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Strengths & weaknesses

Drawn from benchmark results and aggregated community feedback on real use-cases.

Strengths

Most stable vision proposer in our pilot (88% run-identical)Low false-alarm rate (7.1% on the 300-image baseline)Caught blind spots other panel models missed1M-token context variant for long, image-heavy material

Weaknesses

Solo recall 66.9% — a panel is needed to reach 87.5%Reasoning/coding numbers are early (single run, n=1)
Section 06

Capabilities

toolssource: manualvisionjson modepdf inputreasoningjson schemamodel class: mythosprompt cachingmax output tokens: 128000
Section 07

Frequently asked questions

On our 9 June 2026 pilot it was run-identical 88% of the time, flipped its answer only 3.9% of the time, and produced zero false positives — the steadiness you want at the front of a QC pipeline. It was a deliberate product decision based on that pilot.

We would rather hand you the texture — small pilot, one-day baseline, early reasoning numbers — than a clean story that does not survive the next run.

Tokonomix editorial
Section 08

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

Last 30 days

100.0%

n=7

Median response time

13,923ms

n=7

Based on 387 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

7

OK responses (30d)

7

Total calls (7d)

0

OK responses (7d)

0

Image quality control pilot (2026-06-10)

Recall

66.9%

n=300

False-alarm rate

7.1%

n=300

Section 09

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 2 runs
2 correct0 partial0 wrong100% accuracy
claude-sonnet-4-589/100 · 55 runs
47 correct3 partial5 wrong85% accuracy
2026-08-30

Claude Fable 5 shows no benchmark data across two evaluation windows

Claude Fable 5 remains without any benchmark performance data for a second consecutive evaluation window, despite having seven documented capabilities: tools, vision, json_mode, pdf_input, reasoning, json_schema, and prompt_caching. The absence of benchmark results makes it impossible to assess the model's actual performance across standard evaluation metrics or compare it to competing models in the market. Without data on accuracy, reasoning quality, speed, or other measurable outcomes, potential users cannot make informed decisions about whether this model suits their needs. The continued lack of benchmarking is unusual for a model release, particularly one claiming multiple advanced capabilities. Users seeking a model with proven performance should look to alternatives with established benchmark results. Organizations considering Claude Fable 5 for production use would need to conduct their own extensive testing to validate whether the model meets their requirements. Until benchmark data becomes available, the model's practical utility and competitive positioning remain unclear, regardless of its feature set.

Quality

Latency p50

Test runs

0

No benchmark data available Second window without results
Section 10

Full model profile

Claude Fable 5: the steadiest pair of eyes we tested

Anthropic's Claude Fable 5 is a vision-and-reasoning model, and the variant we lean on most — claude-fable-5[1m] — carries a one-million-token context window. We are not going to recite a spec sheet. What we can offer instead is something most write-ups cannot: numbers from our own measurements, with the dates and sample sizes attached, and an honest note about where the evidence runs thin.

Why we made it our default vision proposer

On 9 June 2026 we ran a vision-QC pilot — a set of images, each with a known media-quality flaw to catch, fed to several vision models so we could compare how they behaved. Fable 5 stood out for one quality that matters more than raw cleverness: it was stable.

Across repeated runs on the pilot set it was run-identical 88% of the time, with a flip-rate of 3.9% — meaning it rarely changed its mind about the same image from one pass to the next. On that pilot set it produced zero false positives: it did not invent defects that were not there. It also caught blind spots that other models in the panel missed.

Stability is the property you want from a proposer. A model that flags a real defect today and shrugs at the identical image tomorrow is hard to build a process on. One that answers the same way every time — and does not cry wolf — is one you can put at the front of a pipeline. So we did: Fable 5 is the default vision proposer in our image-consensus panel. (That was a deliberate product decision, made on the strength of the pilot.) It is not in our text-judge pools — its job here is looking at images, not scoring text.

Today's baseline: 300 images, measured live

A pilot is a small thing. So we put the model through a larger run and we publish the result, live, on our vision-QC benchmark.

On 10 June 2026, against the mediaqc-v3-2026-06-10 dataset of 300 images, Fable 5 solo scored:

  • 66.9% recall — the share of real defects it caught. That tied for the best single-model result in the run.
  • 7.1% false-alarm rate — how often it flagged a clean image. For comparison, another strong vision model on the same run sat more than twice as high on false alarms, which is exactly the kind of difference that decides whether a QC step saves time or wastes it.
  • 60.3% class-matched — cases where it not only caught the defect but named the right category of problem.

Recall of two-thirds, solo, is useful but not a finished story — which is the point of running a panel rather than a single model. With Fable 5 sitting in the consensus panel, recall rose to 87.5%. The proposer's steadiness gives the council something reliable to build on; the council closes the gap that one model alone leaves open.

Reasoning and coding: early, honest signal

Vision is where we have the most evidence. On general capability we have less, and we will say so plainly.

In an internal evaluation on 9 June 2026 — a 682-test harness we run across models — Fable 5's results separated sharply from the field: it scored 91% and 73% on the two halves of that harness where the models we compared against scored 3% and 0% on the identical tasks. That is a large gap, and we treat it as a strong signal rather than a settled verdict, because a single harness can flatter or punish a model in ways a second one would correct.

Our first intelligence run on the platform, today, returned a reasoning score of 100 and a coding score of 97. Those are preliminary — a single run, n=1. We are reporting them because hiding early numbers is its own kind of dishonesty, but one run is one run. Watch the leaderboard as the sample grows; that is where these figures will firm up or move.

What the 1M-context variant is for

The claude-fable-5[1m] variant exists for the cases where context is the bottleneck rather than reasoning. A window this wide lets you hold a long document set, a large body of reference material, or an extended interaction in front of the model at once instead of chunking it and hoping nothing important falls out of view. Paired with vision, it suits work where you need the model to reason over a lot of material and look at the images embedded in it — long reports, document sets with figures, screenshots in context.

How to read these numbers

Everything above is measured, dated, and bounded by a sample size we have named. The vision-QC pilot was small; the 300-image baseline is larger but is one dataset on one day; the reasoning and coding figures are early. We would rather hand you that texture than a clean story, because the clean story is usually the one that does not survive the next run.

Claude Fable 5 is available through the Tokonomix gateway and catalogue, and you can put it in front of your own images on our vision-QC benchmark or watch it across tasks on the leaderboard.

Last automated test
Sep 5, 2026 · 08:01 UTC · Speed benchmark
P50 latency
3262 ms
P95 latency
3486 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·June 10, 2026