The answer
Do not switch. Keep Opus 5 at xhigh as the execution seat.
- Use Fable 5.1 for the judgment and planning seats at the effort you set for the session.
- Try Fable 5.1 at low effort only on your own matched tasks, with a search nudge in the prompt.
- That is where the evidence lands, and it is one day of evidence.
The answer at a glance
Seven sentences. Each one links to the evidence that carries it.
- 01 On Claude Max, Fable usage is capped at 50% of the weekly limit and draws that limit faster than other models, with no multiplier published, so Fable 5.1 cannot be the primary execution model whatever the effort setting. seeO1F9the cap
- 02 The one independent cross-model reading puts Fable 5.1 low at index 58 for $0.77 per task against Opus 5 max at 63 for $2.34: cheaper in API dollars, lower scoring, and those dollars are not what a Max plan meters. seeF3O2the chart
- 03 Anthropic's own routing guidance says to start with Opus 5 and raise its effort, and to move to Fable only when evals on Opus 5 at higher effort still fall short. seeF4
- 04 At low effort Fable 5.1 is documented to call search and retrieval tools less often and answer from memory more often, which lands directly on research synthesis and adversarial review. seeF5O3
-
05
Two of the five tips in the post do not hold as stated: the prompt-simplification list is Opus 5 guidance relabelled, refuted 0 to 3, and the mid-conversation effort change is a beta API mechanism that Opus 5 also has, while Claude Code's
/effortstill clears the whole conversation cache. seeS4S5F7F8 - 06 The cheapest true claim in the post, cache reads down to $0.25 per MTok, is an API price, and it buys nothing on a plan that meters usage rather than dollars. seeS3F8
-
07
An open Claude Code issue reports a Fable subagent request being served silently by Sonnet 5 with no error, detectable only by comparing
meta.jsonagainst the model field in the agent transcript, so a Fable seat inside an Opus session is not something the tooling reliably delivers today. seeO4F10
About this research
- Why
- I run a ladder of models: Fable plans and reviews, Opus does the work, Sonnet does the grunt work. The launch tip implied the middle seat could change.
- Trigger
- Lance Martin (Anthropic) on X, launch day: run Fable 5.1 at low effort, cheaper per task than Opus, scores higher.
The five tips, paraphrased
- Low effort: competitive with Opus and Sonnet on cost per task while scoring higher; matches Fable 5 at high effort at a third of the cost (CursorBench 3.2.0).
- Cache reads four times cheaper, down to $0.25 per million tokens.
- Simplify prompts: drop verification rituals, emphasis boosters, scratchpad scaffolds, stale examples, contradictory rules.
- Change effort mid-conversation without breaking the prompt cache.
- Use the Claude Code migration and audit commands.
Each became a claim in the pool (S1 to S6), alongside seven model facts already written into my own rules (R1 to R7).
- When
- Started 02:10 UTC on 2 September 2026, six hours after the launch coverage; finished in 93 minutes. Every practitioner source is one day old.
- Method
- 101 agents, 22 sources, 24 claims, three refute-votes each. Opus 5 at xhigh searched, fetched and verified; Fable 5.1 at xhigh synthesised, then spot-checked five sources against the live pages; Opus 5 built this page.
Tokens by model, and the pipeline
- Opus 5, xhigh, 100 agents, 1,704 turns: 620k output (305k thinking), 17.4M cache writes, 139.1M cache reads. About 157M tokens processed.
- Fable 5.1, xhigh, 1 agent, 10 turns: 21k output (8k thinking), 403k cache writes, 874k cache reads. About 1.3M processed.
- Harness headline: 10.5M tokens (it counts differently from the transcript sums). Not metered: the Fable 5.1 session that framed and reviewed, and the Opus 5 build seat.
How the evidence was produced. The vote step is the one that can kill a claim. - Who else
- Comparing Fable 5.1 to Opus 5 by effort level. On Claude Max and wondering how Fable draws the weekly limit. Routing subagents to Fable in Claude Code. Picking an effort level for research or review work.
- Measure
- Right answers per unit of Max quota, per completed task. Not dollars per token.
Cost, score and the cap
The left panel is the only independent cross-model reading that exists. The right panel is what your plan actually meters.
Score against cost per task
Artificial Analysis Intelligence Index on the vertical axis, cost per index task in API dollars on the horizontal. Read the vertical axis for capability. The horizontal axis is evidence about the vendor's cost story, not about your bill.
The four points as a table, with where each number comes from
| Seat | Index | Cost per task | In the data as |
|---|---|---|---|
| Fable 5.1 low | 58 | $0.77 | Finding F3 and objection O2 |
| Opus 5 max | 63 | $2.34 | Finding F3 and objection O2 |
| Sonnet 5 max | 53-55 | $2.29 | Caveats and objection O2. The range is the spread between the July and September articles |
| Fable 5.1 max | 66 | $3.69 | Objection O3 for the index, S1 verifier evidence for the cost |
One verifier also read Opus 5 high at index 61 and Opus 5 low at index 52 for $0.43 per task on the same index, which would make Opus 5 low the cheapest seat on the board. That figure sits in the S1 verifier evidence only, so it is one reading, not a finding.
Artificial Analysis is a disclosed pre-release evaluation partner of Anthropic, and its Opus 5 and Sonnet 5 cost figures moved between its July and September articles. Treat the horizontal axis as approximate.
What the Max plan meters
Your weekly limit, and the share of it Fable models may take.
Anthropic: Fable models "draw from your plan's regular weekly usage limits and use them faster than other Claude models".
Anthropic: on Max, "You can use up to 50% of your weekly usage limits on Fable models at no extra cost".
- No multiplier is published, and nothing published says whether effort level changes the draw rate.
- The only figure anywhere is in-app Claude Code text quoted on Hacker News in June 2026, about 2x faster than Opus, for Fable 5.
- That is forum grade and predates this model.
The evidence, on screen
Captured from the live pages on 2 September 2026. Click any capture to open it full size.



Claim scorecard
Group S is the five tips in the post plus the subcommand check. Group R is the seven model facts written into the global rules earlier the same day, so a refuted R claim means a rule line is wrong.
S1partly confirmedvote 1-2Claude Fable 5.1 at low effort is often competitive with Claude Opus and Claude Sonnet on cost per task while scoring higher than them.
see alsoF3O2did not survive
Verifier 1refutedhigh confidence
Verifier 2refutedmedium confidence
Verifier 3not refutedmedium confidence
S2confirmedvote 2-1On CursorBench 3.2.0, Claude Fable 5.1 at low effort scores at parity with Claude Fable 5 at high effort, at a third of the cost.
see alsoF2
Verifier 1not refutedhigh confidence
Verifier 2not refutedmedium confidence
Verifier 3refutedhigh confidence
S3confirmedvote 3-0Claude Fable 5.1 prompt-cache reads cost $0.25 per million tokens, down from $1.00 per million tokens, i.e. 4x cheaper.
see alsoF8
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
S4refutedvote 0-3Anthropic's published guidance for Claude Fable 5.1 is to simplify prompts by removing verification rituals, emphasis boosters, scratchpad scaffolds, stale few-shot examples and contradictory rules.
see alsoF7did not survive
Verifier 1refutedhigh confidence
Verifier 2refutedhigh confidence
Verifier 3refutedhigh confidence
S5partly confirmedvote 3-0On Claude Fable 5.1, effort can be changed mid-conversation without invalidating the prompt cache, whereas previously a mid-conversation effort change invalidated the cache.
see alsoF8
Verifier 1not refutedmedium confidence
Verifier 2not refutedmedium confidence
Verifier 3not refutedhigh confidence
S6confirmedvote localThe Claude Code subcommands /claude-api migrate, /claude-api prompt-audit and /claude-api cost-optimize exist.
R1confirmedvote 3-0Claude Fable 5.1 (model ID claude-fable-5-1) is Anthropic's most capable generally available model and the successor to Claude Fable 5, priced at $10 input and $50 output per million tokens, with a 1M-token context window and 128K max output.
see alsoF1
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
R2confirmedvote 3-0As of 2026-09-02, Claude Opus 5 (claude-opus-5) and Claude Sonnet 5 (claude-sonnet-5) are the latest released versions of the Opus and Sonnet tiers; no newer Opus or Sonnet model has been released.
see alsoF1
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
R3confirmedvote 3-0On Claude Fable 5.1, thinking is always on: sending thinking type 'disabled' or a budget_tokens value returns a 400 error, and reasoning depth is controlled only through output_config.effort (low through max).
see alsoF1
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
R4confirmedvote 3-0On Claude Fable 5.1, forced tool use (tool_choice type 'any' or type 'tool') returns a 400 error.
see alsoF1
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
R5partly confirmedvote 3-0On Claude Fable 5.1, editing an earlier conversation turn invalidates the thinking blocks after it ('preserved thinking'), so conversation history must be append-only.
see alsoF1
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
R6confirmedvote 3-0Claude Fable 5.1 is not available on Priority Tier, and is not available under zero data retention unless expressly authorized by Anthropic.
see alsoF1
Verifier 1not refutedhigh confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
R7partly confirmedvote 3-0Anthropic's guidance states that prompts written for prior models are often too prescriptive for Claude Fable 5.1 and reduce its output quality.
see alsoF7
Verifier 1not refutedmedium confidence
Verifier 2not refutedhigh confidence
Verifier 3not refutedhigh confidence
The four objections
Raised against the post before the run started, then tested. Flip a card for the evidence and its sources.
Max cap weighting
Does a Fable 5.1 token draw the Claude Max 5-hour and weekly cap at the same rate as an Opus 5 token, or at a multiple that tracks the API price? If it is roughly 2x, the third-of-the-cost claim does not transfer to this plan at all.
evidence inF9
Max cap weighting
- Objection holds in substance: the API 'third of the cost' figure does not transfer to a Max plan, but the exact multiple is not public.
- Anthropic's support article states Fable models 'draw from your plan's regular weekly usage limits and use them faster than other Claude models' and that on Max 'You can use up to 50% of your weekly usage limits on Fable models at no extra cost'; it publishes no multiplier and says nothing about effort level changing the draw rate.
- The only number is blog and forum grade: in-app Claude Code text quoted on Hacker News in June 2026 said Fable 5 'Uses your limits ~2x faster than Opus', with users reporting steeper than 2x on agentic sessions; that predates Fable 5.1 and is not an Anthropic document.
- Two consequences follow. First, the 50% ceiling alone rules Fable out as the primary execution model on Max whatever the effort level.
- Second, if the meter is token-based with a fixed Fable weight, low effort still helps proportionally (CursorBench shows Fable 5.1 Low at 19,522 tokens/task vs Fable 5 High at 43,747), but against Opus 5 there is no low-effort token comparison, only Snorkel's 58% fewer output tokens at default effort; at a 2x weight that would be roughly 0.84x of Opus draw, which is an inference, not a measurement.
- Pro plans and standard Team seats get no Fable within limits at all (usage credits only).
- Claude Code's model-config page confirms Fable is not the account-type default on any plan and that Fable usage 'can bill to usage credits instead of drawing on your plan's included limits' depending on seat tier.
Benchmark breadth
Is there evidence beyond CursorBench, and beyond parity with Fable 5 at high effort, that Fable 5.1 at low beats Opus 5 at xhigh on non-coding work: reviews, research synthesis, spec writing, long agentic runs?
Benchmark breadth
- Objection holds: no public evidence shows Fable 5.1 at low effort beating Opus 5 at xhigh on non-coding work, and the nearest independent data points the other way.
- Artificial Analysis puts Fable 5.1 low at index 58 for $0.77/task against Opus 5 max at 63 for $2.34/task: cheaper, lower scoring.
- Beyond CursorBench the independent evidence is all coding or code review: Snorkel's Terminal-Bench+ run (Fable 5.1 61.5% pass@1, 58% fewer output tokens, 36% faster, but 18% vs Opus 5's 67% on build-and-dependency tasks, conclusion 'not a strict upgrade', and low effort not tested); CodeRabbit's review harness (Fable 5.1 vs Fable 5 only: latency +48.7%, precision +4.5pp, recall -1.0pp, identical tool-call counts at Low and High); Every's vibe check (subjective, all effort levels 'super good', no Opus comparison).
- Nothing tests PRDs, spec writing, research synthesis or long agentic runs at low effort against Opus 5 xhigh.
- A Hacker News comment citing the Fable 5.1 system card (pp.
- 169-170, FrontierCode Extended) says the model scores best at medium effort, adds unrequested small changes at higher efforts, and its medium is below Fable 5's best at xhigh; that is a single forum comment and was not checked against the system card itself.
- https://artificialanalysis.ai/articles/claude-fable-5-1
- https://artificialanalysis.ai/models/claude-fable-5-1-low
- https://snorkel.ai/blog/fable-5-1-vs-opus-5-coding-benchmark/
- https://www.coderabbit.ai/blog/fable-5-1-model-review
- https://every.to/vibe-check/fable-5-1-vibe-check
- https://news.ycombinator.com/item?id=49526335
- https://cursor.com/cursorbench
Correctness at low effort
Do practitioners report skipped steps, shallower reviews, or more fix cycles in agent loops at low effort?
Correctness at low effort
- Objection partly holds: quality is not intact at low effort, but the documented regressions are specific rather than the 'skips steps, more fix cycles' pattern the objection describes.
- Anthropic documents that at low effort Fable 5.1 'is less likely than Claude Fable 5 to call a search or retrieval tool, and more likely to answer from memory', with the remedy being raising effort for those turns or adding a verification nudge; its own remedy text concedes the failure mode produces out-of-date answers that sound authoritative.
- Artificial Analysis measures an 8-point index drop from Fable 5.1 max (66) to low (58).
- Against that, CodeRabbit found tool-call counts identical at Low and High with no batching regression in their loop, and Every's tester found all effort levels good in practice.
- No practitioner report of extra fix cycles in agent loops, skipped review steps, or shallower reviews at low effort was found, and several verifiers ran out of search budget before sweeping for them, so absence of reports is weak evidence.
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
- https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1
- https://artificialanalysis.ai/models/claude-fable-5-1-low
- https://www.coderabbit.ai/blog/fable-5-1-model-review
- https://every.to/vibe-check/fable-5-1-vibe-check
Runtime mechanics
Can a Claude Code session running as Opus spawn a Fable subagent today, or does the agent model fable error seen on 2026-07-07 still hold? If it holds, Fable low for execution only exists inside a Fable session.
evidence inF10
Runtime mechanics
- Objection holds in effect, with the failure mode changed. anthropics/claude-code issue #82252 (opened 2026-07-29, still open, no staff response, no workaround) reports that a subagent requested with model 'claude-fable-5' or the alias 'fable', directly or via a Workflow agent() helper, is silently served by claude-sonnet-5: meta.json and the agent's self-report show Fable, but every assistant message in the transcript carries model claude-sonnet-5, while a claude-opus-4-8 control was served correctly.
- That is worse than the error observed on 2026-07-07 because nothing surfaces it; the only detection is comparing meta.json against message.model in the agent transcript.
- The issue does not state the session model, so whether a Fable session spawning a Fable subagent works is unverified, and whether the bug persists for claude-fable-5-1 on Claude Code v2.1.255 (Fable 5.1 requires 2.1.250+) is unknown.
- Claude Code's model-config page places no documented restriction on Fable as a subagent model and notes Fable is not the account-type default on any plan.
- Until the issue closes, treat any 'Fable-low for execution' subagent as unverified and check the transcript model field.
The evidence behind the three decisive objections
Captured from the live pages on 2 September 2026. Click any capture to open it full size.



What this means for the two rulings
The post's real claim is that low effort on Fable 5.1 is not a correctness sacrifice. If that held, the effort floor was a correctness rule expressed as a number, and the number should have followed the correctness.
Execution rung is the latest Opus
- Nothing in the evidence moves it, and one thing strengthens it for a reason the ruling did not originally have.
- The ruling was made on where judgment is spent, given that Fable and Opus draw the same Max pool.
- The 50% weekly ceiling adds a plan mechanic: Fable cannot carry the highest-volume seat even if you wanted it to.
- Cycle reviews are the highest-volume role, and the low-effort retrieval regression lands hardest exactly there.
Effort: xhigh default, high is a hard floor
- The evidence does not support dropping the floor, and it does not cleanly condemn low effort either.
- What it shows is narrower: no public measurement compares Fable 5.1 at low effort with Opus 5 at xhigh on non-coding work, and the one documented low-effort regression is retrieval behaviour, not general quality collapse.
- Anthropic's own advice is to step down to medium or low only once your evals show quality holds.
- You have no such evals. Until you run them the floor is the cheaper mistake.
One honest qualification. The floor is a rule about correctness written as an effort number, and the number is not sacred. If a matched test on your own tasks shows Fable 5.1 at low or medium holding the gate, the number should move for those task types, and only those.
The test that would settle it
Nothing public will answer this for your task mix. A small matched eval will, and it costs about a morning.
- Pick two or three real tasks you have already completed well: one cycle review, one PRD prescriptive core, one build step.
- Write the correctness gate before you run anything, and make it mechanical wherever it can be. A gate written afterwards scores the run you got.
- Run each task twice: once as Opus 5 at xhigh, once as Fable 5.1 at low. Run the Fable arm inside a Fable session, so the seat is real.
- Read
/usagebefore and after each run and record the cap percentage consumed, not the token count. - On the Fable low arm, add the search nudge Anthropic recommends. Without it the documented retrieval regression is doing the work rather than the model.
- If any arm carries the Fable seat through a subagent, compare
meta.jsonagainst the model field in the agent transcript before you trust the result. - Score right answers per unit of cap consumed, per seat, per completed task. That is your metric, and no published benchmark reports it.
The 50% weekly ceiling does not depend on the test. Even a clean win for Fable 5.1 at low effort leaves it unable to carry the primary execution seat on this plan.
Anthropic's own routing advice, on screen
Captured from the live model page on 2 September 2026.
Findings
Ten findings survived synthesis, ranked by confidence. The decisive one for this decision is F9, which is medium confidence because Anthropic publishes half of it and nobody publishes the other half.
F1high confidencevote 3-0 on each of R1-R6Model facts and the three hard API blocks
- https://platform.claude.com/docs/en/models/fable-5-1/overview
- https://platform.claude.com/docs/en/about-claude/model-deprecations
- https://platform.claude.com/docs/en/models/fable-5-1/migration-guide
- https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting
- https://platform.claude.com/docs/en/api/service-tiers
- https://platform.claude.com/docs/en/manage-claude/api-and-data-retention
- https://docs.litellm.ai/blog/claude_fable_5_1
F2high confidencevote 2-1The third-of-the-cost figure is Fable against Fable
- https://cursor.com/cursorbench
- https://www.anthropic.com/claude-fable-and-mythos-5-1
- https://llm-stats.com/benchmarks/cursorbench-3.2
answersS2
F4high confidencevote 3-0Anthropic routes you to Opus 5 first
- https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1
- https://platform.claude.com/docs/en/about-claude/models/choosing-a-model
- https://platform.claude.com/docs/en/build-with-claude/effort
- https://platform.claude.com/docs/en/models/fable-5-1/migration-guide
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
F5high confidencevote 3-0, 3-0, 2-1 across three phrasingsLow effort searches less and answers from memory more
- https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
- https://platform.claude.com/docs/en/models/fable-5-1/migration-guide
answersO3
F6high confidencevote docs 3-0; blanket cost claim refuted 1-2Seven documented behaviour changes since Fable 5
- https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
- https://www.coderabbit.ai/blog/fable-5-1-model-review
answersO3
F7high confidencevote S4 0-3; R7 3-0The prompt-simplification list belongs to Opus 5
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
F8high confidencevote 3-0 on S3, S5 and pricing claimCache and effort mechanics
- https://platform.claude.com/docs/en/about-claude/pricing
- https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- https://platform.claude.com/docs/en/build-with-claude/effort
- https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1
- https://code.claude.com/docs/en/prompt-caching
- https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads
F9medium confidencevote n/a (fetched by synthesis for O1; no verifier vote)decisiveWhat a Max plan actually meters
- https://support.claude.com/en/articles/15424964-claude-fable-5-on-your-plan
- https://www.developersdigest.tech/blog/claude-usage-limits-fable-5-explained
- https://code.claude.com/docs/en/model-config
answersO1
F3medium confidencevote S1 1-2; doc-attribution claim 1-2 with verifier evidence saying not refutedThe one independent cross-model reading
- https://artificialanalysis.ai/articles/claude-fable-5-1
- https://artificialanalysis.ai/models/claude-fable-5-1-low
- https://artificialanalysis.ai/models/claude-sonnet-5
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
- https://artificialanalysis.ai/articles/opus-5
F10medium confidencevote n/a (fetched by synthesis for O4; no verifier vote)Claude Code runtime and the silent substitution
- https://code.claude.com/docs/en/model-config
- https://github.com/anthropics/claude-code/issues/82252
- https://support.claude.com/en/articles/15424964-claude-fable-5-on-your-plan
answersO4
What did not survive
Six claims were killed by two or more verifiers. Two are seed claims from the post. Four are claims extracted from the sources during the sweep, labelled E1 to E4. Collapsed by default.
S1refutedvote 1-2seed claim from the postClaude Fable 5.1 at low effort is often competitive with Claude Opus and Claude Sonnet on cost per task while scoring higher than them.
Verifier note 1
Verifier note 2
S4refutedvote 0-3seed claim from the postAnthropic's published guidance for Claude Fable 5.1 is to simplify prompts by removing verification rituals, emphasis boosters, scratchpad scaffolds, stale few-shot examples and contradictory rules.
Verifier note 1
Verifier note 2
Verifier note 3
E1refutedvote 1-2extracted claimFable 5.1 has documented behaviour regressions versus Fable 5 that raise token cost and wall-clock time without lowering answer quality: it may issue one tool call per turn where Fable 5 batched several, it rewrites whole files instead of making targeted edits, it writes fewer user-facing progress updates (especially at higher effort), its prose is denser, it uses less chat formatting, and it more often reproduces source passages in summaries without marking them as quotations.
Verifier note 1
Verifier note 2
E2refutedvote 1-2extracted claimAnthropic's own docs state that at `low` effort, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, and recommend including it wherever you would otherwise run a smaller model at higher effort.
Verifier note 1
Verifier note 2
E3refutedvote 0-3extracted claimFable 5.1 costs $10/$50 per million input/output tokens, double Opus 5's $5/$25; its prompt cache reads are $0.25 per million tokens, half the Opus 5 cache-read rate and a quarter of the Fable 5 rate. Anthropic tells migrators token counts are roughly unchanged, so per-task cost differences come from per-token pricing, not token volume.
Verifier note 1
Verifier note 2
Verifier note 3
E4refutedvote 1-2extracted claimFable 5.1's effort default is `high` with all five levels (low, medium, high, xhigh, max) supported, and Anthropic states its gains over Fable 5 are largest at xhigh and max, at the cost of extra thinking time and time-to-first-response. This directly contradicts any assumption that low effort is the intended default operating point.
Verifier note 1
Verifier note 2
Caveats and open questions
What the run could not establish, and what would settle it.
Caveats
- Verifier disagreement: S1 was refuted 1-2, but a separate verifier found the identical sentence verbatim in Anthropic's Fable 5.1 prompting guide and marked that attribution claim not refuted; treat S1 as Anthropic's own published positioning, vendor-grade, with partial independent support against Sonnet 5 only. S2 and the two low-effort regression claims were 2-1 votes.
- Vendor-only claims: 'gains widest at higher effort', 'medium roughly matches Fable 5', and 'low effort often competitive with Opus and Sonnet on cost per task' have no independent per-effort measurement; Artificial Analysis publishes low-effort figures but is a disclosed pre-release evaluation partner of Anthropic, and its Opus 5 and Sonnet 5 cost-per-task numbers differ between its July and September articles ($2.03 vs $2.34 for Opus 5 max; $1.53 vs $2.29 for Sonnet 5 max), likely index-version drift.
- CursorBench is unpublished and Cursor is a commercial partner.
- Max-plan weighting: the faster-draw statement and 50% cap are official, the ~2x multiplier is a June 2026 blog quoting Hacker News quoting in-app text for Fable 5, and whether effort level changes the draw rate is not published anywhere.
- The system-card finding that Fable 5.1 scores best at medium on FrontierCode Extended comes from one Hacker News comment and was not checked against the system card. The Fable-subagent silent-substitution issue concerns Fable 5 on an earlier Claude Code build and may or may not persist for Fable 5.1 on v2.1.255.
- Several verifiers exhausted their WebSearch budget (200/200) before sweeping for independent corroboration, so R4, R5, and the four migration-guide claims rest on Anthropic docs alone (acceptable for API contract facts, weak for anything about real-world quality).
- Time sensitivity: R2 holds only as of 2026-09-02, with Opus 5.1 and Sonnet 5.1 EAP codenames already leaked; the prompt-cache and effort features cited are beta and header-gated; preserved-thinking enforcement is account-age dependent.
- The refuted 'token counts roughly unchanged, so cost differences come from price not volume' inference should be dropped: Anthropic's own migration guide lists behaviours that move per-task token volume in both directions.
Open questions
- What is the actual Max-plan draw multiplier for Fable 5.1 versus Opus 5, and does low effort reduce the draw proportionally to tokens or is there a per-request weight? No Anthropic source publishes it; the only route is reading the /usage attribution after matched tasks.
- Does the silent Sonnet 5 substitution for a 'fable' subagent (issue #82252) still occur for claude-fable-5-1 on Claude Code v2.1.255, and does a Fable session spawning a Fable subagent work? Checkable by comparing meta.json against message.model in one agent transcript.
- Is there any measurement of Fable 5.1 at low or medium effort against Opus 5 at xhigh on non-coding knowledge work (spec writing, adversarial review, research synthesis)? None exists publicly; the only route is a small matched eval on the user's own PRD and review tasks with the search nudge in the prompt.
- Does the system-card finding that Fable 5.1 scores best at medium effort on FrontierCode Extended, with higher efforts adding unrequested changes, generalise to agentic builds like the WordPress work, and is it confirmed in the system card itself rather than a forum summary?
Sources
Every URL fetched by the run, with its quality tag, the angle it was fetched for, and its publication date where the page states one.
| URL | Quality | Angle | Publication date | Claims |
|---|---|---|---|---|
| platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 | primary | vendor-primary | Not stated on the page (no publish or last-updated date shown); fetched 2026-09-02 | 5 |
| platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompti... | primary | vendor-primary | Not stated on the page (no publish or last-updated date shown); content is current as of retrieval on 2026-09-02 and references an August 31, 2026 account cutoff plus beta headers dated 2026-08-01/08-18/08-21 | 5 |
| platform.claude.com/docs/en/models/fable-5-1/migration-guide | primary | vendor-primary | Not stated on the page; content is dated by internal references (enforcement threshold of 2026-08-31 and an example date of 2026-09-14), so it is current as of at least September 2026. | 5 |
| platform.claude.com/docs/en/build-with-claude/effort | primary | vendor-primary | Not stated on the page (Anthropic platform docs page; no publish or last-updated date shown). Content references a 2026-07-01 beta header and models up to Fable 5.1 / Mythos 5.1, so it is current as of at least mid-2026. | 5 |
| platform.claude.com/docs/en/models/overview | primary | vendor-primary | Not stated on page (no publish or last-updated date rendered); retrieved 2026-09-02. Content is current as of the Fable 5.1 / Opus 5 / Sonnet 5 lineup, with Fable 5.1 retirement commitment "not sooner than September 1, 2027". | 5 |
| support.claude.com/en/articles/15424964-claude-fable-5-on-your-plan | primary | vendor-primary | Not dated on the page; it displays only "Updated" with a relative timestamp (rendered as "today" when fetched on 2026-09-02). | 5 |
| x.com/ArtificialAnlys/status/2094881171066978525 | primary | independent-benchmarks | 2026-09-01 | 5 |
| artificialanalysis.ai/articles/opus-5 | primary | independent-benchmarks | 2026-07-24 | 5 |
| www.vals.ai/benchmarks/terminal-bench-2-1 | primary | independent-benchmarks | 2026-08-31 (page states "Updated 8/31/2026") | 5 |
| composio.dev/content/opus-vs-fable | blog | independent-benchmarks | 2026-07-27 | 5 |
| www.coderabbit.ai/blog/fable-5-1-model-review | primary | independent-benchmarks | 2026-09-01 | 5 |
| every.to/vibe-check/fable-5-1-vibe-check | blog | independent-benchmarks | 2026-09-01 | 5 |
| news.ycombinator.com/item?id=48492210 | forum | practitioner-reports | approx. 2026-06-12 (HN shows "82 days ago" as of 2026-09-02); exact date not displayed | 5 |
| opinionai.substack.com/p/i-tested-opus-5-and-fable-5-the-cheaper | blog | practitioner-reports | 2026-07-28 | 5 |
| chudi.dev/blog/claude-fable-5-vs-opus-4-8 | blog | practitioner-reports | 2026-06-10 (updated 2026-09-01) | 5 |
| www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes | blog | skeptical-regressions | 2026-09-01 | 5 |
| roo.beehiiv.com/p/claude-fable-5-1-mythos-5-1-whats-new | blog | skeptical-regressions | 2026-09-01 | 5 |
| zenn.dev/uhyo/articles/react-profession-bench-13 | blog | skeptical-regressions | 2026-07-17 | 5 |
| www.developersdigest.tech/blog/claude-usage-limits-fable-5-explained | blog | max-plan-weighting | 2026-06-10 | 5 |
| code.claude.com/docs/en/model-config | primary | claude-code-runtime | Not stated on the page (no visible publish or last-updated date); content fetched 2026-09-02 | 5 |
| code.claude.com/docs/en/prompt-caching | primary | claude-code-runtime | Not recorded on the page | 5 |
| github.com/anthropics/claude-code/issues/82252 | forum | claude-code-runtime | 2026-07-29 (issue opened; follow-up comment 2026-08-09, last updated 2026-08-20; still OPEN, label area:agents) | 5 |
22 further URLs carry claims on this page but are not in the run's fetch log
Verifiers read or quoted these while testing claims. They are not in the run's own source table, so their quality tags and publication dates were never recorded.
- artificialanalysis.ai/articles/claude-fable-5-1
- artificialanalysis.ai/models/claude-fable-5-1-low
- artificialanalysis.ai/models/claude-sonnet-5
- cursor.com/cursorbench
- docs.litellm.ai/blog/claude_fable_5_1
- llm-stats.com/benchmarks/cursorbench-3.2
- news.ycombinator.com/item?id=49526335
- platform.claude.com/docs/en/about-claude/model-deprecations
- platform.claude.com/docs/en/about-claude/models/choosing-a-model
- platform.claude.com/docs/en/about-claude/pricing
- platform.claude.com/docs/en/api/service-tiers
- platform.claude.com/docs/en/build-with-claude/preserved-thinking
- platform.claude.com/docs/en/build-with-claude/prompt-caching
- platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5
- platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5
- platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting
- platform.claude.com/docs/en/manage-claude/api-and-data-retention
- platform.claude.com/docs/en/models/fable-5-1/overview
- snorkel.ai/blog/fable-5-1-vs-opus-5-coding-benchmark/
- venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads
- www.anthropic.com/claude-fable-and-mythos-5-1
Built from the run's own output file. Every number, quote and verdict on this page comes from that file or from the run brief. Em dashes in quoted verifier text have been replaced with hyphens, and no other wording was changed. Back to the top.
