VASEOSend a brief

Verified research run, 2026-09-02

Should the execution seat move from Opus 5 at xhigh to Fable 5.1 at low effort?

I run three layers of intelligence. Fable plans and judges, Opus does the work, Sonnet does the grunt work. Anthropic's launch tip said the top layer at low effort could take the middle seat and cost less. Six hours after launch I sent 101 agents at 22 sources to find out, with three verifiers trying to kill every claim. The answer turned on a number the launch coverage never mentioned.

Read this before anything else. Fable 5.1 shipped on 2026-09-01, the day before this run. Every practitioner and benchmark source below is launch-day material, and the vendor documentation has no dated revision history. One day of reports is not a track record, and nothing here should be treated as one.
search, fetch, verify [opus-5/xhigh] synthesis [fable-5-1] build [opus-5/inherit] review [fable-5-1]
22sources fetched6 search angles
110claims extractedtop 12 verified, plus 12 seed claims
24claims verified18 confirmed 6 refuted 0 unverified
101agents100 Opus 5 at xhigh, 1 Fable 5.1 at xhigh
158Mtokens processed0.64M generated, 88% cache re-reads, 93 minutes
Tokens: cached, input and output, by model
Model, effortAgentsTurnsCache readsCache writesUncached inputOutputof which thinkingProcessed
Opus 5, xhigh1001,704139,101,56217,405,4763,408620,137305,228157,130,583
Fable 5.1, xhigh110874,357402,87815,65320,7167,9601,313,604
Total1011,714139,975,91917,808,35419,061640,853313,188158,444,187
What each column means, and the API-price equivalent
  • Cache reads are the same context read again on each turn: the harness system prompt, the tool definitions and the agent's own growing conversation, about 82,000 tokens per turn. They are 88% of the total and are billed at a fraction of the input price.
  • Cache writes are context seen for the first time: each of the 100 agents starts with roughly 170,000 tokens of boilerplate, written once per agent, plus the pages it fetched.
  • Output is what the models actually generated, thinking included. That is the real work: 0.64M tokens.
  • Summed from every agent's transcript. The harness's own headline for the run is 10.5M, which counts differently. Not included: the Opus 5 seat that built this page (289k processed) and the Fable 5.1 session that framed the question and reviewed the result.
  • At API prices this would be about $194 in Opus 5 (reads $0.50, writes $6.25, output $25 per million) and $6 in Fable 5.1. On a Claude Max plan it drew from the weekly quota instead, and that draw is not published per model.
The rest of the run stats
  • 12 seed claims plus 12 extracted claims were verified.
  • 98 lower ranked extracted claims were dropped by the budget cap, 8 claims were dropped for budget inside the sweep, 6 duplicate URLs were filtered, and 10 findings survived synthesis.
  • Seed S6, the existence of the /claude-api subcommands, was checked locally against the bundled skill table before the run started, so no verifier agents were spent on it.

The answer

Do not switch. Keep Opus 5 at xhigh as the execution seat.

  • Use Fable 5.1 for the judgment and planning seats at the effort you set for the session.
  • Try Fable 5.1 at low effort only on your own matched tasks, with a search nudge in the prompt.
  • That is where the evidence lands, and it is one day of evidence.
plan Fable models are capped at 50% of your weekly limit on Max and draw it faster than other models. That alone rules Fable out as the primary execution seat, at any effort level.
quality No public measurement puts Fable 5.1 at low effort against Opus 5 at xhigh on the work you actually run. The one independent cross-model index has Fable low below Opus max on score.
failure mode At low effort Fable 5.1 calls search and retrieval tools less and answers from memory more. That is the wrong failure mode for research and reviews.

The answer at a glance

Seven sentences. Each one links to the evidence that carries it.

  • 01 On Claude Max, Fable usage is capped at 50% of the weekly limit and draws that limit faster than other models, with no multiplier published, so Fable 5.1 cannot be the primary execution model whatever the effort setting. seeO1F9the cap
  • 02 The one independent cross-model reading puts Fable 5.1 low at index 58 for $0.77 per task against Opus 5 max at 63 for $2.34: cheaper in API dollars, lower scoring, and those dollars are not what a Max plan meters. seeF3O2the chart
  • 03 Anthropic's own routing guidance says to start with Opus 5 and raise its effort, and to move to Fable only when evals on Opus 5 at higher effort still fall short. seeF4
  • 04 At low effort Fable 5.1 is documented to call search and retrieval tools less often and answer from memory more often, which lands directly on research synthesis and adversarial review. seeF5O3
  • 05 Two of the five tips in the post do not hold as stated: the prompt-simplification list is Opus 5 guidance relabelled, refuted 0 to 3, and the mid-conversation effort change is a beta API mechanism that Opus 5 also has, while Claude Code's /effort still clears the whole conversation cache. seeS4S5F7F8
  • 06 The cheapest true claim in the post, cache reads down to $0.25 per MTok, is an API price, and it buys nothing on a plan that meters usage rather than dollars. seeS3F8
  • 07 An open Claude Code issue reports a Fable subagent request being served silently by Sonnet 5 with no error, detectable only by comparing meta.json against the model field in the agent transcript, so a Fable seat inside an Opus session is not something the tooling reliably delivers today. seeO4F10

About this research

Why
I run a ladder of models: Fable plans and reviews, Opus does the work, Sonnet does the grunt work. The launch tip implied the middle seat could change.
Trigger
Lance Martin (Anthropic) on X, launch day: run Fable 5.1 at low effort, cheaper per task than Opus, scores higher.
The five tips, paraphrased
  • Low effort: competitive with Opus and Sonnet on cost per task while scoring higher; matches Fable 5 at high effort at a third of the cost (CursorBench 3.2.0).
  • Cache reads four times cheaper, down to $0.25 per million tokens.
  • Simplify prompts: drop verification rituals, emphasis boosters, scratchpad scaffolds, stale examples, contradictory rules.
  • Change effort mid-conversation without breaking the prompt cache.
  • Use the Claude Code migration and audit commands.

Each became a claim in the pool (S1 to S6), alongside seven model facts already written into my own rules (R1 to R7).

When
Started 02:10 UTC on 2 September 2026, six hours after the launch coverage; finished in 93 minutes. Every practitioner source is one day old.
Method
101 agents, 22 sources, 24 claims, three refute-votes each. Opus 5 at xhigh searched, fetched and verified; Fable 5.1 at xhigh synthesised, then spot-checked five sources against the live pages; Opus 5 built this page.
Tokens by model, and the pipeline
  • Opus 5, xhigh, 100 agents, 1,704 turns: 620k output (305k thinking), 17.4M cache writes, 139.1M cache reads. About 157M tokens processed.
  • Fable 5.1, xhigh, 1 agent, 10 turns: 21k output (8k thinking), 403k cache writes, 874k cache reads. About 1.3M processed.
  • Harness headline: 10.5M tokens (it counts differently from the transcript sums). Not metered: the Fable 5.1 session that framed and reviewed, and the Opus 5 build seat.
Search6 anglesOpus 5, xhigh Fetch and extract22 sources, 110 claimsOpus 5, xhigh Verification pool12 seed + 12 extractedthe rest dropped, counted Verify3 verifiers per claimtold to refute; Opus 5, xhigh Synthesisfindings, scorecard, O1 to O4Fable 5.1, xhigh Review and build5 sources spot-checkedFable 5.1 reviews, Opus 5 builds 2 of 3 refutations kill a claim; a failed verifier counts as unverified, never as refuted 18 survived 6 refuted, still shown on the page 0 unverified
How the evidence was produced. The vote step is the one that can kill a claim.
Who else
Comparing Fable 5.1 to Opus 5 by effort level. On Claude Max and wondering how Fable draws the weekly limit. Routing subagents to Fable in Claude Code. Picking an effort level for research or review work.
Measure
Right answers per unit of Max quota, per completed task. Not dollars per token.

Cost, score and the cap

Fable 5.1low effort Fable 5high effort Opus 5xhigh effort what the post measured CursorBench 3.2: 66.2% vs 66.5%, one third of the cost what the decision needs no public measurement on any non-coding work closest independent data: Fable 5.1 low 58 vs Opus 5 max 63 (Artificial Analysis; max, not xhigh)
The launch post compares Fable 5.1 with its own predecessor. The comparison this decision turns on, the dashed arrow, has no public measurement. The nearest independent number points the other way.

The left panel is the only independent cross-model reading that exists. The right panel is what your plan actually meters.

Score against cost per task

Artificial Analysis Intelligence Index on the vertical axis, cost per index task in API dollars on the horizontal. Read the vertical axis for capability. The horizontal axis is evidence about the vendor's cost story, not about your bill.

Intelligence Index against cost per task Fable 5.1 at low effort scores 58 at $0.77 per task. Opus 5 at max effort scores 63 at $2.34. Sonnet 5 at max effort scores 53 to 55 at $2.29. Fable 5.1 at max effort scores 66 at $3.69. 50 55 60 65 70 $0 $1 $2 $3 $4 Cost per task, API dollars Intelligence Index Fable 5.1 low 58 at $0.77 Opus 5 max 63 at $2.34 Sonnet 5 max 53-55 at $2.29 Fable 5.1 max 66 at $3.69
The four points as a table, with where each number comes from
SeatIndexCost per taskIn the data as
Fable 5.1 low58$0.77Finding F3 and objection O2
Opus 5 max63$2.34Finding F3 and objection O2
Sonnet 5 max53-55$2.29Caveats and objection O2. The range is the spread between the July and September articles
Fable 5.1 max66$3.69Objection O3 for the index, S1 verifier evidence for the cost

One verifier also read Opus 5 high at index 61 and Opus 5 low at index 52 for $0.43 per task on the same index, which would make Opus 5 low the cheapest seat on the board. That figure sits in the S1 verifier evidence only, so it is one reading, not a finding.

Artificial Analysis is a disclosed pre-release evaluation partner of Anthropic, and its Opus 5 and Sonnet 5 cost figures moved between its July and September articles. Treat the horizontal axis as approximate.

What the Max plan meters

Your weekly limit, and the share of it Fable models may take.

0%50%100% of the weekly limit

Anthropic: Fable models "draw from your plan's regular weekly usage limits and use them faster than other Claude models".

Anthropic: on Max, "You can use up to 50% of your weekly usage limits on Fable models at no extra cost".

  • No multiplier is published, and nothing published says whether effort level changes the draw rate.
  • The only figure anywhere is in-app Claude Code text quoted on Hacker News in June 2026, about 2x faster than Opus, for Fable 5.
  • That is forum grade and predates this model.

seeO1F9

The evidence, on screen

Captured from the live pages on 2 September 2026. Click any capture to open it full size.

Artificial Analysis article, 1 September 2026: at max effort Fable 5.1 scores 66, ahead of Claude Opus 5 at max on 63; xhigh scores 65 at $2.72 per task against Opus 5 max at $2.34
Artificial Analysis, 1 September 2026. The only independent cross-model index. Fable 5.1 at max scores 66 and at xhigh 65 for $2.72 per task; Opus 5 at max scores 63 for $2.34. Low effort scores 58. The article opens by disclosing that it supported Anthropic's pre-release evaluation.
CursorBench 3.2 score against dollars per task, with the Fable 5.1 and Opus 5 curves across effort levels
CursorBench 3.2, live leaderboard. Score against cost per task; each curve is one model across its effort levels. The launch post's "third of the cost" reads the Fable 5.1 curve against the Fable 5 curve, not against Opus 5. Cursor is a commercial partner and the dataset is unpublished.
Anthropic pricing table: Claude Fable 5.1 cache hits $0.25 per MTok, Fable 5 $1, Opus 5 $0.50
Anthropic pricing page. Cache hits: Fable 5.1 at $0.25 per MTok against Fable 5 at $1 and Opus 5 at $0.50. Base input and output prices unchanged. This is the one launch claim that holds cleanly, and it is an API price, so it does not reach a Max-plan user.

Claim scorecard

Group S is the five tips in the post plus the subcommand check. Group R is the seven model facts written into the global rules earlier the same day, so a refuted R claim means a rule line is wrong.

confirmed partly confirmed refuted Vote reads as kept to refuted across three independent verifiers. Two refutations kill a claim.
Group S: the five tips in the post, plus the subcommand check
S1partly confirmedvote 1-2Claude Fable 5.1 at low effort is often competitive with Claude Opus and Claude Sonnet on cost per task while scoring higher than them.
Verdict noteThe refuting verifiers missed that Lance Martin's sentence is a verbatim quote of Anthropic's Fable 5.1 prompting guide, but the substantive claim is vendor-only except for Artificial Analysis data that supports it against Sonnet 5 max (58 at $0.77 vs 53-55 at $2.29) and refutes 'scoring higher' against Opus 5 max (63 at $2.34).
Run statusrefuted, vote 1-2, 0 verifier errors

see alsoF3O2did not survive

The three verifiers
Verifier 1refutedhigh confidence
REFUTED: no public support, and Anthropic's own docs point the opposite way on three counts. 1) BASELINE SUBSTITUTION (the core defect). Anthropic's launch post (https://www.anthropic.com/claude-fable-and-mythos-5-1) says: "when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost." The comparator is FABLE 5 - the model's own predecessor - not Opus 5 or Sonnet 5. I checked the page specifically for cross-model effort charts: there are none comparing Fable 5.1 low effort to Opus/Sonnet on cost-per-task, and the model benchmark table is not broken down by effort level. The post under review swaps "Fable 5" for "Opus and Sonnet," which changes the claim materially: matching your own prior generation more cheaply is not the same as beating two cheaper models on $/task. 2) ANTHROPIC'S PUBLIC ROUTING ADVICE CONTRADICTS THE IMPLICATION. https://platform.claude.com/docs/en/models/fable-5-1/overview and .../whats-new-fable-5-1 both state verbatim: "For most workloads, start with Claude Opus 5... Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5 at higher effort still fall short." https://platform.claude.com/docs/en/about-claude/models/choosing-a-model says "Most workloads start with Claude Opus 5," listing Fable 5.1 only under "The highest available capability." If low-effort Fable 5.1 were cost-competitive with Opus/Sonnet AND higher-scoring, this guidance would be irrational. Anthropic also advises tuning effort WITHIN a model rather than switching to Fable: "Tuning effort is often a better lever than switching models." 3) "SCORING HIGHER" IS NOT WHAT THE EFFORT DOC SAYS. https://platform.claude.com/docs/en/build-with-claude/effort defines low as "Most efficient. Significant token savings with some capability reduction," use case "Simpler tasks... such as subagents." Fable 5.1 guidance: "Start with high, the default... step down to medium or low for routine or latency-sensitive work ONCE YOUR EVALS SHOW QUALITY HOLDS" - conditional quality PRESERVATION, not superiority. Tellingly, the cross-model boast that exists for Fable 5 ("Lower effort settings on Claude Fable 5 still perform well and often exceed xhigh performance on prior models") is ABSENT from the Fable 5.1 section - and even that one names "prior models," not Opus 5/Sonnet 5, and makes no $/task claim. 4) INDEPENDENT DATA RUNS THE OTHER WAY. Artificial Analysis (https://artificialanalysis.ai/models/claude-fable-5-1) lists Fable 5.1 at Intelligence Index 66 with "Cost per Intelligence Index task: $3.69". Their Opus 5 analysis (https://artificialanalysis.ai/articles/opus-5) reports Opus 5 (max) at $2.03/task vs Fable 5 at $2.75, Sonnet 5 (max) at $1.53 and Opus 4.8 (max) at $1.80 - headline: "Opus 5: Fable 5 level intelligence at a lower cost per task." AA also notes "at low and medium effort Sonnet 5 stays genuinely cheaper." Critically, AA publishes NO low-effort Fable 5.1 cost-per-task figure at all (their Fable 5.1 entry is "Adaptive Reasoning, Max Effort, Default Fallback"), so the specific claim is publicly UNMEASURED, not merely disputed. 5) PRICE ARITHMETIC MAKES IT A DEMANDING CLAIM. Per https://platform.claude.com/docs/en/models/fable-5-1/overview, Fable 5.1 is $10/$50 per MTok vs Opus 5 $5/$25 (2x) and Sonnet 5 $2/$10 (5x). To be "competitive on $/task" with Opus, Fable-low must consume <=~50% of Opus's tokens on the same task; vs Sonnet, <=~20%. No public source establishes those ratios. The one real cost lever - cache reads at $0.25/MTok (0.025x base vs 0.1x on other models) - only bites on long cached agentic sessions that re-read a prefix, which is a narrow regime being presented as general (checklist item 5, cherry-picking). 6) DOCUMENTED LOW-EFFORT REGRESSION ON THE KNOWLEDGE-WORK AXIS. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 lists under "Changed from Claude Fable 5": "Answers from memory more often at low effort. At the lowest effort level the model calls a search or retrieval tool less often." That is a named quality regression at precisely the setting the claim praises - direct counter-evidence to "scoring higher" for research/knowledge work. SOURCE QUALITY / AMBIGUITY: the claim is a vendor employee's X post, marketing-grade per the checklist, with no independent corroboration found for the low-effort-vs-Opus/Sonnet cost-per-task comparison. It is also ambiguous whether "Opus and Sonnet" means the current Opus 5/Sonnet 5 or prior generations (4.8/4.6); on the charitable prior-gen reading it may hold, but then it is not a claim about the current lineup and does not support routing decisions today. NOT REFUTED ON RECENCY - the claim is current (Fable 5.1 released 2026-09-01, per the docs) - it is refuted on absent public support, comparator substitution, and contrary independent cost-per-task measurement. COUNTER-SOURCES: https://platform.claude.com/docs/en/models/fable-5-1/overview | https://platform.claude.com/docs/en/about-claude/models/choosing-a-model | https://platform.claude.com/docs/en/build-with-claude/effort | https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 | https://www.anthropic.com/claude-fable-and-mythos-5-1 | https://artificialanalysis.ai/models/claude-fable-5-1 | https://artificialanalysis.ai/articles/opus-5
Counter sourcehttps://platform.claude.com/docs/en/models/fable-5-1/overview ("For most workloads, start with Claude Opus 5... or when your evals on Claude Opus 5 at higher effort still fall short") and https://artificialanalysis.ai/articles/opus-5 ("Opus 5: Fable 5 level intelligence at a lower cost per task"; Opus 5 max $2.03/task, Sonnet 5 max $1.53/task vs Fable 5.1 $3.69/task per https://artificialanalysis.ai/models/claude-fable-5-1)
Verifier 2refutedmedium confidence
REFUTED AS A GENERAL CLAIM; one sub-case survives. The claim is a faithful restatement of the quote, so it is not a misread - but it fails as stated because "competitive on $/task while scoring higher" than Opus AND Sonnet is true only under one specific, unstated effort pairing, and flips against Opus 5 at other pairings. (1) Anthropic's own public documentation does NOT support the Opus/Sonnet cost comparison at all. On the launch page https://www.anthropic.com/claude-fable-and-mythos-5-1 the cost-vs-score charts (Terminal-Bench 4.0, Terminal-Bench-Science 0.1, Humanity's Last Exam, CursorBench 3.2.0) plot Fable 5.1 only against Fable 5 and Mythos. Opus 5 and Sonnet 5 appear in the raw benchmark score table but are ABSENT from every cost-per-task chart. The only effort-level cost statement Anthropic makes is "When set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost" - Fable 5 as the baseline, not Opus or Sonnet. So the vendor's own docs make a Fable-vs-Fable claim; the post extends it to Opus and Sonnet without vendor backing. (2) Independent data (Artificial Analysis, model pages fetched 2026-09-02) partially contradicts it. Fable 5.1 low effort: index 58, $0.77/task (https://artificialanalysis.ai/models/claude-fable-5-1-low). Opus 5 max: index 63, $2.34/task (https://artificialanalysis.ai/models/claude-opus-5). Opus 5 LOW: index 52, $0.43/task (https://artificialanalysis.ai/models/claude-opus-5-low). Sonnet 5 max: index 55, $1.72/task (https://artificialanalysis.ai/models/claude-sonnet-5). Fable 5.1 max: index 66, $3.76/task (https://artificialanalysis.ai/models/claude-fable-5-1). - vs Sonnet 5 (max): claim HOLDS - 58 @ $0.77 beats 55 @ $1.72 on both axes. - vs Opus 5 (low, matched effort): scores higher (58 vs 52) but costs 1.8x ($0.77 vs $0.43). "Competitive" is a stretch; "cheaper" is false. - vs Opus 5 (max): scores 5 points LOWER (58 vs 63). "Scoring higher" is false. So the claim's truth depends entirely on which Opus configuration is silently chosen, and it is false against the configuration a top-seat agentic-coding user would actually run. (3) The independent headline runs the opposite way from the post's framing. Artificial Analysis's published summary (https://x.com/ArtificialAnlys/status/2094881171066978525, reported at https://cryptobriefing.com/claude-fable-5-1-tops-intelligence-index/) is "Claude Fable 5.1 tops the Intelligence Index but costs 20% MORE per task than Fable 5", with Fable 5.1 max at 1.6x Opus 5 max cost, driven by more verbose output on hard problems. AA's earlier Opus 5 writeup is literally titled "Opus 5: Fable 5 level intelligence at a lower cost per task" (https://artificialanalysis.ai/articles/opus-5). Note AA ran pre-release evals FOR Anthropic, so it is not fully arms-length - but it published cost findings unfavourable to Anthropic, which raises rather than lowers its credibility here. (4) Source quality is insufficient for the claim's strength. Lance Martin is an Anthropic employee posting on X about his employer's launch-day model; per the checklist that is marketing-grade. It generalises ("often", "Opus and Sonnet") from a single effort level on a single composite benchmark, with the comparison effort level for the baselines left unstated - textbook cherry-pick presented as general. (5) Not outdated (Fable 5.1 launched ~2026-09-01), and one further caveat: AA's "$/task" is a benchmark-suite average, not cost-per-COMPLETED-task in agentic work. Fable 5.1's documented behaviour (long turns, always-on thinking, more verbose output) means the metric may not transfer to real agentic loops, and it is irrelevant to Claude Max plan limits, which meter usage windows rather than dollars. Verdict: refuted as worded. Defensible narrower version: "Fable 5.1 at low effort dominates Sonnet 5 on both score and cost per task on the AA Intelligence Index, and outscores Opus 5 at matched low effort for ~1.8x the cost." Do not repeat the Opus half without naming the effort level.
Counter sourcehttps://artificialanalysis.ai/models/claude-opus-5 (Opus 5 max: index 63 @ $2.34/task vs Fable 5.1 low: index 58 @ $0.77/task, https://artificialanalysis.ai/models/claude-fable-5-1-low); https://artificialanalysis.ai/models/claude-opus-5-low (Opus 5 low: index 52 @ $0.43/task - 1.8x cheaper than Fable 5.1 low); https://www.anthropic.com/claude-fable-and-mythos-5-1 (Anthropic's own cost-vs-score charts omit Opus 5 and Sonnet 5 entirely); https://cryptobriefing.com/claude-fable-5-1-tops-intelligence-index/ and https://x.com/ArtificialAnlys/status/2094881171066978525 (AA: Fable 5.1 costs 20% MORE per task than Fable 5, 1.6x Opus 5 at max)
Verifier 3not refutedmedium confidence
SUPPORTED BY PUBLIC EVIDENCE, but materially over-broad on the "scoring higher" half. 1) The claim is verbatim official Anthropic documentation, not just an employee's social post. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#consider-all-effort-levels states: "At `low`, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, so include it in the comparison wherever you'd otherwise run a smaller model at a higher effort level." Lance Martin's paraphrase is faithful; no overreach or misread. (Checklist 1 passes.) 2) Independent corroboration exists, on published per-effort numbers. Artificial Analysis Intelligence Index, all fetched 2026-09-02: - Fable 5.1 low: 58 @ $0.77/task (https://artificialanalysis.ai/models/claude-fable-5-1-low) - Fable 5.1 high: 62 @ $1.43 (https://artificialanalysis.ai/models/claude-fable-5-1-high) - Fable 5.1 max: 66 @ $3.69 (https://artificialanalysis.ai/models/claude-fable-5-1) - Opus 5 low 52 @ $0.43; high 61 @ $1.23; max 63 @ $2.34 (https://artificialanalysis.ai/models/claude-opus-5-low, -high, https://artificialanalysis.ai/models/claude-opus-5) - Sonnet 5 max: 55 @ $1.72 (https://artificialanalysis.ai/models/claude-sonnet-5) Against Sonnet 5 at its TOP setting, Fable 5.1 low scores higher (58 vs 55) at 45% of the cost - the claim's core case, exactly the "smaller model at a higher effort level" framing, independently confirmed. Against Opus 5 low it also scores higher (58 vs 52), though at 1.8x the cost. 3) WHERE IT FAILS, and the caveats a reader needs. Against Opus 5 at high (the Claude Code default) and max, Fable 5.1 at low scores LOWER, not higher: 58 vs 61 and 63. So "scoring higher than Opus" is false at the effort levels most agentic users actually run. Anthropic's own routing guidance points the other way too: https://platform.claude.com/docs/en/about-claude/models/choosing-a-model says "Most workloads start with Claude Opus 5," and https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 positions Fable 5.1 as the escalation target "when your evals on Claude Opus 5 at higher effort still fall short." 4) Vendor-side support for the claim is assertion-only. The launch page (https://www.anthropic.com/claude-fable-and-mythos-5-1) publishes NO cost-per-task comparison against Opus 5 or Sonnet 5 - every cost chart is Fable 5.1 vs Fable 5, and Sonnet 5 appears once, in the safety section. The Effort docs page (https://platform.claude.com/docs/en/build-with-claude/effort) makes no such cross-model cost claim and hedges the low-effort advice as "once your evals show quality holds." The claim rests on the prompting-guide sentence plus AA's numbers, nothing else. 5) Knowledge-work caveat that cuts against "scoring higher": Anthropic's own docs record that at low effort Fable 5.1 "answers from memory more often" and "calls a search or retrieval tool less often" (whats-new-fable-5-1, "Changed from Claude Fable 5"). For research/knowledge work needing fresh information, the low-effort score advantage is not safe to assume. 6) Not outdated (model shipped 2026-09-01, data checked 2026-09-02), but it is ONE benchmark suite, ONE day old, and AA disclosed it "supported @AnthropicAI with pre-release evaluation" - not fully arm's-length, though its headline finding was unflattering (Fable 5.1 max costs 20% more per task than Fable 5 and 1.6x Opus 5 at $3.76 vs $2.34, per https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/), which argues it is not merely repeating marketing. Net: refuted=false because the claim is current, verbatim in public vendor documentation, and independently corroborated for the Sonnet comparison and the Opus-low comparison. It should be quoted only with the qualifier that against Opus 5 at high/max effort, Fable 5.1 at low scores lower, and that the "often" is load-bearing.
Counter sourceArtificial Analysis per-effort Intelligence Index pages (https://artificialanalysis.ai/models/claude-fable-5-1-low vs https://artificialanalysis.ai/models/claude-opus-5-high): Fable 5.1 low scores 58 against Opus 5 high at 61 and Opus 5 max at 63, so the "scoring higher" half is false against Opus 5 at its default and top effort levels. Also Anthropic's own https://platform.claude.com/docs/en/about-claude/models/choosing-a-model ("Most workloads start with Claude Opus 5") and the-decoder/AA finding that Fable 5.1 at max costs $3.76/task vs Opus 5's $2.34.
S2confirmedvote 2-1On CursorBench 3.2.0, Claude Fable 5.1 at low effort scores at parity with Claude Fable 5 at high effort, at a third of the cost.
Verdict noteCursor's live leaderboard shows Fable 5.1 Low 66.2% at $2.90/task vs Fable 5 High 66.5% at $8.77/task (33% of the cost), but this compares Fable to its own predecessor and says nothing about Opus 5 or Sonnet 5.
Run statusconfirmed, vote 2-1, 0 verifier errors

see alsoF2

The three verifiers
Verifier 1not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM, with exact numbers from the benchmark operator (not Anthropic). Primary check - https://cursor.com/cursorbench (CursorBench 3.2, dated "Jul 8, 2026", 61 leaderboard entries). Verbatim rows: - Rank 15 "Fable 5.1 Low" - 66.2%, $2.90/task, 19,522 tokens, 31 steps - Rank 14 "Fable 5 High" - 66.5%, $8.77/task, 43,747 tokens, 48 steps Score delta = -0.3pp (parity, within any reasonable noise band). Cost ratio = 2.90/8.77 = 33.1%, i.e. almost exactly "a third of the cost". Both halves of Lance Martin's sentence check out numerically. Fetched twice with different prompts; numbers stable. Corroboration - https://www.anthropic.com/claude-fable-and-mythos-5-1 states "when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost", plus 73.4% on CursorBench 3.2.0 at max effort vs Fable 5's 70.5%. This is vendor-grade, but it matches the third-party leaderboard rather than exceeding it. Source quality is adequate because the claim is a narrow, checkable numeric comparison and I checked it against the party that runs the benchmark, not against the person who posted it. Currency: fine. CursorBench 3.2 predates the 2026-09-01 Fable 5.1 launch; leaderboard is live as of 2026-09-02. CAVEATS the parent agent should carry (none of which refute the claim as worded): 1. "Parity" is generous framing - Fable 5.1 Low is 0.3pp BELOW Fable 5 High, not equal or above. 2. Not reproducible. CursorBench is Cursor's proprietary suite from real Cursor sessions; llm-stats.com/benchmarks/cursorbench-3.2 lists "Dataset Soon / Code Soon". Cursor is independent of Anthropic but is a major commercial partner, so treat as third-party-but-interested, not academic. 3. Intra-family only. It compares Fable 5.1 to its own predecessor. It says nothing about Opus 5 or Sonnet 5, which is what most of the research question asks. Do not let this claim carry weight it does not have. 4. Field context contradicts the implied cost-efficiency win: on the same leaderboard "Grok 4.6 Extra High" is 70.8% at $2.81 - higher score AND lower cost than Fable 5.1 Low at 66.2%/$2.90. Cherry-pick risk is real if the claim is generalised beyond Anthropic's own lineup. 5. Cost mechanism is mixed: tokens/task fall 43,747 -> 19,522 (~45% of the tokens) but cost falls to ~33%, so the 75% cache-read price cut ($0.25/MTok) does part of the work. The saving is partly a pricing change, not purely effort-driven efficiency, and it is an API-dollar saving that does not map cleanly onto Claude Max plan usage limits.
Counter sourcehttps://cursor.com/cursorbench - leaderboard rows "Fable 5.1 Low" 66.2%/$2.90 and "Fable 5 High" 66.5%/$8.77 (supports, does not counter). Nearest genuine counter on the same page: "Grok 4.6 Extra High" 70.8%/$2.81, beating Fable 5.1 Low on both score and cost, which undercuts any generalised efficiency reading. Independent reproduction is unavailable: https://llm-stats.com/benchmarks/cursorbench-3.2 shows "Dataset Soon / Code Soon".
Verifier 2not refutedmedium confidence
SUPPORTED by public evidence, with caveats. The numbers check out against the benchmark owner's own page, not just the Anthropic post. 1. PRIMARY VERIFICATION (non-Anthropic). https://cursor.com/cursorbench publishes CursorBench 3.2 per-effort rows with score and avg cost/task. Fetched twice with different prompts, both reads agreed: - Fable 5.1: Max 73.4%/$9.64 | XHigh 72.8%/$6.96 | High 69.4%/$4.80 | Med 68.0%/$3.53 | Low 66.2%/$2.90 - Fable 5: Max 70.5%/$17.32 | XHigh 68.4%/$11.73 | High 66.5%/$8.77 | Med 65.2%/$6.80 | Low 62.1%/$4.46 The claim's exact pairing: Fable 5.1 Low 66.2% @ $2.90 vs Fable 5 High 66.5% @ $8.77. $2.90/$8.77 = 33.1%, i.e. "a third of the cost" is accurate to a rounding. Score gap is -0.3pp, so "at parity" rounds a marginal deficit in the model's favour but is within any reasonable noise band. A WebSearch independently surfaced the Fable 5 High 66.5%/$8.77 and Fable 5 Low 62.1%/$4.46 rows; the Fable 5.1 Low row I could only confirm from Cursor's page itself. Benchmark is run by Cursor (Anysphere); cost/task is COMPUTED by applying each model's published per-MTok pricing (input, cache read, cache write, output) to tokens used - modelled spend, not metered spend. 2. ANTHROPIC'S OWN PAGE DOES NOT MAKE THIS CLAIM. https://www.anthropic.com/claude-fable-and-mythos-5-1 carries a "Agentic coding CursorBench 3.2.0 Accuracy vs Cost" chart (axes: Score 60-75%, cost/task $2-$20 log scale, points at low/med/high/xhigh/max) but in text only says "when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost" - no effort-tier pairing, no cost multiple. Its headline framing is "an estimated 25% less than Fable 5 for typical workloads" and "up to approximately 45%" for agentic work. So the specific low-vs-high / third-of-cost formulation is Lance Martin's read off the chart, not a vendor-documented figure - it is checkable, and it checks out. 3. QUALIFIERS THAT STOP IT GENERALISING (why this is not a general cost claim): - https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/ carries an update disputing Anthropic's savings framing, reporting that at MAX effort Fable 5.1 costs roughly 20% MORE per task than Fable 5 on other measures. - https://benchlm.ai/benchmarks/cursorbench tracks CursorBench "display-only because it is a first-party benchmark" - Cursor's own eval, no third-party audit, and it notes Cursor's own Grok contamination caveat. https://llm-stats.com/benchmarks/cursorbench-3.2 lists only 2 models, self-reported, unverified. - The pairing is cross-tier and picked where the ratio is most flattering. Like-for-like at high effort it is 69.4%/$4.80 vs 66.5%/$8.77 (still favourable, but a ~45% cost cut, not a third). On Terminal-Bench-Science 0.1 the analogous figures reported by claudefa.st are Fable 5.1 low ~26% @ ~$11 vs Fable 5 max 24.7% @ ~$44 - a different benchmark, different comparator tier, ~quarter the cost. Different pairings give different multiples. - Much of the $/task delta is driven by the 75% cache-read price cut ($1.00 to $0.25/MTok), so the ratio is workload-dependent: a low-cache-hit workload will not see a third. 4. CURRENCY: fine. Fable 5.1 launched 2026-09-01; checked 2026-09-02. VERDICT: as narrowly stated (one named benchmark, two named effort tiers), the claim is accurate against the benchmark publisher's public table and should not be treated as marketing-only. It should NOT be repeated as a general "Fable 5.1 low is a third the cost of Fable 5" claim, and "parity" is 0.3pp shy of parity.
Counter sourcehttps://benchlm.ai/benchmarks/cursorbench (CursorBench is first-party to Cursor, tracked display-only, no third-party corroboration) and https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/ (update disputing Anthropic's savings claim: ~20% MORE per task at max effort). Also https://www.anthropic.com/claude-fable-and-mythos-5-1, which notably does NOT state the low-vs-high or third-of-cost comparison in text.
Verifier 3refutedhigh confidence
REFUTED: the directional idea has vendor support, but the three specific load-bearing details in the claim (CursorBench + "high effort" comparator + "a third of the cost") are NOT in any public source. 1) "A third of the cost" appears nowhere public. Checked https://www.anthropic.com/claude-fable-and-mythos-5-1 with targeted prompts three times. Anthropic's ONLY cost figures are "Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads" and "For highly agentic work, the savings will often be much larger-up to approximately 45%." "A third of the cost" implies ~67% savings, which EXCEEDS Anthropic's own published ceiling. The page contains no "a third", "1/3" or 3x-reduction wording. No independent source states it either (searched; llm-stats, MarkTechPost, the-decoder, VentureBeat, claudefa.st all restate Anthropic). 2) No cost ratio is published for CursorBench at all. Anthropic's CursorBench chart caption is "CursorBench 3.2.0 by cost (log scale), at each effort level" but reports only accuracy: Fable 5.1 = 73.4% AT MAX EFFORT, Fable 5 = 70.5% with effort level UNLABELLED. There is no stated cost ratio and no statement of low-effort-vs-high-effort parity on this benchmark. 3) The claim appears to import a ratio from a DIFFERENT benchmark and misstate it. The only low-effort-vs-other-effort cost comparison Anthropic publishes is on Terminal-Bench-Science, not CursorBench: Fable 5.1 at low ~$10-11/task vs Fable 5 at MAX (not "high") ~$44-50/task. That is roughly a quarter to a fifth of the cost, not a third, against a different comparator effort, on a different benchmark. Anthropic's actual effort wording is far vaguer: "when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost" - no benchmark, no comparator effort, no ratio. The claim sharpens this into a precise falsifiable statement its source does not make. Note Anthropic uses low/medium/high/xhigh/max, so "high" is a real but distinct level from the "max" its comparison actually uses. 4) Independent evidence cuts against the cost narrative. Per https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/, Artificial Analysis (a pre-release tester) found Fable 5.1 at max effort costs ~20% MORE per task than Fable 5 ($3.76 vs $2.34 per Intelligence Index task) because it emits ~1.7x the output tokens. Cost direction reverses with effort level, so a single low-effort pairing presented as a general cost result is cherry-picked. 5) Source quality insufficient for claim strength. Origin is an Anthropic employee X post = marketing-grade. The only corroboration is Anthropic's own launch page, which lacks the specific figures. https://llm-stats.com/benchmarks/cursorbench-3.2 lists Fable 5.1 at 0.734 with NO effort-level breakdown and no Fable 5 entry, so no independent effort-resolved data exists to check against. CursorBench is Cursor's own internal preference benchmark, not a neutral third-party eval. Claim is current (release 2026-09-01), so not outdated - it fails on unsupported specifics and marketing-only sourcing.
Counter sourcehttps://www.anthropic.com/claude-fable-and-mythos-5-1 (primary vendor source: CursorBench 73.4% is at MAX effort, not low; no cost ratio given for CursorBench; cost savings capped at 25% typical / ~45% agentic, never "a third of the cost"; the low-vs-max cost comparison is on Terminal-Bench-Science at ~$10-11 vs ~$44-50, i.e. ~1/4, not 1/3); https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/ (Artificial Analysis independent testing: Fable 5.1 at max effort costs ~20% MORE per task than Fable 5, $3.76 vs $2.34, ~1.7x output tokens); https://llm-stats.com/benchmarks/cursorbench-3.2 (no effort-level breakdown, no Fable 5 entry, no independent corroboration available)
S3confirmedvote 3-0Claude Fable 5.1 prompt-cache reads cost $0.25 per million tokens, down from $1.00 per million tokens, i.e. 4x cheaper.
Verdict noteAnthropic's pricing table lists Fable 5.1 cache hits at $0.25/MTok (0.025x base) against Fable 5's $1/MTok; only cache reads changed, base input/output and cache writes are unchanged.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF8

The three verifiers
Verifier 1not refutedhigh confidence
CONFIRMED by Anthropic's own primary pricing documentation, not just the vendor post. https://platform.claude.com/docs/en/about-claude/pricing model-pricing table lists "Claude Fable 5.1 | $10 / MTok | $12.50 / MTok | $20 / MTok | $0.25 / MTok | $50 / MTok" and, two rows down, "Claude Fable 5 | $10 / MTok | ... | $1 / MTok | $50 / MTok". Both halves of the claim ($0.25 now, $1.00 before) are on the same table, and $1.00 -> $0.25 is exactly 4x / 75%. A footnote states it explicitly: "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier." The same page's Prompt caching section repeats it: "On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens)", and https://platform.claude.com/docs/en/build-with-claude/prompt-caching carries the same multiplier table. Refutation attempts, all failed: (1) Overreach of the quote - no. The claim is narrowly scoped to cache reads only and is arithmetically exact; it does not assert the model is 4x cheaper overall. (2) Source-quality objection - the seed's "vendor employee post is marketing-grade" concern does not bite here, because a list price is a published contractual rate verifiable at first-party source, not a benchmark result whose transfer to real workloads is in question. (3) Independent corroboration exists anyway: VentureBeat (https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads) reports "$0.25 on input, down from $1.00 for Fable 5"; The New Stack, LiteLLM day-0 notes and OpenRouter's listing agree. (4) Outdated - no; docs fetched 2026-09-02, Fable 5.1 shipped 2026-09-01. (5) Promotional/temporary - no expiry clause anywhere, unlike the Sonnet 5 introductory-pricing note that sits on the same page, which shows Anthropic does flag time-limited rates when they exist. (6) Cherry-picking - not applicable to a single list price. Qualifiers that do NOT refute but should travel with the claim: only cache READS changed - base input/output ($10/$50) and cache writes ($12.50 5m / $20 1h) are identical to Fable 5, so headline workload savings are Anthropic's own ~25% typical and up to ~45% highly-agentic (https://x.com/claudeai/status/2094848588190830982), not 4x. inference_geo:"us" applies a 1.1x multiplier on every token category, making the effective rate $0.275/MTok. Partner clouds (Bedrock, Vertex) set pricing independently per the same docs page, so $0.25 is confirmed for the first-party API and the CCU-billed Anthropic-operated platforms; I did not independently verify partner-cloud rates. Cross-model, the cheaper multiplier does not make Fable cache reads the cheapest in absolute terms - Sonnet 5 cache reads are $0.20/MTok and Opus 5 $0.50/MTok.
Counter sourceNo credible contradicting source found. The nearest thing to a counter is VentureBeat's framing that Fable 5.1 "remains expensive relative to much of the broader market" versus Meta/DeepSeek rates, and its note that the cut likely answers enterprise reluctance to adopt Fable 5 on cost - but that contests the significance of the discount, not the $1.00 -> $0.25 figure, which VentureBeat itself reports verbatim.
Verifier 2not refutedhigh confidence
CONFIRMED by Anthropic's own public documentation, not merely the employee post. https://platform.claude.com/docs/en/about-claude/pricing model-pricing table lists "Claude Fable 5.1 | $10 / MTok base input | $12.50 / MTok 5m cache writes | $20 / MTok 1h cache writes | $0.25 / MTok cache hits and refreshes | $50 / MTok output" and, two rows below, "Claude Fable 5 | $10 / MTok | $12.50 | $20 | $1 / MTok | $50 / MTok". The same page carries an explicit footnote: "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier." $1.00 -> $0.25 is exactly 4x, so all three components of the claim (new price, old price, 4x ratio) check out. Corroborated a second time at https://platform.claude.com/docs/en/build-with-claude/prompt-caching, whose multiplier table reads "Cache read (hit) | 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1)" and states "On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens)." Adversarial checks all failed to dent it: (1) not an overreach of the quote, the numbers match one-for-one; (2) WebSearch for contradicting evidence turned up none, and independent trade press corroborates rather than disputes (VentureBeat, https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads, plus TheNextWeb, both headlining the 75% cache-read cut); (3) source quality is no longer the issue since first-party documentation now carries it, so the marketing-grade concern about the employee post is moot; (4) not outdated, the docs page is live as of 2026-09-02 and still shows Fable 5.1 as current, with Fable 5 listed separately at the old $1; (5) not cherry-picked, this is a posted list price, not a benchmark result at one effort level. Two qualifications that narrow scope without refuting: cache WRITES are unchanged at $12.50/MTok (5m) and $20/MTok (1h), so only the read leg got 4x cheaper, which is exactly what the claim says; and the data-residency 1.1x multiplier for inference_geo "us" applies to cache reads too, making the absolute figure $0.275/MTok in that configuration, though the 4x ratio still holds because the multiplier applies to both sides. Partner-operated platforms (Bedrock, Vertex) price independently and were not verified. Separately, the docs warn that Claude 4.7+ models use a newer tokenizer producing roughly 30% more tokens for the same text, which distorts cost-per-unit-of-text comparisons across generations but not the Fable 5 vs 5.1 comparison, since both use it. Caveat on downstream use: this is an API list price and says nothing about Claude Max plan limits, where usage is metered against rolling 5-hour and weekly caps rather than billed per token, so the 4x cache saving does not transfer to Max plan consumption without separate evidence.
Verifier 3not refutedhigh confidence
CONFIRMED by Anthropic's own public documentation (not just the employee post). Three first-party pages checked, all agreeing exactly: 1. https://platform.claude.com/docs/en/about-claude/pricing - the model pricing table lists "Claude Fable 5.1 | $10 / MTok | $12.50 / MTok | $20 / MTok | $0.25 / MTok | $50 / MTok" and, two rows below, "Claude Fable 5 | ... | $1 / MTok | ...". The same page's prompt-caching section states: "Cache read (hit) | 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1)" and "On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price ($0.25 USD per million tokens)." A footnote reads: "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier." 2. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - Pricing section: "Cache reads (hits and refreshes) cost 0.025 times the base input price on these models, compared with 0.1 on other Claude models. Long agentic sessions that re-read a cached prefix pay a quarter of the Claude Fable 5 rate. Cache writes and the 512-token minimum cacheable prompt length are unchanged." 3. https://platform.claude.com/docs/en/models/fable-5-1/overview - Pricing table: "Cache read | $0.25 / MTok"; overview text: "at the same input and output prices, with cache reads at a quarter of the cost." 4. https://www.anthropic.com/claude-fable-and-mythos-5-1 (announcement) - "Cache reads now cost 75% less, or $0.25 per million tokens." Arithmetic checks out both ways: $1.00 -> $0.25 is exactly 4x cheaper / a 75% reduction, and $0.25 = 0.025 x $10 base input vs Fable 5's $1.00 = 0.1 x $10. Fable 5 and Fable 5.1 share the same tokenizer (both post-Opus-4.7), so the comparison is apples-to-apples with no hidden token-count inflation. ADVERSARIAL CHECKS, all failed to refute: - Quote vs claim: exact match, no overreach. The claim is scoped to cache reads only, which is what the docs say. - Contradicting evidence: none found. Independent trade coverage (VentureBeat, TheNextWeb, MarkTechPost, 9to5Mac, llm-stats.com) all report the same $1.00 -> $0.25 / 75% figure; no source disputes it. - Source quality: this is a published billing schedule, the one claim type where the vendor IS the authoritative source - unlike a benchmark score. It is not marketing-grade, and it is corroborated by the docs' own Fable 5 row showing the $1/MTok baseline. - Currency: Fable 5.1 released 2026-09-01; docs fetched 2026-09-02. Current. - Cherry-picking: not applicable - a flat published rate, not one benchmark or one effort level. SCOPE CAVEATS (qualify downstream inferences, do not refute the claim as stated): - Only the READ leg fell. Cache writes are unchanged at $12.50/MTok (5m) and $20/MTok (1h), so total prompt-caching cost is not 4x cheaper; the write is still 1.25x/2x base input. - The rate is documented for the first-party Claude API. Bedrock and Google Cloud publish independent pricing; regional/multi-region endpoints carry a 10% premium and inference_geo:"us" a 1.1x multiplier that stacks on cache reads. - Anthropic's own headline savings figures ("~25% typical, up to ~45% agentic") are the marketing-grade part, measured by Anthropic on its own August 2026 usage at default effort - treat those as unverified; the $0.25 rate itself is not in doubt. - This is API token billing. It says nothing about Claude Max plan limits, which are not token-invoiced.
Counter sourceNone found. No credible source disputes or qualifies the $1.00 -> $0.25 cache-read figure; searches for contradicting evidence returned only sources corroborating it.
S4refutedvote 0-3Anthropic's published guidance for Claude Fable 5.1 is to simplify prompts by removing verification rituals, emphasis boosters, scratchpad scaffolds, stale few-shot examples and contradictory rules.
Verdict noteThe remove-verification advice is published for Opus 5, not Fable 5.1; the Fable 5.1 guide says existing Fable 5 prompts work without changes and recommends adding a verification nudge at low effort, and three of the five listed items have no published support for any model.
Run statusrefuted, vote 0-3, 0 verifier errors

see alsoF7did not survive

The three verifiers
Verifier 1refutedhigh confidence
REFUTED by public Anthropic docs on two grounds: misattribution to the wrong model, and direct contradiction on the headline item. 1. The dedicated page https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 has 17 sections (effort levels, progress updates, tool-call batching, append-only history, writing density, formatting, quoting sources, finishing the task, compaction summaries, scope/tests, search triggering at low effort, safeguard false positives, targeted edits, long outputs, subagents, vision). NONE is about simplifying prompts. A full-text grep of the page for "simplif", "verification ritual", "emphasis", "ALL CAPS", "scratchpad", "few-shot", "contradict", "conflicting", "self-check", "double-check", "booster", "stale" returns no matching guidance. Its opening line says the opposite of the claim: "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes." Same grep over https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 and https://platform.claude.com/docs/en/models/fable-5-1/migration-guide: no such guidance. 2. The "remove verification instructions" advice IS published, but Anthropic attributes it explicitly to Claude OPUS 5, not Fable 5.1. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 ("Task scope and over-verification"): "Claude Opus 5 verifies its own work without being told to. If your prompt contains explicit verification instructions ... remove them." https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices repeats it as an Opus-5-only carve-out: "Claude Opus 5 is the exception ... When migrating to Claude Opus 5, remove these instructions rather than rewriting them." 3. For the Fable line Anthropic advises the OPPOSITE on verification. Fable 5.1 prompting page, "Search triggering at low effort" (the exact low-effort case in the research question): "In other cases, a prompt nudge toward verification helps" - and it supplies a verification instruction to ADD. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5, "Recommended scaffolding changes": "Make self-verification explicit in long-run prompts ... instruct: 'Establish a method for checking your own work at an interval of [X] as you build.'" 4. Three of the five listed items have no public support at all for any model: scratchpad scaffolds (the only "scratchpad" mention in the best-practices page is POSITIVE - "Using temporary files can improve outcomes particularly for agentic coding use cases"), stale few-shot examples (best practices calls examples "one of the most reliable ways to steer Claude's output format, tone, and structure"), and contradictory rules (no such guidance found). "Emphasis boosters" is real but is general, older guidance dated to Sonnet 4.5, not Fable 5.1 specific. 5. Partial support exists only for a generic "refactor stale prompts" theme, and it is on the Fable 5 page, not 5.1: "Skills developed for prior models are often too prescriptive for Claude Fable 5 ... Review and consider removing older instructions if default performance is better." The Fable 5.1 page does tell you to remove two specific legacy lines (anti-narration "hold all findings for the final response", and anti-formatting rules). That is narrower than, and does not include, the claim's five-item list. Source quality: the claim's only source is a vendor employee's X post (marketing-grade, uncorroborated). The strongest public evidence available contradicts its model attribution rather than corroborating it. The list reads as an Opus 5 guidance summary relabelled as Fable 5.1.
Counter sourcehttps://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 (no simplification guidance; "should perform well ... without changes"; "a prompt nudge toward verification helps" at low effort) and https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 (the remove-verification advice, attributed to Opus 5)
Verifier 2refutedhigh confidence
REFUTED as worded. The claim attributes a specific five-item removal list to "Anthropic's published guidance for Claude Fable 5.1." I read the three governing primary docs in full. That list is not in any of them, and two of its five items are contradicted for the Fable line. (1) https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 - the model-specific page. Grepped the full 52.7KB text for "verification ritual", "emphasis", "ALL CAPS", "scratchpad", "few-shot", "contradict", "booster": zero relevant hits. Its only "remove" advice is two narrow, unrelated items: remove narration-suppression lines ("system prompt lines such as 'hold all findings for the final response.' Remove lines like that before adding anything") and remove anti-formatting rules ("If your prompt contains anti-formatting language, remove it"). Neither is on the claimed list. The page's opening line runs the other way: "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes." Its "Search triggering at low effort" section tells you to ADD a verification nudge, not remove one. (2) https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices - the cross-model reference - pins the remove-verification guidance to a different model and explicitly marks it an exception: "Ask Claude to self-check. Append something like 'Before you finish, verify your answer against [test criteria].' This catches errors reliably... Claude Opus 5 is the exception: it verifies its own work well without explicit instruction... When migrating to Claude Opus 5, remove these instructions rather than rewriting them." Fable 5.1 is not the exception, so the general self-check rule still applies to it. The same page endorses the other two items the claim says to strip: "Examples are one of the most reliable ways to steer Claude's output format, tone, and structure" / "Multishot examples work with thinking"; and on scratchpads, "Claude to use files... as a 'temporary scratchpad'... can improve outcomes particularly for agentic coding use cases." (3) https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 - the Fable-line page - says the opposite of "remove verification rituals": "Make self-verification explicit in long-run prompts. Separate, fresh-context verifier subagents tend to outperform self-critique." It also tells you to BUILD a scratchpad: "Construct a memory system... Provide a place to write notes, as simple as a Markdown file." (4) https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - no prompt-simplification guidance at all; its migration checklist is API-mechanical (tool_choice, append-only history, effort re-tune, tool batching). What IS publicly supported is only the general direction, and only for Fable 5, not 5.1: "Refactor existing prompts and skills. Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions if default performance is better." Plus one specific removal not on the claimed list: strip show-your-thinking instructions, because they trigger the reasoning_extraction refusal category. Source-quality and cherry-picking failures: the seed's only backing is a vendor employee's X post; searches surfaced no Anthropic-published text containing "verification rituals", "emphasis boosters", or the five-item taxonomy. Corroboration comes only from SEO content-farm restatements of the same launch material (mindstudio.ai, knightli.com, kenhuangus.substack.com, aidailycheck.com) - downstream of the vendor, not independent. Search summarisers themselves conflate the models, writing "This principle applies to Fable models as well" as an inference the docs do not make. The post appears to compress Anthropic's Opus 5 guidance and Fable 5's "too prescriptive" line into one punchy list, then label it Fable 5.1 - a paraphrase that inverts published guidance on verification and scratchpads for the very model named.
Counter sourcehttps://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices (verification removal scoped to Opus 5 as "the exception"; few-shot and scratchpad endorsed) and https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 ("Make self-verification explicit in long-run prompts"; "Construct a memory system")
Verifier 3refutedhigh confidence
REFUTED: the five-item list is mis-attributed to Fable 5.1. It matches Anthropic's OPUS 5 guidance, and for the Fable family the docs say the opposite on verification. 1) Verification - CONTRADICTED for Fable 5.1. The "remove verification instructions" advice is published only for Opus 5. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 - "Task scope and over-verification": "Claude Opus 5 verifies its own work without being told to. If your prompt contains explicit verification instructions ... remove them"; "Self-correction": "Avoid instructing re-checks it already performs ('double-check your answer,' 're-verify before responding')". https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices pins this as a single-model exception: "Ask Claude to self-check. Append something like 'Before you finish, verify your answer against [test criteria].' This catches errors reliably... Claude Opus 5 is the exception... When migrating to Claude Opus 5, remove these instructions rather than rewriting them." For Fable, https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 ("Recommended scaffolding changes") says the reverse: "Make self-verification explicit in long-run prompts. Separate, fresh-context verifier subagents tend to outperform self-critique... instruct: 'Establish a method for checking your own work at an interval of [X]...'" And the Fable 5.1 page itself ADDS verification at low effort: "Search triggering at low effort... a prompt nudge toward verification helps" - directly relevant since the research question is about Fable 5.1 at LOW effort. 2) Scratchpads - CONTRADICTED. Best-practices page: temporary files as a "'temporary scratchpad'... can improve outcomes particularly for agentic coding use cases." 3) Few-shot - CONTRADICTED in direction. Same page: "Examples are one of the most reliable ways to steer Claude's output... A few well-crafted examples (known as few-shot or multishot prompting) improve accuracy and consistency," plus "Multishot examples work with thinking." 4) Emphasis boosters and contradictory rules - NO public support found. I read the full Fable 5.1 prompting page (https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1); its 16 sections are effort levels, progress updates, tool-call batching, append-only history, writing density, chat formatting, quoting sources, finishing the task, compaction summaries, scope/tests, search triggering, safeguard false positives, targeted edits, long outputs, subagents, vision. No "simplify your prompts" section, and no occurrence of verification rituals, emphasis boosters, scratchpad, few-shot, or contradictory/conflicting rules. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 contains no prompt-simplification guidance either. Partial kernel that DOES exist (why the post is not pure invention): Fable 5 page - "Refactor existing prompts and skills. Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions"; Fable 5.1 page - "audit your prompt for instructions that suppress narration... Remove lines like that" and "If your prompt contains anti-formatting language, remove it." So a general subtraction theme is real, but it is scoped to narration-suppression, anti-formatting rules and over-prescriptive legacy skills - not to the claim's five named categories. Source quality: an Anthropic employee X post is marketing-grade per the checklist. The only corroboration for the five-item framing is secondary commentary (mindstudio.ai, kenhuangus.substack.com, findskill.ai, aidailycheck.com) paraphrasing the docs, not the docs themselves. Cherry-picking flag: the post generalises Opus-5-specific guidance across the whole Claude 5 line, where Anthropic explicitly warns "Where a technique names a specific model, treat it as measured on that model and re-check it against your own evals before applying it to another."
Counter sourcehttps://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 (recommends making self-verification explicit for Fable) and https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 (the actual home of the remove-verification guidance, Opus 5 only)
S5partly confirmedvote 3-0On Claude Fable 5.1, effort can be changed mid-conversation without invalidating the prompt cache, whereas previously a mid-conversation effort change invalidated the cache.
Verdict noteTrue on the API only via the beta per-message role:system mechanism (header mid-conversation-output-config-2026-07-01), also available on Opus 5 so not a Fable 5.1 novelty, and Claude Code's /effort still invalidates the whole cache with no Fable 5.1 exception.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF8

The three verifiers
Verifier 1not refutedmedium confidence
SUPPORTED by first-party public documentation, but materially narrower than the wording implies. DIRECT CONFIRMATION (both halves): 1. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - section "Change effort mid-conversation (beta)" states verbatim: "On Claude Fable 5.1 you can change the effort level mid-conversation without invalidating the prompt cache. Raise it for a hard step and lower it for routine ones. Per-message effort is in beta: include the `mid-conversation-output-config-2026-07-01` beta header." 2. https://platform.claude.com/docs/en/build-with-claude/effort - "You can run later turns of a conversation at a different effort level in two ways. On Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5, use a per-message effort change, which keeps the prompt cache. On other models, set a new top-level value on the next request, which starts the cache over." 3. The "previously invalidated" half is confirmed on https://platform.claude.com/docs/en/build-with-claude/prompt-caching, which lists effort in its invalidation table: "Changing the `output_config.effort` value always invalidates message blocks, with the same model-specific effect on tool and system caches as thinking parameters. Setting effort explicitly to the model's default is equivalent to omitting it and does not invalidate." FOUR QUALIFICATIONS THAT THE CLAIM OMITS, each from the same primary sources: (a) NOT FABLE-5.1-EXCLUSIVE - and this is the one that matters most for the research question. The effort doc and the whats-new page both name "Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5" as supporting per-message effort. So it is NOT a Fable-5.1-over-Opus-5 differentiator. Anyone reading Martin's post as a reason to choose Fable 5.1 over Opus 5 would be misreading it. The word "previously" is true only in the temporal sense (before the 2026-07-01 beta shipped), not the generational sense. (b) ONLY VIA A SPECIFIC BETA MECHANISM, NOT BY CHANGING EFFORT NORMALLY. Cache preservation requires the `mid-conversation-output-config-2026-07-01` beta header AND a new message form: a `role: "system"` message with empty `content` carrying `output_config.effort`. Changing effort the ordinary way STILL breaks the cache on Fable 5.1 - the effort doc says so explicitly twice: "On Claude Fable 5.1, prefer this form over changing the top-level value between requests. A top-level change restarts the cache and also steers the model less reliably," and in Best Practices: "Hold top-level effort constant within cached conversations: Changing the top-level effort value between requests invalidates prompt caching." Models lacking per-message effort, "including Claude Fable 5, return a 400 error: `output_config.effort requires a model that supports per-turn effort; this model does not`." (c) CLAUDE API ONLY. The whats-new page scopes it: "Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5 support it on the Claude API." No Bedrock / Vertex / Foundry support is stated. (d) FALSE ON THE CLAUDE CODE SURFACE - the strongest counter-evidence I found. https://code.claude.com/docs/en/prompt-caching lists "Changing effort level" under "Actions that invalidate the cache" with NO Fable-5.1 or Opus-5 exception: "**Effort level**: each effort level has its own cache for the same model. Changing effort mid-session recomputes the entire request," and "The cache is keyed by effort level as well as model, so switching with `/effort` means the next request reads the entire conversation history with no cache hits." So for a Claude Code / Claude Max user, changing effort mid-session still costs a full cache rebuild on Fable 5.1 today. The claim does not transfer to that surface. INDEPENDENT CORROBORATION: weak. Secondary write-ups (claudefa.st/blog/models/claude-fable-5-1, digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes, wmedia.es/en/tips/claude-code-fable-5-1-cache-reads) restate the doc language rather than test it; I found no independent empirical measurement of cache hit rates across a per-message effort change on Fable 5.1. Two adjacent Claude Code issues (anthropics/claude-code #61984, #63962) show doc-vs-empirical disagreement about /effort and cache misses on Opus 4.7, i.e. the reliability of effort/cache statements has been contested before, though neither tests Fable 5.1. VERDICT: not refuted. Unlike a benchmark or performance claim, this is a statement about the vendor's own API contract, where the vendor's platform docs are the authoritative primary source, and they state the claim almost word for word. Marked medium rather than high confidence because the claim as worded generalises a beta-gated, API-only, non-exclusive mechanism into an apparent Fable 5.1 capability upgrade, and is contradicted on the Claude Code surface the user actually works on.
Counter sourcehttps://code.claude.com/docs/en/prompt-caching - Anthropic's own Claude Code docs list "Changing effort level" under "Actions that invalidate the cache" with no model exception: "each effort level has its own cache for the same model. Changing effort mid-session recomputes the entire request." Secondary counter-context: https://platform.claude.com/docs/en/build-with-claude/effort ("A top-level change restarts the cache"; per-message effort also supported on Claude Opus 5 and Mythos 5.1, so it is not a Fable 5.1 differentiator), and https://github.com/anthropics/claude-code/issues/61984 / #63962, where documented effort-vs-cache behaviour was empirically disputed on Opus 4.7.
Verifier 2not refutedmedium confidence
SUPPORTED by Anthropic's public technical docs (not a blog or press release), in three independent pages, near-verbatim. 1. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - section "Change effort mid-conversation (beta)" states, almost word for word with the claim: "On Claude Fable 5.1 you can change the effort level mid-conversation without invalidating the prompt cache." It names the required beta header mid-conversation-output-config-2026-07-01 and ships cURL/Python/TS/Go/Java/PHP/Ruby/C#/CLI samples. It is also listed as one of the five additive changes from Fable 5. 2. https://platform.claude.com/docs/en/build-with-claude/effort - "Change effort mid-conversation" section: on Fable 5.1, Mythos 5.1 and Opus 5 a per-message effort change keeps the prompt cache; on other models a new top-level value on the next request starts the cache over. Mechanism: a role:"system" message with empty content carrying output_config.effort, placed anywhere in messages; the new level takes effect from the next user turn and everything before it is unchanged, so the cached prefix still matches. Confirms the "previously" half of the claim: top-level effort shapes the rendered prompt, so changing it between requests does not preserve cached prefixes. Best-practice #5 says the same. 3. https://platform.claude.com/docs/en/build-with-claude/prompt-caching - cache-invalidation table, "Effort setting" row: changing output_config.effort always invalidates message blocks, EXCEPT that on models supporting per-message effort, an effort change carried in a role:"system" message inside messages leaves the cached prefix intact. THREE QUALIFICATIONS the claim as worded omits (caveats, not refutations): (a) Beta-gated. Without the mid-conversation-output-config-2026-07-01 header it does not apply; Fable 5 itself returns a 400 ("requires a model that supports per-turn effort"). (b) Mechanism-specific. Only the per-message role:"system" form preserves the cache. Changing top-level output_config.effort on Fable 5.1 STILL restarts the cache - the effort doc explicitly tells you to prefer the per-message form over a top-level change. Anyone reading the claim as "any mid-conversation effort change is now cache-safe on Fable 5.1" would be wrong. (c) Not a Fable 5.1 exclusive. The same effort doc lists Claude Opus 5 and Mythos 5.1 as supporting it, and Opus 5's own section repeats the cache-preserving line. Framing it as a Fable 5.1 differentiator over-attributes it. SEARCHED FOR CONTRADICTION, found none aimed at this claim. https://github.com/anthropics/claude-code/issues/61984 measured full/near-full cache misses on /effort switches (Opus 4.7, Claude Code v2.1.150, 6 transitions, cold targets read=0) - that is a model with no per-message effort, so it corroborates the "previously invalidated" half; it also shows the caching docs once carried the opposite line ("no effect on the cache") that measurement contradicted, i.e. Anthropic's docs on this exact topic have been wrong before. https://github.com/anthropics/claude-code/issues/63962 (Sonnet 4.6, v2.1.143+) measured a FULL cache hit across a high->xhigh switch, contradicting Claude Code's own warning dialog - messy, but it cuts against absolutism in the "previously" half, not against Fable 5.1. WHY NOT refuted=true: this is a mechanical API-spec claim documented in a first-party API reference with a named beta header and working code, not a performance or quality claim where the marketing-grade discount bites. It is current (Fable 5.1 is live as of 2026-09-02) and not cherry-picked. WHY CONFIDENCE IS ONLY MEDIUM: every supporting source is Anthropic itself; I found no independent empirical test of Fable 5.1 per-message effort cache behaviour (litellm's day-0 post repeats the vendor line rather than measuring it), and issue #61984 establishes that Anthropic's effort/caching documentation has previously misdescribed real behaviour. Treat as documented intent, verified by spec, not by measurement.
Counter sourceNearest counter-evidence, neither of which lands on this claim: github.com/anthropics/claude-code/issues/61984 (Opus 4.7 /effort switches measured full/near-full cache misses against a docs line then claiming no cache effect - proves the effort/caching docs have been wrong before, but concerns a model with no per-message effort and so supports the claim's "previously" half); github.com/anthropics/claude-code/issues/63962 (Sonnet 4.6 measured a full cache hit across an effort change, undercutting "always invalidated" as an absolute). No independent, non-Anthropic measurement of Fable 5.1 per-message effort caching was found; docs.litellm.ai/blog/claude_fable_5_1 restates the vendor claim rather than testing it.
Verifier 3not refutedhigh confidence
SUPPORTED by Anthropic's own public documentation, in near-verbatim wording, on three independent doc pages I fetched. 1. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - section "Change effort mid-conversation (beta)" states verbatim: "On Claude Fable 5.1 you can change the effort level mid-conversation without invalidating the prompt cache. Raise it for a hard step and lower it for routine ones." This is the first half of the claim, word for word. 2. https://platform.claude.com/docs/en/build-with-claude/effort - section "Change effort mid-conversation": "On Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5, use a per-message effort change, which keeps the prompt cache. On other models, set a new top-level value on the next request, which starts the cache over." The sub-section "Top-level effort on the next request" confirms the second half of the claim: "Because top-level effort shapes the rendered prompt, changing it between requests doesn't preserve cached prefixes from earlier turns." Best-practice 5 repeats it: "Changing the top-level effort value between requests invalidates prompt caching." 3. https://platform.claude.com/docs/en/build-with-claude/prompt-caching - the "What invalidates the cache" table: "Changing the output_config.effort value always invalidates message blocks... On models that support per-message effort, an effort change carried in a role: 'system' message inside messages leaves the cached prefix intact." This is a documented API behaviour, not a performance claim, so official docs are the right grade of source; no independent benchmark is needed and none is available. I searched for contradicting evidence and found none for Fable 5.1. FOUR MATERIAL QUALIFICATIONS the bare claim omits, all of which matter if it is repeated as-is: (a) BETA, header required. Per-message effort requires the beta header mid-conversation-output-config-2026-07-01. Without it, no cache-preserving effort change. (b) ONLY via the specific per-message mechanism. The cache survives only when the change is carried in a role: "system" message with empty content and output_config.effort inside messages. A plain top-level output_config.effort change on Fable 5.1 STILL restarts the cache - same effort doc, "Top-level effort on the next request". So "on Fable 5.1 effort can be changed without breaking the cache" is true only of one specific request shape. (c) NOT a Fable 5.1 exclusive, so the "previously" framing is misleading. Claude Opus 5 and Mythos 5.1 support it too (effort doc, quoted above; the what's-new page repeats "Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5 support it on the Claude API"). The beta header is dated 2026-07-01, i.e. it predates Fable 5.1. Claude Fable 5 does not support it and returns a 400: "output_config.effort requires a model that supports per-turn effort; this model does not." So "previously invalidated" is accurate versus Fable 5 and older, but wrong if read as new-with-Fable-5.1. (d) DOES NOT TRANSFER TO CLAUDE CODE, which is the surface Vasso actually uses. https://code.claude.com/docs/en/prompt-caching, section "Changing effort level": "The cache is keyed by effort level as well as model, so switching with /effort means the next request reads the entire conversation history with no cache hits." It lists effort changes under "Actions that invalidate the cache" and states no Fable 5.1 exception. So /effort in Claude Code still costs a full re-read regardless of this API feature. Contradicting evidence checked and dismissed: anthropics/claude-code issues #61984 (docs say no cache effect, empirically full/partial misses, Opus 4.7 on v2.1.150) and #63962 (dialog warns of a miss but cache is preserved). These are mutually contradictory, concern Claude Code + Opus 4.7 rather than the Fable 5.1 API, and neither disputes the documented Fable 5.1 behaviour. Verdict: the claim is true as stated, but should be repeated with (a), (b) and (d) attached - particularly (d), since in Claude Code the practical advice remains "pick effort at the top of a session".
Counter sourcehttps://code.claude.com/docs/en/prompt-caching (Claude Code /effort still fully invalidates the cache, no Fable 5.1 exception); https://platform.claude.com/docs/en/build-with-claude/effort (per-message effort is beta, header-gated, and also on Opus 5 / Mythos 5.1, so not a Fable 5.1 novelty; top-level effort changes still restart the cache); github.com/anthropics/claude-code/issues/61984 and /63962 (contradictory practitioner reports on Claude Code + Opus 4.7 cache behaviour)
S6confirmedvote localThe Claude Code subcommands /claude-api migrate, /claude-api prompt-audit and /claude-api cost-optimize exist.
Verdict noteThe subcommands migrate, prompt-audit and cost-optimize exist. Verified locally against the bundled skill table, so no verifier agents were spent on it and there is no adversarial vote.
Best sourceBundled claude-api skill subcommand table, checked locally on 2026-09-02 before the run started
Run statusconfirmed locally before the run, no agents spent
Group R: the model facts already written into the global rules
R1confirmedvote 3-0Claude Fable 5.1 (model ID claude-fable-5-1) is Anthropic's most capable generally available model and the successor to Claude Fable 5, priced at $10 input and $50 output per million tokens, with a 1M-token context window and 128K max output.
Verdict noteModel ID, $10/$50 pricing, 1M context, 128K max output and successor status all verified on Anthropic's model page and cross-checked on llm-stats.com; 'most capable' is vendor positioning, and the 1M window has no long-context surcharge.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF1

The three verifiers
Verifier 1not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. I did not judge it against the cached skill table; I tested every element against Anthropic's live public docs and one independent source. PRIMARY (https://platform.claude.com/docs/en/models/fable-5-1/overview, fetched 2026-09-02): "Model ID: `claude-fable-5-1`" · "Context window: 1M tokens · Max output: 128K tokens · Input pricing: $10 / MTok · Output pricing: $50 / MTok". Specifications tables repeat all four independently (Model IDs table: Claude API = `claude-fable-5-1`; Pricing table: Input $10/MTok, Output $50/MTok, Cache read $0.25/MTok; Capabilities table: Context window 1M tokens, Max output 128K tokens). Status: "Active (latest)", Released September 1, 2026 - one day before today, so not outdated. SUCCESSOR-TO-FABLE-5: same page - "Claude Fable 5.1 extends Claude Fable 5 at the same input and output prices, with cache reads at a quarter of the cost". Anthropic's pricing page confirms Fable 5 is still served at $10/$50 (differing only on cache read: $1 vs $0.25). So the quote's "same per-token price" detail also holds. MOST-CAPABLE-GENERALLY-AVAILABLE: the qualifier is doing real work and is correct. The docs' comparison table puts Fable 5.1 at the top of the lineup (Fable 5.1 $10/$50 > Opus 5 $5/$25 > Sonnet 5 $2/$10 > Haiku 4.5 $1/$5). The only same-tier peer is Claude Mythos 5.1, which "offers the same capabilities by invitation only, as part of Project Glasswing" - identical specs and price, restricted access, so it does not outrank Fable 5.1 and is not GA. Announcement (https://www.anthropic.com/claude-fable-and-mythos-5-1) calls them "the world's most advanced models for coding and knowledge work" and confirms Mythos 5.1 is "identical to Fable 5.1, but with different levels of safeguards", restricted to vetted users. CROSS-CHECK, INDEPENDENT OF ANTHROPIC (https://llm-stats.com/models/claude-fable-5-1): input $10.00, output $50.00, cached input $0.250, context 1M, output 128K, model ID claude-fable-5-1, released Sep 1 2026. Third-party coverage corroborates the launch date and cache-read cut (venturebeat.com, the-decoder.com, marktechpost.com, macrumors.com, 9to5mac.com, aws.amazon.com/about-aws/whats-new/2026/09/claude-fable-5-1-aws/). I ACTIVELY LOOKED FOR THE OBVIOUS REFUTATION AND IT IS NOT THERE. The likeliest qualification would be a long-context premium tier making "$10/$50" incomplete, as with earlier 1M-context models. https://platform.claude.com/docs/en/about-claude/pricing explicitly closes it: "Claude 4.6 and later models ... include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.)" No >200K surcharge exists for Fable 5.1. CALIBRATION / RESIDUAL QUALIFICATIONS (not enough to refute): 1. The four hard facts (ID, price, context, max output) are administrative facts where the vendor is the definitive authority - a vendor cannot be "marketing-grade" about the price it charges - and they are corroborated by a non-Anthropic aggregator anyway. 2. "Most capable" is the softest element. Anthropic's own announcement does not use that literal phrase; it says "most advanced ... for coding and knowledge work". The claim's phrasing is a fair paraphrase of the docs' positioning, but it is a vendor superlative, not an independently benchmarked result, and should not be carried into any downstream conclusion about capability transfer. 3. Anthropic's own routing advice cuts against reflexive use of it: the overview page says "For most workloads, start with Claude Opus 5 ... Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work". That is a recommendation, not a capability contradiction. 4. "$10/$50" is BASE pricing. `inference_geo: "us"` applies a 1.1x multiplier; Batch API is 50% off. Modifiers, not contradictions. 5. Scope note: this claim says nothing about effort levels, so the research question's Fable-5.1-at-low-effort comparison is untouched by this verification.
Counter sourceNearest qualifiers found, neither refuting: (a) https://platform.claude.com/docs/en/models/fable-5-1/overview advises "For most workloads, start with Claude Opus 5", so Anthropic does not position Fable 5.1 as the default choice despite it topping the lineup; (b) https://www.anthropic.com/claude-fable-and-mythos-5-1 never uses the literal phrase "most capable" - it says "world's most advanced models for coding and knowledge work" - so that specific superlative is a paraphrase of vendor positioning rather than a quoted vendor claim, and is marketing-grade. No source disputing the model ID, $10/$50 pricing, 1M context, 128K max output, or successor status was found.
Verifier 2not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM on every element; I tried to break it and could not. Primary source (Anthropic docs, fetched 2026-09-02): https://platform.claude.com/docs/en/models/fable-5-1/overview states verbatim "Model ID: `claude-fable-5-1`", "Context window: 1M tokens · Max output: 128K tokens · Input pricing: $10 / MTok · Output pricing: $50 / MTok", pricing table Input $10/MTok, Output $50/MTok, Cache read $0.25/MTok, Capabilities table "Context window 1M tokens / Max output 128K tokens", Status "Active (latest)", Released September 1, 2026. Successor relationship confirmed verbatim: "Claude Fable 5.1 extends Claude Fable 5 at the same input and output prices, with cache reads at a quarter of the cost." Fable 5 is listed as a legacy model still available. Corroborated on a second Anthropic page: https://platform.claude.com/docs/en/models/overview lineup table gives Fable 5.1 = 1M context, 128K max output, $10/$50, API ID `claude-fable-5-1`, sitting above Opus 5 ($5/$25), Sonnet 5 ($2/$10), Haiku 4.5. "Most capable generally available model" - this exact phrasing is Anthropic's own on https://www.anthropic.com/claude/fable ("most capable model for coding and knowledge work", "most capable generally available model"). The "generally available" qualifier is precise rather than an overreach: https://www.anthropic.com/claude-fable-and-mythos-5-1 states "Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs." Because "most capable" is vendor-grade framing, I checked independently: https://artificialanalysis.ai/models/claude-fable-5-1 scores it 66 on the Artificial Analysis Intelligence Index and ranks it #1 of 192 models. That is third-party corroboration of the ranking, not just marketing. Qualifications found but NOT contradictions: (a) the docs page itself says "For most workloads, start with Claude Opus 5 ... Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5 at higher effort still fall short" - a cost/latency routing recommendation, not a capability ranking that undercuts the claim; (b) Artificial Analysis flags it as "particularly expensive", slower (66.2 tok/s) and "very verbose" (140M output tokens in evaluation) - relevant to cost-per-task questions elsewhere in the research question, but nothing in the claim asserts efficiency; (c) head-to-head benchmark margins over Opus 5 are vendor-reported (GDPval-AA v2 1,853 vs 1,824; Terminal-Bench 4.0 55.8% vs 52.3%) and should not be treated as independently verified. Not outdated: released 2026-09-01, one day before review, marked "Active (latest)" with retirement not sooner than 2026-09-01. No credible source found disputing any figure.
Counter sourcehttps://platform.claude.com/docs/en/models/fable-5-1/overview (Anthropic's own docs steer most workloads to Opus 5, not Fable 5.1) and https://artificialanalysis.ai/models/claude-fable-5-1 (independent: confirms #1 rank but flags high cost, slow speed, "very verbose" output) - both qualify how the claim should be used, neither contradicts any stated fact.
Verifier 3not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS the claim. I ignored the seed's cached-table source and tested each element against Anthropic's live primary docs. VERIFIED, element by element: 1) Model ID `claude-fable-5-1` - https://platform.claude.com/docs/en/models/fable-5-1/overview ("Model ID: `claude-fable-5-1`") and the lineup table at https://platform.claude.com/docs/en/about-claude/models/overview (Claude API ID row: `claude-fable-5-1`). 2) $10 input / $50 output per MTok - same model page ("Input pricing: $10 / MTok · Output pricing: $50 / MTok") and the pricing table at https://platform.claude.com/docs/en/about-claude/pricing (Fable 5.1 row: $10 base input, $50 output). Anthropic's announcement https://www.anthropic.com/claude-fable-and-mythos-5-1 states the same. 3) 1M context window - confirmed on both the model page and models overview. Not a beta or tier-gate: the pricing page's Long context section says "Claude 4.6 and later models ... include the full 1M token context window at standard pricing." 4) 128K max output - confirmed on both pages. Footnote: this is the synchronous Messages API limit; the 300K batch beta (output-300k-2026-03-24) lists Opus 5, Sonnet 5, Opus 4.8/4.7/4.6, Sonnet 4.6 and NOT Fable 5.1, so 128K is genuinely its ceiling. 5) Successor to Fable 5 - model page: "Claude Fable 5.1 extends Claude Fable 5 at the same input and output prices." Models overview lists Fable 5 under "Legacy models (still available)". Released 1 Sep 2026, status "Active (latest)". 6) "Most capable generally available" - the "generally available" qualifier is load-bearing and correct. Fable 5.1 ships on Claude API, Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS; Claude Mythos 5.1 "offers the same capabilities by invitation only, as part of Project Glasswing" (same specs and price), so Mythos is not a more-capable GA model. Independent corroboration on capability: https://artificialanalysis.ai/models/claude-fable-5-1 scores it 66 on the Intelligence Index, ranked #1 of 192 (median 36). REFUTATION ATTEMPTS THAT FAILED: (a) Mythos 5.1 as a more capable model - docs say same capabilities, and it is invitation-only, so the GA qualifier holds; (b) 1M context as beta/gated - explicitly standard pricing, no long-context premium; (c) outdated - released 1 Sep 2026, one day before today, and the docs mark it "Latest"; (d) price changed at the .1 release - docs confirm same $10/$50, only cache reads dropped ($1 -> $0.25/MTok). QUALIFICATIONS WORTH CARRYING (do not refute the claim, but temper it): Anthropic's own docs do NOT position Fable 5.1 as the blanket default - "start with Claude Opus 5 for most workloads ... Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5 at higher effort still fall short." Fable 5.1 is listed as "Slower" latency at 2x Opus 5's price, and Artificial Analysis independently flags it as "particularly expensive" and "slower than average and very verbose". So "most capable" is defensible on intelligence rankings but is a specialist top-end, not a recommended default. Also noted: the seed source described its model table as cached 2026-06-24, which predates the 1 Sep 2026 release - the table was evidently newer than its stated cache date, but that provenance wobble does not affect the facts, which all check out against live docs.
R2confirmedvote 3-0As of 2026-09-02, Claude Opus 5 (claude-opus-5) and Claude Sonnet 5 (claude-sonnet-5) are the latest released versions of the Opus and Sonnet tiers; no newer Opus or Sonnet model has been released.
Verdict noteThe model-status table lists claude-opus-5 and claude-sonnet-5 as the newest Active rows of their tiers as of 2026-09-02; Opus 5.1 / Sonnet 5.1 exist only as leaked EAP codenames and rumour, so this expires the day a successor ships.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF1

The three verifiers
Verifier 1not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. Three Anthropic-operated pages fetched live on 2026-09-02 agree, and an adversarial search for successors found none. 1) https://platform.claude.com/docs/en/about-claude/models/overview - the "Compare models" table lists exactly four current models: Claude Fable 5.1 (`claude-fable-5-1`), Claude Opus 5 (`claude-opus-5`), Claude Sonnet 5 (`claude-sonnet-5`), Claude Haiku 4.5. Its "Legacy models" line names Fable 5, Opus 4.8/4.7/4.6/4.5, Sonnet 4.6/4.5 - i.e. every other Opus and Sonnet is OLDER than 5. No Opus or Sonnet above 5 appears anywhere on the page. 2) https://platform.claude.com/docs/en/about-claude/model-deprecations - the "Model status" table is the authoritative catalogue of every model Anthropic serves (Active/Legacy/Deprecated/Retired). The newest Opus row is `claude-opus-5` (Active, retirement "not sooner than July 24, 2027"); the newest Sonnet row is `claude-sonnet-5` (Active, "not sooner than June 30, 2027"). No `claude-opus-5-1`, `claude-opus-6`, `claude-sonnet-5-1` or similar exists in the table. Both exact model IDs in the claim are confirmed verbatim. 3) https://www.anthropic.com/news - Claude Opus 5 announced Jul 24, 2026. The most recent model announcement (Sep 1, 2026) is Claude Fable 5.1 and Claude Mythos 5.1, which are SEPARATE tiers above Opus, not Opus/Sonnet releases; corroborated by 9to5mac.com/2026/09/01/anthropic-upgrades-claude-with-new-fable-5-1-model-details-here/ and by Anthropic's own framing that "Haiku, Sonnet, and Opus are the small, medium, and large models below Fable". So the newest release does not disturb the Opus/Sonnet tier heads. REFUTATION ATTEMPT (failed): searching "Claude Opus 5.1" OR "Claude Sonnet 5.1" returns only third-party speculation (cometapi.com/claude-opus-5-1/, cellcog.ai/blog/claude-opus-5-1-release-date/, orcarouter.ai) discussing unreleased EAP codenames `claude-marshmallow-eap` / `claude-melon-eap` from ~Aug 21-24 screenshots, with an expected release date that slipped past Aug 28 and Aug 31 with no announcement. These sources explicitly confirm nothing newer has shipped - they corroborate rather than contradict. CAVEATS (do not change the verdict): (a) The cited source's own "cached 2026-06-24" label is internally inconsistent, since Opus 5 launched Jul 24 and Fable 5.1 on Sep 1 - both after that date; the cache stamp is unreliable even though its content happens to be right. (b) Source-quality is adequate here because this is a catalogue fact about the vendor's own products, not a performance claim where vendor bias transfers. (c) The claim is inherently time-sensitive: an Opus/Sonnet successor is actively rumoured and could land within weeks, so this holds as of 2026-09-02 only.
Counter sourcecometapi.com/claude-opus-5-1/ and cellcog.ai/blog/claude-opus-5-1-release-date/ - speculate about imminent Opus 5.1 / Sonnet 5.1 based on leaked EAP codenames (claude-marshmallow-eap, claude-melon-eap), but both confirm no official release has occurred; they do not contradict the claim, only flag that it may expire soon.
Verifier 2not refutedhigh confidence
VERDICT: NOT REFUTED. Public, first-party evidence fetched live on 2026-09-02 supports the claim on both tiers. 1) https://platform.claude.com/docs/en/about-claude/models/overview (fetched 2026-09-02). The "Compare models" table lists exactly four current models: Claude Fable 5.1 (`claude-fable-5-1`), Claude Opus 5 (`claude-opus-5`), Claude Sonnet 5 (`claude-sonnet-5`), Claude Haiku 4.5 (`claude-haiku-4-5-20251001`). Everything older is explicitly demoted: "Legacy models (still available): Claude Fable 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Sonnet 4.6, Claude Sonnet 4.5." No Opus or Sonnet model above 5 appears anywhere on the page. Guidance line: "start with Claude Opus 5 for most workloads." 2) https://platform.claude.com/docs/en/about-claude/model-deprecations (fetched 2026-09-02). The Model status table is the strongest test, because it enumerates EVERY model with a lifecycle state. Opus rows: claude-opus-5 (Active, retire not sooner than July 24 2027), then 4-8, 4-7, 4-6, 4-5-20251101, and retired 4-1 / 4. Sonnet rows: claude-sonnet-5 (Active, not sooner than June 30 2027), then 4-6, 4-5-20250929, retired sonnet-4 and 3-7. There is no claude-opus-5-1 and no claude-sonnet-5-1 row. A newer release would have to appear here to receive a retirement commitment, so its absence is meaningful, not just an un-updated marketing page. 3) https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions (fetched 2026-09-02) confirms the ID convention and enumerates the Opus/Sonnet family without any 5.x entry: "For example: claude-sonnet-4-6, claude-sonnet-5, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8, and claude-opus-5". It also confirms dateless IDs are pinned snapshots, not evergreen pointers, so `claude-opus-5` cannot be silently pointing at a newer model. 4) Independent/third-party corroboration of release dates and of the negative: web search surfaced Sonnet 5 released 2026-06-30 and Opus 5 released 2026-07-24 (anthropic.com/news/claude-sonnet-5, anthropic.com/news/claude-opus-5, plus AWS ML blog and Thurrott coverage), and a non-Anthropic tracker (https://cellcog.ai/blog/claude-opus-5-1-release-date/) states Opus 5.1 / Sonnet 5.1 had NOT been announced or released as of 2026-08-30, with only unconfirmed codename sightings ("claude-marshmallow-eap", "claude-melon-eap") behind the rumours. Community codenames are not releases. Scope check (the one place this could have been an overreach): Claude Fable 5.1 and Claude Mythos 5 are newer models overall, but they sit in the Mythos-class tier ABOVE Opus, not in the Opus or Sonnet tiers, so they do not contradict a claim scoped to those two tiers. Mythos 5 is also restricted-access (Project Glasswing), not a general Opus/Sonnet release. Currency: this is not a stale check. All three docs pages were fetched today, 2026-09-02, the same date the claim is scoped to.
Counter sourceNo source contradicts the claim. Two weaknesses worth flagging, neither fatal: (a) the seed's stated provenance is internally inconsistent - a claude-api skill table "cached 2026-06-24" could not have known about Sonnet 5 (released 2026-06-30) or Opus 5 (released 2026-07-24), so either the cache date is wrong or the table was refreshed later; the claim survives only because live public docs independently confirm it, not because the cached table is trustworthy. (b) Unconfirmed rumours of Opus 5.1 / Sonnet 5.1 circulate from leaked EAP identifiers (claude-marshmallow-eap, claude-melon-eap, ~2026-08-21 to 08-24, per cellcog.ai); these are speculation, not releases, but they mean the claim is a point-in-time fact with a plausibly short shelf life - it should be re-verified rather than cached.
Verifier 3not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM (not just the cached seed source). 1) https://platform.claude.com/docs/en/about-claude/models/overview - the "Compare models" table of the CURRENT lineup contains exactly four models: Claude Fable 5.1 (`claude-fable-5-1`), Claude Opus 5 (`claude-opus-5`), Claude Sonnet 5 (`claude-sonnet-5`), Claude Haiku 4.5. The page's own "Legacy models (still available)" line reads: Fable 5, Opus 4.8, Opus 4.7, Opus 4.6, Opus 4.5, Sonnet 4.6, Sonnet 4.5. No Opus 5.x above 5, no Sonnet 5.x above 5, anywhere on the page. 2) https://platform.claude.com/docs/en/about-claude/model-deprecations - the full lifecycle table lists every active API model name: claude-fable-5-1, claude-fable-5, claude-opus-5, claude-opus-4-8, -4-7, -4-6, -4-5-20251101, claude-sonnet-5, claude-sonnet-4-6, claude-sonnet-4-5-20250929, claude-haiku-4-5-20251001. `claude-opus-5` is the highest Opus entry (retirement "not sooner than July 24, 2027") and `claude-sonnet-5` the highest Sonnet entry ("not sooner than June 30, 2027"). Nothing newer in either tier exists in Anthropic's own canonical status table. 3) https://platform.claude.com/docs/en/about-claude/pricing - model pricing table lists Opus 5 as the top Opus row and Sonnet 5 as the top Sonnet row. It does introduce two models absent from the seed quote (Claude Mythos 5.1 and Mythos 5, "limited availability", anthropic.com/glasswing), but Mythos is a separate line, not an Opus or Sonnet release, so it does not contradict the claim as worded. 4) Contradiction search: WebSearch for "Anthropic 'Opus 5.1' OR 'Sonnet 5.1' release announcement" returns only speculative third-party SEO blogs (cometapi.com, cellcog.ai, kie.ai, orcarouter.ai) built on unverified "claude-marshmallow-eap / claude-melon-eap" test-string screenshots circulated 21-24 Aug 2026. Those sources themselves state Anthropic has made no official Opus 5.1 or Sonnet 5.1 announcement. No credible source disputes the claim. CAVEATS (do not reach refutation but qualify it): - The seed's own cited source is inadequate for the claim's date. The bundled skill table is cached 2026-06-24, ~10 weeks stale, and is demonstrably drifted in other respects (the public anthropics/skills models.md still shows Fable 5, not Fable 5.1, as the top model). The claim is true because the live docs say so today, not because the cached table said so in June. - Short shelf life: the same speculative reporting places an Opus-tier point release around 5 Sep 2026 (three days from now). The claim is a snapshot fact, not a durable one, and should be re-checked before being written into any rule. - Scope precision: "latest of the Opus and Sonnet tiers" is correct; "latest Anthropic models overall" would be false (Fable 5.1 and Mythos 5.1 sit above both).
Counter sourcehttps://www.cometapi.com/claude-opus-5-1/ and https://cellcog.ai/blog/claude-opus-5-1-release-date/ - speculative third-party posts about an imminent Opus 5.1/Sonnet 5.1 based on unverified "claude-marshmallow-eap"/"claude-melon-eap" test strings; both concede no official Anthropic announcement exists, so they qualify the claim's shelf life rather than refute it.
R3confirmedvote 3-0On Claude Fable 5.1, thinking is always on: sending thinking type 'disabled' or a budget_tokens value returns a 400 error, and reasoning depth is controlled only through output_config.effort (low through max).
Verdict noteBoth thinking disabled and budget_tokens return 400 at any effort level, effort defaults to high with all five levels accepted; thinking type adaptive is still accepted, so 'any thinking parameter' is not rejected, only those two.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF1

The three verifiers
Verifier 1not refutedhigh confidence
Public Anthropic platform docs support the claim directly and independent third parties corroborate it, so it is NOT refuted. 1) Per-model rejection table, https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting - row "Claude Fable 5.1 | Adaptive only | Always on | Rejected with 400: `"enabled"`, `"disabled"`". Same page: "Models marked `Always on` cannot turn thinking off", and it gives the exact server error text for `"thinking.type.disabled" is not supported for this model`, naming Fable 5.1 among the models that reject it. Since `budget_tokens` is only valid inside `thinking: {type: "enabled"}`, and `"enabled"` is 400-rejected, the budget_tokens half of the claim follows and is stated explicitly elsewhere (below). 2) https://platform.claude.com/docs/en/models/fable-5-1/migration-guide - "Adaptive thinking is always on, unchanged from Claude Fable 5... No `thinking` configuration is required. Both `thinking: {type: "disabled"}` and manual extended thinking (`thinking: {type: "enabled", budget_tokens: N}`) return a 400 error." And: "on `claude-fable-5-1` ... `thinking: {type: "disabled"}` returns a 400 error at any effort level." 3) Effort ladder confirmed on the same migration guide: "The effort parameter default is `high`, and all five levels are supported" and "Only the named levels are accepted (`low`, `medium`, `high`, `xhigh`, `max`)". https://platform.claude.com/docs/en/build-with-claude/thinking adds "steer thinking depth with `effort` instead of `budget_tokens`." 4) Independent corroboration (not vendor marketing): LiteLLM's Fable 5 day-0 post and OpenRouter/olostep developer guides both state explicit thinking budgets are rejected with 400 and that depth is steered by effort only. Adversarial checks: (a) not an overreach of the seed quote - public docs say more, not less; (b) no contradicting source found in two searches; (c) source quality is appropriate - this is an API contract claim about Anthropic's own endpoint, where vendor reference docs are authoritative, unlike a benchmark/performance claim; (d) current - the migration guide is live as of 2026-09-02, days after the Fable 5.1 release; (e) not cherry-picked - it is a hard request-validation rule, not a benchmark result. Two qualifications that narrow but do not refute the claim: (i) "controlled ONLY through output_config.effort" is slightly strong - the docs also list prompt/system-prompt steering and a beta per-message `role: "system"` output_config message (`mid-conversation-output-config-2026-07-01` header) as ways to change depth, and `max_tokens` as the hard ceiling; effort is the only top-level API depth knob, which is what the claim means. (ii) `thinking: {type: "adaptive"}` (with `display`) IS accepted on Fable 5.1 - only `"disabled"` and `"enabled"` are 400s - so "sending a thinking type" is not universally rejected, just those two.
Counter sourceNone found. Nearest thing to a counterpoint: Claude Opus 5 (not Fable 5.1) DOES accept thinking {type:"disabled"} at effort high or below, per the same troubleshooting table - a contrast that could be misread as contradicting the claim but actually reinforces the Fable-5.1-specific rule.
Verifier 2not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. Verified against three separate Anthropic public docs pages (not the cached skill table). (1) https://platform.claude.com/docs/en/models/fable-5-1/migration-guide states verbatim: "Adaptive thinking is always on, unchanged from Claude Fable 5. The model decides when and how much to think. No `thinking` configuration is required. Both `thinking: {type: \"disabled\"}` and manual extended thinking (`thinking: {type: \"enabled\", budget_tokens: N}`) return a 400 error." Same page: "On `claude-fable-5-1` and `claude-mythos-5-1`, adaptive thinking is always on, and `thinking: {type: \"disabled\"}` returns a 400 error at any effort level." (2) https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting per-model table row: "Claude Fable 5.1 | Adaptive only | Always on | Rejected with 400: \"enabled\", \"disabled\"" and gives the exact server messages, e.g. "\"thinking.type.enabled\" is not supported for this model. Use \"thinking.type.adaptive\" and \"output_config.effort\" to control thinking behavior." Note this is a stronger rejection than the claim: `type: "enabled"` is rejected outright, with or without budget_tokens. (3) https://platform.claude.com/docs/en/build-with-claude/effort confirms the level range: "Claude Fable 5.1 supports all five effort levels" (low, medium, high, xhigh, max), set at `output_config.effort`, and "Where adaptive thinking is available, effort is the recommended way to control thinking depth." TWO CAVEATS, neither fatal: (a) "only through output_config.effort" is a mild overreach. https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost calls effort "the primary control" but documents two further non-parameter levers - system-prompt guidance and per-message prompting ("Please think hard before responding.") - plus `max_tokens` as the hard cost cap. So effort is the only API *parameter* controlling depth, not the only mechanism. (b) The claim omits that `thinking: {type: "adaptive"}` IS accepted (with `display: "summarized"/"omitted"/"updates"`), so "any thinking config errors" would be wrong; the claim as written names only disabled and budget_tokens, which is correct. Source-quality check: this is a documented API error-surface fact, where the vendor's own reference docs are the authoritative primary source rather than marketing-grade material - unlike the benchmark/quality half of the wider research question. No contradicting source found; the only near-miss in search results (GitHub anthropics/claude-code#76689, a 400 on Opus 4.8 with xhigh + thinking disabled) is about a different model and is consistent with the documented Opus 5-and-later restriction. Current as of 2026-09-02, one day after the Fable 5.1 launch.
Counter sourceNo credible contradicting source found. Nearest qualification: platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost, which shows system-prompt and per-message prompting as additional (non-parameter) thinking-depth levers, softening the word "only" in the claim.
Verifier 3not refutedhigh confidence
SUPPORTED by first-party public documentation, and I could not refute it. (1) https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - under "Unchanged from Claude Fable 5" it states verbatim: "Adaptive thinking is always on. `thinking: {\"type\": \"enabled\"}` with `budget_tokens` and `thinking: {\"type\": \"disabled\"}` both return a 400 error. Omit `thinking` or send `{\"type\": \"adaptive\"}`." The same page's spec block says: "Thinking: adaptive thinking is always on. Use the effort parameter to control thinking depth." It also gives thinking-always-on as the stated reason forced tool use is rejected ("Thinking is always on for these models, and a forced tool call would skip it"). (2) https://platform.claude.com/docs/en/build-with-claude/effort - confirms the parameter path is `output_config.effort`, that it is supported on `claude-fable-5-1` with no beta header, that the ladder is exactly low/medium/high/xhigh/max ("Claude Fable 5.1 supports all five effort levels. Start with `high`, the default"), and that "Where adaptive thinking is available, effort is the recommended way to control thinking depth." (3) Independent corroboration that the 400 is real in production, not just doc text: langchainjs issue #11052, "ChatAnthropic sends thinking: {\"type\":\"disabled\"} by default - rejected with 400 by claude-fable-5" (https://github.com/langchain-ai/langchainjs/issues/11052), plus AWS's own Bedrock adaptive-thinking page. Adversarial checks: not outdated (docs current for Fable 5.1 on 2026-09-02); not marketing-grade - this is an API contract fact in the vendor's own reference docs, which is the authoritative source for its own error codes, and it is corroborated by a third-party bug report; not cherry-picked. ONE NUANCE, not enough to refute: the word "only" in "controlled only through output_config.effort" is slightly stronger than the docs. Anthropic calls effort the "recommended" depth control and separately documents prompt steering (https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost) and advisory task budgets (https://platform.claude.com/docs/en/build-with-claude/task-budgets) as additional levers. As a statement about request parameters ("no other parameter sets thinking depth") the claim holds; read as "nothing else influences how much it thinks", it overstates marginally. Also note `thinking.display` (omitted/summarized/updates) remains configurable - it controls visibility, not depth, so it does not contradict "thinking is always on".
Counter sourceNo contradicting source found. Nearest qualification: platform.claude.com/docs/en/build-with-claude/effort frames effort as the "recommended" (not sole) depth control and documents prompting and task budgets as additional steering, so the claim's "only" is precise for API parameters but slightly absolute otherwise.
R4confirmedvote 3-0On Claude Fable 5.1, forced tool use (tool_choice type 'any' or type 'tool') returns a 400 error.
Verdict notetool_choice any and tool return 400 invalid_request_error on Messages, Batches and token-counting endpoints, independently corroborated by LiteLLM; strict:true is unavailable in CMEK orgs so the workaround is instruction-only there.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF1

The three verifiers
Verifier 1not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM; it is not merely a cached-tool-table artifact. I did not judge it against its own source. (1) PRIMARY, first-party public doc - https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1, section "Breaking changes > Forced tool use is not supported", fetched 2026-09-02. Verbatim: "Claude Fable 5.1 and Claude Mythos 5.1 don't support forced tool use. `tool_choice` set to `{\"type\": \"any\"}` or `{\"type\": \"tool\", \"name\": \"...\"}` returns a 400 `invalid_request_error`:" followed by the literal error string "tool_choice: type \"tool\" and \"any\" are not supported for this model." The same page adds: "`tool_choice: {\"type\": \"auto\"}` (the default) and `{\"type\": \"none\"}` are unchanged. The same validation applies to the token counting endpoint." Rationale given: "Thinking is always on for these models, and a forced tool call would skip it." (2) SECOND first-party public page - https://platform.claude.com/docs/en/models/fable-5-1/migration-guide, "Breaking changes" item 1, fetched 2026-09-02. Verbatim: "Forced tool choice is not supported: Claude Fable 5 accepts `tool_choice` `auto`, `none`, `any`, and `tool`. On `claude-fable-5-1`, `{type: \"any\"}` and `{type: \"tool\", name: \"...\"}` return a 400 `invalid_request_error`" with the identical error string, plus "The check applies on the Messages API, the Message Batches API, and the token counting endpoint." It carries before/after code samples across 7 SDKs showing the migration from `tool_choice: {type: "tool"}` to `auto` + strict:true. (3) INDEPENDENT, non-Anthropic corroboration - https://docs.litellm.ai/blog/claude_fable_5_1 (LiteLLM day-0 support post). Verbatim: "Forced tool use returns a 400. Thinking is always on, and a forced call would skip it. LiteLLM maps OpenAI's `tool_choice: \"required\"` to Anthropic's `any`, so a request that worked on Fable 5 fails here unchanged." This is a third party reporting the failure in their own integration, i.e. behaviour observed in practice, not a vendor assertion. CHECKLIST DISPOSITION: 1. Overreach? No. The claim is a strict subset of the quoted doc text; it names both rejected forms and the 400. 2. Contradicting evidence? None found across searches for tool_choice/400/Fable 5.1. Multiple independent search paths returned the same restriction. 3. Source quality vs claim strength? Adequate and then some. This is an API contract fact (what HTTP status the vendor's own endpoint returns), the one class of claim where first-party docs ARE the authority - unlike a benchmark or quality claim, there is no "does the vendor claim transfer" gap. It is additionally corroborated by an independent SDK vendor. Note the seed's cited source (a cached skill table dated 2026-06-24) is NOT what I relied on; the live public docs are. 4. Outdated? No. Fable 5.1 is the current top model as of 2026-09-02 and the doc pages are live and current; this is a launch-time breaking change, not a deprecated behaviour. 5. Cherry-picked? Not applicable. This is deterministic request validation, not a benchmark or effort-level-dependent result - it fires regardless of effort, prompt, or workload.
Counter sourceNo contradicting source found. The nearest thing to a qualification is scope-widening rather than contradiction: Anthropic's migration guide states the rejection applies not only to the Messages API but also to the Message Batches API and the token-counting endpoint, and the same restriction is documented for Claude Mythos 5.1 (Project Glasswing only), so the claim as stated is narrower than reality, not broader. A separate CMEK caveat affects the recommended workaround (structured outputs / strict:true unavailable on Fable models in CMEK organizations), not the 400 behaviour itself.
Verifier 2not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS the claim, verbatim and from the vendor's own public docs (not just the cached skill table). 1) PRIMARY - https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 has a "Breaking changes" section headed "Forced tool use is not supported": "Claude Fable 5.1 and Claude Mythos 5.1 don't support forced tool use. `tool_choice` set to `{\"type\": \"any\"}` or `{\"type\": \"tool\", \"name\": \"...\"}` returns a 400 `invalid_request_error`:" followed by the literal error string `tool_choice: type "tool" and "any" are not supported for this model.` It adds that `auto` (default) and `none` are unchanged, that the same validation applies to the token counting endpoint, and gives the stated reason (thinking is always on; a forced call would skip it and push working-out into tool arguments). The remedies named match the seed quote exactly: keep `auto` + `strict: true` (strict tool use), or move the schema to structured outputs, or name the tool in the prompt. 2) PRIMARY #2 - https://platform.claude.com/docs/en/models/fable-5-1/migration-guide, breaking-change item 1: "Forced tool choice is not supported: Claude Fable 5 accepts `tool_choice` `auto`, `none`, `any`, and `tool`. On `claude-fable-5-1`, `{type: \"any\"}` and `{type: \"tool\", name: \"...\"}` return a 400 `invalid_request_error`" with the same error string, and "The check applies on the Messages API, the Message Batches API, and the token counting endpoint." Migration checklist step 1 on the what's-new page: "Remove any `tool_choice` of type `any` or `tool`." ADVERSARIAL CHECKS - all pass: - Quote overreach: none. The seed quote is a faithful compression of the doc text, including the two remedies. It is a binary API-validation fact, not a benchmark or performance claim, so there is no effort-level or cherry-picking dimension to misrepresent. - Contradicting evidence: searched for any source saying forced tool use still works (including a possible partner-platform carve-out, the most plausible qualification). Found none. AWS Bedrock coverage runs the same way - forced tool use unsupported on Fable 5.1/Mythos 5.1 across InvokeModel, InvokeModelWithResponseStream and CountTokens - so the restriction is not first-party-API-only. - Independent (non-marketing) corroboration: LiteLLM's day-0 page https://docs.litellm.ai/blog/claude_fable_5_1 confirms it from an integrator's pain point: "Forced tool use returns a 400. Thinking is always on... LiteLLM maps OpenAI's `tool_choice: \"required\"` to Anthropic's `any`, so a request that worked on Fable 5 fails here unchanged." A third-party writeup (https://roo.beehiiv.com/p/claude-fable-5-1-mythos-5-1-whats-new) says the same and notes no platform exceptions. This is the class of vendor claim that transfers cleanly - it is a validation error anyone can reproduce, not a quality claim. - Currency: Fable 5.1 shipped ~2026-09-01; today is 2026-09-02. Docs are the current release docs, one day old, not stale. - Source strength vs claim strength: the claim is narrow and mechanical; two vendor doc pages plus AWS docs plus an independent SDK/proxy vendor is more than sufficient. Only nuance found (does not affect the claim): in a CMEK organization structured outputs and `strict: true` are unavailable on Fable models, so the fallback there is the prompt instruction alone. That qualifies the workaround, not the 400.
Counter sourceNo contradicting source found. Searched for platform carve-outs (Bedrock/Vertex/Foundry) and for practitioner reports of forced tool_choice succeeding on claude-fable-5-1; AWS Bedrock documentation reinforces rather than contradicts the restriction, and LiteLLM's integration notes independently reproduce the 400.
Verifier 3not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. Confirmed verbatim on three independent first-party Anthropic doc pages fetched live 2026-09-02, not from the cached skill table. (1) https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - section "Forced tool use is not supported": "Claude Fable 5.1 and Claude Mythos 5.1 don't support forced tool use. `tool_choice` set to `{\"type\": \"any\"}` or `{\"type\": \"tool\", \"name\": \"...\"}` returns a 400 `invalid_request_error`:" followed by the literal error string `tool_choice: type \"tool\" and \"any\" are not supported for this model.` It is listed as one of three breaking changes vs Claude Fable 5. `auto` (default) and `none` are explicitly unchanged. Stated rationale: thinking is always on and a forced tool call would skip it, so the model writes its working-out into the tool arguments, lowering argument quality. (2) https://platform.claude.com/docs/en/models/fable-5-1/migration-guide - breaking change #1: "Claude Fable 5 accepts `tool_choice` `auto`, `none`, `any`, and `tool`. On `claude-fable-5-1`, `{type: \"any\"}` and `{type: \"tool\", name: \"...\"}` return a 400 `invalid_request_error`" with the same error string, plus a scope detail the seed quote omits: "The check applies on the Messages API, the Message Batches API, and the token counting endpoint." Includes before/after code showing the migration to `tool_choice: {\"type\": \"auto\"}` + naming the tool in the instruction + `strict: true`. (3) https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools - the "Forcing tool use" restriction table has a dedicated row: "Claude Fable 5.1 and Claude Mythos 5.1 | `any` and `tool` return a 400 error | `auto` with strict tool use to guarantee schema-valid tool inputs, or structured outputs when you need a response in a fixed JSON shape. Prompting still influences which tool `auto` picks. `none` is also supported." Adversarial checks all pass. (a) No overreach: the seed quote's three workarounds (auto + prompt instruction, strict:true, structured outputs) match the docs exactly. (b) No contradicting source found; independent third-party writeups (docs.litellm.ai day-0 Fable 5.1 post, digitalapplied.com, claudefa.st "3 breaking changes") report the same restriction. (c) Source-quality objection does not apply here: this is a hard API contract, not a benchmark or performance claim, and the vendor is the authoritative source for its own endpoint's validation behaviour - the "vendor claims may not transfer" concern is about quality claims, not 400 responses. (d) Current as of today's fetch. (e) Not cherry-picked: binary API behaviour, no effort level or benchmark involved. Two qualifications worth carrying forward, neither of which refutes the claim: the restriction extends beyond the Messages API to the Message Batches API and token counting endpoint (broader than stated); and in a CMEK organisation, structured outputs and `strict: true` are not available on Claude Fable models, so one of the three recommended workarounds does not exist there and the prompt instruction is the only route.
Counter sourceNone found. Searched for contradicting evidence on tool_choice/400/Fable 5.1; every source located (Anthropic platform docs x3, LiteLLM day-0 post, third-party migration writeups) agrees. The only near-miss is a separate, distinct restriction on the same docs page: manual extended thinking (thinking: {type: "enabled"}) also blocks any/tool on other models, while adaptive thinking models such as Claude Opus 5 DO support forced tool use - so the restriction is specific to Fable/Mythos 5.1, not general to all thinking-on models.
R5partly confirmedvote 3-0On Claude Fable 5.1, editing an earlier conversation turn invalidates the thinking blocks after it ('preserved thinking'), so conversation history must be append-only.
Verdict noteThe mechanism is exactly as documented, but the 400 is enforced by default only for accounts created on or after 2026-08-31, drop_block avoids it, and Claude Code, claude.ai and Agent SDK users are explicitly unaffected, so 'must be append-only' bites only on hand-rolled Messages API loops.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF1

The three verifiers
Verifier 1not refutedhigh confidence
SUPPORTED by Anthropic's own PUBLIC primary documentation, not just the cached skill table. I fetched https://platform.claude.com/docs/en/build-with-claude/preserved-thinking (page title "Preserved thinking"). Opening line: "On Claude Fable 5.1, changing prior turns in the conversation (the `system` prompt, the `tools`, or any earlier message) affects the API response. By default, it makes the API reject the request with an error." The page's "What counts as an edit" table lists "Edit, reorder, or delete any earlier `user`, `assistant`, or `system` message" -> later thinking blocks "Invalid", and separately "Remove a `thinking` block from the middle of the history and keep later ones" -> "Invalid for every later thinking block". The append-only half is Anthropic's own wording, not an inference: "You should check your code and ensure that the `messages` array is treated as append-only", and the page checklist repeats "Consecutive request bodies are byte-identical in `system`, `tools`, and the shared `messages` prefix." The 400 text is quoted verbatim in the doc: "The block is bound to a different conversation." Corroborated by Anthropic's help-centre article https://support.claude.com/en/articles/16761192-preserved-thinking-... and by two INDEPENDENT third-party posts (https://claudefa.st/blog/models/claude-fable-5-1 and https://www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes), which describe it as a breaking change and quote the same error. This is a spec claim about the vendor's own API surface, so vendor docs are the authoritative source rather than marketing-grade; the independent posts corroborate anyway. Current, not outdated: enforcement began 2026-08-31 00:00 UTC, two days before today. ADVERSARIAL ATTEMPTS THAT FAILED, and the qualifications they surfaced (these narrow the "must", they do not refute the mechanism): (1) Scope of enforcement. "Who is affected" says the check is enforced by default ONLY for accounts created on or after 2026-08-31 00:00 UTC; older accounts are not enforced unless they set `thinking.block_binding.prefix_mismatch_behavior`. So on Fable 5.1 today, an older-account harness that edits history does NOT get a 400. The doc adds "Later models will enforce the check for all users" - which is exactly what the cached skill quote says, so the quote is accurate rather than overreaching. (2) "Must be append-only" is not absolute. `prefix_mismatch_behavior: "drop_block"` (beta header `thinking-binding-controls-2026-08-01`) lets the request succeed with the affected blocks stripped, and simple client-side compaction (replay a summary and NO earlier thinking blocks) is explicitly still fine: "This check doesn't prohibit client-side compaction. The rule is narrower: don't keep a thinking block behind a prefix you've rewritten." (3) Relevance to this user. The doc says users of official products/SDKs - Claude Code, claude.ai, Claude Managed Agents, the Claude Agent SDK - are unaffected ("These keep the prefix intact for you"; checklist: "stop here"). The help-centre article says the same. So the operational instruction "make every harness append-only" only bites on hand-rolled Messages API loops, not on Claude Code sessions. (4) I searched specifically for contradiction/dispute ("developers complaint", "not required", "older accounts") and found no credible source disputing the mechanism - only sources restating the account-age carve-out. Net: mechanism claim = correct and primary-sourced; the word "must" should carry the caveats in (1)-(3).
Counter sourcehttps://platform.claude.com/docs/en/build-with-claude/preserved-thinking (section "Who is affected" - enforced by default only for accounts created on/after 2026-08-31; `prefix_mismatch_behavior: "drop_block"` avoids the error; Claude Code / claude.ai / Agent SDK users unaffected) and https://claudefa.st/blog/models/claude-fable-5-1 ("Older accounts get the mismatch recorded but not acted on unless the request sets thinking.block_binding.prefix_mismatch_behavior")
Verifier 2not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. It is not judged against its own source (the cached skill table); I found a dedicated first-party doc page. 1) Primary: https://platform.claude.com/docs/en/build-with-claude/preserved-thinking - "On Claude Fable 5.1, changing prior turns in the conversation (the `system` prompt, the `tools`, or any earlier message) affects the API response. By default, it makes the API reject the request with an error". Its "What counts as an edit" table lists "Edit, reorder, or delete any earlier `user`, `assistant`, or `system` message" -> later thinking blocks "Invalid"; likewise editing `system`, editing `tools`, or "Remove a `thinking` block from the middle of the history and keep later ones" -> "Invalid for every later thinking block". The append-only wording is the doc's own: "ensure that the `messages` array is treated as append-only", and the checklist requires consecutive request bodies be "byte-identical in `system`, `tools`, and the shared `messages` prefix". Exact 400: "messages.1.content.0: Invalid `signature` in `thinking` block. The block is bound to a different conversation." 2) First-party corroboration: https://support.claude.com/en/articles/16761192-preserved-thinking-changing-how-the-messages-api-handles-thinking-blocks-to-protect-against-distillation - the API "will verify that a thinking block is sent back with the same system prompt, tools, and messages that produced it, and will return an error if they don't match"; applies to "new API accounts created after August 31, 2026 12:00:00 AM UTC"; "Only Claude Fable 5.1 is impacted ... Preserved thinking will apply to all users for future models." 3) Independent corroboration (not vendor): https://www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes - "Modifying the system prompt, the tool list, or any earlier message before a thinking block produces a 400 on the next request"; "Enforced for API accounts created on or after August 31, 2026; recorded but not acted on for older accounts unless you opt in." Practitioner reports of the same class of error exist at https://github.com/anthropics/claude-code/issues/22278 (400 "`thinking` ... blocks in the latest assistant message cannot be modified"). ADVERSARIAL CHECKS RUN, none fatal: (a) Not an overreach - the seed's own quote carries the account-age qualifier, and "append-only" is literally the docs' term. (b) Not outdated - the page describes betas dated 2026-08-01/2026-08-21 and today is 2026-09-02. (c) Source quality is adequate: this is an API contract about the vendor's own endpoint, where first-party docs are authoritative rather than marketing-grade, and it is independently corroborated. (d) Not cherry-picked (no benchmark/effort-level dependence). TWO QUALIFICATIONS a harness author should carry, which narrow but do not refute the claim: (i) enforcement is account-gated today - Fable 5.1 enforces by default only for accounts created on/after 2026-08-31 00:00 UTC; older accounts must opt in via `thinking.block_binding.prefix_mismatch_behavior` (beta header `thinking-binding-controls-2026-08-01`), though "Later models will enforce the check for all users." (ii) "Must be append-only" is the practical rule for histories that replay thinking blocks, not an absolute ban on client-side compaction: `prefix_mismatch_behavior: "drop_block"` drops the offending blocks and succeeds, and the docs explicitly bless "simple compaction" that replays no earlier thinking blocks, plus removing thinking blocks from the front of the history. Also note Claude Code / claude.ai / the Agent SDK "keep the prefix intact for you", so the rule bites on hand-rolled Messages API loops.
Counter sourcehttps://www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes (independent; corroborates rather than contradicts). Nearest thing to a counter-point is the first-party https://support.claude.com/en/articles/16761192 plus https://platform.claude.com/docs/en/build-with-claude/preserved-thinking, which qualify the "must": enforcement is default-on only for API accounts created on/after 2026-08-31, and `prefix_mismatch_behavior: "drop_block"` lets an edited history succeed with blocks silently dropped instead of erroring.
Verifier 3not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. Anthropic's own public docs carry a dedicated page named exactly "Preserved thinking" (https://platform.claude.com/docs/en/build-with-claude/preserved-thinking), which states: "On Claude Fable 5.1, changing prior turns in the conversation (the `system` prompt, the `tools`, or any earlier message) affects the API response. By default, it makes the API reject the request with an error"; "You should check your code and ensure that the `messages` array is treated as append-only"; and lists as prefix-invalidating "Trimming or dropping older turns", "Injecting a reminder into an earlier turn and removing it on the next request", "Rebuilding the `system` prompt each request", "Adding or removing entries in `tools` mid-session". Its "What counts as an edit" table marks "Edit, reorder, or delete any earlier `user`, `assistant`, or `system` message" as Invalid, and "Remove a `thinking` block from the middle of the history and keep later ones" as "Invalid for every later thinking block". The 400 text is quoted: "messages.1.content.0: Invalid `signature` in `thinking` block. The block is bound to a different conversation." Model binding is confirmed too: "A block is readable by the model that produced it and by later models, not by earlier ones." Enforcement scope confirmed verbatim: "The check is enforced by default for new accounts created on or after August 31, 2026, 00:00 UTC... Later models will enforce the check for all users." Anthropic's support article (https://support.claude.com/en/articles/16761192-preserved-thinking-changing-how-the-messages-api-handles-thinking-blocks-to-protect-against-distillation) independently repeats the account cutoff and extends it to Bedrock, Vertex and Azure Foundry. INDEPENDENT CORROBORATION (not Anthropic-authored): LiteLLM's day-0 page (https://docs.litellm.ai/blog/claude_fable_5_1) states "Editing earlier turns invalidates thinking blocks. Rebuilding `system` or `tools` mid-conversation errors with `The block is bound to a different conversation`"; and https://www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes states "Modifying the system prompt, the tool list, or any earlier message before a thinking block produces a 400 on the next request" and advises treating conversations as "append-only". CHECKLIST: (1) claim is a faithful summary of its quote, no overreach - every element (model binding, invalidation on edit, "preserved thinking" name, 2026-08-31 new-account cutoff, later-models-for-everyone, append-only harnesses) appears verbatim in the primary doc; (2) no contradicting source found; (3) source quality is adequate BECAUSE this is an API-contract claim, not a performance claim - the vendor's own API reference is ground truth for its own error behaviour, and it is corroborated by two non-Anthropic parties anyway; (4) current, not outdated - Fable 5.1 shipped ~2026-09-01, docs live as of 2026-09-02; (5) not cherry-picked - it is a documented API rule, not a benchmark result. The only correction worth carrying forward is the scope qualification in counterSource: "append-only" is the recommended harness discipline, while the enforced rule is narrower (no thinking block left behind a rewritten prefix), and enforcement is default-on only for accounts created on/after 2026-08-31 unless a request opts in via prefix_mismatch_behavior.
Counter sourceNo credible contradicting source found. Closest thing to a counter is a scope qualification inside the same primary doc: https://platform.claude.com/docs/en/build-with-claude/preserved-thinking states the check "is enforced by default for new accounts created on or after August 31, 2026, 00:00 UTC" (older accounts unaffected until "later models will enforce the check for all users"), that `thinking.block_binding.prefix_mismatch_behavior: "drop_block"` (beta header `thinking-binding-controls-2026-08-01`) turns the 400 into a silent drop, and that official surfaces (Claude Code, claude.ai, Claude Agent SDK, Managed Agents) already keep the prefix intact so their users need do nothing. It also narrows "append-only" slightly: "This check doesn't prohibit client-side compaction. The rule is narrower: don't keep a thinking block behind a prefix you've rewritten" - simple compaction (replace history with a summary, carry no thinking blocks) and removing thinking blocks from the FRONT of history both stay valid. These qualify the blast radius, not the claim itself.
R6confirmedvote 3-0Claude Fable 5.1 is not available on Priority Tier, and is not available under zero data retention unless expressly authorized by Anthropic.
Verdict noteBoth halves verbatim in Anthropic docs; Priority Tier also excludes Opus 5 and Sonnet 5 and is closed to new purchase, and the launch post says eligible customers may get ZDR pending Enterprise Frontier Safeguards.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF1

The three verifiers
Verifier 1not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS BOTH HALVES, verbatim, from Anthropic's own platform docs (checked 2026-09-02, one day after Fable 5.1's 2026-09-01 release). (1) ZDR half - https://platform.claude.com/docs/en/manage-claude/api-and-data-retention states, near-word-for-word: "Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5: These models require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic." The same page's Model-specific section repeats it and adds the mechanism: the four are "designated Covered Models ... ZDR is therefore not available for any of them unless expressly authorized by Anthropic. On the Claude API, requests to Claude Fable 5 from an organization whose data retention configuration does not meet this requirement return a `400 invalid_request_error`". So the cached skill table's 400-error detail is also corroborated. (2) Priority Tier half - https://platform.claude.com/docs/en/api/service-tiers, "Supported models": "Priority Tier is supported on all available Claude models except Claude Fable 5.1, Claude Mythos 5.1, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5, and Claude Sonnet 5." Checklist results. (a) No overreach: the claim is a plain restatement of the quote, and the quote matches the public docs. (b) Searched for contradiction; found none. Independent write-ups (jetstream.security, listiak.dev, synthorai.io, constellationr.com, digitalapplied.com) all report the same 30-day-retention/no-ZDR posture and treat it as a live enterprise problem, so the policy is corroborated outside Anthropic's own pages. (c) Source quality is adequate BECAUSE this is not a performance or marketing claim - it is a policy/availability fact about the vendor's own API surface, where the vendor's docs are the authoritative source rather than marketing-grade material. (d) Not outdated: model page shows "Released September 1, 2026", status "Active (latest)". (e) Not cherry-picked: no benchmark or effort level is involved. Two qualifications the claim omits, neither of which contradicts it: - The launch announcement (https://www.anthropic.com/claude-fable-and-mythos-5-1) says "eligible customers will be able to use Fable 5.1 with zero data retention" as an interim measure until Enterprise Frontier Safeguards ships "in phases, beginning later this fall". That is an authorized exception, i.e. exactly the "unless expressly authorized" carve-out - it widens who may qualify, it does not falsify the rule. The retention doc also documents a workspace-level override: a ZDR org can enable 30-day retention on one workspace and keep ZDR everywhere else. - The Priority Tier fact is real but less distinguishing than it sounds: the same page carries a warning that "Priority Tier capacity commitments are no longer available for purchase", and the exclusion list also covers Claude Opus 5 and Claude Sonnet 5 - so this is not a Fable-5.1-specific penalty.
Counter sourceNo contradicting source found. Nearest qualifier: https://www.anthropic.com/claude-fable-and-mythos-5-1 ("eligible customers will be able to use Fable 5.1 with zero data retention" pending Enterprise Frontier Safeguards), plus https://platform.claude.com/docs/en/api/service-tiers noting Priority Tier commitments are closed to new purchase and that Opus 5 / Sonnet 5 are excluded too.
Verifier 2not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. Both halves are confirmed near-verbatim in Anthropic's own public platform docs (not just the cached skill table), each on a separate page. (1) ZDR half - https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 (Availability section) states: "Claude Fable 5.1 and Claude Mythos 5.1 carry 30-day data retention and aren't available under zero data retention unless expressly authorized by Anthropic. Both are Covered Models, like Claude Fable 5 and Claude Mythos 5." Corroborated on a second page, https://platform.claude.com/docs/en/manage-claude/api-and-data-retention: "Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5: These models require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic," plus "Covered Models are not available under ZDR regardless of feature eligibility." That page also carries the 400 error the seed quote cites: "In order to access this model, your organization or workspace must have data retention enabled" (type invalid_request_error). (2) Priority Tier half - https://platform.claude.com/docs/en/api/service-tiers, "Supported models": "Priority Tier is supported on all available Claude models except Claude Fable 5.1, Claude Mythos 5.1, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5, and Claude Sonnet 5." ADVERSARIAL CHECKS RUN, none refuting: - Contradiction hunt: the launch post https://www.anthropic.com/claude-fable-and-mythos-5-1 says "Until EFS [Enterprise Frontier Safeguards] is available, eligible customers will be able to use Fable 5.1 with zero data retention." That reads as a partial softening but is consistent, not contradictory - "eligible customers" is the same gate as "expressly authorized by Anthropic," and the docs' hard rule stands for everyone else. - Source quality: this is Anthropic's own technical documentation, so it is vendor-sourced - but the claim is about Anthropic's own product policy, where the vendor IS the authoritative source. This is not a transferability-of-benchmark-claims situation, so marketing-grade discounting does not apply. - Currency: docs checked 2026-09-02, the day after the 2026-09-01 Fable 5.1 launch; the pages describe 5.1 explicitly. - Not cherry-picked: two independent doc pages plus the launch post agree. TWO QUALIFICATIONS worth carrying forward (neither refutes): (a) the "no Priority Tier" fact is partly moot - the same service-tiers page opens with "Priority Tier capacity commitments are no longer available for purchase," so no new customer could buy Priority Tier for any model; (b) the doc's worked 400-error example is written against Claude Fable 5 specifically, though the requirement text names Fable 5.1 alongside it, and a ZDR-authorized org can instead enable 30-day retention on a single workspace to use Covered Models there while other workspaces stay ZDR.
Counter sourcehttps://www.anthropic.com/claude-fable-and-mythos-5-1 - "Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention." Softens the absoluteness but does not contradict: "eligible customers" is the express-authorization carve-out. Secondary: https://platform.claude.com/docs/en/api/service-tiers notes Priority Tier is closed to new purchase entirely, making the Priority Tier exclusion largely moot in practice.
Verifier 3not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM. Both halves verified verbatim against Anthropic's live first-party platform documentation (fetched 2026-09-02, one day after the 2026-09-01 Fable 5.1 release). (1) PRIORITY TIER - https://platform.claude.com/docs/en/api/service-tiers, section "Existing Priority Tier commitments > Supported models": "Priority Tier is supported on all available Claude models except Claude Fable 5.1, Claude Mythos 5.1, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5, and Claude Sonnet 5." Fable 5.1 is named explicitly in the exclusion list. Exact match. (2) ZDR - https://platform.claude.com/docs/en/manage-claude/api-and-data-retention, section "Model-specific data retention requirements": "Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5 are designated Covered Models ... and require 30-day data retention; ZDR is therefore not available for any of them unless expressly authorized by Anthropic." The same page's ZDR-scope section repeats it: "These models require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic." The seed quote's wording ("ZDR only if expressly authorized by Anthropic; no Priority Tier") is a near-verbatim restatement, not an overreach. The 400 invalid_request_error detail also checks out, though the doc's error example names Claude Fable 5 specifically while the 30-day rule is stated generally for all Covered Models. ADVERSARIAL CHECKS RUN: - Contradiction hunt: one apparent contradiction surfaced - a WebSearch snippet of the service-tiers page listing exclusions as only "Mythos 5, Mythos Preview, Opus 5, Sonnet 5" (i.e. Fable supported). Fetching the live page shows this is a STALE CACHE predating the 5.1 release; the current page adds Fable 5.1 and Mythos 5.1. Not a real contradiction. - Source quality: this is a policy/availability statement about Anthropic's own service, so Anthropic's platform docs are the authoritative primary source, not marketing-grade. The marketing-transfer concern applies to benchmark/performance claims, not to "which service tiers does our product support". Independently corroborated in kind by TechCrunch (2026-09-01), Constellation Research and Unite.AI coverage of the release, all of which report the retention restriction. - Currency: current as of today; the Fable 5.1 model page (https://platform.claude.com/docs/en/models/fable-5-1/overview) confirms Released September 1, 2026, status "Active (latest)". - Cherry-picking: N/A - categorical policy statement, not a benchmark or single-effort-level result. TWO QUALIFICATIONS (do not refute, but the claim is thinner than it sounds): (a) The Priority Tier exclusion is close to moot. The same page carries a warning: "Priority Tier capacity commitments are no longer available for purchase. Organizations with an existing commitment can continue to use Priority Tier through their contract end date." So Fable 5.1 is excluded from a programme already closed to new customers - this is not a Fable-specific penalty relative to what a new buyer could otherwise get. It matters only to orgs with a live pre-existing commitment. (b) "Not available under ZDR" is not absolute. The retention doc documents a workspace-level override: "Organizations with a ZDR arrangement can make these models available in a specific workspace by enabling 30-day retention for that workspace only. Other workspaces in the organization keep zero data retention." Separately, Anthropic's announcement (https://www.anthropic.com/claude-fable-and-mythos-5-1) states eligible customers can currently use Fable 5.1 with zero data retention until Enterprise Frontier Safeguards (EFS) ships "later this fall", with EFS then giving "the same as a zero data retention policy" via customer-controlled cloud storage. Both are consistent with the "unless expressly authorized" carve-out rather than contradicting it.
Counter sourceNo credible counter-source found. The only near-miss was a stale search-index snapshot of https://platform.claude.com/docs/en/api/service-tiers listing the Priority Tier exclusions without Fable 5.1; the live page (fetched 2026-09-02) includes it, so the snippet reflects a pre-release cache rather than a dispute.
R7partly confirmedvote 3-0Anthropic's guidance states that prompts written for prior models are often too prescriptive for Claude Fable 5.1 and reduce its output quality.
Verdict noteThe 'too prescriptive, can degrade output quality' sentence is published for Fable 5, not 5.1; the 5.1 guide says Fable 5 prompts need no changes and only names two legacy lines to remove, so the rule's attribution to 5.1 is a transfer, and no independent measurement of the effect exists.
Run statusconfirmed, vote 3-0, 0 verifier errors

see alsoF7

The three verifiers
Verifier 1not refutedmedium confidence
PUBLIC EVIDENCE SUPPORTS THE SUBSTANCE, WITH A REAL VERSION-SCOPING CAVEAT. 1) Near-verbatim confirmation on Anthropic's public docs (not the cached tool table). https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 - section "Recommended scaffolding changes", bullet titled "Refactor existing prompts and skills": "Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions if default performance is better." So the guidance exists publicly and is not merely a vendor social post. 2) Corroborated model-agnostically on a page that explicitly names Fable 5.1 as in scope. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices ("This is the reference for prompt engineering with current Claude models, including Claude Fable 5.1..."): "Prefer general instructions over prescriptive steps. A prompt like 'think thoroughly' often produces better reasoning than a hand-written step-by-step plan. Claude's reasoning frequently exceeds what a human would prescribe." The same page's Opus 5 note applies the identical logic to a different behaviour: "verification instructions carried over from prompts tuned for earlier models can cause over-verification... remove these instructions rather than rewriting them." 3) The 5.1-specific guide repeats the direction in narrower, concrete form. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 - "Second, audit your prompt for instructions that suppress narration. Some earlier models were eager to give updates while working, which led to system prompt lines such as 'hold all findings for the final response.' Remove lines like that before adding anything." And: "Earlier models overused bullets and bold in chat, and many prompts carry anti-formatting rules written to hold that down... If your prompt contains anti-formatting language, remove it." QUALIFICATIONS I COULD NOT DISMISS (why confidence is medium, not high): a) VERSION LABEL. The "too prescriptive / degrade output quality" sentence names Claude Fable 5, not 5.1. I grepped the full text of both 5.1-specific docs - https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 and https://platform.claude.com/docs/en/models/fable-5-1/migration-guide - and neither contains "prescriptive", neither contains the degrade-quality statement, the 5.1 prompting guide never uses the word "skill" at all, and neither links back to the Fable 5 prompting guide. The Fable 5.1 overview (https://platform.claude.com/docs/en/models/fable-5-1/overview) points only to "Prompting Claude Fable 5.1" as its prompting resource. b) A DIRECTLY QUALIFYING LINE. The 5.1 guide opens: "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes, but a handful of behavioral differences are worth knowing about." Fable 5 is a prior model relative to 5.1, so read literally, the seed claim is broader than Anthropic's 5.1 guidance. The claim is best read as "prompts written for pre-Fable-5 models", which is what the source doc actually says. c) WORDING. Source doc says "Skills", claim says "prompts"; doc says "can degrade", claim says "reduce". The bullet heading ("Refactor existing prompts and skills") makes the prompts/skills widening defensible, but it is a widening. d) NO INDEPENDENT CORROBORATION of the underlying performance effect. Everything I found tracing the "too prescriptive" line back (mindstudio.ai, kenhuangus.substack.com, alphasignalai.substack.com, findskill.ai, a garrytan/gstack GitHub issue) is downstream commentary quoting the same Anthropic sentence, not independent measurement. If the claim is read as an empirical performance assertion rather than a claim about what Anthropic's guidance says, it is vendor-only. e) PROVENANCE ANACHRONISM in the seed (does not affect the verdict, but flag it): the cited source is a tool table cached 2026-06-24, yet Fable 5.1 released 2026-09-01 (https://platform.claude.com/docs/en/models/fable-5-1/overview, Bloomberg/TechCrunch 2026-09-01). A June cache could not have carried 5.1-specific guidance; the guidance it carries is the Fable 5 guidance. VERDICT: not refuted. The claim is a claim about the content of Anthropic's public guidance, and that guidance is verifiable verbatim on platform.claude.com, so vendor-doc sourcing matches claim strength here. The only defect is that the sentence is scoped to Fable 5 and to "skills"; applying it to Fable 5.1 is an unrestated but consistent extension (5.1 is documented as extending Fable 5, and 5.1's own guide gives two concrete "remove instructions written for earlier models" cases). Correct restatement: Anthropic's Fable 5 prompting guide states that skills developed for prior models are often too prescriptive and can degrade output quality; the 5.1 guide does not restate this and says Fable 5 prompts carry over unchanged.
Counter sourcehttps://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 - the Fable 5.1-specific prompting guide contains no "too prescriptive"/degrade-quality statement, never mentions skills, and opens with "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes", which qualifies the claim's attribution of that guidance to 5.1. The verbatim sentence lives only on the Fable 5 page: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5
Verifier 2not refutedhigh confidence
SUPPORTED by public Anthropic documentation, with two precision caveats. DIRECT SUPPORT. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 (fetched 2026-09-02), section "Recommended scaffolding changes", bullet "Refactor existing prompts and skills": "Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions if default performance is better." This is near-verbatim the seed quote, so the cached skill table is reporting real published guidance, not inventing it. CAVEAT 1 (model naming). The published sentence names Claude Fable 5, not Fable 5.1. I fetched the 5.1-specific guide, https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1, and grepped its full text: it does NOT contain "prescriptive" or "degrade", has no "Recommended scaffolding changes" section, and never links back to the Fable 5 prompting guide. It instead opens: "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes." So the claim transfers a Fable-5 statement onto 5.1. That transfer is nonetheless sound on three independent grounds: (a) The same 5.1 page twice instructs removing prior-model instructions, just narrower ones: "Some earlier models were eager to give updates while working, which led to system prompt lines such as 'hold all findings for the final response.' Remove lines like that before adding anything"; and "Earlier models overused bullets and bold in chat, and many prompts carry anti-formatting rules written to hold that down... If your prompt contains anti-formatting language, remove it". (b) The cross-model page https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices explicitly states it covers Fable 5.1 and says: "Prefer general instructions over prescriptive steps. A prompt like 'think thoroughly' often produces better reasoning than a hand-written step-by-step plan. Claude's reasoning frequently exceeds what a human would prescribe." Same page: "Tune anti-laziness prompting: If your prompts previously encouraged the model to be more thorough or use tools more aggressively, dial back that guidance." (c) Since 5.1 inherits Fable 5 prompts unchanged per the 5.1 guide, guidance about pre-Fable-5 prompts applies transitively. CAVEAT 2 (dropped hedge). Docs say "can degrade output quality"; the claim states "reduce its output quality", dropping the modal. Mild overstatement of certainty, though "often" is preserved in both. PARTIAL COUNTER-EVIDENCE, checked and judged non-fatal. The best-practices page's "Migration considerations" section pushes the opposite direction on specificity: "Be specific about desired behavior", "Frame your instructions with modifiers... 'Include as many relevant features and interactions as possible'", "Request specific features explicitly". This is not a contradiction: it concerns specifying the desired OUTPUT, whereas the claim concerns prescribing the PROCEDURE. Anthropic's own replacement pattern (goal, reason, boundaries, self-verification, not step-by-step scripts) reconciles the two. SOURCE-QUALITY ASSESSMENT. The claim is a claim ABOUT what Anthropic's guidance states, not about the underlying effect, so first-party docs are the authoritative source and the marketing-grade discount does not bite. Searches for dissent returned only restatements of Anthropic's own doc (kenhuangus.substack.com, alphasignalai.substack.com, medium.com/@mehmet.ozel2701, mindstudio.ai, github.com/garrytan/gstack issue 1994) - these are echoes, not independent corroboration. No credible source disputes or qualifies the guidance. IMPORTANT LIMIT: I found NO independent controlled benchmark demonstrating that prescriptive prompts actually reduce Fable 5.1 output quality. If the claim were restated as "prescriptive prompts do reduce quality" rather than "Anthropic's guidance states X", it would be vendor-asserted only and I would refute it. CURRENCY. Both docs live and current as of 2026-09-02; Fable 5.1 released ~2026-09-01, so the guidance is days old, not stale.
Counter sourcehttps://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 (the 5.1-specific guide omits the "too prescriptive / degrade output quality" sentence entirely and states existing Fable 5 prompts need no changes); and https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices "Migration considerations" items 1-3, which advise MORE specificity ("Be specific about desired behavior", "Frame your instructions with modifiers", "Request specific features explicitly"). Both examined; neither refutes the claim, since the first is superseded transitively and reinforced by the same page's remove-earlier-model-instructions guidance, and the second concerns output specification rather than procedural prescription.
Verifier 3not refutedhigh confidence
PUBLIC EVIDENCE SUPPORTS THE CLAIM (found, not absent). Primary: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 - section "Recommended scaffolding changes", bullet headed "Refactor existing prompts and skills": "Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions if default performance is better." This is near-verbatim the seed quote and is on a first-party Anthropic docs page, not a social post or a cached tool table. TWO CAVEATS I tried to break the claim on, both survive: (1) Version mismatch. The verbatim sentence names Fable 5, not 5.1. I fetched https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 and grepped the full page: it contains NO occurrence of "prescriptive", "degrade", "prior model" or "scaffold". Its lead reads "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes" - i.e. 5.1 inherits Fable 5's guidance for anything older; it does not contradict it. The transfer to 5.1 is documented rather than inferred by the cross-model page https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices, whose opening line explicitly scopes it to "current Claude models, including Claude Fable 5.1", and which states: "Prefer general instructions over prescriptive steps. A prompt like 'think thoroughly' often produces better reasoning than a hand-written step-by-step plan. Claude's reasoning frequently exceeds what a human would prescribe." That same page gives 5.1-SPECIFIC instances of legacy instructions actively hurting output: "Claude Fable 5.1 already formats less than earlier models, so on that model a block like this can suppress structure the content needs. Remove it"; and "remove any instruction telling it to keep that text brief." https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 adds "remove any prompt line that tells it to hold findings for the final response." (2) Wording drift. The doc sentence says "Skills", the claim says "prompts"; the doc says "can degrade", the claim says "reduce". The bullet's own heading is "Refactor existing prompts and skills", so prompts are in scope; the hedge loss ("can" -> flat) is minor and "often" is preserved. SOURCE-QUALITY TEST: the claim is explicitly about what Anthropic's guidance STATES, not about an independently validated effect size, so first-party docs are the correct evidence class and the marketing-grade discount does not bite. If it were restated as a validated behavioural fact, it would be weaker - see counterSource. NO CONTRADICTING EVIDENCE FOUND. Searches for practitioners disputing it returned only restatements (kenhuangus.substack.com/p/claude-fable-5-what-changed-and-how; github.com/garrytan/gstack issue #1994 "audit skills against Anthropic's new prompting guide"), which corroborate rather than dispute. Not outdated: the Fable 5.1 docs are live as of 2026-09-02 and the best-practices page still carries the principle. Not cherry-picked: it is standing prompting guidance, not a single benchmark run.
Counter sourcePartial counter, not fatal: the Fable 5.1 prompting page (https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1) does NOT itself repeat the "too prescriptive / degrade output quality" line - it appears only on the Fable 5 page - and its own migration advice is "your existing Claude Fable 5 prompts should perform well ... without changes". So anyone citing this as a 5.1-branded statement is citing the 5.1-inherited version of a Fable 5 sentence. Second: every source is Anthropic's own documentation. I found no independent benchmark or third-party A/B test quantifying how much prescriptive scaffolding actually costs on Fable 5.1 - third-party writeups (Ken Huang Substack, MindStudio, garrytan/gstack #1994) restate the doc rather than test it. The claim about what the guidance says is verified; the underlying quality effect is vendor-asserted only.

The four objections

Raised against the post before the run started, then tested. Flip a card for the evidence and its sources.

O1objection partly holdsdecisive

Max cap weighting

Does a Fable 5.1 token draw the Claude Max 5-hour and weekly cap at the same rate as an Opus 5 token, or at a multiple that tracks the API price? If it is roughly 2x, the third-of-the-cost claim does not transfer to this plan at all.

evidence inF9

O1objection partly holds

Max cap weighting

  • Objection holds in substance: the API 'third of the cost' figure does not transfer to a Max plan, but the exact multiple is not public.
  • Anthropic's support article states Fable models 'draw from your plan's regular weekly usage limits and use them faster than other Claude models' and that on Max 'You can use up to 50% of your weekly usage limits on Fable models at no extra cost'; it publishes no multiplier and says nothing about effort level changing the draw rate.
  • The only number is blog and forum grade: in-app Claude Code text quoted on Hacker News in June 2026 said Fable 5 'Uses your limits ~2x faster than Opus', with users reporting steeper than 2x on agentic sessions; that predates Fable 5.1 and is not an Anthropic document.
  • Two consequences follow. First, the 50% ceiling alone rules Fable out as the primary execution model on Max whatever the effort level.
  • Second, if the meter is token-based with a fixed Fable weight, low effort still helps proportionally (CursorBench shows Fable 5.1 Low at 19,522 tokens/task vs Fable 5 High at 43,747), but against Opus 5 there is no low-effort token comparison, only Snorkel's 58% fewer output tokens at default effort; at a 2x weight that would be roughly 0.84x of Opus draw, which is an inference, not a measurement.
  • Pro plans and standard Team seats get no Fable within limits at all (usage credits only).
  • Claude Code's model-config page confirms Fable is not the account-type default on any plan and that Fable usage 'can bill to usage credits instead of drawing on your plan's included limits' depending on seat tier.
Sources
O2objection holds

Benchmark breadth

Is there evidence beyond CursorBench, and beyond parity with Fable 5 at high effort, that Fable 5.1 at low beats Opus 5 at xhigh on non-coding work: reviews, research synthesis, spec writing, long agentic runs?

evidence inF3F4

O2objection holds

Benchmark breadth

  • Objection holds: no public evidence shows Fable 5.1 at low effort beating Opus 5 at xhigh on non-coding work, and the nearest independent data points the other way.
  • Artificial Analysis puts Fable 5.1 low at index 58 for $0.77/task against Opus 5 max at 63 for $2.34/task: cheaper, lower scoring.
  • Beyond CursorBench the independent evidence is all coding or code review: Snorkel's Terminal-Bench+ run (Fable 5.1 61.5% pass@1, 58% fewer output tokens, 36% faster, but 18% vs Opus 5's 67% on build-and-dependency tasks, conclusion 'not a strict upgrade', and low effort not tested); CodeRabbit's review harness (Fable 5.1 vs Fable 5 only: latency +48.7%, precision +4.5pp, recall -1.0pp, identical tool-call counts at Low and High); Every's vibe check (subjective, all effort levels 'super good', no Opus comparison).
  • Nothing tests PRDs, spec writing, research synthesis or long agentic runs at low effort against Opus 5 xhigh.
  • A Hacker News comment citing the Fable 5.1 system card (pp.
  • 169-170, FrontierCode Extended) says the model scores best at medium effort, adds unrequested small changes at higher efforts, and its medium is below Fable 5's best at xhigh; that is a single forum comment and was not checked against the system card itself.
Sources
O3objection partly holds

Correctness at low effort

Do practitioners report skipped steps, shallower reviews, or more fix cycles in agent loops at low effort?

evidence inF5F6

O3objection partly holds

Correctness at low effort

  • Objection partly holds: quality is not intact at low effort, but the documented regressions are specific rather than the 'skips steps, more fix cycles' pattern the objection describes.
  • Anthropic documents that at low effort Fable 5.1 'is less likely than Claude Fable 5 to call a search or retrieval tool, and more likely to answer from memory', with the remedy being raising effort for those turns or adding a verification nudge; its own remedy text concedes the failure mode produces out-of-date answers that sound authoritative.
  • Artificial Analysis measures an 8-point index drop from Fable 5.1 max (66) to low (58).
  • Against that, CodeRabbit found tool-call counts identical at Low and High with no batching regression in their loop, and Every's tester found all effort levels good in practice.
  • No practitioner report of extra fix cycles in agent loops, skipped review steps, or shallower reviews at low effort was found, and several verifiers ran out of search budget before sweeping for them, so absence of reports is weak evidence.
Sources
O4objection partly holds

Runtime mechanics

Can a Claude Code session running as Opus spawn a Fable subagent today, or does the agent model fable error seen on 2026-07-07 still hold? If it holds, Fable low for execution only exists inside a Fable session.

evidence inF10

O4objection partly holds

Runtime mechanics

  • Objection holds in effect, with the failure mode changed. anthropics/claude-code issue #82252 (opened 2026-07-29, still open, no staff response, no workaround) reports that a subagent requested with model 'claude-fable-5' or the alias 'fable', directly or via a Workflow agent() helper, is silently served by claude-sonnet-5: meta.json and the agent's self-report show Fable, but every assistant message in the transcript carries model claude-sonnet-5, while a claude-opus-4-8 control was served correctly.
  • That is worse than the error observed on 2026-07-07 because nothing surfaces it; the only detection is comparing meta.json against message.model in the agent transcript.
  • The issue does not state the session model, so whether a Fable session spawning a Fable subagent works is unverified, and whether the bug persists for claude-fable-5-1 on Claude Code v2.1.255 (Fable 5.1 requires 2.1.250+) is unknown.
  • Claude Code's model-config page places no documented restriction on Fable as a subagent model and notes Fable is not the account-type default on any plan.
  • Until the issue closes, treat any 'Fable-low for execution' subagent as unverified and check the transcript model field.
Sources

The evidence behind the three decisive objections

Captured from the live pages on 2 September 2026. Click any capture to open it full size.

Anthropic support article: on Max plans you can use up to 50% of your weekly usage limits on Fable models, and they draw from the regular weekly limits faster than other Claude models
O1. Anthropic's plan page. The highlighted sentence: Max plans can use "up to 50% of your weekly usage limits" on Fable models, which "draw from your plan's regular weekly usage limits and use them faster than other Claude models". No multiplier is published. The article says Fable 5 and 5.1 work the same way.
Anthropic's Prompting Claude Fable 5.1 guide: existing Fable 5 prompts should perform well without changes; a listed behaviour is answering from memory instead of searching at low effort
O3. Anthropic's Fable 5.1 prompting guide. Two things on one screen: "existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes" (which is why the prompt-simplification tip does not hold as stated), and the listed behaviour "answers from memory instead of searching at low effort".
GitHub issue 82252 on anthropics/claude-code, open, area:agents: subagent model override claude-fable-5 or fable silently served by claude-sonnet-5
O4. Claude Code issue #82252, still open. A subagent requested as Fable was served by Sonnet 5 with no error; the metadata and the agent's own self-report both said Fable. Filed for Fable 5 on 29 July 2026; no staff reply. Whether it persists for Fable 5.1 is unknown.

What this means for the two rulings

A task arrives Judgment-heavy? diagnosis, architecture, PRD core, flagged review Mechanical only? reformat, extract, boilerplate Fable 5.1, top seat plans, rules, reviews the hard cases Max: at most 50% of the weekly limit low effort: less search, more memory Opus 5, xhigh: the execution seat implementation + fresh-context review the seat the launch tip wanted to move Sonnet 5, xhigh yes no no yes cannot take this seat: the 50% cap
The ladder as it stands after the evidence. Fable stays where a wrong answer costs a rebuild. Opus keeps the work, because the Max cap alone rules Fable out of the execution seat at any effort, before quality is even considered.

The post's real claim is that low effort on Fable 5.1 is not a correctness sacrifice. If that held, the effort floor was a correctness rule expressed as a number, and the number should have followed the correctness.

Ruling 1 stands

Execution rung is the latest Opus

  • Nothing in the evidence moves it, and one thing strengthens it for a reason the ruling did not originally have.
  • The ruling was made on where judgment is spent, given that Fable and Opus draw the same Max pool.
  • The 50% weekly ceiling adds a plan mechanic: Fable cannot carry the highest-volume seat even if you wanted it to.
  • Cycle reviews are the highest-volume role, and the low-effort retrieval regression lands hardest exactly there.

carried byO1F9F4F5

Ruling 2 stands

Effort: xhigh default, high is a hard floor

  • The evidence does not support dropping the floor, and it does not cleanly condemn low effort either.
  • What it shows is narrower: no public measurement compares Fable 5.1 at low effort with Opus 5 at xhigh on non-coding work, and the one documented low-effort regression is retrieval behaviour, not general quality collapse.
  • Anthropic's own advice is to step down to medium or low only once your evals show quality holds.
  • You have no such evals. Until you run them the floor is the cheaper mistake.

One honest qualification. The floor is a rule about correctness written as an effort number, and the number is not sacred. If a matched test on your own tasks shows Fable 5.1 at low or medium holding the gate, the number should move for those task types, and only those.

carried byO2O3F4F5S1

The test that would settle it

Nothing public will answer this for your task mix. A small matched eval will, and it costs about a morning.

  1. Pick two or three real tasks you have already completed well: one cycle review, one PRD prescriptive core, one build step.
  2. Write the correctness gate before you run anything, and make it mechanical wherever it can be. A gate written afterwards scores the run you got.
  3. Run each task twice: once as Opus 5 at xhigh, once as Fable 5.1 at low. Run the Fable arm inside a Fable session, so the seat is real.
  4. Read /usage before and after each run and record the cap percentage consumed, not the token count.
  5. On the Fable low arm, add the search nudge Anthropic recommends. Without it the documented retrieval regression is doing the work rather than the model.
  6. If any arm carries the Fable seat through a subagent, compare meta.json against the model field in the agent transcript before you trust the result.
  7. Score right answers per unit of cap consumed, per seat, per completed task. That is your metric, and no published benchmark reports it.

The 50% weekly ceiling does not depend on the test. Even a clean win for Fable 5.1 at low effort leaves it unable to carry the primary execution seat on this plan.

Anthropic's own routing advice, on screen

Captured from the live model page on 2 September 2026.

Claude Fable 5.1 overview page: for most workloads, start with Claude Opus 5; use Fable 5.1 for demanding reasoning or when your evals on Opus 5 at higher effort still fall short
Anthropic's Fable 5.1 overview. "For most workloads, start with Claude Opus 5", and move to Fable 5.1 "when your evals on Claude Opus 5 at higher effort still fall short". The vendor's own default agrees with keeping Opus in the execution seat.

Findings

Ten findings survived synthesis, ranked by confidence. The decisive one for this decision is F9, which is medium confidence because Anthropic publishes half of it and nobody publishes the other half.

high confidence
F1high confidencevote 3-0 on each of R1-R6Model facts and the three hard API blocks
The claimFable 5.1 is Anthropic's top generally available model (claude-fable-5-1, released 2026-09-01) at $10/$50 per MTok, 1M context, 128K max output, and Opus 5 / Sonnet 5 remain the newest models of their tiers as of 2026-09-02. Three API breaking changes are hard-blocked: thinking cannot be disabled or budgeted (400 at any effort), forced tool_choice any/tool returns 400 invalid_request_error, and assistant prefill returns 400. Priority Tier and default ZDR are unavailable.
EvidenceEvery fact verified verbatim on live Anthropic docs and cross-checked against llm-stats.com and LiteLLM's day-0 post; these are API contract and catalogue facts where the vendor is the authority. The rules written into CLAUDE.md this session (R1-R7) are accurate, with two attribution softenings: preserved-thinking 400s are enforced by default only for accounts created on or after 2026-08-31 and never affect Claude Code users, and the 'too prescriptive' guidance is published for Fable 5, not 5.1.
Sources

answersR1R2R3R4R5R6

F2high confidencevote 2-1The third-of-the-cost figure is Fable against Fable
The claimThe 'third of the cost' claim is real but intra-family: on CursorBench 3.2 Fable 5.1 Low scores 66.2% at $2.90/task against Fable 5 High at 66.5% and $8.77/task. It compares Fable to its own predecessor, is not reproducible (dataset unpublished), and on the same board Grok 4.6 Extra High beats Fable 5.1 Low on both score and cost.
EvidenceNumbers read directly from Cursor's live leaderboard on 2026-09-02, stable across two fetches; tokens per task fall to about 45% (43,747 to 19,522) while cost falls to 33%, so the 75% cache-read price cut does part of the work. Cursor is a commercial partner, and the claim carries no information about Opus 5.
Sources

answersS2

F4high confidencevote 3-0Anthropic routes you to Opus 5 first
The claimAnthropic's own routing and effort guidance argues against replacing Opus 5 xhigh with Fable 5.1 low: 'For most workloads, start with Claude Opus 5', move to Fable 5.1 only 'when your evals on Claude Opus 5 at higher effort still fall short'; Fable 5.1's default is high, 'gains over Claude Fable 5 are largest at xhigh and max', and stepping down to medium or low is advised only 'once your evals show quality holds'. Effort names do not map across models, so any sweep tuned on Fable 5 or Opus must be re-run.
EvidenceAll quotes verbatim across four Anthropic pages fetched 2026-09-02. This is a cost-default heuristic, not a capability ranking; the same docs list Fable 5.1 as the highest available capability. Snorkel's independent coding run supports task-specific selection rather than a strict hierarchy either way and did not test low effort.
Sources

answersO2O3

F5high confidencevote 3-0, 3-0, 2-1 across three phrasingsLow effort searches less and answers from memory more
The claimLow effort on Fable 5.1 has a documented regression that lands on this user's core work: it 'calls a search or retrieval tool less often' and 'answers from memory more often', with Anthropic's remedies being to raise effort for the affected turns (per-message, beta) or add a verification nudge to the system prompt. It is narrow and behavioural, not a general quality collapse, but it is the failure mode that matters most for research synthesis and reviews.
EvidenceDocumented verbatim on three Anthropic pages; a vendor disclosing a limitation of its newest model is the least marketing-suspect class of vendor statement. The same page says low effort is 'often competitive' with Opus and Sonnet on cost per task, so the regression must not be generalised into 'low effort is worse than Fable 5 overall'. No independent measurement of the magnitude exists.
Sources

answersO3

F6high confidencevote docs 3-0; blanket cost claim refuted 1-2Seven documented behaviour changes since Fable 5
The claimFable 5.1 changes default behaviour versus Fable 5 in seven documented ways: two raise cost and time (one tool call per turn in loops where the next read is only implied; whole-file rewrites for small edits), three reduce output volume (fewer progress updates, denser prose, less chat formatting), and two are quality-relevant (unmarked quotations in summaries; less retrieval at low effort). Each has a one-line prompting fix. Independent measurement (CodeRabbit) confirms the wall-clock cost (+48.7% per review task) but found output volume down, precision up 4.5pp, recall down 1.0pp, and no batching regression in its harness.
EvidenceThe seven behaviours are verbatim bullet headings in Anthropic's docs; the blanket claim that they all raise cost without touching quality was refuted because three cut output and two affect correctness. CodeRabbit is one harness and compares to Fable 5, not Opus 5. The token-volume half of a refuted claim matters here: Anthropic's 'token counts are roughly unchanged' is tokenizer parity, and the same guide says extra turns and always-on thinking can raise per-task tokens, so re-baseline cost on your own workloads.
Sources

answersO3

F7high confidencevote S4 0-3; R7 3-0The prompt-simplification list belongs to Opus 5
The claimThe prompt-simplification list in the X post (verification rituals, emphasis boosters, scratchpad scaffolds, stale few-shot examples, contradictory rules) is Opus 5 guidance relabelled as Fable 5.1. Anthropic's Fable 5.1 guide says existing Fable 5 prompts 'should perform well on Claude Fable 5.1 without changes', names only two legacy lines to remove (anti-narration and anti-formatting rules), and at low effort recommends adding a verification nudge. The Fable 5 guide, not the 5.1 guide, says skills for prior models 'are often too prescriptive' and 'can degrade output quality'.
EvidenceFull-text search of the Fable 5.1, what's-new and migration pages found no simplification guidance; the remove-verification sentence is on the Opus 5 page and the best-practices page calls Opus 5 'the exception'. For the CLAUDE.md rule that deletes self-check scaffolding: it is correctly grounded for Opus 5 seats; for Fable seats the docs point the other way at low effort and in long-run prompts.
Sources

answersS4R7

F8high confidencevote 3-0 on S3, S5 and pricing claimCache and effort mechanics
The claimCache and effort mechanics: Fable 5.1 cache reads are $0.25/MTok (0.025x base, a quarter of the 0.1x every other Claude model uses; Opus 5 is $0.50, Sonnet 5 $0.20), with base prices, cache writes and the 512-token minimum unchanged. Mid-conversation effort changes preserve the cache only via the beta per-message role:system message (header mid-conversation-output-config-2026-07-01), which Opus 5 also supports; a top-level effort change still restarts the cache, and Claude Code's /effort invalidates the entire conversation cache with no Fable 5.1 exception.
EvidencePricing verified on Anthropic's table and reported identically by VentureBeat; the cache-read cut is the concrete API lever for long cached agentic sessions and is worth nothing on a Max plan where usage is metered by plan limits. The practical Claude Code advice stays: pick effort at the top of a session.
Sources

answersS3S5

medium confidence
F9medium confidencevote n/a (fetched by synthesis for O1; no verifier vote)decisiveWhat a Max plan actually meters
The claimOn Claude Max, Fable models draw from the regular weekly limit faster than other Claude models and are capped at 50% of the weekly limit at no extra cost; Anthropic publishes no multiplier and says nothing about effort changing the rate. The only figure is in-app Claude Code text quoted on Hacker News in June 2026 ('Uses your limits ~2x faster than Opus', for Fable 5), with users reporting steeper draw on agentic sessions. Pro plans and standard Team seats get Fable only on usage credits.
EvidenceThe faster-draw statement and the 50% ceiling are official (support article, fetched 2026-09-02); the 2x number is blog and forum grade and predates Fable 5.1. The 50% ceiling is decisive on its own: Fable cannot be the primary execution model on Max whatever the effort level, and the API 'third of the cost' figure has no Max-plan equivalent.
Sources

answersO1

F3medium confidencevote S1 1-2; doc-attribution claim 1-2 with verifier evidence saying not refutedThe one independent cross-model reading
The claimOn the one independent cross-model measurement, Fable 5.1 at low effort beats Sonnet 5 at max on both score and cost (index 58 at $0.77/task vs 53-55 at $2.29) but scores below Opus 5 at max (63 at $2.34/task). Lance Martin's sentence is a verbatim quote of Anthropic's Fable 5.1 prompting guide; 'scoring higher' holds against Sonnet and fails against Opus 5.
EvidenceArtificial Analysis is the sole independent source and is a disclosed pre-release evaluation partner of Anthropic, though its headlines are consistently critical of Anthropic's cost story (Fable 5.1 max costs about 20% more per task than Fable 5). Its Opus 5 cost-per-task figures differ between its July and September articles ($2.03 vs $2.34), likely index-version drift. Verifiers split 1-2 on S1 largely because one refuter did not find the verbatim doc sentence another verifier did.
Sources

answersS1O2

F10medium confidencevote n/a (fetched by synthesis for O4; no verifier vote)Claude Code runtime and the silent substitution
The claimClaude Code runtime: default effort is high on every model except Opus 4.7 (xhigh); Fable is not the account-type default on any plan; Fable 5.1 needs Claude Code 2.1.250 or later. Issue #82252 (open since 2026-07-29, no staff response, no workaround) reports a 'fable' or 'claude-fable-5' subagent request being silently served by claude-sonnet-5 while meta.json and the agent's self-report claim Fable; a claude-opus-4-8 control was served correctly.
EvidenceModel-config facts are from Anthropic's Claude Code docs; the silent-substitution report is a single open GitHub issue about Fable 5 and does not state the session model, so its status for claude-fable-5-1 on v2.1.255 is unknown. Any Fable subagent's transcript model field should be checked until the issue closes.
Sources

answersO4

What did not survive

Six claims were killed by two or more verifiers. Two are seed claims from the post. Four are claims extracted from the sources during the sweep, labelled E1 to E4. Collapsed by default.

S1refutedvote 1-2seed claim from the postClaude Fable 5.1 at low effort is often competitive with Claude Opus and Claude Sonnet on cost per task while scoring higher than them.
Where it came fromX post by Lance Martin (@RLanceMartin, Anthropic), screenshot supplied by the user on 2026-09-02
Why it was killed
Verifier note 1
REFUTED: no public support, and Anthropic's own docs point the opposite way on three counts. 1) BASELINE SUBSTITUTION (the core defect). Anthropic's launch post (https://www.anthropic.com/claude-fable-and-mythos-5-1) says: "when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost." The comparator is FABLE 5 - the model's own predecessor - not Opus 5 or Sonnet 5. I checked the page specifically for cross-model effort charts: there are none comparing Fable 5.1 low effort to Opus/Sonnet on cost-per-task, and the model benchmark table is not broken down by effort level. The post under review swaps "Fable 5" for "Opus and Sonnet," which changes the claim materially: matching your own prior generation more cheaply is not the same as beating two cheaper models on $/task. 2) ANTHROPIC'S PUBLIC ROUTING ADVICE CONTRADICTS THE IMPLICATION. https://platform.claude.com/docs/en/models/fable-5-1/overview and .../whats-new-fable-5-1 both state verbatim: "For most workloads, start with Claude Opus 5... Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5 at higher effort still fall short." https://platform.claude.com/docs/en/about-claude/models/choosing-a-model says "Most workloads start with Claude Opus 5," listing Fable 5.1 only under "The highest available capability." If low-effort Fable 5.1 were cost-competitive with Opus/Sonnet AND higher-scoring, this guidance would be irrational. Anthropic also advises tuning effort WITHIN a model rather than switching to Fable: "Tuning effort is often a better lever than switching models." 3) "SCORING HIGHER" IS NOT WHAT THE EFFORT DOC SAYS. https://platform.claude.com/docs/en/build-with-claude/effort defines low as "Most efficient. Significant token savings with some capability reduction," use case "Simpler tasks... such as subagents." Fable 5.1 guidance: "Start with high, the default... step down to medium or low for routine or latency-sensitive work ONCE YOUR EVALS SHOW QUALITY HOLDS" - conditional quality PRESERVATION, not superiority. Tellingly, the cross-model boast that exists for Fable 5 ("Lower effort settings on Claude Fable 5 still perform well and often exceed xhigh performance on prior models") is ABSENT from the Fable 5.1 section - and even that one names "prior models," not Opus 5/Sonnet 5, and makes no $/task claim. 4) INDEPENDENT DATA RUNS THE OTHER WAY. Artificial Analysis (https://artificialanalysis.ai/models/claude-fable-5-1) lists Fable 5.1 at Intelligence Index 66 with "Cost per Intelligence Index task: $3.69". Their Opus 5 analysis (https://artificialanalysis.ai/articles/opus-5) reports Opus 5 (max) at $2.03/task vs Fable 5 at $2.75, Sonnet 5 (max) at $1.53 and Opus 4.8 (max) at $1.80 - headline: "Opus 5: Fable 5 level intelligence at a lower cost per task." AA also notes "at low and medium effort Sonnet 5 stays genuinely cheaper." Critically, AA publishes NO low-effort Fable 5.1 cost-per-task figure at all (their Fable 5.1 entry is "Adaptive Reasoning, Max Effort, Default Fallback"), so the specific claim is publicly UNMEASURED, not merely disputed. 5) PRICE ARITHMETIC MAKES IT A DEMANDING CLAIM. Per https://platform.claude.com/docs/en/models/fable-5-1/overview, Fable 5.1 is $10/$50 per MTok vs Opus 5 $5/$25 (2x) and Sonnet 5 $2/$10 (5x). To be "competitive on $/task" with Opus, Fable-low must consume <=~50% of Opus's tokens on the same task; vs Sonnet, <=~20%. No public source establishes those ratios. The one real cost lever - cache reads at $0.25/MTok (0.025x base vs 0.1x on other models) - only bites on long cached agentic sessions that re-read a prefix, which is a narrow regime being presented as general (checklist item 5, cherry-picking). 6) DOCUMENTED LOW-EFFORT REGRESSION ON THE KNOWLEDGE-WORK AXIS. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 lists under "Changed from Claude Fable 5": "Answers from memory more often at low effort. At the lowest effort level the model calls a search or retrieval tool less often." That is a named quality regression at precisely the setting the claim praises - direct counter-evidence to "scoring higher" for research/knowledge work. SOURCE QUALITY / AMBIGUITY: the claim is a vendor employee's X post, marketing-grade per the checklist, with no independent corroboration found for the low-effort-vs-Opus/Sonnet cost-per-task comparison. It is also ambiguous whether "Opus and Sonnet" means the current Opus 5/Sonnet 5 or prior generations (4.8/4.6); on the charitable prior-gen reading it may hold, but then it is not a claim about the current lineup and does not support routing decisions today. NOT REFUTED ON RECENCY - the claim is current (Fable 5.1 released 2026-09-01, per the docs) - it is refuted on absent public support, comparator substitution, and contrary independent cost-per-task measurement. COUNTER-SOURCES: https://platform.claude.com/docs/en/models/fable-5-1/overview | https://platform.claude.com/docs/en/about-claude/models/choosing-a-model | https://platform.claude.com/docs/en/build-with-claude/effort | https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 | https://www.anthropic.com/claude-fable-and-mythos-5-1 | https://artificialanalysis.ai/models/claude-fable-5-1 | https://artificialanalysis.ai/articles/opus-5
Verifier note 2
REFUTED AS A GENERAL CLAIM; one sub-case survives. The claim is a faithful restatement of the quote, so it is not a misread - but it fails as stated because "competitive on $/task while scoring higher" than Opus AND Sonnet is true only under one specific, unstated effort pairing, and flips against Opus 5 at other pairings. (1) Anthropic's own public documentation does NOT support the Opus/Sonnet cost comparison at all. On the launch page https://www.anthropic.com/claude-fable-and-mythos-5-1 the cost-vs-score charts (Terminal-Bench 4.0, Terminal-Bench-Science 0.1, Humanity's Last Exam, CursorBench 3.2.0) plot Fable 5.1 only against Fable 5 and Mythos. Opus 5 and Sonnet 5 appear in the raw benchmark score table but are ABSENT from every cost-per-task chart. The only effort-level cost statement Anthropic makes is "When set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5's at a much lower cost" - Fable 5 as the baseline, not Opus or Sonnet. So the vendor's own docs make a Fable-vs-Fable claim; the post extends it to Opus and Sonnet without vendor backing. (2) Independent data (Artificial Analysis, model pages fetched 2026-09-02) partially contradicts it. Fable 5.1 low effort: index 58, $0.77/task (https://artificialanalysis.ai/models/claude-fable-5-1-low). Opus 5 max: index 63, $2.34/task (https://artificialanalysis.ai/models/claude-opus-5). Opus 5 LOW: index 52, $0.43/task (https://artificialanalysis.ai/models/claude-opus-5-low). Sonnet 5 max: index 55, $1.72/task (https://artificialanalysis.ai/models/claude-sonnet-5). Fable 5.1 max: index 66, $3.76/task (https://artificialanalysis.ai/models/claude-fable-5-1). - vs Sonnet 5 (max): claim HOLDS - 58 @ $0.77 beats 55 @ $1.72 on both axes. - vs Opus 5 (low, matched effort): scores higher (58 vs 52) but costs 1.8x ($0.77 vs $0.43). "Competitive" is a stretch; "cheaper" is false. - vs Opus 5 (max): scores 5 points LOWER (58 vs 63). "Scoring higher" is false. So the claim's truth depends entirely on which Opus configuration is silently chosen, and it is false against the configuration a top-seat agentic-coding user would actually run. (3) The independent headline runs the opposite way from the post's framing. Artificial Analysis's published summary (https://x.com/ArtificialAnlys/status/2094881171066978525, reported at https://cryptobriefing.com/claude-fable-5-1-tops-intelligence-index/) is "Claude Fable 5.1 tops the Intelligence Index but costs 20% MORE per task than Fable 5", with Fable 5.1 max at 1.6x Opus 5 max cost, driven by more verbose output on hard problems. AA's earlier Opus 5 writeup is literally titled "Opus 5: Fable 5 level intelligence at a lower cost per task" (https://artificialanalysis.ai/articles/opus-5). Note AA ran pre-release evals FOR Anthropic, so it is not fully arms-length - but it published cost findings unfavourable to Anthropic, which raises rather than lowers its credibility here. (4) Source quality is insufficient for the claim's strength. Lance Martin is an Anthropic employee posting on X about his employer's launch-day model; per the checklist that is marketing-grade. It generalises ("often", "Opus and Sonnet") from a single effort level on a single composite benchmark, with the comparison effort level for the baselines left unstated - textbook cherry-pick presented as general. (5) Not outdated (Fable 5.1 launched ~2026-09-01), and one further caveat: AA's "$/task" is a benchmark-suite average, not cost-per-COMPLETED-task in agentic work. Fable 5.1's documented behaviour (long turns, always-on thinking, more verbose output) means the metric may not transfer to real agentic loops, and it is irrelevant to Claude Max plan limits, which meter usage windows rather than dollars. Verdict: refuted as worded. Defensible narrower version: "Fable 5.1 at low effort dominates Sonnet 5 on both score and cost per task on the AA Intelligence Index, and outscores Opus 5 at matched low effort for ~1.8x the cost." Do not repeat the Opus half without naming the effort level.
S4refutedvote 0-3seed claim from the postAnthropic's published guidance for Claude Fable 5.1 is to simplify prompts by removing verification rituals, emphasis boosters, scratchpad scaffolds, stale few-shot examples and contradictory rules.
Where it came fromX post by Lance Martin (@RLanceMartin, Anthropic), screenshot supplied by the user on 2026-09-02
Why it was killed
Verifier note 1
REFUTED by public Anthropic docs on two grounds: misattribution to the wrong model, and direct contradiction on the headline item. 1. The dedicated page https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 has 17 sections (effort levels, progress updates, tool-call batching, append-only history, writing density, formatting, quoting sources, finishing the task, compaction summaries, scope/tests, search triggering at low effort, safeguard false positives, targeted edits, long outputs, subagents, vision). NONE is about simplifying prompts. A full-text grep of the page for "simplif", "verification ritual", "emphasis", "ALL CAPS", "scratchpad", "few-shot", "contradict", "conflicting", "self-check", "double-check", "booster", "stale" returns no matching guidance. Its opening line says the opposite of the claim: "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes." Same grep over https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 and https://platform.claude.com/docs/en/models/fable-5-1/migration-guide: no such guidance. 2. The "remove verification instructions" advice IS published, but Anthropic attributes it explicitly to Claude OPUS 5, not Fable 5.1. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 ("Task scope and over-verification"): "Claude Opus 5 verifies its own work without being told to. If your prompt contains explicit verification instructions ... remove them." https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices repeats it as an Opus-5-only carve-out: "Claude Opus 5 is the exception ... When migrating to Claude Opus 5, remove these instructions rather than rewriting them." 3. For the Fable line Anthropic advises the OPPOSITE on verification. Fable 5.1 prompting page, "Search triggering at low effort" (the exact low-effort case in the research question): "In other cases, a prompt nudge toward verification helps" - and it supplies a verification instruction to ADD. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5, "Recommended scaffolding changes": "Make self-verification explicit in long-run prompts ... instruct: 'Establish a method for checking your own work at an interval of [X] as you build.'" 4. Three of the five listed items have no public support at all for any model: scratchpad scaffolds (the only "scratchpad" mention in the best-practices page is POSITIVE - "Using temporary files can improve outcomes particularly for agentic coding use cases"), stale few-shot examples (best practices calls examples "one of the most reliable ways to steer Claude's output format, tone, and structure"), and contradictory rules (no such guidance found). "Emphasis boosters" is real but is general, older guidance dated to Sonnet 4.5, not Fable 5.1 specific. 5. Partial support exists only for a generic "refactor stale prompts" theme, and it is on the Fable 5 page, not 5.1: "Skills developed for prior models are often too prescriptive for Claude Fable 5 ... Review and consider removing older instructions if default performance is better." The Fable 5.1 page does tell you to remove two specific legacy lines (anti-narration "hold all findings for the final response", and anti-formatting rules). That is narrower than, and does not include, the claim's five-item list. Source quality: the claim's only source is a vendor employee's X post (marketing-grade, uncorroborated). The strongest public evidence available contradicts its model attribution rather than corroborating it. The list reads as an Opus 5 guidance summary relabelled as Fable 5.1.
Verifier note 2
REFUTED as worded. The claim attributes a specific five-item removal list to "Anthropic's published guidance for Claude Fable 5.1." I read the three governing primary docs in full. That list is not in any of them, and two of its five items are contradicted for the Fable line. (1) https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 - the model-specific page. Grepped the full 52.7KB text for "verification ritual", "emphasis", "ALL CAPS", "scratchpad", "few-shot", "contradict", "booster": zero relevant hits. Its only "remove" advice is two narrow, unrelated items: remove narration-suppression lines ("system prompt lines such as 'hold all findings for the final response.' Remove lines like that before adding anything") and remove anti-formatting rules ("If your prompt contains anti-formatting language, remove it"). Neither is on the claimed list. The page's opening line runs the other way: "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes." Its "Search triggering at low effort" section tells you to ADD a verification nudge, not remove one. (2) https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices - the cross-model reference - pins the remove-verification guidance to a different model and explicitly marks it an exception: "Ask Claude to self-check. Append something like 'Before you finish, verify your answer against [test criteria].' This catches errors reliably... Claude Opus 5 is the exception: it verifies its own work well without explicit instruction... When migrating to Claude Opus 5, remove these instructions rather than rewriting them." Fable 5.1 is not the exception, so the general self-check rule still applies to it. The same page endorses the other two items the claim says to strip: "Examples are one of the most reliable ways to steer Claude's output format, tone, and structure" / "Multishot examples work with thinking"; and on scratchpads, "Claude to use files... as a 'temporary scratchpad'... can improve outcomes particularly for agentic coding use cases." (3) https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 - the Fable-line page - says the opposite of "remove verification rituals": "Make self-verification explicit in long-run prompts. Separate, fresh-context verifier subagents tend to outperform self-critique." It also tells you to BUILD a scratchpad: "Construct a memory system... Provide a place to write notes, as simple as a Markdown file." (4) https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - no prompt-simplification guidance at all; its migration checklist is API-mechanical (tool_choice, append-only history, effort re-tune, tool batching). What IS publicly supported is only the general direction, and only for Fable 5, not 5.1: "Refactor existing prompts and skills. Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions if default performance is better." Plus one specific removal not on the claimed list: strip show-your-thinking instructions, because they trigger the reasoning_extraction refusal category. Source-quality and cherry-picking failures: the seed's only backing is a vendor employee's X post; searches surfaced no Anthropic-published text containing "verification rituals", "emphasis boosters", or the five-item taxonomy. Corroboration comes only from SEO content-farm restatements of the same launch material (mindstudio.ai, knightli.com, kenhuangus.substack.com, aidailycheck.com) - downstream of the vendor, not independent. Search summarisers themselves conflate the models, writing "This principle applies to Fable models as well" as an inference the docs do not make. The post appears to compress Anthropic's Opus 5 guidance and Fable 5's "too prescriptive" line into one punchy list, then label it Fable 5.1 - a paraphrase that inverts published guidance on verification and scratchpads for the very model named.
Verifier note 3
REFUTED: the five-item list is mis-attributed to Fable 5.1. It matches Anthropic's OPUS 5 guidance, and for the Fable family the docs say the opposite on verification. 1) Verification - CONTRADICTED for Fable 5.1. The "remove verification instructions" advice is published only for Opus 5. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 - "Task scope and over-verification": "Claude Opus 5 verifies its own work without being told to. If your prompt contains explicit verification instructions ... remove them"; "Self-correction": "Avoid instructing re-checks it already performs ('double-check your answer,' 're-verify before responding')". https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices pins this as a single-model exception: "Ask Claude to self-check. Append something like 'Before you finish, verify your answer against [test criteria].' This catches errors reliably... Claude Opus 5 is the exception... When migrating to Claude Opus 5, remove these instructions rather than rewriting them." For Fable, https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 ("Recommended scaffolding changes") says the reverse: "Make self-verification explicit in long-run prompts. Separate, fresh-context verifier subagents tend to outperform self-critique... instruct: 'Establish a method for checking your own work at an interval of [X]...'" And the Fable 5.1 page itself ADDS verification at low effort: "Search triggering at low effort... a prompt nudge toward verification helps" - directly relevant since the research question is about Fable 5.1 at LOW effort. 2) Scratchpads - CONTRADICTED. Best-practices page: temporary files as a "'temporary scratchpad'... can improve outcomes particularly for agentic coding use cases." 3) Few-shot - CONTRADICTED in direction. Same page: "Examples are one of the most reliable ways to steer Claude's output... A few well-crafted examples (known as few-shot or multishot prompting) improve accuracy and consistency," plus "Multishot examples work with thinking." 4) Emphasis boosters and contradictory rules - NO public support found. I read the full Fable 5.1 prompting page (https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1); its 16 sections are effort levels, progress updates, tool-call batching, append-only history, writing density, chat formatting, quoting sources, finishing the task, compaction summaries, scope/tests, search triggering, safeguard false positives, targeted edits, long outputs, subagents, vision. No "simplify your prompts" section, and no occurrence of verification rituals, emphasis boosters, scratchpad, few-shot, or contradictory/conflicting rules. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 contains no prompt-simplification guidance either. Partial kernel that DOES exist (why the post is not pure invention): Fable 5 page - "Refactor existing prompts and skills. Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions"; Fable 5.1 page - "audit your prompt for instructions that suppress narration... Remove lines like that" and "If your prompt contains anti-formatting language, remove it." So a general subtraction theme is real, but it is scoped to narration-suppression, anti-formatting rules and over-prescriptive legacy skills - not to the claim's five named categories. Source quality: an Anthropic employee X post is marketing-grade per the checklist. The only corroboration for the five-item framing is secondary commentary (mindstudio.ai, kenhuangus.substack.com, findskill.ai, aidailycheck.com) paraphrasing the docs, not the docs themselves. Cherry-picking flag: the post generalises Opus-5-specific guidance across the whole Claude 5 line, where Anthropic explicitly warns "Where a technique names a specific model, treat it as measured on that model and re-check it against your own evals before applying it to another."
E1refutedvote 1-2extracted claimFable 5.1 has documented behaviour regressions versus Fable 5 that raise token cost and wall-clock time without lowering answer quality: it may issue one tool call per turn where Fable 5 batched several, it rewrites whole files instead of making targeted edits, it writes fewer user-facing progress updates (especially at higher effort), its prose is denser, it uses less chat formatting, and it more often reproduces source passages in summaries without marking them as quotations.
Why it was killed
Verifier note 1
SPLIT VERDICT - the six enumerated behaviours are verbatim-documented and survive; the umbrella clause binding them ("raise token cost and wall-clock time without lowering answer quality") is an overreach and is contradicted for most of the list. WHAT CHECKS OUT (primary: https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1, "Behavior differences > Changed from Claude Fable 5", fetched 2026-09-02). All six items appear verbatim as bullet headings: "Parallel tool calling is more variable" ("may issue one tool call per turn where Claude Fable 5 batched several ... The extra turns cost tokens, round trips, and wall-clock time but don't reduce answer quality"); "Fewer progress updates during long tool runs" ("less user-facing text between tool calls, especially at higher effort"); "Denser prose in places"; "Less formatting in chat"; "Unmarked quotations in summaries"; "Whole-file rewrites for small changes" ("more likely to rewrite the entire file than make a targeted edit ... costs more output tokens and time"). Each is re-stated on https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 with a prompting fix. Source quality is fine here: this is the vendor documenting its own model's downsides (admission against interest), and the claim is only about what is DOCUMENTED. Current, not cherry-picked (it is the doc's full list, minus one item - see below). WHY REFUTED ANYWAY - three defects: 1. OVERREACH ON THE QUOTE. The supporting quote licenses "raises cost/time, quality unaffected" for exactly ONE item (tool batching); the doc adds a weaker "The result is usually the same" for whole-file rewrites. The claim generalises it across all six. It does not hold: fewer progress updates, denser prose and less chat formatting all REDUCE output tokens, they do not raise them. The doc never says these leave quality intact - it warns the opposite for formatting ("anti-formatting rules ... can suppress structure the content needs") and treats unmarked quotations as an attribution defect requiring a corrective example in the system prompt, which is a quality/correctness problem, not a cost one. 2. INDEPENDENT DATA CONTRADICTS THE COST HALF. CodeRabbit's own harness benchmark (https://www.coderabbit.ai/blog/fable-5-1-model-review) corroborates the wall-clock half (18:38 vs 12:32 per review task, +48.7%) but measured output volume DOWN, not up - 87 fewer final comments and 70.2% fewer nitpicks vs Fable 5 - and found NO batching regression in their loop ("completed the same number of tasks with the same number of calls", ~2.0 review-file calls per task at both Low and High). Their quality result is also mixed rather than flat: precision +4.5pp (32.8%->37.3%), recall -1.0pp (61.9%->61.0%). So "raises token cost" is unsupported in the one independent measurement available, and the batching regression is harness-dependent, not universal - the doc itself scopes it to loops "where the next independent calls are implied" and says explicitly "Requests naming several things to fetch still run in parallel." 3. MATERIAL OMISSION THAT BREAKS "WITHOUT LOWERING ANSWER QUALITY". The same doc bullet list contains a seventh change the claim drops: "Answers from memory more often at `low` effort ... the model calls a search or retrieval tool less often." That is a straightforward accuracy regression, and it is the item most directly on point for this research question (Fable 5.1 at LOW effort). Its omission is what makes the blanket quality-unaffected framing possible. Minor: "regressions" is a strong label for what the docs class as "Behavior differences", each with a one-line prompting fix, alongside "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes." USABLE REFORMULATION: "Anthropic documents seven default-behaviour changes in Fable 5.1 vs Fable 5. Two raise cost/time (harness-dependent one-call-per-turn batching; whole-file rewrites) and are documented as quality-neutral. Three reduce output volume (fewer progress updates, denser prose, less formatting). Two are quality-relevant, not cost-relevant (unmarked quotations in summaries; less search/retrieval at low effort). All have documented prompting fixes."
Verifier note 2
VERDICT: the enumerated facts check out verbatim; the claim's framing sentence is an overreach. Refuted on checklist item 1 (overreach), not on source quality or currency. WHAT SURVIVES. All six behaviours are verbatim in the "Changed from Claude Fable 5" section of https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 - "Parallel tool calling is more variable" (one call per turn where Fable 5 batched several), "Fewer progress updates during long tool runs" ("especially at higher effort"), "Denser prose in places", "Less formatting in chat", "Unmarked quotations in summaries", "Whole-file rewrites for small changes". The supporting quote is accurate. Source is current (Fable 5.1 shipped ~1 Sep 2026; checked 2 Sep 2026) and is primary API reference documentation, not a press release. It is also a vendor disclosing NEGATIVE behaviour about its own model, i.e. a statement against interest - the strongest grade of vendor evidence, so the checklist's "marketing-grade" discount largely does not bite here. Independent corroboration on the time cost exists: https://www.coderabbit.ai/blog/fable-5-1-model-review measured 18m38s vs 12m32s per review task (+48.7% latency) with essentially flat recall (61.0% vs 61.9%). WHY IT STILL FAILS. (1) The framing "regressions ... that raise token cost and wall-clock time" is true of only 2 of the 6 listed items. Anthropic attributes token/round-trip/wall-clock cost ONLY to tool batching ("The extra turns cost tokens, round trips, and wall-clock time") and whole-file rewrites ("a rewrite costs more output tokens and time", https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1). Denser prose, less formatting, and unmarked quotations carry no cost claim at all, and FEWER progress updates would if anything reduce output tokens - the claim attributes a cost property to items where the source asserts the opposite direction or nothing. (2) "Without lowering answer quality" is asserted by Anthropic for exactly one item - batching ("This doesn't affect answer quality") - and is HEDGED for rewrites ("The resulting file is usually the same", "usually" doing real work). It is contradicted within the claim's own list: "Unmarked quotations in summaries" is an attribution/correctness defect, not a cost defect. The prompting guide's fix for it is a full worked example whose rationale demands "one short marked phrase from one source; every other claim is reworded" - that is a quality remedy, not a token-saving one. (3) Selective omission that matters for this research question. The same docs list a SEVENTH change the claim drops: "Answers from memory more often at `low` effort. At the lowest effort level the model calls a search or retrieval tool less often." Since the question is specifically Fable 5.1 AT LOW EFFORT, omitting the one documented change that degrades factual grounding while asserting "without lowering answer quality" is a material misrepresentation. (4) Materially missing qualifier: each behaviour has a documented one-line prompting fix, and the guide states the edit instruction "brings Claude Fable 5.1 back in line with Claude Fable 5" and that "Your existing Claude Fable 5 prompts should perform well on Claude Fable 5.1 without changes". Anthropic frames these as "Behavior differences", never as regressions. The cost increase is default-behaviour-only and mitigable - the claim presents it as an unqualified property of the model. (5) The vendor-favourable half ("quality unchanged") has only one independent data point (CodeRabbit, a single code-review domain, and that same test did not record token totals: "this evaluation did not record reliable input and output token totals"). No independent source I found measured the token-cost claim at all. SALVAGE: the list of six behaviours is safe to state as documented fact with the URL. The sentence wrapping it is not - cost/time applies to tool batching and whole-file rewrites only, quality-neutrality is Anthropic's assertion for batching alone, and the low-effort search-triggering change must be included rather than dropped.
E2refutedvote 1-2extracted claimAnthropic's own docs state that at `low` effort, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, and recommend including it wherever you would otherwise run a smaller model at higher effort.
Why it was killed
Verifier note 1
QUOTE IS REAL AND VERBATIM. I fetched https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 and the sentence exists word-for-word in the "Consider all effort levels" section: "At `low`, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, so include it in the comparison wherever you'd otherwise run a smaller model at a higher effort level." So checklist item 1 (overreach vs quote) mostly passes, and item 4 (outdated) fails - the page is current. The claim is refuted on items 2, 3 and 5. (A) MATERIAL OMISSION OF THE SURROUNDING CONDITIONALITY. The docs sentence is not a standalone recommendation; it sits inside an explicitly eval-gated paragraph: "Start at the default effort level, `high`, then test the other levels against your own evals... step down to `medium` or `low` where your evals show quality holds." And the operative verb is "include it in the COMPARISON" - i.e. add Fable 5.1 as a candidate in your own sweep, not "it wins." The claim drops "in the comparison" and drops the eval gate, converting a hypothesis-to-test into a general recommendation. The same page also warns "effort level names don't correspond to the same amount of thinking across models," which makes the cross-model low-effort comparison non-portable by Anthropic's own admission. (B) CHERRY-PICK / SUPPRESSED CAVEAT ON THE SAME PAGE. Two paragraphs later the identical doc states that at `low`, Fable 5.1 "calls search and retrieval tools less often" (its own "Search triggering at low effort" section). That is a documented behavioural regression at exactly the setting being recommended, and it bears directly on agentic and research work - the use case in the research question. Recommending low effort while omitting this is a one-sided read. (C) VENDOR-ONLY; NO INDEPENDENT CORROBORATION EXISTS FOR THE LOW-EFFORT CLAIM. Artificial Analysis (https://artificialanalysis.ai/articles/claude-fable-5-1), which ran Anthropic's own pre-release eval, publishes Fable 5.1 scores across all five effort levels (58 at low -> 66 at max, 11x output-token spread, 13.1M->143.7M) but publishes NO cost-per-task figure for low, medium or high. The specific assertion is therefore untested by the leading independent per-task cost benchmark. (D) INDEPENDENT DATA POINTS THE OTHER WAY AT EVERY LEVEL ACTUALLY PRICED. Same AA article: "At xhigh effort Fable 5.1 scores 65 at $2.72 per task... but still above Claude Opus 5 (max, 63) at $2.34"; Fable 5.1 max is 66 at $3.76, i.e. ~60% more per task than Opus 5. AA's Opus 5 write-up (https://artificialanalysis.ai/articles/opus-5) is titled "Opus 5: Fable 5 level intelligence at a lower cost per task" and reports Opus 5 (max) $2.03/task vs Fable 5 $2.75, Sonnet 5 (max) $1.53 - the exact inverse of the direction Anthropic's docs claim. On scores, AA puts Fable 5.1 at low = 58 versus Opus 5 max = 63, so "while scoring higher" is FALSE against Opus 5 on the leading independent index; it survives only against Sonnet 5 (53). The claim's "Opus and Sonnet models" is an unspecified comparison class doing a lot of work. (E) THE ONE INDEPENDENT AGENTIC-CODING BENCHMARK QUALIFIES IT FURTHER. Snorkel AI (https://snorkel.ai/blog/fable-5-1-vs-opus-5-coding-benchmark/) tested NO effort levels at all, and its verdict cuts against the quality half: Fable 5.1 posts 61.5% pass@1; on the matched set both solved 18, "Opus alone solved 5, Fable alone solved 2"; conclusion quoted verbatim - "Fable is not a strict upgrade in this evaluation. It is faster and more token-efficient, with clear wins on specific tasks, but Opus remains more robust across the full matched set." Cheaper per successful run, but loses more attempts outright. VERDICT: the quote is authentic, but it is a marketing-grade vendor assertion published with zero supporting numbers, stripped of its eval gate and of the low-effort search-regression caveat, uncorroborated by any independent source, and contradicted in direction by the only independent cost-per-task benchmark (Fable 5.1 costs MORE per task than Opus 5 at both effort levels AA priced) and in the "scoring higher" half (Fable 5.1 low = 58 vs Opus 5 = 63). Safe to say "Anthropic's docs claim X"; not safe to state X as established.
Verifier note 2
QUOTE IS REAL BUT THE CLAIM OVERREACHES IT, AND THE SUBSTANCE IS VENDOR-ONLY AND PARTLY CONTRADICTED. 1. Quote verified verbatim. Fetched https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 - the sentence exists exactly as given, in the "Consider all effort levels" section: "At `low`, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, so include it in the comparison wherever you'd otherwise run a smaller model at a higher effort level." So the attribution half is accurate. 2. OVERREACH: the claim drops "in the comparison." Anthropic recommends adding Fable-5.1-at-low to your EVAL COMPARISON SET, not deploying it there. The claim's "recommend including it wherever you would otherwise run a smaller model at higher effort" reads as a routing recommendation. The same paragraph is explicitly eval-conditional: "test the other levels against your own evals" and "step down to `medium` or `low` where your evals show quality holds." The claim strips that conditionality. 3. SAME-SOURCE CONTRADICTION of the implied routing. https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 states the opposite default: "For most workloads, start with Claude Opus 5... Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5 at higher effort still fall short." 4. SAME-PAGE CAVEAT OMITTED, at exactly the effort level and workload in question. The prompting page's "Search triggering at low effort" section: at `low`, Fable 5.1 "is less likely than Claude Fable 5 to call a search or retrieval tool, and more likely to answer from memory" - a documented regression for agentic/research work, with the recommended fix being to RAISE effort. 5. INDEPENDENT DATA CONTRADICTS THE "SCORING HIGHER" HALF. Artificial Analysis (https://artificialanalysis.ai/articles/claude-fable-5-1): Fable 5.1's five effort settings "score from 58 to 66" on the Intelligence Index, i.e. 58 at low. Claude Opus 5 at max scores 63 (~$2.03-$2.34/task) and Sonnet 5 at max is ~$1.53/task. On that independent index, Fable-at-low scores BELOW Opus 5 at max - precisely the "smaller model at a higher effort level" comparison the doc invites. AA publishes no low-effort cost-per-task figure, so the cost half is not independently verified either way. 6. NEAREST INDEPENDENT CODING EVAL QUALIFIES RATHER THAN CORROBORATES. Snorkel AI Terminal-Bench+ (https://snorkel.ai/blog/fable-5-1-vs-opus-5-coding-benchmark/) tests NO effort levels and reports NO cost per task; it finds Fable 5.1 uses 58% fewer output tokens and is 36% faster but "loses more attempts outright" - Opus 5 alone solved 5 extra tasks vs Fable's 2, and Build/Dependency Management was 18% (Fable) vs 67% (Opus). 7. NO INDEPENDENT REPRODUCTION EXISTS of the specific low-effort-vs-Opus/Sonnet cost-per-task comparison. https://www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes quotes the same Anthropic line but states there is no independent effort-level breakdown and treats it as a vendor statement needing customer validation. The widely circulated "~26% at ~$11/task at low vs Fable 5 max 24.7% at ~$44" figure is an INTRA-FAMILY comparison (Fable 5.1 vs Fable 5), not against Opus or Sonnet, so it does not support the claim. Not outdated (Fable 5.1 is current as of 2026-09-02). Fails on source strength (vendor performance/cost claim, no independent corroboration), on paraphrase overreach, and on selective omission of the same source's own low-effort caveat and default-to-Opus-5 guidance.
E3refutedvote 0-3extracted claimFable 5.1 costs $10/$50 per million input/output tokens, double Opus 5's $5/$25; its prompt cache reads are $0.25 per million tokens, half the Opus 5 cache-read rate and a quarter of the Fable 5 rate. Anthropic tells migrators token counts are roughly unchanged, so per-task cost differences come from per-token pricing, not token volume.
Why it was killed
Verifier note 1
PRICING NUMBERS: fully confirmed, no refutation possible. https://platform.claude.com/docs/en/about-claude/pricing model table gives Fable 5.1 = $10/MTok input, $50/MTok output, cache hits $0.25/MTok; Opus 5 = $5/MTok, $25/MTok, cache hits $0.50/MTok; Fable 5 = $10/$50 with cache hits $1/MTok. So "double Opus 5", "half the Opus 5 cache-read rate" and "a quarter of the Fable 5 rate" are all exactly right, with the footnote "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier." The migration guide (https://platform.claude.com/docs/en/models/fable-5-1/migration-guide) repeats it in item 6 of the Opus-5 section and, in the Fable-5 section, "a quarter of the Claude Fable 5 rate" (the part the supplied quote did NOT cover, but the doc does). Currency is fine: the pricing page carries a note dated to the 1 Sep 2026 Sonnet 5 price decision, so it is live as of 2026-09-02. Vendor-grade sourcing is not a defect here, since Anthropic is definitionally authoritative on its own list prices. REFUTATION is confined to the claim's final sentence, the causal inference "so per-task cost differences come from per-token pricing, not token volume." That is the claim's own reasoning, not Anthropic's, and the SAME migration guide contradicts it on exactly the workload the research question asks about: (1) Behavior changes, item 1: "In long-running loops where the next independent reads are only implied by the task (custom coding agents, bash-and-editor harnesses, computer use), Claude Fable 5.1 may issue one tool call per turn. Each extra turn costs tokens, a round trip, and wall-clock time." Anthropic states a per-task token VOLUME increase for agentic coding and prescribes a batching instruction to mitigate it, which would be pointless if volume were fixed. (2) Opus 5 migration item 1: Opus 5 accepts thinking:{type:"disabled"} at effort high or lower; Fable 5.1 returns 400 at any effort level, and the checklist says to "Control token spend with lower effort levels, and revisit max_tokens" for "workloads that ran with thinking disabled." A thinking-disabled Opus 5 workload emits strictly more output tokens on Fable 5.1: volume, not price. (3) Scope misread: line 41 of the guide says the "tokenizer" matches Fable 5, and the pricing page separately notes 4.7-and-later models use a tokenizer producing "approximately 30% more tokens for the same text." "Token counts are roughly unchanged" therefore reads as tokenizer/accounting parity for the same text, not a promise that a completed task consumes the same tokens. The claim upgrades tokenizer parity into a per-task cost conclusion. (4) The checklist sentence is an instruction to MEASURE ("Re-baseline cost on your own workloads"), which the claim inverts into a reason volume can be ignored. Anthropic also tells migrators to run a fresh effort sweep starting at high rather than carrying over a tuned setting, and effort directly governs thinking-token volume. Note the volume effects run both ways, so the claim is not merely optimistic but simply wrong that volume is constant: behaviour changes 2 and 3 (fewer progress messages, shorter agentic coding summaries, fewer search/retrieval calls at low effort, the low-effort case in the research question) push volume DOWN, while change 1 pushes it UP. Separately, "double Opus 5" is an incomplete headline for a cached agent loop: per-token, Fable 5.1 cache reads are HALF Opus 5's ($0.25 vs $0.50), so the effective multiple ranges roughly 0.5x to 2x depending on cache-hit mix, which again supports re-baselining rather than the claim's conclusion. Caveat on method: this session's WebSearch budget was exhausted (200/200), so no independent/practitioner corroboration was obtainable. It was not needed for the refutation, which rests on Anthropic's own document contradicting the claim's inference, and the numeric half is vendor-authoritative. Recommended repair: keep the four pricing figures verbatim, drop the "so ... not token volume" clause, and replace it with Anthropic's actual instruction to re-baseline on your own workloads, flagging one-tool-call-per-turn and always-on thinking as agentic-coding volume risks.
Verifier note 2
SPLIT VERDICT - the pricing numbers are fully verified; the causal inference in the final sentence is contradicted by the same primary source. Refuted on the inference, which is the operationally load-bearing part for "cost per completed task". CONFIRMED (independent of the quoted page). https://platform.claude.com/docs/en/about-claude/pricing model-pricing table: Claude Fable 5.1 = $10/MTok input, $50/MTok output, cache hits/refreshes $0.25/MTok; Claude Opus 5 = $5/MTok, $25/MTok, cache hits $0.50/MTok; Claude Fable 5 cache hits $1/MTok. Footnote: "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier." So $10/$50 is exactly double $5/$25; $0.25 is exactly half Opus 5's $0.50 and a quarter of Fable 5's $1.00. All four numeric assertions hold, and are dual-sourced (pricing page + migration guide line 1592, which matches the supporting quote verbatim). No long-context surcharge undercuts the "double": "Claude 4.6 and later models ... include the full 1M token context window at standard pricing." Source quality is adequate here - for list prices the vendor is the authority, not marketing-grade. REFUTED - the clause "so per-task cost differences come from per-token pricing, not token volume" is contradicted by three passages in the very migration guide cited, all inside the Opus 5 -> Fable 5.1 section the quote comes from (https://platform.claude.com/docs/en/models/fable-5-1/migration-guide): (1) "What changed" item 1: Opus 5 accepts thinking: {type: "disabled"} at effort high or lower; on claude-fable-5-1 adaptive thinking is always on and disabling "returns a 400 error at any effort level. Remove the field, control token spend with lower effort levels, and revisit max_tokens for workloads that ran with thinking disabled." Any workload that ran Opus 5 with thinking off necessarily emits thinking tokens on Fable 5.1, billed as output at $50/MTok. That is a token-VOLUME increase on the exact migration path the claim describes. (2) Agent-loop behavior delta 1: "Fewer parallel tool calls in long agent loops: ... Claude Fable 5.1 may issue one tool call per turn. Each extra turn costs tokens, a round trip, and wall-clock time." The doc states outright that extra turns cost tokens - a per-task volume change in agentic coding, mitigated only by an added batching instruction. (3) Agent-loop behavior delta 3: "Fewer search and retrieval calls at low effort: At low effort Claude Fable 5.1 answers from memory more often than Claude Fable 5 instead of calling a search or retrieval tool." Directly relevant to the low-effort framing of the research question, and another volume change. MISREAD OF THE QUOTE. "Token counts are roughly unchanged" appears as a migration-checklist bullet adjacent to "Re-evaluate effort with a fresh sweep, starting at high, and re-baseline cost and latency" (line 1564) and "Re-baseline cost on your own workloads" (line 1604). Read in context it asserts tokenizer/counting equivalence for the same text - corroborated by the pricing page note that the newer tokenizer applies to "Claude 4.7 and later models", so Opus 5 and Fable 5.1 share it - not equivalence of per-task token consumption. The claim upgrades a tokenizer statement into a total-cost-of-task conclusion. Anthropic never says per-task cost scales purely with price; it says the opposite by telling migrators to re-sweep effort and revisit max_tokens. Practical consequence: budgeting "2x price, same tokens = 2x cost" understates real cost for thinking-disabled migrations and for long agent loops where Fable 5.1 serialises tool calls. CURRENCY: both pages current as of 2026-09-02 (pricing page carries the note that Sonnet 5's $2/$10 became standard rather than rising on 2026-09-01, and lists Fable 5.1). Not outdated. LIMITATION: this session's WebSearch budget (200/200) was exhausted before I could seek independent practitioner or third-party corroboration, so no non-Anthropic source was checked. That gap does not affect the numeric half (vendor list prices are self-authoritative) but means the refutation rests on Anthropic's own documentation contradicting the claim's inference - which is the strongest available ground anyway.
Verifier note 3
SPLIT VERDICT - the two pricing sentences are exact; the third sentence's conclusion is contradicted by the very section it quotes. CONFIRMED (https://platform.claude.com/docs/en/about-claude/pricing, model pricing table, fetched 2026-09-02): Fable 5.1 = $10/MTok input, $50/MTok output, cache hits $0.25/MTok. Opus 5 = $5/$25, cache hits $0.50/MTok. Fable 5 = $10/$50, cache hits $1/MTok. So $10/$50 IS double Opus 5's $5/$25; $0.25 IS half Opus 5's $0.50 and a quarter of Fable 5's $1.00. Footnote confirms the mechanism: "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier." The migration guide states both ratios verbatim (item 6, "half the Claude Opus 5 rate"; line 30, "a quarter of the Claude Fable 5 rate"). Vendor docs are definitionally authoritative for list prices, so the marketing-grade concern does not bite here. REFUTED - sentence 3, "so per-task cost differences come from per-token pricing, not token volume": (1) CONTRADICTED BY THE SAME SECTION. The quote is from the Opus 5 -> Fable 5.1 checklist. Item 1 of that section's "What changed" reads: "Claude Opus 5 accepts thinking: {type: 'disabled'} at an effort level of high or lower. On claude-fable-5-1 ... adaptive thinking is always on, and thinking: {type: 'disabled'} returns a 400 error at any effort level. Remove the field, control token spend with lower effort levels, and revisit max_tokens for workloads that ran with thinking disabled." Any Opus 5 workload running thinking-disabled generates MORE tokens on Fable 5.1 - a token-volume change on precisely the migration path the claim describes, on top of the 2x per-token rate. The guide also states (Fable 5 -> 5.1, item 1) that Fable 5.1 "may issue one tool call per turn" in long agent loops and "Each extra turn costs tokens" - Anthropic's own words that agentic behaviour moves token volume. (2) INVERTS THE QUOTE'S OWN INSTRUCTION. The quoted line leads with "Re-baseline cost on your own workloads." That instruction only exists because per-task cost is NOT derivable from the price ratio. "Token counts are roughly unchanged" is a statement about the TOKENIZER - line 41 lists "per-token pricing, tokenizer" among things that "match Claude Fable 5", i.e. the same text tokenizes to the same count. It is not a claim that a task consumes the same tokens end-to-end. (3) SCOPE OVERREACH ON "MIGRATORS". The unqualified "Anthropic tells migrators" is false. The phrase appears only in the Fable 5 -> 5.1 and Opus 5 -> 5.1 checklists (same tokenizer). The Opus 4.8-or-earlier checklist omits it, saying only "Per-token pricing differs." The pricing page notes: "Claude 4.7 and later models ... use a newer tokenizer ... approximately 30% more tokens for the same text. Claude Sonnet 4.6 and earlier models use the previous tokenizer." A migrator from Sonnet 4.6/Haiku 4.5 sees ~30% more tokens. MINOR QUALIFICATION on "double": Opus 5 fast mode bills at $10/$50 (pricing page, Fast mode section) - identical to Fable 5.1. "Double" holds for standard Opus 5 only. RELEVANT TO THE LOW-EFFORT RESEARCH QUESTION: the guide warns that "At low effort Claude Fable 5.1 answers from memory more often than Claude Fable 5 instead of calling a search or retrieval tool" - fewer tokens but a quality risk, so cost per COMPLETED task diverges further from the price ratio. NOT OUTDATED: pricing page fetched live 2026-09-02; Sonnet 5's $2/$10 is confirmed as now-standard. LIMITATION: WebSearch budget was exhausted (200/200) before independent third-party corroboration could run. This does not change the verdict - the refutation rests on Anthropic's own documentation contradicting the claim's conclusion, not on an absent external source.
E4refutedvote 1-2extracted claimFable 5.1's effort default is `high` with all five levels (low, medium, high, xhigh, max) supported, and Anthropic states its gains over Fable 5 are largest at xhigh and max, at the cost of extra thinking time and time-to-first-response. This directly contradicts any assumption that low effort is the intended default operating point.
Why it was killed
Verifier note 1
SPLIT VERDICT - the two factual halves check out verbatim; the load-bearing inference is contradicted by Anthropic's own docs. VERIFIED (accurate, quote is verbatim): 1) https://platform.claude.com/docs/en/models/fable-5-1/migration-guide - I fetched the page today (2026-09-02). The supporting quote is word-for-word correct, under "Start at `high` effort and sweep": "The effort parameter default is `high`, and all five levels are supported. Keep the Claude Fable 5 guidance: `high` for most work, and `medium` as a cost control worth testing. Claude Fable 5.1's gains over Claude Fable 5 are largest at `xhigh` and `max`, but those levels also add thinking time and time-to-first-response, so step up to them for the most capability-sensitive tasks and where your evals show the gain." 2) Corroborated at https://platform.claude.com/docs/en/build-with-claude/effort - "Recommended effort levels for Claude Fable 5.1": "Claude Fable 5.1 supports all five effort levels. Start with `high`, the default." Plus "The API defaults to `high`" and the levels table (low/medium/high/xhigh/max). Vendor-source concern does not bite on these two points: an API parameter default and a "what Anthropic states" attribution are facts about the vendor's own surface, where the vendor is the definitive authority, and the claim correctly hedges the gains statement as "Anthropic states". REFUTED (the third sentence, which is the part that would change a decision): "This directly contradicts any assumption that low effort is the intended default operating point" is an overreach on two counts. (a) It is not supported by the quote: the quote establishes a *parameter default*, which says nothing about whether low effort is a sound operating point for a given workload. (b) It is contradicted by Anthropic's own Fable 5.1 prompting guide, which endorses precisely the comparison the parent research is asking about - https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 § "Consider all effort levels": "At `medium`, results roughly match Claude Fable 5 at lower cost, so step down to `medium` or `low` where your evals show quality holds. At `low`, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, so include it in the comparison wherever you'd otherwise run a smaller model at a higher effort level." The same effort doc reinforces this: Best practices item 2 is "Use low for speed-sensitive or simple tasks", and the Fable 5.1 recommendation itself says "step down to `medium` or `low` for routine or latency-sensitive work once your evals show quality holds." The migration guide also treats low effort as a normal supported mode, documenting a low-effort-specific behavior change (item 3: "Fewer search and retrieval calls at low effort") rather than warning against the level. CHERRY-PICKING CHECK: the claim quotes the one migration-guide bullet that leans high/xhigh/max and omits the same doc set's explicit low-effort cost-per-task guidance. It also conflates "API default" with "intended operating point" - a conflation the cited source does not make. CURRENCY: current. All three pages fetched live on 2026-09-02. NOTE: I could not run WebSearch for independent/practitioner corroboration - this session had exhausted its 200-call web search budget - so this verification rests on primary vendor documentation only. That is sufficient for parts 1-2 (vendor is definitive on its own API default) but means the practitioner-transfer question in the research prompt is untested here. USABLE RESIDUE for the parent: keep "default is high, all five levels supported" and "Anthropic says gains over Fable 5 are largest at xhigh/max, costing thinking time and TTFR" - both are solid and citable. Drop the inference that this argues against low effort; Anthropic explicitly recommends benchmarking Fable 5.1 at low against Opus/Sonnet on cost per task.
Verifier note 2
SPLIT VERDICT - the factual half verifies, the load-bearing conclusion is refuted by the same vendor's own docs. VERIFIED (fetched https://platform.claude.com/docs/en/models/fable-5-1/migration-guide, item 4 "Start at `high` effort and sweep"): the supporting quote is verbatim and accurate - "The effort parameter default is `high`, and all five levels are supported... Claude Fable 5.1's gains over Claude Fable 5 are largest at `xhigh` and `max`, but those levels also add thinking time and time-to-first-response". Independently corroborated on https://platform.claude.com/docs/en/build-with-claude/effort: "Claude Fable 5.1 supports all five effort levels. Start with `high`, the default." and "By default, Claude uses high effort". REFUTED - sentence 3 ("This directly contradicts any assumption that low effort is the intended default operating point"): 1. CONTRADICTED BY PRIMARY SOURCE. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 ("Consider all effort levels") states the opposite of the claim's implication for exactly the research question: "At `medium`, results roughly match Claude Fable 5 at lower cost, so step down to `medium` or `low` where your evals show quality holds. At `low`, Claude Fable 5.1 is often competitive with Claude Opus and Claude Sonnet models on cost per task while scoring higher, so include it in the comparison wherever you'd otherwise run a smaller model at a higher effort level." Anthropic explicitly endorses Fable-5.1-at-low-effort as a candidate operating point against Opus 5 on cost per task. The effort page's own table lists `low` as the level for "subagents", and Best Practice 2 reads "Use low for speed-sensitive or simple tasks." 2. CATEGORY ERROR. The claim conflates the API parameter's default value (`high`, a request-shape fact) with "the intended default operating point" (a workload-tuning judgment). The docs settle the first and reject the second as a blanket rule: "Set effort explicitly: The API defaults to `high`, but the right starting point depends on your model and workload." 3. CHERRY-PICKED. The claim quotes half of one bullet and drops its own hedges - the same bullet says "`medium` as a cost control worth testing" and "Run a fresh sweep on your own evals rather than carrying over a setting tuned for Claude Fable 5." The gains-at-xhigh/max sentence is also stated conditionally in the docs ("step up... where your evals show the gain"), not as a general recommendation. 4. SOURCE-QUALITY GAP on the performance half. "Gains over Fable 5 are largest at `xhigh` and `max`" is an uncorroborated vendor performance claim - no benchmark numbers, no methodology, no independent replication located. (WebSearch budget for this session was exhausted at 200/200 before I could sweep for independent benchmarks, so I could not affirmatively confirm or find contradicting third-party measurement; per the brief this counts as marketing-grade-without-independent-support, not as verified.) The default-is-`high` and five-levels facts, by contrast, are API-surface facts where vendor docs are the correct authority and do not need independent support. NOT OUTDATED: pages current as of 2026-09-02 and describe the shipping model. Net: safe to assert is "the API default is `high` and all five levels are supported." Not safe to assert is that this "directly contradicts" running Fable 5.1 at low effort as a deliberate operating point - Anthropic's own prompting guide recommends putting exactly that configuration into the comparison against Opus/Sonnet.

Caveats and open questions

What the run could not establish, and what would settle it.

Caveats

  • Verifier disagreement: S1 was refuted 1-2, but a separate verifier found the identical sentence verbatim in Anthropic's Fable 5.1 prompting guide and marked that attribution claim not refuted; treat S1 as Anthropic's own published positioning, vendor-grade, with partial independent support against Sonnet 5 only. S2 and the two low-effort regression claims were 2-1 votes.
  • Vendor-only claims: 'gains widest at higher effort', 'medium roughly matches Fable 5', and 'low effort often competitive with Opus and Sonnet on cost per task' have no independent per-effort measurement; Artificial Analysis publishes low-effort figures but is a disclosed pre-release evaluation partner of Anthropic, and its Opus 5 and Sonnet 5 cost-per-task numbers differ between its July and September articles ($2.03 vs $2.34 for Opus 5 max; $1.53 vs $2.29 for Sonnet 5 max), likely index-version drift.
  • CursorBench is unpublished and Cursor is a commercial partner.
  • Max-plan weighting: the faster-draw statement and 50% cap are official, the ~2x multiplier is a June 2026 blog quoting Hacker News quoting in-app text for Fable 5, and whether effort level changes the draw rate is not published anywhere.
  • The system-card finding that Fable 5.1 scores best at medium on FrontierCode Extended comes from one Hacker News comment and was not checked against the system card. The Fable-subagent silent-substitution issue concerns Fable 5 on an earlier Claude Code build and may or may not persist for Fable 5.1 on v2.1.255.
  • Several verifiers exhausted their WebSearch budget (200/200) before sweeping for independent corroboration, so R4, R5, and the four migration-guide claims rest on Anthropic docs alone (acceptable for API contract facts, weak for anything about real-world quality).
  • Time sensitivity: R2 holds only as of 2026-09-02, with Opus 5.1 and Sonnet 5.1 EAP codenames already leaked; the prompt-cache and effort features cited are beta and header-gated; preserved-thinking enforcement is account-age dependent.
  • The refuted 'token counts roughly unchanged, so cost differences come from price not volume' inference should be dropped: Anthropic's own migration guide lists behaviours that move per-task token volume in both directions.

Open questions

  1. What is the actual Max-plan draw multiplier for Fable 5.1 versus Opus 5, and does low effort reduce the draw proportionally to tokens or is there a per-request weight? No Anthropic source publishes it; the only route is reading the /usage attribution after matched tasks.
  2. Does the silent Sonnet 5 substitution for a 'fable' subagent (issue #82252) still occur for claude-fable-5-1 on Claude Code v2.1.255, and does a Fable session spawning a Fable subagent work? Checkable by comparing meta.json against message.model in one agent transcript.
  3. Is there any measurement of Fable 5.1 at low or medium effort against Opus 5 at xhigh on non-coding knowledge work (spec writing, adversarial review, research synthesis)? None exists publicly; the only route is a small matched eval on the user's own PRD and review tasks with the search nudge in the prompt.
  4. Does the system-card finding that Fable 5.1 scores best at medium effort on FrontierCode Extended, with higher efforts adding unrequested changes, generalise to agentic builds like the WordPress work, and is it confirmed in the system card itself rather than a forum summary?

Sources

Every URL fetched by the run, with its quality tag, the angle it was fetched for, and its publication date where the page states one.

URLQualityAnglePublication dateClaims
platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1primaryvendor-primaryNot stated on the page (no publish or last-updated date shown); fetched 2026-09-025
platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompti...primaryvendor-primaryNot stated on the page (no publish or last-updated date shown); content is current as of retrieval on 2026-09-02 and references an August 31, 2026 account cutoff plus beta headers dated 2026-08-01/08-18/08-215
platform.claude.com/docs/en/models/fable-5-1/migration-guideprimaryvendor-primaryNot stated on the page; content is dated by internal references (enforcement threshold of 2026-08-31 and an example date of 2026-09-14), so it is current as of at least September 2026.5
platform.claude.com/docs/en/build-with-claude/effortprimaryvendor-primaryNot stated on the page (Anthropic platform docs page; no publish or last-updated date shown). Content references a 2026-07-01 beta header and models up to Fable 5.1 / Mythos 5.1, so it is current as of at least mid-2026.5
platform.claude.com/docs/en/models/overviewprimaryvendor-primaryNot stated on page (no publish or last-updated date rendered); retrieved 2026-09-02. Content is current as of the Fable 5.1 / Opus 5 / Sonnet 5 lineup, with Fable 5.1 retirement commitment "not sooner than September 1, 2027".5
support.claude.com/en/articles/15424964-claude-fable-5-on-your-planprimaryvendor-primaryNot dated on the page; it displays only "Updated" with a relative timestamp (rendered as "today" when fetched on 2026-09-02).5
x.com/ArtificialAnlys/status/2094881171066978525primaryindependent-benchmarks2026-09-015
artificialanalysis.ai/articles/opus-5primaryindependent-benchmarks2026-07-245
www.vals.ai/benchmarks/terminal-bench-2-1primaryindependent-benchmarks2026-08-31 (page states "Updated 8/31/2026")5
composio.dev/content/opus-vs-fableblogindependent-benchmarks2026-07-275
www.coderabbit.ai/blog/fable-5-1-model-reviewprimaryindependent-benchmarks2026-09-015
every.to/vibe-check/fable-5-1-vibe-checkblogindependent-benchmarks2026-09-015
news.ycombinator.com/item?id=48492210forumpractitioner-reportsapprox. 2026-06-12 (HN shows "82 days ago" as of 2026-09-02); exact date not displayed5
opinionai.substack.com/p/i-tested-opus-5-and-fable-5-the-cheaperblogpractitioner-reports2026-07-285
chudi.dev/blog/claude-fable-5-vs-opus-4-8blogpractitioner-reports2026-06-10 (updated 2026-09-01)5
www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changesblogskeptical-regressions2026-09-015
roo.beehiiv.com/p/claude-fable-5-1-mythos-5-1-whats-newblogskeptical-regressions2026-09-015
zenn.dev/uhyo/articles/react-profession-bench-13blogskeptical-regressions2026-07-175
www.developersdigest.tech/blog/claude-usage-limits-fable-5-explainedblogmax-plan-weighting2026-06-105
code.claude.com/docs/en/model-configprimaryclaude-code-runtimeNot stated on the page (no visible publish or last-updated date); content fetched 2026-09-025
code.claude.com/docs/en/prompt-cachingprimaryclaude-code-runtimeNot recorded on the page5
github.com/anthropics/claude-code/issues/82252forumclaude-code-runtime2026-07-29 (issue opened; follow-up comment 2026-08-09, last updated 2026-08-20; still OPEN, label area:agents)5
22 further URLs carry claims on this page but are not in the run's fetch log

Built from the run's own output file. Every number, quote and verdict on this page comes from that file or from the run brief. Em dashes in quoted verifier text have been replaced with hyphens, and no other wording was changed. Back to the top.