Which Claude Model Should You Use? (Most of You Don't Need Fable)
An honest, cost-first guide to choosing a Claude model after Opus 5.5: what each costs, why the coder seat just got cheaper, and why most people still don't need the expensive model — from an agency that runs Claude in production daily.
The map got redrawn twice this summer, and most guides are still on the first draft. Claude Opus 5.5 shipped on September 22, 2026 at $4 / $20 per million tokens. That is 20% below the Opus 5 it replaces and 40% of Fable 5.1's $10 / $50, and Anthropic says it performs at Fable's level on most work while costing about 40% less to run than Opus 5 on typical workloads. So the old "just use the best model" reflex is now wrong twice over: the hard, deep work you used to reach Fable for has a cheaper home, and the cheaper home just got cheaper. Our routing in one line: default to Sonnet 5 ($2 / $10) for most work, drop to Haiku 4.5 ($1 / $5) for high-volume simple tasks, send hard coding and analysis to Opus 5.5 ($4 / $20), and reserve Fable 5.1 ($10 / $50) for fast, high-context strategy and orchestration where its speed and decisiveness still earn the premium. The real trick isn't picking the smartest model. It's treating the four like a team with different temperaments and giving each the work it was built for. We run production systems on Claude every day and keep the bill boring by doing exactly that. Here's the map, the math, the temperaments, and an honest take on who still needs the expensive model.
Every "which Claude model" guide online tells you what each model is good at. Almost none tell you what each one costs you, whether you're overpaying, or — the part nobody says out loud — that most people reaching for the top model don't need it. This is that version, updated the week Opus 5.5 landed, from people who watch the meter.
What each Claude model costs
Claude has four current models, and the price gap between the cheapest and the most expensive is 10×, which makes "just use the best one" a quietly expensive habit. Here are the API rates, per million tokens, as listed on Anthropic's pricing page on September 24, 2026. These are also the rates you pay once you exceed a subscription plan's limits and switch on usage credits — so they matter even if you never touch the raw API.
| Model | Input / M tokens | Output / M tokens | Built for |
|---|---|---|---|
| Haiku 4.5 | $1 | $5 | Fast, high-volume, mechanical work |
| Sonnet 5 | $2* | $10* | The balanced daily driver |
| Opus 5.5 | $4 | $20 | Deep, thorough coding, hard reasoning, high-stakes review |
| Fable 5.1 | $10 | $50 | Fast frontier strategy, orchestration, very long context |
*Sonnet 5 launched in June at an introductory $2 / $10 that was slated to step up to $3 / $15 after August 31, 2026. As of September 24, Anthropic's pricing page still lists $2 / $10 with no introductory note, so that is the rate we use here. If it moves, the math below moves with it. Opus 5 remains available at $5 / $25.
To make the spread real: a month of 100 million input tokens and 50 million output tokens runs about $350 on Haiku, $700 on Sonnet, $1,400 on Opus 5.5, and $3,500 on Fable 5.1. Same volume. The model choice alone is a 10× swing from bottom to top, and Fable costs two and a half Opus 5.5s. That single fact is the whole of model economics, and it's why the smartest teams treat model selection as a cost decision, not just a capability one.
What changed on September 22: Opus 5.5 landed at 40% of Fable
Claude Opus 5.5 shipped September 22, 2026 at $4 / $20 per million tokens, 20% below the Opus 5 it replaces, and Anthropic's own line is that it performs at the level of Claude Fable 5.1 on most work while costing about 40% less to run than Opus 5. On Anthropic's published benchmarks it leads in agentic coding, computer use and knowledge work, and it beats Fable 5.1 on every board they show (Terminal-Bench 4.0: 66.4% vs 55.8%; FrontierCode: 54.4% vs 50.3%; GDPval-AA: 1846 vs 1735). Anthropic adds a caveat we'd have added ourselves: at this level the benchmark margins have become a less reliable guide to real differences, and in their own use the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest.
Two things about it matter more than the leaderboard. First, it's cheaper per task, not just per token. Anthropic reports it uses fewer tokens to finish the same work, generates output more than 30% faster than Opus 5, and cache reads dropped to $0.20 per million (from $0.50). Their internal example: a 200,000-line codebase audit that took Opus 5 over 20 hours and 2.5× the tokens, done in under three hours. Second, they say it writes more clearly. Verbose, hard-to-follow output was the most common complaint about Opus 5, and 5.5 is pitched as putting the important thing first and dropping the jargon. That claim is the one we care about most, because a big part of the case for Fable was the prose. We haven't re-run our own writing test on 5.5 yet, and until we do the honest position is "Anthropic says it closed the gap; we'll report what we find."
The temperament question is open too. Opus 5 was thorough to a fault: slow to start, token-hungry, prone to over-engineering a one-line change. Anthropic says 5.5 makes fewer, more complete edits and stops retrying, and that it's noticeably less likely to take hard-to-reverse actions or step outside the boundaries it's given, which matters for anyone running it unattended overnight. One quirk to know: Opus 5.5 always thinks. There is no setting to switch extended thinking off.
The Fable question: what changed, and what didn't
As of July 7, 2026, Claude Fable is billed as metered usage credits at $10 input / $50 output per million tokens (Fable 5.1 since September 1), two and a half times Opus 5.5 and ten times Haiku, now that its subsidized period on paid plans has ended. Through early July, Pro, Max, Team, and select Enterprise plans included Fable for up to about half of your weekly usage limit. That included allowance is gone; if you keep using Fable, you're paying API rates on top of your subscription.
The part that catches people: if you haven't enabled usage credits in the Claude Console, Fable access simply stops — no silent bill, but no access either. And if you have enabled credits without a spending cap, a few long Fable sessions on a big context can quietly dwarf your monthly subscription, because it's the most expensive model Anthropic sells. Nobody overspends here on purpose. They overspend by leaving "use the newest, best model" as the default and not noticing the meter turned on. The two-minute fix is to open the Console, decide whether you even need Fable enabled, and set a monthly cap before you forget.
Here's what Opus 5.5 does to that decision: it took the hardest half of Fable's job and now does it for 40 cents on the dollar. The deep coding, debugging, and high-stakes review that used to be the honest case for Fable have a cheaper home. What's left for Fable is narrower and more specific, and it's a real band, which we'll get to.
The honest take: most of you still don't need Fable
Fable 5.1 is an excellent model, and most people reaching for it are wrong to — it's overqualified for the work the average knowledge worker actually does, and now Opus 5.5 covers most of what remained. The daily work of most knowledge workers — drafting, summarizing, answering, analyzing a normal document, writing and reviewing ordinary code — sits comfortably inside what Sonnet already does well, faster and at a fraction of the price. The hard, deep work that genuinely needs more horsepower increasingly belongs to Opus 5.5, at 40% of Fable's rate. Paying ten times Haiku for headroom you never touch isn't buying quality; it's buying idle horsepower.
One band we'd add since 5.1 shipped: writing that has to sound like a person. We broke that down in Fable 5.1 vs Opus 5, and Opus 5.5's clearer-writing claim is exactly the thing that test has to re-check.
So where does Fable still earn its rate? In a narrow, real band: fast, high-context strategic work where you want a model to hold a lot of moving pieces, reason several steps ahead, and commit with confidence — setting direction, orchestrating other models and tools, making the judgment call at the top of a workflow — and genuinely long-context, long-output tasks that need its window. That's roughly the list. Notice the contrast with Opus: where Opus is the careful specialist who double-checks, Fable moves with a decisiveness that skips the second-guessing — it behaves like it already knows, which is precisely what you want from the model steering the ship and precisely what you don't want reviewing your auth code.
Which model for which job: a team with four temperaments
Stop thinking of the four models as a ladder and start thinking of them as a team, each with a rank and a personality — and assign work by the temperament the job needs, not by prestige. Here's the staff.
Haiku 4.5 — the fast operator. Classification, tagging, extraction, format conversion, routing, first-pass triage, simple boilerplate. Anything high-volume and rule-shaped. At half of Sonnet's price and a tenth of Fable's, Haiku is where cost-conscious throughput lives — and a common pattern is to run Haiku first on incoming work and escalate only the items it flags as hard.
Sonnet 5 — the reliable generalist. This is what most work should run on: the majority of coding, content drafting, data analysis, and everyday reasoning. It outperforms earlier Opus models on many benchmarks while running much faster and cheaper, and at $2 / $10 it's an easy default. Start here; escalate deliberately.
Opus 5.5 — the meticulous senior engineer, now with a shorter commute. Reach for Opus 5.5 when the work is genuinely hard and being thorough pays: intricate debugging, architectural planning, large-codebase changes, security and high-stakes review, long multi-step agent loops. It holds a big codebase in its head, recovers from its own mistakes, and (per Anthropic) finishes in fewer steps and fewer tokens than Opus 5 did, with output that's clearer to check. It's the seat where a missed edge case costs you, so let it be careful. Just don't put it on live chat or one-liners; that's still Sonnet's job.
Fable 5.1 — the confident principal. Reserve it for fast, high-context strategy and orchestration: scoping a complex project, holding the whole board while it reasons several moves ahead, directing other models, and the frontier long-context/long-output jobs that need its window. It's the most expensive seat and it moves with a decisiveness the others don't. Opus 5.5 narrowed its band by taking the coding half — but at the top of a hard workflow, where speed × confidence × context all matter at once, it still earns the premium.
What we actually run
Here's our real division of labor: Fable 5.1 is our strategist, designer, and orchestrator; Opus does the actual coding, and as of this update that seat is still Opus 5, with 5.5 in trial. Fable scopes the work, sets the design direction, and directs the run — fast and decisive, because that's the seat where hesitation costs you. Then Opus grinds the implementation, thorough and self-checking, because that's the seat where a missed edge case costs you. It's the planner-does-the-thinking, specialist-does-the-building split — except when the building itself is hard, the builder isn't Haiku, it's Opus. We'll move the seat to 5.5 once it has earned it on our own work, not on a launch page, and say so here when it does.
Our five production loops — including the one that runs this site's SEO every day — cost about $200/month in fixed subscription plus a few dollars of data APIs, and a big reason is that recurring operational work never touches the frontier tier. Scheduled, bounded work on Sonnet- and Haiku-tier models is cheap and entirely good enough for the job: pull the search data, check the answer engines, draft the fixes, reconcile the books. We escalate to Opus for the genuinely hard builds and to Fable for the strategic passes — and almost never for routine automation.
That discipline is the same one behind the honest agency economics we've written about, and it's exactly why business loops stay affordable as you add more of them — the marginal cost of the next scheduled run on Sonnet is close to nothing. The teams that get surprised by an AI bill are almost always the ones defaulting to the biggest model for work the workhorse would have handled.
Four levers that cut the bill without downgrading the work
Beyond picking the right model, three settings cut model cost dramatically: prompt caching, the Batch API, and disciplined routing. A fourth, fast mode, buys speed when a human is waiting. We use the first three every day, and together they're why an agency running autonomous systems all day pays predictable, boring bills.
| Lever | What it does | Use it when |
|---|---|---|
| Prompt caching | Cuts repeated-context input cost by 90 to 97% (Fable 5.1 cached input drops from $10 to $0.25 per million; Opus 5.5 from $4 to $0.20; Sonnet 5 from $2 to $0.20) | You re-send the same system prompt, tool definitions, or file context every turn, i.e. any long session or agent loop |
| Batch API | Halves both input and output rates (Fable → roughly Opus-priced, Opus 5.5 → Sonnet-priced, Sonnet → Haiku-priced) | The work doesn't need to happen in real time: overnight analysis, bulk generation, scheduled reports |
| Fast mode | Runs Opus 5.5 at up to 2.5× speed for $8 / $40 per million (Claude API only) | The task is hard enough to need Opus and a human is waiting on the answer |
| Model routing | Runs each task on the smallest capable model; reserves Opus 5.5 and Fable for jobs that need them | Always. This is the discipline that separates predictable bills from billing surprises |
Prompt caching is the highest-leverage of the four for anyone running long sessions, and it's essentially a switch. It matters even more now that a cache hit on Opus 5.5 costs a nickel on the dollar. Batch pricing is free money for any workload nobody is watching in real time, which describes most business automation. Fast mode is the honest answer to "Opus is slow": double the rate for up to 2.5× the speed is still cheaper than reaching for Fable. And routing is the habit that makes the other three matter: cache and batch a Fable workload that should have been an Opus or Sonnet workload, and you've still overpaid.
So which should you actually pick?
If you use Claude in the app, the model choice is mostly handled for you — just don't manually default everything to Fable; if you build on the API, route deliberately and set a cap. For a chat or Cowork user, stay on Sonnet for daily work, let the product escalate when it needs to, and pick Opus 5.5 by hand when you've got a genuinely hard coding or analysis task in front of you. For an API builder or anyone running automations, the routing map above is the difference between a predictable line item and a scary one — default to Sonnet, drop to Haiku for volume, escalate to Opus 5.5 for hard problems, reserve Fable for fast strategic and orchestration passes, cache and batch what you can, and keep the frontier tier on a leash with a Console spending cap.
And if your usage is heavy or spiky enough that you're weighing plans against the API, that's a related but separate decision — we broke down the subscription tiers (Pro, Max, Team, Enterprise) and when each one beats metered usage in our Claude plans breakdown. If you'd rather just get a read on where your own operation is overspending, that's what our Revenue Audit is for.
Frequently asked questions
Which Claude model should I use?
Default to Sonnet 5 for most work: coding, writing, analysis. Drop to Haiku 4.5 for high-volume simple tasks like classification and extraction. Escalate to Opus 5.5 for genuinely hard coding, debugging, planning, or high-stakes review. Reserve Fable 5.1 for fast, high-context strategy and orchestration, or jobs that need its very large context and long output. The rule of thumb: start with Sonnet and escalate only when a task earns it.
Is Claude Opus 5 worth it, and what is it best at?
Opus 5 has been superseded. Claude Opus 5.5 (shipped September 22, 2026, $4 / $20 per million tokens) is the model for hard coding and long agent loops: Anthropic reports it performs at Fable 5.1's level on most work, leads on agentic coding and knowledge-work benchmarks, finishes in fewer tokens and generates output over 30% faster than Opus 5, and costs about 40% less to run. Opus 5 is still available at $5 / $25, but there is no longer a price or capability reason to pick it for new work. Use 5.5 for genuinely hard work, not real-time chat or trivial tasks.
Opus 5 vs Fable 5: which should I use for coding?
For coding, use Opus 5.5. It leads Fable 5.1 on Anthropic's coding benchmarks (66.4% vs 55.8% on Terminal-Bench 4.0; 54.4% vs 50.3% on FrontierCode) at 40% of the token price ($4 / $20 vs $10 / $50), and it's built for error recovery and navigating large codebases. Keep Fable 5.1 for the top of the workflow: fast, high-context strategy, design direction, and orchestrating other models, where its speed and decisiveness matter more than raw coding depth. A common split is Fable to plan and direct, Opus 5.5 to build.
Do I actually need Claude Fable 5?
For most people, no, and less than ever. Fable 5.1 is Anthropic's most capable generally available model and its most expensive, overqualified for typical daily knowledge work. Since September 22, 2026, Opus 5.5 covers the hard coding and deep reasoning that used to be Fable's main justification at 40% of the price, and Anthropic says its writing has closed most of the clarity gap too. What's left for Fable 5.1 is a narrow band: fast, high-context strategy and orchestration, genuinely long-context autonomous tasks, and prose that has to sound like a person (pending our re-test). If you're unsure whether you're in that group, you're almost certainly not, and Sonnet plus Opus 5.5 will cover you.
What does the Claude API cost per token?
Per million tokens (input / output): Haiku 4.5 is $1 / $5, Sonnet 5 is $2 / $10, Opus 5.5 is $4 / $20, and Fable 5.1 is $10 / $50, per Anthropic's pricing page on September 24, 2026. A sample month of 100M input and 50M output tokens runs roughly $700 on Sonnet, $1,400 on Opus 5.5, and $3,500 on Fable. The model choice alone is a 5× swing.
How much does Claude Fable 5 cost now that the subsidized credits ended?
As of July 7, 2026, Fable is billed as metered usage credits at $10 per million input tokens and $50 per million output tokens (Fable 5.1 since September 1, 2026), about two and a half times Opus 5.5 and ten times Haiku 4.5. Before that date, paid plans included Fable for up to roughly half of weekly usage limits. To keep using it you must enable usage credits in the Claude Console; if you don't, Fable access stops rather than silently billing you.
Should I default to Sonnet or Opus 5?
Default to Sonnet. Sonnet 5 handles the roughly 90% of everyday tasks (coding, analysis, writing, standard research) at half of Opus's price, and it's faster. Escalate to Opus 5.5 only for genuinely hard, multi-step reasoning or coding that Sonnet visibly struggles with. Opus 5.5 is twice Sonnet's price on both input and output, and it always runs with extended thinking on, so Sonnet-first with deliberate escalation is the cost-correct default for all but the hardest work.
How do I lower my Claude API bill?
Three levers, plus one. Enable prompt caching to cut repeated-context input cost by 90 to 97% (Fable 5.1 cached input drops from $10 to $0.25 per million; Opus 5.5 from $4 to $0.20). Use the Batch API for non-real-time work to halve both rates. Route by task: most work on Sonnet, bulk work on Haiku, hard work on Opus 5.5, Fable for fast strategic and orchestration passes. And if Opus is too slow for a human who's waiting, fast mode at $8 / $40 is cheaper than reaching for Fable. Set a monthly spending cap in the Console as a backstop.
What is the cheapest Claude model?
Haiku 4.5 is the cheapest at $1 per million input tokens and $5 per million output tokens, half of Sonnet's price and a tenth of Fable's. It's built for fast, high-volume, mechanical work like classification, extraction, formatting, and simple questions, and it is not suited to deep reasoning or nuanced tasks. A common cost pattern is to run Haiku first and escalate only what it flags as hard.
Published July 2026; updated September 24, 2026 for the launch of Claude Opus 5.5. All model rates were verified against Anthropic's pricing page on the update date; the Opus 5.5 performance and efficiency figures are Anthropic's published claims, not our own benchmarks. Sonnet 5.5 followed on September 28, 2026 at the same $2 / $10 list price as Sonnet 5 (Anthropic says it runs 30%+ faster and costs up to 30% less for most work); Haiku 5.5 is still to come, and we'll update the routing once the lineup settles. Claude pricing and models change frequently. Confirm current numbers before you buy or enable usage credits.
Related: Claude plans: Pro vs Max vs Team vs Enterprise · Claude Cowork pricing · The economics of an AI automation agency · Business loops · Claude loops for business
Written by Joseph Darnell, founder of Automaton. He runs the agency's autonomous SEO and AEO program and builds AI systems for clients.