Which Claude Model Should You Use? (Most of You Don't Need Fable)
An honest, cost-first guide to choosing a Claude model: what each costs, what the July 7 Fable change means, and why most people don't need the expensive model — from an agency that runs Claude in production daily.
Two things just redrew this map, and most guides haven't caught up. Claude Opus 5 shipped on July 24, 2026 at $5 / $25 per million tokens — the same price as the Opus it replaced, and exactly half of Fable 5 — while landing within 0.5% of Fable's peak coding score. That makes the old "just use the best model" reflex doubly wrong: for the hard, deep work you used to reach Fable for, Opus 5 now gets you there for half the price. Our routing in one line: default to Sonnet 5 ($3/$15, intro $2/$10 through August) for most work, drop to Haiku 4.5 ($1/$5) for high-volume simple tasks, send hard coding and analysis to Opus 5 ($5/$25), and reserve Fable 5 ($10/$50) for fast, high-context strategy and orchestration where its speed and decisiveness earn the premium. The real trick isn't picking the smartest model — it's treating the four like a team with different temperaments and giving each the work it's actually built for. We run production systems on Claude every day and keep the bill boring by doing exactly that. Here's the map, the math, the temperaments, and an honest take on who still needs the expensive model.
Every "which Claude model" guide online tells you what each model is good at. Almost none tell you what each one costs you, whether you're overpaying, or — the part nobody says out loud — that most people reaching for the top model don't need it. This is that version, updated the week Opus 5 landed, from people who watch the meter.
What each Claude model costs
Claude has four models, and the price gap between the cheapest and the most expensive is 10×, which makes "just use the best one" a quietly expensive habit. Here are the current API rates, per million tokens, verified on the publish date. These are also the rates you pay once you exceed a subscription plan's limits and switch on usage credits — so they matter even if you never touch the raw API.
| Model | Input / M tokens | Output / M tokens | Built for |
|---|---|---|---|
| Haiku 4.5 | $1 | $5 | Fast, high-volume, mechanical work |
| Sonnet 5 | $3 ($2 intro*) | $15 ($10 intro*) | The balanced daily driver |
| Opus 5 | $5 | $25 | Deep, thorough coding, hard reasoning, high-stakes review |
| Fable 5 | $10 | $50 | Fast frontier strategy, orchestration, very long context |
*Sonnet 5 launched at an introductory $2/$10 per million tokens through August 31, 2026, then moves to its standard $3/$15.
To make the spread real: a month of 100 million input tokens and 50 million output tokens runs about $1,050 on Sonnet, $1,750 on Opus 5, and $3,500 on Fable. Same volume — the model choice alone more than triples the bill. That single fact is the whole of model economics, and it's why the smartest teams treat model selection as a cost decision, not just a capability one.
What changed on July 24: Opus 5 landed at half of Fable
Claude Opus 5 shipped July 24, 2026 at the same $5 / $25 as the Opus 4.8 it replaced — but it comes, in Anthropic's own words, "close to the frontier intelligence of Fable 5 at half the price." On CursorBench it lands within 0.5% of Fable 5's peak score for roughly half the cost per task. It's meaningfully better than the Opus before it at recovering from errors, navigating large codebases, and completing long multi-step agent loops. If you write code or run long autonomous jobs, this is the upgrade that matters this quarter.
But it has a temperament, and it's worth knowing before you point it at everything. Opus 5 is thorough to a fault. It takes time to clarify the goal, weighs several credible approaches before committing, and documents its reasoning heavily — which is exactly what you want on a gnarly refactor and exactly what you don't want on a one-line change. Independent reviews clock it reading about 50% more and writing about 65% more per call than the frontier baseline, with time-to-first-token around 68 seconds at maximum effort versus a class median under 3 seconds. It can over-check and over-engineer simple requests. In our own use it visibly second-guesses itself — which is a feature when the cost of a missed edge case is high, and a tax when the task was never hard. It is not the model for real-time chat or trivial work; it's the model for work that rewards being slow and careful.
The Fable question: what changed, and what didn't
As of July 7, 2026, Claude Fable 5 is billed as metered usage credits at $10 input / $50 output per million tokens — roughly double Opus 5 and ten times Haiku — now that its subsidized period on paid plans has ended. Through early July, Pro, Max, Team, and select Enterprise plans included Fable for up to about half of your weekly usage limit. That included allowance is gone; if you keep using Fable, you're paying API rates on top of your subscription.
The part that catches people: if you haven't enabled usage credits in the Claude Console, Fable access simply stops — no silent bill, but no access either. And if you have enabled credits without a spending cap, a few long Fable sessions on a big context can quietly dwarf your monthly subscription, because it's the most expensive model Anthropic sells. Nobody overspends here on purpose. They overspend by leaving "use the newest, best model" as the default and not noticing the meter turned on. The two-minute fix is to open the Console, decide whether you even need Fable enabled, and set a monthly cap before you forget.
Here's what Opus 5 changed about that decision: it took the hardest half of Fable's job and did it for half the price. The deep coding, debugging, and high-stakes review that used to be the honest case for Fable now has a cheaper home. What's left for Fable is narrower and more specific — and it's a real band, which we'll get to.
The honest take: most of you still don't need Fable
Fable 5 is an excellent model, and most people reaching for it are wrong to — it's overqualified for the work the average knowledge worker actually does, and now Opus 5 covers most of what remained. The daily work of most knowledge workers — drafting, summarizing, answering, analyzing a normal document, writing and reviewing ordinary code — sits comfortably inside what Sonnet already does well, faster and at a fraction of the price. The hard, deep work that genuinely needs more horsepower increasingly belongs to Opus 5, at half Fable's rate. Paying ten times Haiku for headroom you never touch isn't buying quality; it's buying idle horsepower.
So where does Fable still earn its rate? In a narrow, real band: fast, high-context strategic work where you want a model to hold a lot of moving pieces, reason several steps ahead, and commit with confidence — setting direction, orchestrating other models and tools, making the judgment call at the top of a workflow — and genuinely long-context, long-output tasks that need its window. That's roughly the list. Notice the contrast with Opus 5: where Opus is the careful specialist who double-checks, Fable moves with a decisiveness that skips the second-guessing — it behaves like it already knows, which is precisely what you want from the model steering the ship and precisely what you don't want reviewing your auth code.
Which model for which job: a team with four temperaments
Stop thinking of the four models as a ladder and start thinking of them as a team, each with a rank and a personality — and assign work by the temperament the job needs, not by prestige. Here's the staff.
Haiku 4.5 — the fast operator. Classification, tagging, extraction, format conversion, routing, first-pass triage, simple boilerplate. Anything high-volume and rule-shaped. At a fifth of Sonnet's price and a tenth of Fable's, Haiku is where cost-conscious throughput lives — and a common pattern is to run Haiku first on incoming work and escalate only the items it flags as hard.
Sonnet 5 — the reliable generalist. This is what most work should run on: the majority of coding, content drafting, data analysis, and everyday reasoning. It outperforms earlier Opus models on many benchmarks while running much faster and cheaper, and its introductory pricing through August makes it an especially easy default right now. Start here; escalate deliberately.
Opus 5 — the meticulous senior engineer. Reach for Opus 5 when the work is genuinely hard and being thorough pays: intricate debugging, architectural planning, large-codebase changes, security and high-stakes review, long multi-step agent loops. It recovers from its own mistakes, holds a big codebase in its head, and documents as it goes. The flip side is the temperament — it's slow to start, token-hungry, and will over-engineer a task that didn't need it. Give it the hard problems and let it be careful; don't put it on live chat or one-liners.
Fable 5 — the confident principal. Reserve it for fast, high-context strategy and orchestration: scoping a complex project, holding the whole board while it reasons several moves ahead, directing other models, and the frontier long-context/long-output jobs that need its window. It's the most expensive seat and it moves with a decisiveness the others don't. Opus 5 narrowed its band by taking the coding half — but at the top of a hard workflow, where speed × confidence × context all matter at once, it still earns the premium.
What we actually run
Here's our real division of labor: Fable 5 is our strategist, designer, and orchestrator; Opus 5 does the actual coding. Fable scopes the work, sets the design direction, and directs the run — fast and decisive, because that's the seat where hesitation costs you. Then Opus 5 grinds the implementation, thorough and self-checking, because that's the seat where a missed edge case costs you. It's the planner-does-the-thinking, specialist-does-the-building split — except when the building itself is hard, the builder isn't Haiku, it's Opus 5. Two premium seats, two temperaments, matched to two very different kinds of work.
Our five production loops — including the one that runs this site's SEO every day — cost about $200/month in fixed subscription plus a few dollars of data APIs, and a big reason is that recurring operational work never touches the frontier tier. Scheduled, bounded work on Sonnet- and Haiku-tier models is cheap and entirely good enough for the job: pull the search data, check the answer engines, draft the fixes, reconcile the books. We escalate to Opus 5 for the genuinely hard builds and to Fable for the strategic passes — and almost never for routine automation.
That discipline is the same one behind the honest agency economics we've written about, and it's exactly why business loops stay affordable as you add more of them — the marginal cost of the next scheduled run on Sonnet is close to nothing. The teams that get surprised by an AI bill are almost always the ones defaulting to the biggest model for work the workhorse would have handled.
Three levers that cut the bill without downgrading the work
Beyond picking the right model, three settings cut model cost dramatically: prompt caching, the Batch API, and disciplined routing. We use all three, and together they're why an agency running autonomous systems all day pays predictable, boring bills.
| Lever | What it does | Use it when |
|---|---|---|
| Prompt caching | Cuts repeated-context input cost by ~90% (Fable cached input drops from $10 to ~$1/M; Opus 5 from $5 to ~$0.50/M) | You re-send the same system prompt, tool definitions, or file context every turn — i.e. any long session or agent loop |
| Batch API | Halves both input and output rates (Fable → Opus-priced, Opus 5 → Sonnet-priced, Sonnet → near-Haiku) | The work doesn't need to happen in real time: overnight analysis, bulk generation, scheduled reports |
| Model routing | Runs each task on the smallest capable model; reserves Opus 5 and Fable for jobs that need them | Always. This is the discipline that separates predictable bills from billing surprises |
Prompt caching is the highest-leverage of the three for anyone running long sessions, and it's essentially a switch — which matters more with Opus 5, since its thoroughness means longer contexts and more turns. Batch pricing is free money for any workload nobody is watching in real time — which describes most business automation. And routing is the habit that makes the other two matter: cache and batch a Fable workload that should have been an Opus 5 or Sonnet workload, and you've still overpaid.
So which should you actually pick?
If you use Claude in the app, the model choice is mostly handled for you — just don't manually default everything to Fable; if you build on the API, route deliberately and set a cap. For a chat or Cowork user, stay on Sonnet for daily work, let the product escalate when it needs to, and pick Opus 5 by hand when you've got a genuinely hard coding or analysis task in front of you. For an API builder or anyone running automations, the routing map above is the difference between a predictable line item and a scary one — default to Sonnet, drop to Haiku for volume, escalate to Opus 5 for hard problems, reserve Fable for fast strategic and orchestration passes, cache and batch what you can, and keep the frontier tier on a leash with a Console spending cap.
And if your usage is heavy or spiky enough that you're weighing plans against the API, that's a related but separate decision — we broke down the subscription tiers (Pro, Max, Team, Enterprise) and when each one beats metered usage in our Claude plans breakdown. If you'd rather just get a read on where your own operation is overspending, that's what our Revenue Audit is for.
FAQ
Published July 2026; updated July 29, 2026 for the launch of Claude Opus 5. All model rates, the Sonnet 5 introductory window, the Fable 5 usage-credit change (effective July 7, 2026), and the Opus 5 launch (July 24, 2026) were verified against Anthropic's pricing pages, documentation, and independent benchmark reviews on the update date. Claude pricing and models change frequently — confirm current numbers before you buy or enable usage credits.
Related: Claude plans: Pro vs Max vs Team vs Enterprise · Claude Cowork pricing · The economics of an AI automation agency · Business loops · Claude loops for business
Written by Joseph Darnell, founder of Automaton. He runs the agency's autonomous SEO and AEO program and builds AI systems for clients.