Which Claude model I suppose use? Price guide
Compare Claude Opus 4.8, Sonnet 5, and Haiku 4.5 rates for July 2026. See how much 100,000 jobs go cost you and which settings dey increase your bill.
Which Claude model I suppose use?
Short answer for which Claude model you suppose use: start with Claude Opus 4.8, and no move away from am unless you get one specific reason. Anthropic guidance na the same thing. "If you no sure which model to use, start with Claude Opus 4.8 for complex agentic coding and enterprise work." Move down to Claude Sonnet 5 when the task clear well-well and you dey run am many times every day. Move down again to Claude Haiku 4.5 for mechanical, high-volume work wey you don already know wetin correct answer suppose look like. Claude Fable 5 dey top all of dem, for long-running agents.
The choice get price wey follow am. These na the rates for Claude API (application programming interface) as of 23 July 2026, per million tokens, wey the docs write as MTok.
- Claude Fable 5 (
claude-fable-5): $10 / MTok in, $50 / MTok out. 1M context. - Claude Opus 4.8 (
claude-opus-4-8): $5 in, $25 out. 1M context. - Claude Opus 4.7 (
claude-opus-4-7): $5 in, $25 out. 1M context. - Claude Sonnet 5 (
claude-sonnet-5): $2 in, $10 out on introductory pricing through 31 August 2026. The standard $3 in, $15 out go start on 1 September 2026. 1M context. - Claude Haiku 4.5 (
claude-haiku-4-5): $1 in, $5 out. 200K context.
Those identifiers complete as dem write dem. Nothing dey add for dem.
Size itself no get extra charge for the 1M-token models: "A 900k-token request na the same per-token rate as a 9k-token request."
Wetin every model dey do
Anthropic explain wetin every model dey do inside one line, and those lines better pass any leaderboard.
- Claude Fable 5: "Next-generation intelligence for long-running agents." Na him slow pass for latency among the four.
- Claude Opus 4.8: "For complex agentic coding and enterprise work." Latency dey middle.
- Claude Sonnet 5: "The best combination of speed and intelligence." E fast.
- Claude Haiku 4.5: "The fastest model with near-frontier intelligence."
Haiku 4.5 get some limits wey go help you know if e fit for your work before you look for price. Im context window na 200K tokens instead of 1M, so large repository or long agent transcript no go fit inside. Im maximum output for synchronous Messages API na 64K tokens, while the others na 128K, and im knowledge cutoff na February 2025, while the other three na January 2026.
Wetin the choice go cost me for real workload?
Every model for dis lineup charge for output at five times wetin dem charge for input. Opus 4.8 na $5 for input and $25 for output. Haiku 4.5 na $1 for input and $5 for output. Dis ratio dey stay same for all models, so di model wey you pick go matter pass for work wey dey generate plenty output.
Agentic work na dat kind work, because dem charge thinking tokens as output tokens and dem dey count dem for max_tokens even when di text no reach you. For Fable 5, Opus 4.8, Opus 4.7 and Sonnet 5, dem no dey send reasoning summary by default, so di thinking field go come back empty. Di billing no go change: "Either way the block is billed the same and passed back the same in multi-turn conversations." Wetin really dey fill Claude token bill explain dis matter well-well.
Whether thinking dey run or no dey depend on di models wey you dey compare, and if you miss dis one, your cost test go spoil. For Sonnet 5 and Fable 5, thinking don dey on already and e no need configuration. For Opus 4.8 and Opus 4.7, e dey off until you set thinking: {type: "adaptive"} for di request. Make sure say configuration for both sides dey match before you start to look di numbers.
Why the cheapest model fit cost you more money
Take one coding task wey dey stand alone. The request carry 60,000 tokens of context, and the model go produce 8,000 tokens of output including thinking. If you no use caching, the math go clear.
On Opus 4.8: 0.06 MTok of input at $5 na $0.30, and 0.008 MTok of output at $25 na $0.20. The attempt cost $0.50. On Haiku 4.5, the same attempt na $0.06 plus $0.04, so $0.10.
Haiku cheap five times per attempt, wey fit look like say e get plenty space. But e go change when you see wetin one failed attempt dey do. You go read wrong answer, and that one go waste your time. You go re-send that answer as context when you retry, so every new attempt go big pass the one before am. And if the third attempt still fail, you go use better model anyway: $0.30 of Haiku plus $0.50 of Opus na $0.80, wey mean 60% more than if you run Opus once.
So the real question na how cheap you fit check the answer. If you fit see wrong answer inside one second, the small model na bargain. But if to see error means say you must read diff carefully, the small model go cost you time wey no go show for your invoice.
Wey small model go win outright
Haiku 4.5 na correct answer for these cases:
- Mechanical subagent work. Subagent wey dey rename files or collect search output no need heavy reasoning. Na standard pattern be that once you build an AI agent with Claude and give am helpers.
- Log triage. To decide if one line na noise or if e worth human attention na narrow judgement wey get obvious failure mode.
- Classification against a fixed label set. Output dey short and you fit measure accuracy for sample.
- High-volume production calls. At 100,000 runs a day, the per-token gap no be small thing again. Na how AI workflows wired into n8n dey work normally.
- Anything wey response speed matter to the user. Haiku 4.5 na fastest model for the lineup.
One limit na for Haiku alone. Its minimum cacheable prompt na 4,096 tokens, while Opus 4.8 and Sonnet 5 na 1,024. If e below that minimum, "any requests to cache fewer than this number of tokens will be processed without caching, and no error is returned". One 1,500-token instruction block wey dey cache for Sonnet 5 no go cache for Haiku 4.5. You go know am when cache_creation_input_tokens and cache_read_input_tokens both dey zero, wey be one of the ways wey keeping an always-on agent's costs down dey cover.
The levers that change the answer
Each of these moves a bill further than the model name does.
Effort. output_config.effort controls how much work the model does before it answers. The levels are low, medium, high, xhigh and max, and high is the default: "Setting effort to high produces exactly the same behavior as omitting the effort parameter entirely." It nests inside output_config rather than sitting at the top level of the request:
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=4096,
output_config={"effort": "medium"},
messages=messages,
)Effort touches every token in the response: "It can affect all token spend including tool calls. For example, lower effort would mean Claude makes fewer tool calls." That compounds on an agentic loop. It is not a token cap: "Effort is a behavioral signal, not a strict token budget." Anthropic publishes no cost multiplier per level and tells you to test your own use case, so a Sonnet 5 run at high against an Opus 4.8 run at low is a real comparison you can run this afternoon. Haiku 4.5 is absent from the list of models that support effort.
Prompt caching. A cache read costs 0.1x the base input rate, and the number that decides a model choice is what that does across tiers. A cache hit on Opus 4.8 costs $0.50 / MTok, and uncached input on Haiku 4.5 costs $1 / MTok, so a well-cached Opus prefix is cheaper per input token than a cold Haiku prompt. Session hygiene beats model choice on any workload that re-sends a large stable prefix.
Batches. If nothing is waiting on the answer, the Message Batches API runs the same models with "a 50% discount on both input and output tokens". It halves every row of the comparison below, and it works across tiers: a batched Opus 4.8 job costs less than a synchronous Sonnet 5 job at the standard price from 1 September 2026, trading latency for capability at the same money.
One cost comparison for work
Workload: classify 100,000 support emails. Every call go use 2,000 input tokens and 300 output tokens. Dis mean 200 MTok for input and 30 MTok for output.
- Haiku 4.5: 200 x $1 = $200 in, 30 x $5 = $150 out. Total $350.
- Sonnet 5, introductory rate: 200 x $2 = $400 in, 30 x $10 = $300 out. Total $700.
- Sonnet 5, from 1 September 2026: 200 x $3 = $600 in, 30 x $15 = $450 out. Total $1,050.
- Opus 4.8: 200 x $5 = $1,000 in, 30 x $25 = $750 out. Total $1,750.
You fit change four numbers if you wan use tomorrow prices. De adjustments below go change de result pass de model name:
- Batching go cut every line by half: $175, $350, $525 and $875. No thing dey wait for nightly classification run, so no need to skip am.
- Caching de shared prefix. Suppose say 1,200 of de 2,000 input tokens na de same instruction block every time. For Sonnet 5 at de introductory rate, de cost na $0.20 / MTok as cache hits instead of $2 / MTok. So 120 MTok go cost $24 instead of $240. Add 80 MTok of fresh input at $2, which na $160, plus $300 of output, and Sonnet 5 go land near $484 instead of $700. Dis same trick no go work for Haiku 4.5, because 1,200 tokens small pass de 4,096-token minimum.
- Retries. Suppose say you reject Haiku answer for 8% of de emails and you re-run dem for Opus 4.8. Add 0.08 x $1,750 = $140 to de $350, so e go be $490. Check your own rejection rate instead of to use mine.
- Tool definitions, if de job use dem. De system prompt dem generate na 290 tokens for Opus 4.8 when
tool_choiceset toauto, 354 for Sonnet 5 and 496 for Haiku 4.5. De cheapest model carry de biggest fixed overhead.
Haiku with escalation path at $490 and cached Sonnet 5 at $484 cost de same, and one of dem fit do de job for single pass. De model no be de whole matter.
If I change model mid-session, I go save money?
E no dey happen as the sticker prices dey show, because the saving na per token, and wetin you dey risk na the whole cached prefix.
Anthropic don write down wetin dey spoil cache. Dem dey build prefixes for the order tools, then system, then messages, and "Changes at each level invalidate that level and all subsequent levels." Two request settings dey for that list too: "The thinking configuration and the resolved effort level dey inside the prompt itself, so if you change any of dem, e go start new cache prefix." For effort, the docs talk say: "vary effort across workloads rather than within a conversation that relies on cache hits".
The model no dey for that documented list, so no assume anything. Read cache_read_input_tokens for the first request after you change model, make the number answer you. For 150,000-token Opus 4.8 prefix, cache read cost about $0.08 and fresh write cost about $0.94, which plenty pass the per-token gap wey you dey chase.
So change these things between tasks, no do am inside one task. For Claude Code, that means /clear first, when dem dey discard the cache anyway, then /model or /effort.
How to test the answer on your own workload
No use token count from another model for size your prompt. The token-counting endpoint free to use and e dey count with the tokenizer of any model wey you name, so name the one wey you plan to call. Opus 4.7 and later, Fable 5 and Sonnet 5 use new tokenizer wey "produces approximately 30% more tokens for the same text", so if you measure budget for old model, e go low when you use new model even if per-token rate no change.
Then read wetin dem return. input_tokens na only the uncached remainder, so the true prompt size na that field plus cache_creation_input_tokens plus cache_read_input_tokens. A first Claude API app on a VPS na the smallest honest place to run that measurement, and which Claude plan fits your usage na different decision about whether you go pay per token at all.
FAQ
Which Claude model best for coding?
Claude Opus 4.8 na him dem document say na him be starting point for complex agentic coding, at $5 per million input tokens and $25 per million output as of July 2026. Claude Sonnet 5 na him be best combination of speed and intelligence, and at $2 / $10 through 31 August 2026, e cost less than half of the other one. Run the same task on both, keep the thinking configuration constant, and compare total token spend from response.usage.
Claude Haiku 4.5 cheap enough to replace Sonnet 5?
Per token, e easy: $1 / $5 against $2 / $10 for Sonnet 5 on introductory pricing, or $3 / $15 from 1 September 2026. The limits na dem decide am. Haiku 4.5 get 200K token context window instead of 1M, 64K maximum output on the synchronous Messages API, February 2025 reliable knowledge cutoff, no output_config.effort support, and 4,096-token minimum cacheable prompt wey dey quietly disable caching on short system prompts.
Switching to a cheaper Claude model mid-session save money?
E no dey happen as much as e look like. Anthropic document say cache invalidators na changes to the tools, system or messages prefix, changes to the thinking configuration, and changes to output_config.effort. Dem no document wetin model switch dey do to existing cache, so treat am as unknown and read cache_read_input_tokens on the first request afterwards. The stake worth knowing: on a 150,000-token Opus 4.8 prefix, a cache read cost about $0.08 and a fresh write cost about $0.94, wey pass the money wey you go save for several turns of per-token saving. Switch between tasks, after /clear in Claude Code, when the cache dey throw away regardless.
Wetin output_config.effort dey do to my bill?
E dey change how many tokens the model dey spend for text, tool calls and thinking. Lower effort make fewer tool calls, wey dey compound on agentic loops because every tool result dey re-sent for later turns. The levels na low, medium, high, xhigh and max, with high as the default. Anthropic no publish cost multiplier per level, dem dey call effort as behavioural signal instead of token budget, so measure am for your own task.
How much Claude Batches API save?
50% on both input and output tokens, in exchange for asynchronous delivery. As of July 2026, e put Opus 4.8 at $2.50 in / $12.50 out, Sonnet 5 at $1 / $5 on introductory pricing, and Haiku 4.5 at $0.50 / $2.50. The comparison wey worth noticing run across tiers: a batched Opus 4.8 job cost less than a synchronous Sonnet 5 job at the standard rate from 1 September 2026, so batching dey give you stronger model for the same money.