Claude Opus 5.5: price, plans and fast mode
Opus 5.5 costs $4 per million input tokens and $20 output, 20% below Opus 5. See which Claude plans include it and when fast mode is worth double the rate.
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on the Claude API, as of October 2026. Fast mode doubles both rates and gives you faster output in return. Opus 5.5 is included in the Pro, Max and Team plans. The Free plan gets Sonnet and Haiku, not Opus.
The prices in this guide come from the Claude pricing page and from Anthropic's Opus 5.5 announcement of 22 September 2026. They were checked on 3 October 2026. Prices change, so recheck the pricing page before you plan a budget around them.
Two terms first. A token is a small piece of text, often part of a word, and Claude bills by the token. MTok means one million tokens, and every price here is per MTok. To get an idea of how much text that is, read how much text one million Claude tokens really holds.
Opus 5.5 API prices: input, output and prompt caching
The data behind this chart
[
{
"label": "Input",
"usd_per_mtok": 4
},
{
"label": "Output",
"usd_per_mtok": 20
},
{
"label": "Cache write",
"usd_per_mtok": 5
},
{
"label": "Cache read",
"usd_per_mtok": "0.20"
}
]Input tokens are everything you send: your prompt plus everything sent with it, such as attached files and earlier turns of the conversation. Output tokens are the text the model writes back. Output costs five times as much as input here. A task that writes long answers therefore costs far more than a task that reads a lot and answers briefly.
The two cache rows matter most for coding agents. Prompt caching lets you mark the long, unchanging start of a prompt, such as a system prompt or a set of project files, so later requests can reuse it. The first request writes that prefix to the cache at $5 per MTok, a little more than normal input. Every later request that reuses it reads it at $0.20 per MTok, which is one twentieth of the normal input rate. Claude Code sends your project context again on every turn, so much of its input is billed as cache reads. That is why a long coding session costs less than its raw token count suggests. How Claude Code spends tokens explains where those tokens go.
How does Opus 5.5 compare with Opus 5, Sonnet 5.5 and Fable 5.1?
The data behind this chart
[
{
"label": "Opus 5.5",
"input_usd_per_mtok": 4,
"output_usd_per_mtok": 20
},
{
"label": "Opus 5",
"input_usd_per_mtok": 5,
"output_usd_per_mtok": 25
},
{
"label": "Sonnet 5.5",
"input_usd_per_mtok": 2,
"output_usd_per_mtok": 10
},
{
"label": "Fable 5.1",
"input_usd_per_mtok": 10,
"output_usd_per_mtok": 50
},
{
"label": "Opus 5.5 fast mode",
"input_usd_per_mtok": 8,
"output_usd_per_mtok": 40
}
]Opus 5 lists at $5 input and $25 output per MTok. Opus 5.5 is 20% lower on both. Sonnet 5.5, at $2 and $10, costs half as much per token as Opus 5.5. Fable 5.1, at $10 and $50, costs 2.5 times as much.
The last row is Opus 5.5 in fast mode, at $8 and $40. It is the same model, so the answers do not get better. You pay for speed only. Even so, it costs less per token than Fable 5.1.
A lower price does not make a model the right one. For a simple task, Sonnet 5.5 at half the price will often do the job just as well. For the hardest work you have, a stronger model can finish in fewer attempts, and fewer attempts means fewer tokens. Choosing between Opus, Sonnet and Haiku covers that decision. Claude Fable 5 pricing and the jobs it suits covers the model above Opus.
Is Opus 5.5 really 40% cheaper than Opus 5?
Not per token. Two different figures describe Opus 5.5, and they measure different things.
The list price fell 20%. You can check this in the chart above: every Opus 5.5 rate is four fifths of the Opus 5 rate.
Anthropic's announcement also says Opus 5.5 is "40% less to run than Opus 5". That is Anthropic's claim about the cost to finish a task. It is not a price. The cost of a task is the tokens it uses multiplied by the price per token. If the price fell 20% and the cost per task fell 40%, the rest of the saving comes from Opus 5.5 using fewer tokens on Anthropic's tasks, or a cheaper mix of input and output tokens. Your tasks are not Anthropic's tasks. To see what you save, run the same job on both models and compare the token counts each one reports.
Which Claude plans include Opus 5.5?
The announcement names Pro, Max and Team as the plans that include Opus. The Free plan gets Sonnet and Haiku only. If you are on Free and need Opus, the cheapest upgrade is Pro, described in Claude Pro's price and usage limits. If you already hit Pro's limits, the comparison of the two Claude Max tiers shows what each one adds.
Anthropic also raised the five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans with this release. The five-hour limit is the window that decides how much you can use before you have to wait for it to reset. The announcement does not publish a message count for any plan, so this guide does not give one either. A message count would mislead you anyway, because one message with large files attached uses far more of your limit than a short question. How Claude's usage limits work explains the window in detail.
The announcement also gives no context window for Opus 5.5. The context window is the largest amount of text the model can read in one request. Do not assume it matches an earlier model. Check the model page in the Claude documentation before you design around very long inputs.
Where else can you run Opus 5.5?
Outside Anthropic's own API, Opus 5.5 is available on AWS, Google Cloud and Azure. This is useful if your company already buys cloud services from one of them, because the usage is charged on the same bill. Two details differ from the Anthropic API. The model ID is often written differently on each cloud, so copy it from your provider's model catalogue instead of reusing claude-opus-5-5. The price you pay is the one on the provider's own pricing page, which this guide does not quote.
How to use Opus 5.5 in Claude Code
Claude Code is Anthropic's coding agent for the terminal. If it is not installed yet, follow the guide to running Claude on the Linux desktop and in the CLI, then come back here.
Start a session pinned to Opus 5.5 with the full model ID:
claude --model claude-opus-5-5Inside a running session, type /model to open the model picker and switch without restarting. Type /status to confirm which model the session uses. It should name Opus 5.5. If it names another model, the switch did not take effect, so run /model again and pick Opus 5.5 from the list.
To make Opus 5.5 the default for every session, set it in your user settings file at ~/.claude/settings.json:
{
"model": "claude-opus-5-5"
}You can also set it for one shell with an environment variable:
export ANTHROPIC_MODEL=claude-opus-5-5
claudeUse the full ID, not the short alias opus. The alias follows whichever model Anthropic currently ships as Opus. When a newer Opus is released, the alias moves, and your costs move with it. The full ID stays on Opus 5.5.
Note the hyphens. The ID is claude-opus-5-5, not claude-opus-5.5. A dotted ID is not a valid model name, so the API rejects the request. The API check further down shows the exact error.
Turn fast mode on and off
Type /fast in a Claude Code session to toggle fast mode. Anthropic describes it as "up to 2.5x speed" for the same Opus 5.5 model, at the doubled rate shown in the chart above. Fast mode is also offered on the Claude Platform. Leave it off by default. Turn it on for the work where you are actively waiting for the answer.
When does fast mode pay for itself?
Fast mode doubles both the input rate and the output rate. A fast session therefore always costs exactly twice the same session at standard speed, whatever the mix of input and output. The only question is whether the time you save is worth that second bill.
Here is one worked example, computed from the list prices only. Take a single agent task that sends 2 million input tokens and receives 200,000 output tokens. The example ignores caching, because the prices quoted here do not list cache rates for fast mode. Real Claude Code sessions use the cache heavily, so your standard-speed cost will usually be lower than this.
The data behind this chart
[
{
"label": "Standard Opus 5.5",
"input_cost_usd": 8,
"output_cost_usd": 4,
"total_cost_usd": 12
},
{
"label": "Fast mode",
"input_cost_usd": 16,
"output_cost_usd": 8,
"total_cost_usd": 24
}
]At standard speed the task costs $8 for input plus $4 for output, $12 in total. In fast mode it costs $24. The extra charge is the same as the whole standard bill, $12.
Now add time. Assume the task takes 60 minutes at standard speed. At the full 2.5x speed-up it would take 24 minutes, which saves 36 minutes for the extra charge. You are paying about $20 for each hour you save (12 divided by 0.6). If you only get 1.5x, the task takes 40 minutes and you save 20 minutes, so each saved hour now costs $36. "Up to 2.5x" is a ceiling. Plan with a lower figure until you have timed your own work.
So fast mode pays for itself when you sit and wait for the result, and an hour of your waiting time is worth more than that figure. Short interactive loops fit that description, such as fixing a failing test while a release waits on it. Long jobs you start and leave do not fit, such as overnight refactors or batch runs. Nobody is waiting, so the speed buys nothing.
On a subscription the picture is different. The per-token rates above are what you pay on the API and the Claude Platform. If you run Claude Code on a Pro or Max plan, check the pricing page for how fast mode is charged on your plan, because the figures in this guide do not cover it. Comparing API billing with a subscription helps you decide which way of paying fits how you work.
Check the model ID and your token use with one API call
If you pay through the API, one request confirms that your key can reach Opus 5.5. It also shows exactly what the request was billed for. Set your key in ANTHROPIC_API_KEY first.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "claude-opus-5-5", "max_tokens": 64, "messages": [{"role": "user", "content": "Reply with one word."}]}'A healthy reply is a JSON object whose model field reads claude-opus-5-5 and whose usage field lists input_tokens and output_tokens. Multiply those counts by the rates in the first chart and you have the cost of the call. The usage field also reports cache_creation_input_tokens and cache_read_input_tokens. Those are billed at the cache write and cache read rates.
Send the dotted ID claude-opus-5.5 instead, and the API answers with an error like this:
{"type":"error","error":{"type":"not_found_error","message":"model: claude-opus-5.5"}}not_found_error means the API has no model by that name. The fix is the hyphenated ID. An authentication_error with HTTP status 401 is a different problem: the key is missing or wrong. Check that ANTHROPIC_API_KEY is set in the same shell that runs curl.
FAQ
How much does Claude Opus 5.5 cost per million tokens?
As of October 2026, Claude Opus 5.5 lists at $4 per million input tokens and $20 per million output tokens. Cache writes cost $5 and cache reads $0.20 per million tokens. That is 20% below Opus 5. Recheck the Claude pricing page before you budget, because prices change.
Can I use Opus 5.5 on the free Claude plan?
No. Anthropic's announcement lists Opus on the Pro, Max and Team plans, and the Free plan gets Sonnet and Haiku. Upgrading to Pro is the cheapest way to get Opus 5.5 in the Claude apps. The API is a separate option, billed per token.
What does fast mode cost in Claude Code?
Fast mode bills Opus 5.5 at $8 per million input tokens and $40 per million output tokens. That is double the standard rate, for what Anthropic calls "up to 2.5x speed". The model is the same, so the answers are not better. Toggle it with /fast, and use it only when you are waiting for the result.
What is the context window of Claude Opus 5.5?
Anthropic's announcement does not state one, and this guide does not guess. Check the model page in the Claude documentation before you plan work that sends very long inputs in one request.
Is Opus 5.5 really 40% cheaper than Opus 5?
The per-token list price is 20% lower. The "40% less to run" figure is Anthropic's claim about the cost to finish a task, and that cost also depends on how many tokens the model uses. To find your real saving, run the same task on both models and compare the token counts each one reports.