SSD Nodes Learn Hosting plans →
How to do am Matt ConnorBy Matt Connor · Updated 2026-08-29

Why Claude output tokens cost five times pass input

Claude output tokens cost five times more because decoding happens one token at a time. See how this asymmetry affects your monthly agent bills based on your specific workload.

Why output tokens cost more than input tokens

Output tokens dey cost five times wetin input tokens dey cost for every Claude model wey dey the current catalogue. The reason na how the computation dey work. To read prompt na one pass over the model. To write reply na one pass per token, and every pass must wait for the one wey come before am.

That ratio dey the same for every row for the price list, so the model wey you pick no go change how much of your bill na output. Na your workload shape dey decide that. One agent step wey read 60,000 tokens and answer with 800 tokens no go spend almost anything for output. One drafting job wey read 2,000 tokens and write 12,000 tokens no go spend almost anything for input. We don work both cases below against the rates wey Anthropic publish for August 2026.

Prefill dey run one time, decoding dey run per token

Inference server dey handle request for two phases wey get different cost. Prefill dey read the prompt. Decoding dey write the reply.

Prefill dey take the whole prompt one time. Every prompt token dey enter the network for the same forward pass, so the attention and feed-forward work dey turn to small number of big matrix multiplications wey cover thousands of tokens at once. One read of the model weights from memory dey serve the whole prompt. The accelerator matrix units dey busy, wey mean say prefill na compute-bound: the limit na how fast the chip fit multiply.

Decoding no fit work like that, because token 2 dey depend on token 1. The token wey the model just produce dey become part of the input for the next step, so the steps no fit run at the same time. Each output token get im own forward pass, and each of those passes dey read the full set of model weights from high-bandwidth memory to produce one single token. That one dey make decoding memory-bound: the limit na how fast weights fit move, no be how fast dem fit multiply. The same weight traffic wey consume whole prompt during prefill na im go give you one token during decoding.

Serving systems dey fight back by batching. Plenty requests dey decode together, so one read of the weights dey produce one token for each request inside the batch. That na why decoding dey affordable at all. The limit na memory again. Every request wey dey active dey hold KV cache (key/value cache, the stored attention state for each token so far), that cache dey grow with every token wey dem generate, and when e fill the accelerator, the batch no fit grow again.

None of this one give you exact number, and you no suppose read 5x as measured hardware ratio. E be price, wey Anthropic set, based on that asymmetry. Wetin you fit check by yourself na the direction, and e dey take about one minute.

Measure the input and output gap yourself

Install the tools for any Ubuntu machine:

sudo apt update && sudo apt install -y curl jq moreutils

Now, stream one short prompt wey ask for long answer, and put time stamp for every line as e dey arrive.

curl -sN https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","max_tokens":1000,"stream":true,
       "messages":[{"role":"user","content":"Count from 1 to 300, one number per line."}]}' \
  | ts -s '%.s'

ts -s dey add the seconds wey don pass since the command start for front of every line. Two things dey wey you suppose notice for that output. The first content_block_delta line na your time to first token, and all the prefill happen inside that time. Every line wey follow na one small step of decoding, and the stamps go continue to increase until message_stop arrive.

Now, make you turn the shape upside down. Put long document inside the prompt and limit the answer to small amount of tokens.

curl -sN https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d "$(jq -n --rawfile doc ./long-document.txt \
       '{model:"claude-sonnet-5", max_tokens:16, stream:true,
         messages:[{role:"user", content:("Answer in one word. Is this document about networking?\n\n" + $doc)}]}')" \
  | ts -s '%.s'

The first delta take longer pass the one wey you do with short prompt, because prefill get plenty text to read. After e arrive, the response finish almost immediately, because only few tokens remain to decode. Tens of thousands of tokens enter and the clock barely move. Only few hundred come out and the clock run the whole time.

Every non-streaming response dey end with the numbers wey dem go bill you.

{
  "usage": {
    "input_tokens": 41283,
    "output_tokens": 6,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}

Log all the four fields for every request. output_tokens include extended thinking, so model wey dey think before e answer dey bill that thinking at the output rate. To know the price of prompt before you send am, POST /v1/messages/count_tokens accept the same request body, return {"input_tokens": N} without running the model, and e dey free. No be only that part of the API dey free, and which parts of the Claude API you are never billed for na something wey you suppose check before you plan your first project budget.

How much Claude dey charge per million tokens as of August 2026

ChartClaude API list rates, US dollars per million tokens, August 2026
The data behind this chart
[
  {
    "label": "Haiku 4.5",
    "input_usd": 1,
    "output_usd": 5,
    "output_multiple": 5
  },
  {
    "label": "Sonnet 5 (to 31 Aug)",
    "input_usd": 2,
    "output_usd": 10,
    "output_multiple": 5
  },
  {
    "label": "Sonnet 5 (from 1 Sep)",
    "input_usd": 3,
    "output_usd": 15,
    "output_multiple": 5
  },
  {
    "label": "Opus 5",
    "input_usd": 5,
    "output_usd": 25,
    "output_multiple": 5
  },
  {
    "label": "Fable 5",
    "input_usd": 10,
    "output_usd": 50,
    "output_multiple": 5
  }
]

The last column na output wey dem divide by input, and e dey show 5 for every row. Haiku 4.5 dey bill $1 for input and $5 for output. Opus 5 dey bill $5 and $25. Fable 5, wey be the most expensive one, dey bill $10 and $50, and wetin those Fable 5 rates give you na better tin to read before you commot that top row for your mind. As you dey move up the range, e dey multiply both sides by the same factor, so e dey change your total but e no dey change how input and output split.

Sonnet 5 appear twice because the introductory rate go expire. Up to 31 August 2026, e dey bill $2 and $10. From 1 September 2026, the standard rate of $3 and $15 go apply, wey be 50% more for both sides. Every example wey we work below use the August rate.

Rates dey change, and no be this page you suppose check dem. claude.com/pricing na the place wey get the correct info. Wetin dey survive price change na the method.

One warning wey the price list no show. Anthropic documentation talk say Claude 4.7 and later models dey use new tokenizer wey dey produce roughly 30% more tokens for the same text pass the tokenizer wey dey Sonnet 4.6 and earlier ones. If you compare two models based on price per million tokens alone, the new one go look cheaper, because the same document go get more tokens for am. Compare based on cost per finished task, and count your real prompts against the model wey you plan to use. The same trap dey between providers, because their tokenizers different pass that one, so how to calculate cost for real job on both Claude and ChatGPT go tell you more pass just to put the two rate cards side by side. wetin one million Claude tokens mean for real text explain how that volume look for real life.

When output tokens come dey cost pass input tokens?

Because dem price output 5 times pass input, e easy to calculate when you go start pay more. Make we call your input tokens I and your output tokens O. Input cost na I. Output cost na 5 times O. Output go don pass half of your bill once 5 times O big pass I, wey mean say the ratio na 5 input tokens to 1 output token.

So, if your prompt long pass five times your reply, na input tokens dey cost pass. If your prompt no reach that length, na output tokens dey cost pass.

ChartShare of spend by input to output token ratio, at 5x output pricing
The data behind this chart
[
  {
    "label": "100:1",
    "input_share_pct": 95.2,
    "output_share_pct": 4.8
  },
  {
    "label": "75:1",
    "input_share_pct": 93.75,
    "output_share_pct": 6.25
  },
  {
    "label": "20:1",
    "input_share_pct": 80,
    "output_share_pct": 20
  },
  {
    "label": "10:1",
    "input_share_pct": 66.7,
    "output_share_pct": 33.3
  },
  {
    "label": "5:1",
    "input_share_pct": 50,
    "output_share_pct": 50
  },
  {
    "label": "1:1",
    "input_share_pct": 16.7,
    "output_share_pct": 83.3
  },
  {
    "label": "1:6",
    "input_share_pct": 3.2,
    "output_share_pct": 96.8
  }
]

If the ratio na 100 to 1, output na 4.8% of your bill, and na only to reduce prompt size go make sense. If the ratio na 5 to 1, the two sides dey equal. If the ratio na 1 to 6, output na 96.8% and the prompt cost no even matter again. Plenty pipo dey guess their ratio wrong, so check your logs first before you try optimize anything.

Workload wey agent dey do: long context for inside, short answer for outside

Make we look one step wey retrieval agent dey take: 60,000 input tokens of documents wey dem retrieve plus conversation history, come add 800 token answer. That one na 75 to 1 ratio, wey be normal thing for anything wey must read before e write.

ChartOne agent step, 60,000 input and 800 output tokens, US dollars per call
The data behind this chart
[
  {
    "label": "Haiku 4.5",
    "input_cost": 0.06,
    "output_cost": 0.004,
    "total_cost": 0.064
  },
  {
    "label": "Sonnet 5 (Aug)",
    "input_cost": 0.12,
    "output_cost": 0.008,
    "total_cost": 0.128
  },
  {
    "label": "Opus 5",
    "input_cost": 0.3,
    "output_cost": 0.02,
    "total_cost": 0.32
  },
  {
    "label": "Fable 5",
    "input_cost": 0.6,
    "output_cost": 0.04,
    "total_cost": 0.64
  }
]

Output na 6.25% of that call for every model, sake of say the ratio dey fixed across the whole price list. The call cost $0.32 for Opus 5, $0.128 for Sonnet 5 for the August rate, and $0.064 for Haiku 4.5. Two hundred of those steps every day for Opus 5 na $64 per day.

The way to save money dey clear once you see how the split dey. If you cut the answer from 800 tokens go 400, you go save like 3% of the call. If you cut 20,000 tokens of old context comot from the prompt, you go save like one-third of the cost. To try reduce output length for agent wey dey read plenty things no really get sense. where a coding agent's tokens actually go explain wetin dey full that prompt for the first place.

Workload wey get big output: short prompt, long draft

Now make we turn di tin upside down. One 2,000 token brief, one 12,000 token draft, ratio of 1 to 6.

ChartOne draft, 2,000 input and 12,000 output tokens, US dollars per draft
The data behind this chart
[
  {
    "label": "Haiku 4.5",
    "input_cost": 0.002,
    "output_cost": 0.06,
    "total_cost": 0.062,
    "batch_total_cost": 0.031
  },
  {
    "label": "Sonnet 5 (Aug)",
    "input_cost": 0.004,
    "output_cost": 0.12,
    "total_cost": 0.124,
    "batch_total_cost": 0.062
  },
  {
    "label": "Opus 5",
    "input_cost": 0.01,
    "output_cost": 0.3,
    "total_cost": 0.31,
    "batch_total_cost": 0.155
  },
  {
    "label": "Fable 5",
    "input_cost": 0.02,
    "output_cost": 0.6,
    "total_cost": 0.62,
    "batch_total_cost": 0.31
  }
]

Output na 96.8% of dis bill. Opus 5 cost $0.31 per draft against $0.062 wey Haiku 4.5 go charge. Dat five times difference mostly come from di output side, wey be exactly wia cheaper model dey save you pass.

Di last column na di same job wey pass through Batch API, wey dey cut 50% comot from input and output. Opus 5 come drop go $0.155 per draft. Batch dey return result within 24 hours instead of immediately, so e fit work for report wey you wan generate for night or bulk classification. E no fit work for any tin wey person dey wait make e finish sharp-sharp.

Model routing dey pay for dis side for way wey no dey happen for agent step. If di long-long part of di job na just mechanical work, like to reformat text or expand outline wey you don already approve, di cheap model go produce dose tokens for one-fifth of di price. how to choose between Opus, Sonnet and Haiku explain wia di quality line actually dey.

Caching dey give discount for input, and na only input

Prompt caching dey store part of your prompt for server and e dey charge small fraction of the input rate to read am again. As of August 2026, the multipliers na 1.25x the base input rate to write 5 minute cache, 2x to write 1 hour cache, and 0.1x to read one hit.

Output no dey part of this deal. No cache dey for output. Every token wey the model write, you go pay full output rate for am, every time, no matter how much of the prompt come back as cache hit.

Make we look at the same agent step for Opus 5, wey get 55,000 out of the 60,000 input tokens wey come from warm cache.

ChartThe same Opus 5 agent step, with and without a warm 55,000 token cache, US dollars
The data behind this chart
[
  {
    "label": "No cache",
    "input_cost": 0.3,
    "output_cost": 0.02,
    "total_cost": 0.32
  },
  {
    "label": "55k prefix cache read",
    "input_cost": 0.0525,
    "output_cost": 0.02,
    "total_cost": 0.0725
  }
]

The call price go drop from $0.32 go $0.0725. The output line no go move: $0.02 before, $0.02 after. Caching dey cut the bill and e dey change how the cost dey look. Output na 6.25% of that call before. Now e don pass one quarter, and this one dey change which part you suppose focus on next.

The first call dey pay for the write. 5 minute cache write dey cost 1.25x base input, so e go pay for itself after just one hit. 1 hour write dey cost 2x, so e need two hits. the write and read multipliers, and where caching stops paying explain that calculation well.

Four levers wey you fit control

  1. Set max_tokens base on your p95 output length, no be base on the model maximum.
  2. Route the long-long steps go one cheaper model.
  3. Batch anything wey nobody dey wait for.
  4. Delete the instructions wey dey make reply too long.

max_tokens na hard ceiling, and to set am high no dey cost anything by itself, because na only tokens wey you produce dem dey bill you, no be the ceiling. Wetin one generous cap dey do na to remove the limit for reply wey go wrong. Pull the output_tokens distribution comot from your logs, set the cap small bit above the 95th percentile, and handle stop_reason: "max_tokens" for code by continuing the response or retrying. Truncation wey you detect dey cost less than 4,000 token wey you go pay for and come throway. Extended thinking dey land for output_tokens too, so set that budget base on the same evidence.

Routing dey work when the expensive part of one step na volume, no be judgement. Make the strong model do the decision, and give the typing work to something cheaper. Measure the routed version for your own evaluation set first, because one cheap model wey need two attempts dey cost pass one expensive attempt.

Batching na the only lever wey dey give discount for output. 50% off for both sides, results go show within 24 hours, and anything wey get schedule qualify.

The last lever na the one wey people dey skip. Phrases like "be thorough" and "explain your reasoning" dey set your output length for every call wey you go ever make. Replace dem with the shape wey you want: "Answer in at most three sentences", or "Return only the JSON object, with no preamble". One system prompt wey add 300 tokens to every reply dey cost five times wetin the same 300 tokens cost for the prompt. how to keep agent costs under control cover the monitoring side, and whether the API or a flat subscription dey cheaper for your pattern na something wey you suppose settle before you spend one week dey tune per-token spend wey subscription for don cover. For one developer, e mostly depend on whether Claude Pro $20 per month and the usage limits wey come with am cover the work wey you for dey meter. If you don dey hit those limits mid-session, to find out which window you dey wait for na the first thing, because the fix from there na smaller model, lighter context, extra usage credits, or to move that work go the metered API. If the metered API come turn out to be the cheaper place for that work, to drop to a smaller plan or cancel am no go affect the month wey you don already pay for, so to switch no dey cost you anything. If the plan wey you dey weigh Pro against na ChatGPT own instead of the metered API, the two subscription ladders wey dem place side by side go show which one come out cheaper for coding work. If na team dey ask this question, no forget say Claude Enterprise dey pair per seat fee with tokens wey dem meter at these same API rates, so every lever for this page still apply to the metered half of that bill.

FAQ

Why output tokens dey cost pass input tokens?

To generate dem dey take more accelerator time per token. Dem dey process prompt for one single forward pass, so one read of model weights dey cover thousands of tokens and hardware dey limited by multiply throughput. Dem dey produce reply one token at a time, and every token need im own forward pass wey go read all model weights again, so hardware dey limited by memory bandwidth instead. Anthropic dey price output five times the input across all current catalogue, from Haiku 4.5 go reach Fable 5.

Prompt caching dey make output tokens cheap?

No. Prompt caching dey apply to input only. As of August 2026, cache read dey cost 0.1x of base input rate, and cache write dey cost 1.25x for 5 minute duration or 2x for 1 hour duration. Dem dey bill output at full rate for every call, no matter wetin cache do. Na why caching dey change the shape of your bill and the size: once input side reduce, output come be the part wey you need watch.

High max_tokens dey cost me money if the reply come short?

No. Dem dey bill you for the tokens wey the model actually produce, so max_tokens na ceiling, e no be reservation. E still matter, because na the only hard limit for reply wey dey run away. Set am small bit above the 95th percentile of your observed output_tokens, then handle stop_reason: "max_tokens" for your code instead of make you send answer wey dem cut short without notice.

How I go find my own input to output token ratio?

Log input_tokens, output_tokens, cache_read_input_tokens and cache_creation_input_tokens from the usage object of every response, then divide the totals over one week. If the ratio pass 5 input to 1 output, your money dey the prompt, so cache the part wey stable and trim the rest. If e dey below that, your money dey the reply, so cap the length and move the steps wey dey generate most of am go cheaper model or go Batch API.