Claude cloud session credits explained
A cloud session runs Claude Code on Anthropic's machines. Learn which pool it draws from, what the credit balance really is, and what happens when it empties.
Why a cloud session spends a credit when you already pay for a plan
Cloud session credits are a prepaid balance that a cloud session draws down before it touches the usage your plan already includes. A cloud session is a Claude Code session that runs on Anthropic's machines instead of on your laptop. It still belongs to your subscription, and Anthropic's documentation is explicit that there is no separate compute charge for the virtual machine (VM) the session runs on. What the credit changes is the order in which the meter is read: while a credit balance is attached to your account, cloud sessions spend that balance first, and only after it is empty do they start counting against the same pool your terminal sessions and your chats already share.
So the answer to "why is this charging a credit when I already pay for Pro or Max" is that the credit is an extra pool sitting in front of your plan, not a replacement for it. Nothing about your subscription changed. A second meter appeared ahead of it.
Is this session running on your machine or in the cloud?
Everything else depends on this distinction, and the interface does not always make it loud. A session runs in the cloud when you start it from claude.ai/code in a browser, from the Code tab in the Claude mobile app, from the Desktop app with Cloud selected instead of Local, from the terminal with claude --cloud, or from a scheduled routine. A session runs on your own machine when you start it in your terminal, in your IDE, or in the Desktop app with Local selected.
One case looks like an exception and is not. Remote Control lets you watch and steer a session from your phone or browser while that session executes on your own machine. It is a local session, so it never touches a cloud credit, because the work is running on hardware you already own.
Cloud sessions are available on Pro, Max, and Team plans, and to Enterprise users with premium seats or with Chat + Claude Code seats.
Where does a cloud session's usage come from?
From your account's normal pool, once any credit in front of it is gone. The cloud sessions documentation states it plainly: "cloud sessions share rate limits with all other Claude and Claude Code usage within your account. Running multiple tasks in parallel consumes more rate limits proportionately. There is no separate compute charge for the cloud VM."
Read the second sentence twice, because that is the one that surprises people. Each claude --cloud command creates its own session with its own context window. Start four tasks at once and four sessions send their own requests at the same time, so the draw on your account is roughly four times what one session would use. The VM is free. The tokens are not.
The word "credit" means two different things
The interface uses one word for two separate mechanisms, and that overloading is most of the confusion behind this question.
Usage credits are the permanent mechanism. They are a prepaid balance you fund yourself in Settings > Usage, available on Pro, Max 5x, and Max 20x. They do nothing at all until you reach your plan's included limit. After that, if credits are on and you have funds, your work continues and, in Anthropic's words, "your subsequent usage will be billed at standard API pricing rates". API means application programming interface, and standard API rates are charged per token, so this is metered spending rather than an allowance. You can set a monthly spend cap, and /usage shows the month's spend against it. Usage credits cover both Claude conversations and Claude Code. For the whole picture of that balance, including the second unrelated sense the word carries, read what usage credits are and where the word gets overloaded.
A cloud session credit is a promotional balance. It is granted rather than purchased, it applies only to cloud sessions, and it is applied automatically when a cloud session starts. It is spent before your plan usage, which is why Anthropic's framing of the offer is that if you hit a limit locally you can keep going in the cloud until the credit runs out.
Here the honest answer is that the two are documented to different standards, so check your own account rather than trust any summary, including this one. As of 27 September 2026, the cloud sessions page at code.claude.com covers rate limit sharing and says nothing about a promotional credit, while the usage credits help article covers the prepaid balance and says nothing about cloud sessions. The promotional terms live in the offer itself. Reported terms for the September 2026 offer include a claim deadline in early October 2026 and an expiry date in early November 2026 for any unspent balance, but those are the offer's own dates: read them on the claim page inside your account before you plan around them.
What happens when the pool is empty?
Three different things, depending on which pool emptied.
When a promotional cloud session credit runs out, nothing stops. Cloud sessions go back to counting against your normal plan limits. This is the case worth watching, because the credit is the reason cloud work felt free, and the switch back is silent.
When your plan's included limit is reached and usage credits are on with funds available, work continues and is billed per token at standard API rates until you reach your monthly spend limit. At that point Claude Code reports a message naming the limit it hit, such as "You've hit your monthly spend limit", and on Pro and Max it offers to raise or remove that cap without leaving the CLI.
When your plan's limit is reached and usage credits are off or unfunded, work stops. You see "You've hit your session limit" or "You've hit your weekly limit", along with the time that window resets. Those limits cover every model, so /model will not rescue you. A model-specific message behaves differently: after "You've hit your Opus limit", switching to another family with /model sonnet does keep you working. The options at the moment a limit lands and how the reset windows are counted are both easier to read before you need them than during.
Nothing on this path quietly downgrades your model or trims your context to keep going. The meter either has something to draw on or it stops.
Where do you see what a cloud session consumed?
Two places, and one of them misleads you if you read it wrong.
Settings > Usage on claude.ai is the authoritative view. It holds your plan usage, whether usage credits are on, your balance, this month's spend, and your monthly spend limit. From a terminal, /usage-credits opens that same page in your browser.
/usage inside a session shows your plan usage bars, and on Pro, Max, Team, and Enterprise it adds a breakdown of where recent usage went: skills, subagents, plugins, and individual MCP servers, each as a percentage of the total. It also flags behaviors that account for a large share of recent usage, such as long context or cache misses. Press d or w to switch between the last day and the last week.
Now the trap. Those day and week figures are computed from session history on the machine you typed the command on. The documentation says so directly: "usage from other devices or claude.ai is not included." So running /usage in your terminal after an afternoon of heavy cloud work shows you a quiet day, because the cloud work happened somewhere else. Run /usage inside the cloud session, or open Settings > Usage, when you want the cloud figures.
What makes a cloud session expensive?
The cost of any Claude Code session is the size of the context it sends multiplied by how many times it sends it. Cloud sessions tend to inflate both halves, for reasons that have nothing to do with the VM.
A long background run resends the whole conversation on every turn, and each tool call adds another request carrying that batch of results. A session left working for an hour therefore pays for its own history many times over. That is the same mechanism described in what actually consumes tokens in a Claude Code session, and it applies identically to local work. The difference is that nobody is watching a cloud session, so it runs much longer before someone stops it.
Re-reading the repository is the other big one. The VM clones the repository fresh and starts with no memory of whatever you explored yesterday. A vague instruction like "clean up the error handling" makes Claude search widely and read many files just to find where the error handling lives. A specific instruction like "add the retry wrapper from src/http/retry.ts to the three callers in src/sync/" sends a fraction of the tokens, because it removes the search.
/clear does not exist in a cloud session. Locally, clearing between unrelated tasks costs nothing and keeps the next task's context small. In a cloud session you have to start a new session from the sidebar instead, so a thread you keep reusing carries everything that came before into every later request.
Idle gaps cost money too. A cloud session stops after a period of inactivity and its VM is reclaimed. Reopening the session from claude.ai/code provisions a fresh VM with your conversation history restored, though background work that was still running when the VM went away is not restored. Your first message after that gap also misses the prompt cache, which means the full context is reprocessed at full price instead of being read back cheaply, and the cache lifetime is shorter while you are drawing on usage credits than while you are inside your plan.
Controls that work inside a cloud session: /context to see what is in the window, /compact keep the test output to summarize with a focus you choose, and /model sonnet or /effort so you stop paying Opus rates for mechanical work. Note that /compact reads the conversation it summarizes, so compacting a very large context is itself a large request. Subagents keep verbose output in a separate context window, but each subagent sends its own requests, so they shrink your main context rather than your total. The shape of that overhead is covered in what an agent harness adds to every request.
Running the same job locally
The local route for this work is the ordinary CLI, and the useful pattern is to split the job rather than pick one side.
Plan on your machine, where exploration is cheap to interrupt:
claude --permission-mode planIn plan mode Claude reads files and runs commands to explore, then proposes an approach without editing source. Commit the plan, push it, and hand the execution to a cloud session:
claude --cloud "Execute the migration plan in docs/migration-plan.md"One detail catches everyone once. The cloud VM clones your repository's GitHub remote at your current branch, not your local working copy, so anything you have not pushed does not exist as far as that session is concerned. Push first.
To finish a cloud session on your own machine, pull it down:
claude --teleportTeleport fetches and checks out the session's branch, then loads the full conversation history into your terminal. It requires a clean working tree and a checkout of the same repository, not a fork. After that the terminal has its own copy of the session: new work there stays local and does not appear back on claude.ai.
Running locally does not make a single token cheaper. What it buys you is control: /clear between tasks, /resume to come back later, no VM expiry, and the ability to press Escape the second Claude heads somewhere wrong. For work you will sit next to anyway, that control is usually worth more than the convenience. For a test suite that takes twenty minutes, or a task you want progressing while your laptop is shut, the cloud earns its place. Keeping the context small pays off in both places, which is the subject of managing the context window deliberately.
When changing plan is the real answer
Credits are the expensive way to buy capacity, because they are billed per token at standard rates while a subscription is a flat fee. If you top up every week, the arithmetic has already turned against you.
Two signals say the plan is the problem rather than your habits. The first is hitting the limit during ordinary work, before you have attempted anything unusual. The second is a credit balance that empties on the same schedule every month. Tighter prompts fix neither.
What to compare, and where the plans actually differ, belongs on the plan pages rather than in a paragraph here. Start with which plan matches the way you work, and if you are already on Max, the difference between the two Max tiers. Decide with your own figures from Settings > Usage open in front of you.
FAQ
Do cloud sessions cost extra on top of my Claude plan?
No. The cloud sessions documentation states that "there is no separate compute charge for the cloud VM" and that cloud sessions "share rate limits with all other Claude and Claude Code usage within your account". The machine is included. What you spend is tokens, drawn from the same pool your terminal and your chats use, unless a credit balance sits in front of that pool and is spent first.
Why does my local /usage not show what my cloud session did?
Because its day and week figures are computed from session history on the machine you ran the command on. The documentation is explicit that "usage from other devices or claude.ai is not included", so cloud work never appears there. Run /usage inside the cloud session itself, or open Settings > Usage on claude.ai, which is the authoritative view for both plan usage and credit spend.
Does my cloud session stop when the credit runs out?
It depends which credit. When a promotional cloud session credit is exhausted, cloud sessions keep running and go back to counting against your normal plan limits. When your plan limit is reached and usage credits are on and funded, work continues billed per token at standard API rates until your monthly spend limit. When your plan limit is reached with no credits available, work stops with a message such as "You've hit your session limit" that names the reset time.
Is Claude Code cheaper to run locally than in the cloud?
Per token, no: the same request costs the same wherever the session executes, and the VM is free either way. In practice local sessions often use less, because /clear between unrelated tasks costs nothing, you can stop a run that is going wrong, and no VM expiry forces a restart that reprocesses your history. A cloud session gets expensive when it is left to work unattended on a vague instruction.