Caveman Skill: Make Coding Agents Say Less
Caveman shortens what a coding agent writes, and its proxy shrinks what the agent reads. Install it pinned to a release, then measure your own sessions.
What the Caveman skill does
The Caveman skill makes a coding agent answer in short, plain sentences, so each reply costs fewer output tokens. The same project ships a second, separate piece: a local proxy that compresses logs, JSON, CSV and test output before the model reads them. The two halves save different tokens. Judge each one on its own numbers, and measure your own sessions before you trust either.
Caveman is an open-source project by Julius Brussee, dual licensed under Apache-2.0 and MIT. As of October 2026 the latest release is v3.1.0, tagged on 2026-10-04. It installs into Claude Code, Codex, Gemini CLI, Cursor and many other agents. If skills are new to you, what an agent skill is and how an agent loads one covers the format Caveman builds on.
Two halves that save different tokens
A coding agent session spends tokens in four places:
- Output: the text the model writes back to you. The Caveman skill works here.
- Tool results: file reads, command output, logs and API responses that the agent pulls into its context. The Caveman proxy works here.
- Harness overhead: the system prompt, tool definitions, rules files and memory files that the agent sends with every request. Neither half of Caveman touches this.
- History: every earlier turn, sent again on each new request until the context is compacted.
The split matters because the halves pull on different parts of the bill. Output tokens cost several times more per token than input tokens on most model price lists. But a session reads far more than it writes, because every request sends the whole context again. So a modest cut in what the agent reads can be worth more than a large cut in what it writes. In a chat-heavy session with little tool use, the reverse is true.
Caveman sits beside two tools we have covered. Ponytail, the lazy senior dev agent works on the same output side as the Caveman skill. Tura, a harness built to send fewer tokens works on the harness overhead line, which Caveman leaves alone. To find out how large that line is before you choose, read how much a harness adds to every request.
Half one: the skill changes what the agent writes
The skill is a set of writing rules that the agent follows while it is active. The README lists them:
- Answer first, with no greeting and no recap.
- One idea per sentence, at most 20 words.
- Keep the meaning: negations and numbers stay intact.
- Keep the payload verbatim: code, commands, paths and error messages are copied character for character.
- Stay quiet between tool runs instead of narrating each step.
In v3.1.0 the skill has three modes, which replaced the older intensity levels. /caveman is the standard mode. /ultracave is the aggressive mode, and it is where most of the measured output saving comes from. /megacave is the third mode. /caveman off turns the style off, and /caveman status shows which mode is active.
The plugin also adds task commands. /caveman-commit writes a one-line Conventional Commits message. /caveman-review returns code review findings, one per line. /caveman-compress <file> shrinks a memory file such as CLAUDE.md. /caveman-stats shows the token use of the current session, and /caveman-help prints the full command reference.
The skill only changes output. The project's own notes put its input reduction at 0%, because it is an output-style instruction. Its rules are also input: the model reads them on every request where the style is active.
Half two: the proxy compresses what the agent reads
The proxy is a separate program, installed through npm. It runs on your machine and sits between the agent and the model provider. It uses your own API keys, or your Claude Pro or Max login. Each request passes through it, and it compresses the bulky parts before they go to the provider.
The proxy checks what kind of payload each piece is, and sends it to a matching compressor. Logs keep errors and stack traces and lose progress lines. Code keeps imports and function signatures, which it finds with tree-sitter (a parsing library that reads source code as a syntax tree). JSON keeps its structure plus any subtree that holds an error. The original files stay on disk, and the agent can read them again in full.
The proxy does not change what the agent writes. Run it without the skill and replies stay as long as before. The two halves are independent, so you can install either one alone.
How to install the Caveman skill in Claude Code
The README gives a plugin marketplace route for Claude Code. Add #v3.1.0 to the marketplace name, so you get that release and not whatever lands on the default branch next. That is the right default for a tool that changes the text your agent produces.
claude plugin marketplace add JuliusBrussee/caveman#v3.1.0
claude plugin install caveman@caveman
claude plugin listclaude plugin list should show caveman@caveman with an enabled status. The #<ref> suffix pins the marketplace catalog to that branch or tag. The README line leaves the tag off, which means your copy tracks the default branch. Check the releases page on GitHub for a newer tag before you copy this, and replace v3.1.0 with it.
Start a session and type /caveman. The next reply should open with the answer and no greeting. Type /caveman status to confirm the mode.
To move to a newer release later, remove the marketplace and add it again with the new tag. Removing a marketplace also uninstalls its plugins, so run the install line again afterwards.
claude plugin marketplace remove caveman
claude plugin marketplace add JuliusBrussee/caveman#v3.1.0
claude plugin install caveman@cavemanIf plugins and marketplaces are new to you, how Claude Code plugins are packaged and installed explains the parts.
Install with npx for Codex and other agents
The README's universal route uses the skills CLI, which installs a skill into the supported agents it finds:
npx skills add JuliusBrussee/caveman -gThe -g flag installs into your user directory, so every project sees the skill. Without it, the skill goes into the current project only. npx skills list shows what is installed. This line follows the default branch, like the unpinned marketplace line.
For a pinned install, the README gives the project's own installer at a release tag:
npx --allow-git=root -y github:JuliusBrussee/caveman#v3.1.0npm 12 and newer block packages installed straight from a git repository, which is why the --allow-git=root flag is there. On an older npm, drop that flag. The installer detects the agents on your machine and runs each agent's native setup, such as a plugin or a skills folder. For Claude Code it also installs hooks and a status line badge, so the style is on without typing /caveman each time. It also accepts --dry-run, which prints what it would write without writing anything. How Claude Code hooks run covers what those hook entries do.
Check the Claude Code side after the installer finishes:
cat ~/.claude/.caveman-activeThe file should print caveman. If cat reports No such file or directory, the installer did not set up Claude Code. Run it again and read its output for the Claude Code step.
To remove everything the installer added:
npx -y github:JuliusBrussee/caveman -- --uninstallOn npm 12 and newer, add --allow-git=root here too.
Set up the Caveman proxy
The proxy ships as the @caveman-ai/cli npm package. Check the current version first, so you know which release you are running:
npm view @caveman-ai/cli version
npm install -g @caveman-ai/cli
caveman setup --install
caveman claudeTo pin the CLI, install @caveman-ai/cli@<version> with the version that npm view printed. If npm install -g fails with EACCES: permission denied, your global npm prefix is owned by root. Use a Node version manager, or set a user-owned prefix, rather than running npm as root. caveman setup --install configures the proxy and your provider login. caveman claude starts Claude Code with its traffic going through the proxy. The same pattern works for other agents, for example caveman codex. If you run Codex on a ChatGPT plan, using Codex with a ChatGPT subscription covers that login.
You can also compress the output of a single command without routing a whole session through the proxy. The README's example:
caveman shrink -- pnpm testOn a VPS, check where the proxy listens. Its documentation says it binds to the loopback address 127.0.0.1:8787 by default. It has no inbound authentication unless you set CAVEMAN_AUTH_TOKEN, so anything that can reach the port can send requests through it.
ss -tln | grep 8787The line should show 127.0.0.1:8787. If it shows 0.0.0.0:8787 or your public address, the proxy is open to the network. Fix that before you do anything else. The same care applies to every tool you run beside the agent, and the CLI tools a coding agent needs on a VPS covers the rest of that setup.
The CLI sends usage telemetry by default. The README lists what it sends: commands run, token counts, a random install ID, the OS and your IP address. It says it never sends prompts, code or file paths. Turn it off with caveman telemetry off, or set DO_NOT_TRACK=1 in the environment.
What the Caveman benchmarks claim, and under what conditions
Every number in this section is the project's own claim, copied from its v3.1.0 README and its docs/HONEST-NUMBERS.md page. We have not reproduced them. Each one was measured under narrow conditions, and the conditions matter more than the headline.
The data behind this chart
[
{
"category": "CSV",
"before_tokens": "28,041",
"after_tokens": "314",
"reduction_pct": 98.9
},
{
"category": "Logs",
"before_tokens": "22,810",
"after_tokens": "348",
"reduction_pct": 98.5
},
{
"category": "JSON",
"before_tokens": "18,837",
"after_tokens": "281",
"reduction_pct": 98.5
},
{
"category": "YAML",
"before_tokens": "20,447",
"after_tokens": "178",
"reduction_pct": 99.1
},
{
"category": "All six file types, including HTML",
"before_tokens": "130,611",
"after_tokens": "22,994",
"reduction_pct": 82.4
}
]The CSV file went from 28,041 tokens to 314, a 98.9% cut. These are per-file best cases. Each row is one large file of one type, passed through the compressor that suits it. A real session mixes such files with source code, short command output and conversation, which compress far less. The combined row covers all six types the project tested, and it falls to 82.4%. Part of that gap is HTML: the README says HTML has no compressor yet, so HTML adds work without a saving. Test output was also tested, and the README reports 98.9% for it without the raw token counts.
The session figure is the one closest to a bill:
The data behind this chart
[
{
"config": "Claude Code, no proxy",
"input_tokens": "885,793",
"reduction_pct": 0
},
{
"config": "Claude Code, with proxy",
"input_tokens": "591,673",
"reduction_pct": 33.2
}
]The project describes this as 54 Claude Code runs on its own tasks, with 18 of 18 answers right. Input went from 885,793 to 591,673 tokens, which is 33.2% fewer. That is whole-session input, so it includes the system prompt and every part the proxy cannot compress. It is a far more honest number than the per-file rows. It is still a measurement of their tasks, not yours.
The skill figures are smaller:
The data behind this chart
[
{
"config": "Answer concisely (control)",
"output_tokens": "4,334",
"median_reduction_pct": 0
},
{
"config": "/caveman",
"output_tokens": "4,119",
"median_reduction_pct": 3
},
{
"config": "/ultracave",
"output_tokens": "2,693",
"median_reduction_pct": 35
}
]The control here is a plain one-line instruction, "Answer concisely." Against that, standard /caveman saved a median of 3%, and /ultracave saved 35%. So the standard mode is a small step past an instruction you could write yourself, and the aggressive mode is where the output saving is. The run was ten questions, one run each. It counted output length with the tiktoken o200k tokenizer, an OpenAI tokenizer used as an estimate and not Claude's own count. The project's notes say plainly that token counts "do not prove semantic or technical equivalence": a shorter answer is not shown to be an equally correct one.
Measure your own session before you trust a number
The project's own advice is the right one: compare provider-billed totals on the same task, with and without Caveman, and turn it off for any workload where it costs more. Compare billed totals and not raw token counts. Cached input is billed at a fraction of the normal input rate, so a cut in raw input tokens does not turn into the same cut in cost.
- Pick a real task you can repeat, such as fixing one failing test in your own repository.
- Run it with Caveman off, and record the billed tokens and whether the fix was right.
- Run it again with the skill only.
- Run it again with the skill and the proxy.
- Compare the billed totals, and compare the quality of the result.
The CLI has a helper for this: caveman trial -- claude runs an A/B comparison on real work, and caveman stats shows your token history. For a view that does not depend on Caveman's own counting, read how Claude Code counts input, output and cache tokens and use one of the spend-tracking tools for Claude Code. On a subscription plan, the number that matters is how fast you reach your usage limit, not a dollar figure, and API billing versus a subscription explains why those two behave differently.
The project's notes record cases where Caveman produced a net loss: terse coding question-and-answer sessions, and per-request billing. In a short session with few tool calls there is little for the proxy to compress, and the skill's own rules are extra input on every request.
When terse output and compressed input hurt
Terse output hurts reviews and handoffs. The agent writes for you, now, with the whole session in its context. A teammate who reads the same reply tomorrow does not have that context. A one-line commit message from /caveman-commit says what changed but not why, and the why is what a reviewer needs six months later. A /caveman-review list of one-line findings is fine for the person who asked, and thin for a junior teammate who has to act on it. Turn the style off before you ask for a pull request description or a handoff note: type /caveman off, then ask.
Proxy compression can hide the one log line that mattered. The compressor decides what to keep by the shape of each line. Errors and stack traces match that shape and survive. A warning that does not look like an error, or a note that three tests were skipped, can be dropped as noise. The agent then reasons from the shortened version, and it does not know what it never saw. When a failure makes no sense, run the raw command yourself without caveman shrink or the proxy, or ask the agent to read the original file, which the proxy keeps on disk.
Smaller replies are not always cheaper sessions. If a terse answer leaves out a step and you ask a follow-up, that second request sends the whole context again. One extra turn can cost more input than the skill saved in output. This is one more reason to measure whole tasks and not single replies.
FAQ
Does the Caveman skill reduce input tokens?
No. The project's own notes put the skill's input reduction at 0%, because it is an output-style instruction. Its rules are themselves extra input while the style is active. Input savings come only from the Caveman proxy, which is a separate install through the @caveman-ai/cli npm package.
Will I see the 98% compression figures in my own sessions?
Probably not. The 98% figures are per-file best cases: one large CSV, log, JSON or YAML file passed through its matching compressor. The project's whole-session figure is 33.2% fewer input tokens, measured on its own tasks. Your result depends on how much of your session is bulky tool output, so measure the same task with and without the proxy and compare billed totals.
How do I pin Caveman to a release in Claude Code?
Add the release tag to the marketplace name: claude plugin marketplace add JuliusBrussee/caveman#v3.1.0, then claude plugin install caveman@caveman. Without the tag your copy tracks the default branch. Check the project's releases page for the newest tag first. For other agents, the pinned installer is npx --allow-git=root -y github:JuliusBrussee/caveman#v3.1.0.
How do I turn Caveman off for a handoff or a review?
Type /caveman off in the session, then ask for the pull request description or handoff note. /caveman status confirms the current mode. To remove Caveman completely, run npx -y github:JuliusBrussee/caveman -- --uninstall, adding --allow-git=root on npm 12 and newer, or remove the marketplace with claude plugin marketplace remove caveman.