JuliusBrussee/caveman

β˜… 105,636β‘‚ 0

πŸͺ¨ why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.

105,636Star
0Fork
0Watch
0Issue
GoLanguage
-License
Created Β· last push Β· repository size 0 KB Β· default branch -

README

https://github.com/JuliusBrussee/caveman/blob/HEAD/Caveman

why use many token when few do trick

Your AI coding agent bills by the word and writes like it knows that. Caveman make it stop.

https://github.com/JuliusBrussee/caveman/blob/HEAD/ThePrimeagen reacts to Caveman: No way this actually works

▢️ ThePrimeagen reacts: "No way this actually works"

https://github.com/JuliusBrussee/caveman/blob/HEAD/GitHub stars https://github.com/JuliusBrussee/caveman/blob/HEAD/npm downloads https://github.com/JuliusBrussee/caveman/blob/HEAD/30+ agents https://github.com/JuliusBrussee/caveman/blob/HEAD/10 native wrap profiles https://github.com/JuliusBrussee/caveman/blob/HEAD/License https://github.com/JuliusBrussee/caveman/blob/HEAD/skills.sh

πŸ† #1 on GitHub Trending Β· July 2026  Β·  πŸ₯‡ #1 Repository of the Day on Trendshift Β· April 2026

#1 on Hacker News Β· 904 points Β· 366 comments  Β·  #8 Product of the Day on Product Hunt

πŸ“„ Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style output cutting cost 1.4 to 2.4Γ—, up to 3Γ—  Β·  πŸ§ͺ Tested by JetBrains on 86 real coding tasks: "costs you nothing measurable in quality"

https://github.com/JuliusBrussee/caveman/blob/HEAD/Caveman - why use many token when few do trick | Product Hunt https://github.com/JuliusBrussee/caveman/blob/HEAD/JuliusBrussee%2Fcaveman | Trendshift

⚑ One command, no account, no API key. npx skills add JuliusBrussee/caveman -g β†’ Quick Start

---

See it Β· Quick Start Β· The Numbers Β· In the Wild Β· The Skill Β· The Proxy Β· Wrap Β· When to Skip Β· Docs

---

πŸͺ¨ See it

πŸ—£οΈ Normal agent Β· 69 tokens https://github.com/JuliusBrussee/caveman/blob/HEAD/ Caveman agent Β· 19 tokens
The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.

New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo.

Same diagnosis. Same fix. Same useMemo. The only thing that died was the throat-clearing.

Code, commands, file paths, and exact error messages never get cavemanned. Only the prose around them does. Security warnings and "are you sure?" confirmations come back in full sentences on their own, then caveman resumes.

Caveman no make brain smaller. Caveman make mouth smaller.

Half the fun is that your agent talks like it just discovered fire. The other half is that it is still right.

---

🌍 Why this exists

A token is what AI billing counts, roughly three quarters of a word. Your agent pays for every token it writes and every token it reads. Most agents write like a cover letter and read like a firehose.

Caveman attacks both ends:

Started as a joke on a Friday in April 2026. Hit 4,000 stars in a week. Now past 100,000, with a research paper, a JetBrains lab test, and a Primeagen reaction video. The joke got serious. The voice did not.

---

⚑ Quick Start

Caveman come in two sizes. Start small.

Small rock: the skill

A rule file that makes your agent answer in caveman. MIT, free forever, works in 30+ agents (Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot, more). One command:

npx skills add JuliusBrussee/caveman -g

Type /caveman if your agent doesn't wake up on its own. That the whole install. One rock.

Big rock: the proxy

Runs on your machine, between your agent and the AI provider, and shrinks what the agent reads before every call. MIT CLI, BSL-1.1 runtime:

npm install -g @caveman-ai/cli && caveman setup --install
caveman claude        # or codex Β· gemini Β· aider Β· kilo Β· qwen Β· opencode Β· hermes Β· openclaw Β· pi

They stack. Most people start with the small rock and graduate.

More doors into the cave Β· full installer, Windows, single agents, uninstall


The full installer wires up Claude Code hooks and the statusline badge, finds every supported agent on your machine, and skips agents you no have. Safe to re-run. Needs Node.js 22.13+.

curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/v2.7.0/install.sh | bash

Windows, PowerShell 5.1+:

irm https://raw.githubusercontent.com/JuliusBrussee/caveman/v2.7.0/install.ps1 | iex

Just one agent:

# Claude Code
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman

Gemini CLI

gemini extensions install https://github.com/JuliusBrussee/caveman

Qwen Code CLI, then its Caveman wrapper

npm i -g @qwen-code/qwen-code caveman qwen

Codex, Cursor, Windsurf, Cline, and other skills-compatible agents

npx skills add JuliusBrussee/caveman --skill '*' -a codex --yes -g # replace codex with your agent profile

Install broke? Open your agent in this repo and say: "Read CLAUDE.md and INSTALL.md, install caveman for me." Agent read repo, agent fix own brain. Snake eat tail.

Changed your mind: npx -y github:JuliusBrussee/caveman -- --uninstall

The full 30+ agent matrix, dry runs, flags, and verification live in INSTALL.md.

πŸ• The first five minutes

Small rock. The skill, right after npx skills add:

1. Ask it something. Any coding question. Watch the preamble vanish and the answer stay. 2. Turn the dial. /caveman lite for tight-but-polite. /caveman ultra for grunts. /caveman wenyan for classical Chinese, because someone asked. 3. Commit like a caveman. /caveman-commit writes a Conventional Commit in one line. 4. Review like a caveman. /caveman-review gives one finding per line: L42: πŸ”΄ null deref. Guard it. 5. Shrink your memory files. /caveman-compress CLAUDE.md cuts the prose, keeps every heading, path, and command, and backs up the original. 6. Come home. Say stop caveman. Normal prose returns. No hard feelings.

Big rock. The proxy, right after npm install -g @caveman-ai/cli:

1. Find out where your tokens go. caveman learn reads months of agent history already on your disk, locally, and ranks your token sinks worst-first with a one-line fix behind each. Do this before anything else. It is the most useful five minutes in this README. 2. Let it fix them. caveman learn implement hands each fix to Claude Code or Codex one diff at a time, applied only on your yes, and reverts anything that did not lower tokens per turn. 3. Wrap your agent. caveman claude (or codex, gemini, aider, opencode, pi, …) puts the proxy in front of it. Logs, test output, JSON, and diffs get shrunk before the provider sees them. Originals stay on disk, and the agent can pull any of them back. 4. Shrink the noisy stuff. caveman shrink -- pnpm test compresses command output. caveman browse gives the agent a compressed view of a web page instead of a 15,000-token accessibility dump. 5. Prove it on your own work. caveman trial -- claude runs a real session with and without caveman, then caveman trial report shows the difference. That A/B outranks every number on this page. 6. Shrink caveman itself. caveman convert --dry-run shows which installed skills get cheaper as PNG pages the model reads as an image. Convert the profitable ones, revert byte-for-byte any time. 7. Watch the bill. caveman stats for history and estimates. /caveman-stats inside Claude Code for that session.

---

πŸ“Š The Numbers

Every number below is either from a committed run in this repo or from a named third party. Nothing rounded up. Where a number is small, it says so. Where a row is red, it stays red.

What the skill saves (writing less)

| Who measured | What they measured | Result | |---|---|---| | Adobe Research (CAVEWOMAN, arXiv 2606.24083) | Eight models, five datasets, five compression levels | Output-side caveman style cuts realized cost 1.4 to 2.4Γ— per model, up to 3Γ— in the best case | | JetBrains | 86 real coding tasks, paired A/B, Claude Code 2.1.200. Skill only, no proxy (July 2026, before the proxy existed) | 8.5% fewer output tokens, about 10% cost. No detectable quality change (sign test p = 0.82) | | This repo (committed eval snapshot) | Ten dev questions, skill vs a plain Answer concisely. control, claude-opus-4-6 | 50% fewer output tokens at the median on top of the terse control. Length only, not correctness |

Read those three together and you get the honest picture. Chat-style Q&A: big cut. Agentic coding sessions, where most tokens are code and tool calls that the skill never touches: high single digits on output, quality flat.

The JetBrains number is why the proxy exists. They measured the skill alone, in July 2026, before the proxy shipped. Their finding was that an agent's bill is mostly reading, not writing, and no talking style fixes that. So we built the thing that shrinks the reading. The table below is what that changed.

The Adobe paper's other finding matters too: compressing the human's prompt into caveman-speak makes models answer longer and worse. Caveman never rewrites your prompts. Only the agent's mouth.

The rules add input tokens on every call, and whether shorter output pays for them depends on your agent, caching, and billing. Full accounting: docs/HONEST-NUMBERS.md.

No reviewed API benchmark result is published here yet. Run uv run python benchmarks/run.py to generate a new result, then review its raw response pairs and quality before publishing the generated table.

What the proxy saves (reading less)

Your agent rereads logs, test output, diffs, and half your repo all day. The proxy shrinks that stream before it reaches the provider. Pinned 54-run Claude Code benchmark, provider-reported input tokens, three runs per case, every answer checked against an exact oracle:

| Case | Direct Claude Code | Through caveman | Change | | ---------------------- | -----------------: | --------------: | ---------: | | CSV outlier hunt | 165,823 | 74,484 | -55.1% | | Log needle in haystack | 148,807 | 74,068 | -50.2% | | YAML config drift | 132,124 | 71,027 | -46.2% | | Test output failure | 150,377 | 108,514 | -27.8% | | Deployment JSON drift | 147,975 | 108,939 | -26.4% | | Dashboard HTML alert | 140,687 | 154,641 | +9.9% | | Total | 885,793 | 591,673 | -33.2% |

18 of 18 answer checks passed. Case-clustered 95% interval: 14.6% to 48.5%. In the same suite, Headroom's wrap saved 6.7% and failed 3 of 18 checks. Method, provenance hashes, and limits: docs/WRAP-BENCHMARK.md. Raw harness artifacts are not in this checkout, so treat it as a pinned report, not a public reproduction.

Maintainer note. The HTML row is red and it stays red. That case had no compression transform, so caveman paid its own overhead and won nothing back. The day I hide a red row is the day you should stop trusting the green ones.

Everything else caveman shrinks

| Surface | Measured | Number | |---|---|---| | Browser pages | Focused question against a 200-row table, vs the Playwright ARIA snapshot | 121 tokens vs 15,704. 129.8Γ— smaller. Tiny forms lose 2.3Γ—; the benchmark says so | | Memory files (/caveman-compress) | Five real CLAUDE.md-style fixtures | 46% smaller on average, headings, code, paths, and URLs verified intact | | The skill itself (pixel mode) | Rendered to PNG pages the model reads as an image | 1,069 to 415 estimated tokens, a 61% cut | | Your harness prefix (subagent-tax) | What every subagent re-sends before doing any work | On one real machine, 219k of a 267k-char request was tool schemas. Run it on yours |

---

πŸ“£ In the Wild

ThePrimeagen Β· "No way this actually works"
Full reaction on The PrimeTime β†’

Adobe Research Β· CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
Adeyemi, Rossi, Dernoncourt Β· arXiv, June 2026 Β· cites this repo. The style is now a benchmarked register.

JetBrains Β· Speaking to AI Agents like Cavemen Saves 65% of Tokens. We Test.
The most rigorous outside A/B so far, run on the skill alone before the proxy existed. Their verdict: "Use it if you like it. It is fun, and it costs you nothing measurable in quality." Their 8.5% is the number that made us build the proxy.

Hacker News Β· #1, 904 points, 366 comments

The New Stack Β· Getting Claude Code to grunt in Caveman-speak might not save as many tokens as you think
Fair headline. We link it anyway. See The Numbers.

GitHub Trending Β· #1 overall, July 2026
Trendshift Β· #1 Repository of the Day (April 2026) Β· #1 JavaScript repo of the month (April) Β· #1 Go repo of the month (July)

Product Hunt Β· #8 Product of the Day

Star History Chart

---

πŸ’¬ The skill, unpacked

One rule file, one talking style, plus a small toolbox. /caveman lite|full|ultra|wenyan-lite|wenyan-full|wenyan-ultra sets intensity. /caveman off or normal mode turns it off.

| Level | Same question: "Why does my React component re-render?" | |---|---| | lite | Your component re-renders because you create a new object reference each render. Wrap it in useMemo. | | full (default) | New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo. | | ultra | Inline obj prop, new ref, re-render. useMemo. | | wenyan-full | 每ηΉͺζ–°η”Ÿε°θ±‘εƒη…§οΌŒζ•…ι‡ηΉͺοΌ›δ»₯ useMemo εŒ…δΉ‹ε‰‡ε…γ€‚ |

Three things the skill will never do: shorten your code, paraphrase an error message, or grunt through a security warning. It drops to full sentences for anything irreversible, then picks the club back up.

Everything in the box Β· commit messages, reviews, subagents, work patterns


| Tool / command | What you get | | ----------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | /caveman [lite\|full\|ultra\|wenyan-lite\|wenyan-full\|wenyan-ultra\|off] | Shorter replies at the intensity you choose. | | cavecrew-investigator, cavecrew-builder, cavecrew-reviewer | Compressed subagent presets for locating, editing, and reviewing code. | | /caveman-commit | Terse Conventional Commit messages. | | /caveman-review | One-line, actionable review findings. | | /caveman-compress | Smaller Markdown memory files, with the original backed up. | | /caveman-stats | Recorded Claude Code token usage; savings unknown without a measured comparison. | | /caveman-help | One-screen reminder of every mode and command. | | investigate-first, lean-build, surgical-patch, safe-refactor, migration, verify-and-stop | Work patterns that write less code, so the agent bills fewer tokens. Your agent picks these up on its own when a task fits. | | /caveman-setup, /caveman-discover, /caveman-learn, /caveman-manage, /caveman-optimize, /caveman-explore, /caveman-evidence-review | Drive the caveman engine and proxy: set it up, find where tokens go, act on what it finds. |

---

πŸ”§ The proxy, unpacked

One local process. Your agent talks to it, it talks to your provider. No Caveman server in the path, and your Claude Pro/Max login passes through to Anthropic untouched. Originals of everything it compresses sit in a SQLite file on your machine with a recovery handle, so the agent can always ask for the full version back.

 Your agent  (Claude Code Β· Codex Β· Gemini Β· Aider Β· opencode Β· Pi Β· …)
      β”‚   tool output Β· logs Β· JSON Β· diffs Β· search results
      β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚  caveman proxy   (your machine, your keys)          β”‚
 β”‚  detect() β†’ json Β· log Β· code Β· diff Β· search Β· textβ”‚
 β”‚  originals β†’ local SQLite, recovery handle returned β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚   smaller prompt, same answer
      β–Ό
 Your provider  (Anthropic Β· OpenAI Β· Google Β· Bedrock Β· Vertex Β· Azure Β· OpenRouter)

Whole team? One container. Same proxy in your VPC, one shared token, keys stay on server. Deploy it β†’

https://github.com/JuliusBrussee/caveman/blob/HEAD/Terminal demo: caveman compress reads a large JSON payload and emits a much smaller compressed version, byte-exact recoverable

What the engine keeps, by payload type Β· and the wrap stack diagram


https://github.com/JuliusBrussee/caveman/blob/HEAD/coding agent talks to a local caveman proxy that forwards upstream to the provider with auth passed through byte-exact; a CCR store below the proxy keeps the original bytes and returns a recovery handle to the agent; an MCP toolkit side-channel gives the agent caveman_retrieve, toon encode/decode, and browse

detect() types each payload and routes it to a compressor that keeps what answers depend on:

| Detected type | Keeps | Target savings | | --------------- | ---------------------------------------------------------------------- | -------------- | | json | keys, structure, error/message subtrees; collapses repetitive arrays | 70-90% | | log | errors, stack traces, first/last lines; drops INFO and progress noise | 85-95% | | code | imports, signatures, types; elides function bodies, syntax stays valid | 40-70% | | diff | file/hunk headers and changed lines; elides repeated context | 60-80% | | search-result | top/bottom hits plus diagnostic/security hits | 80-95% | | text / HTML | headings, opening/closing context, important sections | 50-80% |

contextwindow.Pack() additionally fits candidate context into a token budget by BM25 relevance, recency, and error signal, returned in original order so chronology survives.

Any MCP host gets the same powers through five tools: caveman_compress, caveman_retrieve, caveman_stats, caveman_toon_encode, caveman_toon_decode.

Where your tokens go

Months of your agent history already sit on your disk. caveman learn reads it, locally, read-only, no account, and ranks your token sinks worst-first with a one-line fix behind each.

caveman learn             # Claude Code + Codex + Gemini CLI + opencode; aider via CAVEMAN_AIDER_ROOT
caveman learn implement   # hand the fixes to Claude Code or Codex, one diff at a time, applied only on your yes

https://github.com/JuliusBrussee/caveman/blob/HEAD/Caveman Learn report: TLDR summary and savings cards on the left; ranked token sinks with an expanded fix and a session context depth histogram on the right

implement re-measures after every change and reverts anything t

More AI Agent Skills Trending projects

1

affaan-m / ECC

JavaScriptβ˜… 258,555β‘‚ 0
β†’
2

NousResearch / hermes-agent

Pythonβ˜… 245,605β‘‚ 0
β†’
3

deepseek-ai / deepseek-harness

TypeScriptβ˜… 224,514β‘‚ 0
β†’
4

firecrawl / firecrawl

TypeScriptβ˜… 180,546β‘‚ 0
β†’
5

anthropics / skills

Pythonβ˜… 176,366β‘‚ 0
β†’
6

langchain-ai / langchain

Pythonβ˜… 146,352β‘‚ 0
β†’