Of all the ways to make an AI coding assistant cheaper, "grunt at it" is not the one most engineers would have predicted. Yet that is the premise of Caveman, a viral skill and proxy that ranked among GitHub's trending repositories this week with roughly 108,900 stars — 271 of them added in a single day — and a README that states the thesis in broken English: "why use many token when few token do trick".

The mechanism is straightforward. Left to itself, a coding agent spends a large share of its output on preambles, restatements and polite summaries — the connective prose around the code. Caveman instructs the model to drop it and answer in compressed, telegraphic form, claiming to cut output tokens by around 65% while keeping code and technical claims intact. Version 2.0 pushed further, compressing the input side of the conversation as well; version 3.0, released in recent weeks, put the entire repository — engine, proxy, browser tooling, an MCP server, a compression module and the "cavemem" memory core — under Apache 2.0.

The idea has spread far beyond one repository. GitHub trending now carries a whole genre of context-economy tools: proxies that shrink prompts, pre-indexed knowledge graphs that stop an agent re-reading a codebase, and skills that trim session output. Several of this week's fastest-rising repositories were built on exactly that promise.

The scepticism is equally visible. A recurring critique holds that output tokens are a small slice of a typical session's cost — some estimates put visible prose at roughly 1 to 10% — so aggressive trimming of the model's prose can leave the bill almost unchanged even as the transcript looks dramatically shorter. One widely shared write-up was bluntly titled "Caveman promises 65% fewer tokens — my bill didn't move".

That tension is the interesting part. Caveman works in the narrow sense that it changes model behaviour exactly as advertised; whether it saves what a developer actually pays depends on where their tokens were going in the first place. As agents grow more autonomous and their loops longer, the appetite for tools that shrink that loop shows no sign of fading.