Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Cut LLM output tokens 40–70% by stripping grammatical scaffolding while preserving every fact — telegraphic output modes, when they pay (pipelines, long sessions) and when they don't (single shots, human-facing prose), with the mode lines to switch on demand. Use when asked make the model respond tersely, cut output token costs, caveman mode, or compress agent-to-agent messages. Produces the diet-mode instruction block ready to paste, the three compression levels with examples, the economics of
.claude/skills/mohitagw15856-token-diet/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -51% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -27% | 0% |
Most of an LLM's output is grammatical scaffolding the reader's brain (or the next model in the pipeline) reconstructs for free: articles, hedges, pleasantries, "it's worth noting that." Strip it and the facts survive in 30–60% of the tokens — output reads like a telegram, and models parse telegrams fine. But the diet has real economics: output tokens cost 3–5× input, so dieting output pays disproportionately — while in single-shot calls the mode instruction itself costs more than it saves, and human-facing prose dieted into fragments just transfers the reading cost to a person. This skill installs the three levels, the switch lines, and the judgment about when each pays.
Ask for these if not provided:
> The instruction text, e.g. L2: "Respond in compressed prose: short declaratives, no filler, no hedges, no restating the question. Facts and actions only. Full grammar where ambiguity threatens."]
One realistic paragraph at L0/L1/L2/L3 with token counts — the trade made visible]
This use case's volume × the level's reduction × output pricing — worth-it verdict in one sentence; measure with token-cost]
The exclusions relevant to this user's context, named]
The output-compression register pattern — telegraphic prompting for output-token reduction (as in Caveman and the caveman-compression method) — systematized here into levels, economics, and exclusions.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 53,172 | 20,280 | -62% | 1 | 1 | 0% | 6,934 | 3,411 | -51% | 0 | 0 | — |
case-02 | pass→pass | 26,872 | 18,063 | -33% | 1 | 1 | 0% | 3,351 | 3,294 | -2% | 0 | 0 | — |
case-03 | fail→fail | 60,714 | 25,677 | -58% | 1 | 1 | 0% | 4,936 | 4,129 | -16% | 0 | 0 | — |
case-04 | fail→pass | 32,817 | 20,845 | -36% | 1 | 1 | 0% | 4,764 | 4,014 | -16% | 0 | 0 | — |
case-05 | fail→pass | 54,949 | 25,766 | -53% | 1 | 1 | 0% | 8,239 | 3,980 | -52% | 0 | 0 | — |
case-06 | fail→fail | 34,755 | 21,716 | -38% | 1 | 1 | 0% | 3,986 | 3,464 | -13% | 0 | 0 | — |
case-07 | fail→pass | 40,339 | 15,972 | -60% | 1 | 1 | 0% | 4,818 | 3,600 | -25% | 0 | 0 | — |
case-08 | fail→fail | 53,527 | 21,907 | -59% | 1 | 1 | 0% | 8,234 | 3,684 | -55% | 0 | 0 | — |
case-09 | fail→pass | 40,481 | 23,578 | -42% | 1 | 1 | 0% | 6,118 | 4,469 | -27% | 0 | 0 | — |
case-10 | fail→pass | 31,517 | 21,250 | -33% | 1 | 1 | 0% | 4,789 | 3,653 | -24% | 0 | 0 | — |
case-11 | fail→pass | 39,677 | 15,705 | -60% | 1 | 1 | 0% | 5,730 | 3,275 | -43% | 0 | 0 | — |
case-12 | fail→pass | 32,373 | 22,124 | -32% | 1 | 1 | 0% | 4,119 | 3,437 | -17% | 0 | 0 | — |
case-13 | fail→fail | 42,349 | 23,276 | -45% | 1 | 1 | 0% | 6,923 | 3,519 | -49% | 0 | 0 | — |
case-14 | fail→pass | 37,447 | 16,801 | -55% | 1 | 1 | 0% | 4,531 | 3,220 | -29% | 0 | 0 | — |
case-15 | fail→pass | 32,682 | 15,105 | -54% | 1 | 1 | 0% | 2,944 | 3,599 | +22% | 0 | 0 | — |
case-16 | fail→pass | 38,814 | 20,830 | -46% | 1 | 1 | 0% | 6,688 | 4,043 | -40% | 0 | 0 | — |
case-17 | pass→pass | 33,634 | 13,847 | -59% | 1 | 1 | 0% | 3,972 | 3,138 | -21% | 0 | 0 | — |
case-18 | fail→fail | 31,287 | 27,527 | -12% | 1 | 1 | 0% | 4,480 | 4,175 | -7% | 0 | 0 | — |
case-19 | fail→fail | 27,659 | 23,282 | -16% | 1 | 1 | 0% | 4,748 | 3,721 | -22% | 0 | 0 | — |
case-20 | fail→pass | 19,000 | 23,060 | +21% | 1 | 1 | 0% | 3,272 | 5,383 | +65% | 0 | 0 | — |
case-21 | fail→fail | 20,781 | 33,750 | +62% | 1 | 1 | 0% | 3,013 | 6,338 | +110% | 0 | 0 | — |
case-22 | pass→pass | 20,678 | 13,294 | -36% | 1 | 1 | 0% | 2,622 | 3,672 | +40% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.