Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when answering at a decision point - recommending an approach, reporting a verification result, finishing a task, proposing a fix, or assessing whether something is ready - so the judgment leads and the reader can evaluate it instead of only accepting it. Triggers on "brief mode", "bluf this", "give me the bottom line", "brief it", "stop burying the answer".
.claude/skills/escoffier-labs-brief/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 202% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 52% | 0% |
Before service the chef briefs the line: what is on, what is 86'd, what changed since yesterday. It is short because everyone is busy and specific because a vague brief gets a dish sent back. It is not a tour of the walk-in. It is what you need to work from, said first.
This skill is that brief applied to an agent's answer. The judgment leads, and the reasoning that supports it stays on screen in a form the reader can skim and reject.
Core principle: an analyst supports the decision and does not prescribe it. Handing over a command makes the decision for the reader. Handing over a judgment, its basis, and the runner-up supports theirs.
The usual fix for a buried answer is to cut everything except the action. That reads fast and quietly removes the reader's ability to disagree. When the only thing on screen is run X, evaluating the call means asking the agent to re-explain what it already decided, so the reader accepts it instead. Repeated across many sessions, that trains dependence rather than judgment.
Reasoning does not have to be long to be auditable. Prose hides reasoning inside sentences. A short labeled block exposes it as a list. Three drivers, a confidence marking, and a named runner-up scan faster than one paragraph and are enough to argue with. Same word budget, restructured.
Which shape fires depends on whether the agent is asserting or reporting. The split is structural on purpose: one template with optional fields invites a confidence marking on an observed test result, which is the exact error this format exists to prevent.
Fires when the agent is making a call.
BLUF: <the answer>. <probability term>, confidence <High|Moderate|Low> - <basis, one line>.
Why:
- <driver>
- <driver>
Alternative: <runner-up> - better if <indicator>.
Assuming: <the assumption that flips this if wrong>
Next: <one action, under 2 min>Two to four drivers, one line each. Assuming is omitted when no load-bearing assumption exists. Next is always present.
Fires when the agent is stating what happened.
<What happened, as fact. No probability term.>
Evidence: <the command and its actual output line>
Next: <one action>Observed results never carry a probability term or a confidence marking. Forecasts and recommendations always carry both. Probability describes how likely the judgment is. Confidence describes how good the basis for judging is. They are separate axes, and a high-confidence unlikely call is a coherent thing to say.
Term definitions, the probability ladder, the banned ambiguous phrasings, and the fact/assumption/judgment test live in references/estimative-language.md. Read it before marking anything.
Enumerated rather than described, because an agent asked to judge "does this matter" either fires constantly or never.
Fires, recommendation shape:
Fires, report shape:
Does not fire:
check governs whether a claim of done, fixed, or passing may be made at all: it requires fresh verification evidence in the same reply. brief governs how that evidence is presented once check is satisfied.
The report shape's Evidence: line is where check's evidence lands. Its brevity is a presentation constraint and never licence to shorten, skip, or paraphrase the verification. A compact report of a command you did not run is still a false report.
The format creates its own failure modes, and each one produces output that looks correct.
Alternative line invites a strawman when only one real path exists. In that case the line reads No real alternative: <why>. A fabricated runner-up is worse than none, because it implies a choice was weighed when it was not.rm -rf, force push, schema migration, dropping data: confirm before acting. Safety outranks brevity.This skill governs structure, not vocabulary. Word choice, banned phrasing, and slop are governed by the project's own writing rules where they exist. Do not restate them here. A second copy drifts from the first. brief adds no vocabulary rules of its own beyond the estimative terms.
Alternative with an option nobody would choose, to avoid leaving the line out.The persistence model, the no-preamble-no-closer rule, the one-concrete-next-action rule, and the pre-send check come from i-have-adhd (MIT). No text is copied. The analytic layer that distinguishes this skill - probability and confidence as separate axes, the fact/assumption/judgment separation, the mandatory runner-up and indicator - comes from intelligence-writing tradecraft.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 10,437 | 9,236 | -12% | 1 | 1 | 0% | 1,711 | 2,870 | +68% | 0 | 0 | — |
case-02 | fail→pass | 8,633 | 12,404 | +44% | 1 | 1 | 0% | 1,173 | 3,537 | +202% | 0 | 0 | — |
case-03 | pass→pass | 7,193 | 6,393 | -11% | 1 | 1 | 0% | 1,095 | 2,499 | +128% | 0 | 0 | — |
case-04 | fail→pass | 13,816 | 18,076 | +31% | 1 | 1 | 0% | 2,023 | 2,727 | +35% | 0 | 0 | — |
case-05 | fail→fail | 11,222 | 9,069 | -19% | 1 | 1 | 0% | 1,695 | 2,950 | +74% | 0 | 0 | — |
case-06 | fail→pass | 15,764 | 8,803 | -44% | 1 | 1 | 0% | 2,070 | 3,212 | +55% | 0 | 0 | — |
case-07 | fail→pass | 13,931 | 8,295 | -40% | 1 | 1 | 0% | 1,874 | 2,856 | +52% | 0 | 0 | — |
case-08 | pass→pass | 13,816 | 17,361 | +26% | 1 | 1 | 0% | 2,098 | 4,276 | +104% | 0 | 0 | — |
case-09 | fail→pass | 12,344 | 12,770 | +3% | 1 | 1 | 0% | 1,911 | 3,640 | +90% | 0 | 0 | — |
case-10 | fail→pass | 11,296 | 9,503 | -16% | 1 | 1 | 0% | 1,699 | 3,011 | +77% | 0 | 0 | — |
case-11 | fail→pass | 7,783 | 5,010 | -36% | 1 | 1 | 0% | 1,189 | 2,334 | +96% | 0 | 0 | — |
case-12 | fail→pass | 8,903 | 4,505 | -49% | 1 | 1 | 0% | 1,486 | 2,341 | +58% | 0 | 0 | — |
case-13 | pass→pass | 6,954 | 3,945 | -43% | 1 | 1 | 0% | 1,250 | 2,237 | +79% | 0 | 0 | — |
case-14 | pass→pass | 5,804 | 5,257 | -9% | 1 | 1 | 0% | 1,107 | 2,283 | +106% | 0 | 0 | — |
case-15 | fail→pass | 7,518 | 8,696 | +16% | 1 | 1 | 0% | 1,052 | 2,025 | +92% | 0 | 0 | — |
case-16 | pass→pass | 6,460 | 6,866 | +6% | 1 | 1 | 0% | 977 | 2,342 | +140% | 0 | 0 | — |
case-17 | pass→pass | 23,308 | 16,045 | -31% | 1 | 1 | 0% | 3,733 | 3,726 | -0% | 0 | 0 | — |
case-18 | fail→pass | 9,969 | 7,785 | -22% | 1 | 1 | 0% | 1,353 | 2,517 | +86% | 0 | 0 | — |
case-19 | pass→pass | 20,616 | 14,582 | -29% | 1 | 1 | 0% | 3,176 | 3,792 | +19% | 0 | 0 | — |
case-20 | pass→pass | 2,150 | 2,867 | +33% | 1 | 1 | 0% | 185 | 1,894 | +924% | 0 | 0 | — |
case-21 | pass→fail | 21,618 | 8,878 | -59% | 1 | 1 | 0% | 3,239 | 3,250 | +0% | 0 | 0 | — |
case-22 | fail→fail | 5,383 | 19,043 | +254% | 1 | 1 | 0% | 711 | 4,610 | +548% | 0 | 0 | — |
case-23 | fail→fail | 3,620 | 6,377 | +76% | 1 | 1 | 0% | 401 | 2,479 | +518% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +43 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.