Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Close defects and incidents with a BUG RECEIPT and VERIFIED, PARTIAL, or BLOCKED status after diagnosis, repair, or recovery.
.claude/skills/github-bug-receipt/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 134% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 447% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 225% | 0% |
For every bug or incident closeout decision, return the complete receipt below as the entire user-facing result, even when the user requests a concise reply or does not name this format. Concision shortens field values; it never removes or renames a row. Do not replace the receipt with prose.
textBUG RECEIPT · VERIFIED | PARTIAL | BLOCKED Problem <observed defect and intended behavior> Baseline <failing interaction or command and decisive result; or not run> Root cause <proven mechanism; or unproven hypothesis> Change <responsible change; or none> Proof <supplied or executed check: result; include every decisive layer> Gaps <none; or exact missing proof and single next experiment/package> Source executed now | supplied | mixed
Use not run, unproven, or none explicitly. Never omit a row to make the receipt look complete.
Before editing, record the observed problem, intended behavior, strongest direct check, and evidence source: executed now, supplied, or mixed. Never imply that supplied evidence was executed in the current run.
Keep evidence privacy-minimal. Redact credentials, tokens, cookies, personal data, private URLs, and sensitive payloads; preserve only the identifiers and excerpts needed to reproduce or correlate the failure.
Reproduce the failure with the narrowest safe check when possible. If reproduction is unavailable, preserve the evidence obtained and cap the result at PARTIAL or BLOCKED.
Do not convert a plausible patch, stale log, source read, or passing build into proof of the user-visible behavior.
Run only checks required by the affected contract:
Use these decisive boundaries:
| Surface | Required direct proof | | --- | --- | | Logic or failing test | Original failing input or focused test now passes | | UI behavior | Real interaction plus relevant console and network observation | | API or integration | Request, response, and responsible service behavior | | Persistence | Write/read or reload round trip through the real owner path | | Race or lifecycle | Repeated concurrent trigger; zero-or-one success; affected-row and transaction evidence; final invariant | | Cross-system blocker | One sanitized failing request/response with timestamp or request ID, edge and application logs, and identity-provider logs when the trace reaches that owner |
VERIFIED: observed baseline, concrete cause, responsible change, all declared checks passed, no material gap.PARTIAL: useful evidence exists, but a required proof layer is missing or inconclusive.BLOCKED: a specific external condition prevents reproduction, repair, or proof.For PARTIAL or BLOCKED, name the single minimal experiment or correlated evidence package that closes the decisive gap. Never invent a command, observation, count, location, or result.
For a machine-readable receipt or CI integration, read references/receipt-contract.md and conform to its JSON fields, evidence-source marker, compatibility rule, and status invariants.
When a JSON artifact is requested, start from assets/receipt.template.json, write it to a task-owned path, and validate it with node scripts/validate-receipt.mjs <receipt.json> from this skill directory. Do not commit the generated receipt unless the user requests it.
Originally published at https://github.com/lMysticl/bug-receipt under the MIT License.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,048 | 20,107 | +34% | 1 | 1 | 0% | 1,677 | 3,929 | +134% | 0 | 0 | — |
case-02 | fail→pass | 8,838 | 19,233 | +118% | 1 | 1 | 0% | 1,524 | 2,856 | +87% | 0 | 0 | — |
case-03 | fail→pass | 21,557 | 12,904 | -40% | 1 | 1 | 0% | 2,938 | 2,427 | -17% | 0 | 0 | — |
case-04 | fail→pass | 2,202 | 11,633 | +428% | 1 | 1 | 0% | 398 | 2,179 | +447% | 0 | 0 | — |
case-05 | fail→fail | 6,953 | 21,613 | +211% | 1 | 1 | 0% | 1,255 | 4,158 | +231% | 0 | 0 | — |
case-06 | pass→pass | 11,991 | 11,615 | -3% | 1 | 1 | 0% | 1,324 | 3,161 | +139% | 0 | 0 | — |
case-07 | pass→pass | 7,833 | 11,816 | +51% | 1 | 1 | 0% | 1,443 | 2,402 | +66% | 0 | 0 | — |
case-08 | pass→pass | 5,410 | 16,785 | +210% | 1 | 1 | 0% | 893 | 3,146 | +252% | 0 | 0 | — |
case-09 | fail→pass | 11,301 | 13,944 | +23% | 1 | 1 | 0% | 1,062 | 3,453 | +225% | 0 | 0 | — |
case-10 | fail→pass | 11,660 | 14,799 | +27% | 1 | 1 | 0% | 1,681 | 2,653 | +58% | 0 | 0 | — |
case-11 | fail→pass | 10,523 | 15,052 | +43% | 1 | 1 | 0% | 1,634 | 2,847 | +74% | 0 | 0 | — |
case-12 | fail→fail | 9,587 | 36,468 | +280% | 1 | 1 | 0% | 935 | 7,987 | +754% | 0 | 0 | — |
case-13 | fail→fail | 6,718 | 9,144 | +36% | 1 | 1 | 0% | 1,094 | 2,548 | +133% | 0 | 0 | — |
case-14 | pass→pass | 8,002 | 25,185 | +215% | 1 | 1 | 0% | 1,364 | 4,792 | +251% | 0 | 0 | — |
case-15 | pass→fail | 18,328 | 22,014 | +20% | 1 | 1 | 0% | 3,133 | 4,124 | +32% | 0 | 0 | — |
case-16 | pass→pass | 27,616 | 15,259 | -45% | 1 | 1 | 0% | 3,100 | 3,441 | +11% | 0 | 0 | — |
case-21 | fail→pass | 12,589 | 16,377 | +30% | 1 | 1 | 0% | 2,023 | 4,248 | +110% | 0 | 0 | — |
case-17 | fail→pass | 9,276 | 6,706 | -28% | 1 | 1 | 0% | 1,442 | 2,135 | +48% | 0 | 0 | — |
case-18 | fail→pass | 5,187 | 6,066 | +17% | 1 | 1 | 0% | 887 | 1,999 | +125% | 0 | 0 | — |
case-19 | fail→pass | 13,384 | 8,114 | -39% | 1 | 1 | 0% | 2,060 | 2,479 | +20% | 0 | 0 | — |
case-20 | fail→pass | 6,149 | 4,855 | -21% | 1 | 1 | 0% | 1,140 | 2,063 | +81% | 0 | 0 | — |
case-22 | fail→pass | 10,222 | 12,408 | +21% | 1 | 1 | 0% | 1,660 | 3,231 | +95% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.