Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn a process description into an operational quality-control checklist — binary pass/fail checks, acceptance criteria, critical-vs-routine tiers, and failure modes with catch-points. Use when the user says "build a QC checklist", "standardize how we review X", "acceptance criteria for this process", or describes quality varying by who does the work.
.claude/skills/sgharlow-qc-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 54% | 0% |
Convert "experienced people just know" into checks anyone can run. A good check is binary (pass/fail), observable (no judgment calls hidden inside), and placed at the point where the failure it catches actually happens.
what has actually gone wrong before. Real past failures anchor the checklist; ask for 2–3 if none are offered.
detected, and the cost of missing it. Every checklist item must trace to a failure mode — a check that catches nothing is process theater.
are never skippable and each states its acceptance criterion measurably.
Phrase every item as a verifiable pass/fail question with the evidence to look at — "Invoice total matches PO total (compare field X to field Y)", not "check invoice".
who signs off, where results are recorded, and the escalation when a critical check fails. A checklist without a recorded result is a suggestion.
failure the user named must be caught by some item; items that catch nothing on the dry-run get justified or cut. Then hand over with a 90-day review note — checklists drift as processes change.
"appropriate", "sufficient", or "high quality".
rather than shipping a wall nobody completes under time pressure.
Full walkthrough, examples, and variations: recipes/Recipe-048-Quality-Control-Checklists-Standards.md. For enforcing checks on AI-assisted code specifically, the same philosophy is implemented in ai-control-framework.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,941 | 32,306 | +91% | 1 | 1 | 0% | 2,673 | 3,281 | +23% | 0 | 0 | — |
case-02 | fail→pass | 17,712 | 14,656 | -17% | 1 | 1 | 0% | 2,648 | 2,860 | +8% | 0 | 0 | — |
case-03 | fail→pass | 16,978 | 20,318 | +20% | 1 | 1 | 0% | 2,759 | 3,949 | +43% | 0 | 0 | — |
case-04 | fail→pass | 18,924 | 16,175 | -15% | 1 | 1 | 0% | 2,977 | 3,014 | +1% | 0 | 0 | — |
case-05 | pass→pass | 20,043 | 21,066 | +5% | 1 | 1 | 0% | 3,266 | 3,843 | +18% | 0 | 0 | — |
case-06 | fail→pass | 14,289 | 19,038 | +33% | 1 | 1 | 0% | 2,365 | 3,648 | +54% | 0 | 0 | — |
case-07 | fail→fail | 18,751 | 19,371 | +3% | 1 | 1 | 0% | 2,704 | 3,444 | +27% | 0 | 0 | — |
case-08 | fail→pass | 14,552 | 17,712 | +22% | 1 | 1 | 0% | 2,493 | 3,543 | +42% | 0 | 0 | — |
case-09 | fail→fail | 15,735 | 18,978 | +21% | 1 | 1 | 0% | 2,390 | 3,769 | +58% | 0 | 0 | — |
case-10 | fail→pass | 17,436 | 16,673 | -4% | 1 | 1 | 0% | 2,684 | 3,089 | +15% | 0 | 0 | — |
case-11 | fail→pass | 16,469 | 21,225 | +29% | 1 | 1 | 0% | 2,420 | 3,799 | +57% | 0 | 0 | — |
case-12 | fail→pass | 13,861 | 15,945 | +15% | 1 | 1 | 0% | 2,235 | 3,024 | +35% | 0 | 0 | — |
case-13 | fail→pass | 17,272 | 20,862 | +21% | 1 | 1 | 0% | 2,843 | 3,730 | +31% | 0 | 0 | — |
case-14 | fail→pass | 17,708 | 21,338 | +20% | 1 | 1 | 0% | 2,744 | 3,857 | +41% | 0 | 0 | — |
case-15 | fail→fail | 20,948 | 23,270 | +11% | 1 | 1 | 0% | 3,286 | 4,193 | +28% | 0 | 0 | — |
case-16 | fail→pass | 17,087 | 19,617 | +15% | 1 | 1 | 0% | 2,796 | 3,447 | +23% | 0 | 0 | — |
case-17 | fail→pass | 16,782 | 21,072 | +26% | 1 | 1 | 0% | 2,535 | 3,848 | +52% | 0 | 0 | — |
case-18 | fail→fail | 15,360 | 18,744 | +22% | 1 | 1 | 0% | 2,321 | 3,309 | +43% | 0 | 0 | — |
case-19 | fail→pass | 16,302 | 18,630 | +14% | 1 | 1 | 0% | 2,520 | 3,268 | +30% | 0 | 0 | — |
case-20 | pass→fail | 12,350 | 22,445 | +82% | 1 | 1 | 0% | 1,903 | 4,199 | +121% | 0 | 0 | — |
case-21 | pass→fail | 16,750 | 23,665 | +41% | 1 | 1 | 0% | 2,620 | 3,932 | +50% | 0 | 0 | — |
case-22 | pass→fail | 13,790 | 16,058 | +16% | 1 | 1 | 0% | 2,302 | 3,200 | +39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.