Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Stop hook that blocks Claude from finishing until quality checks pass. Detects rationalization patterns (surface text heuristics), stale learning logs (filesystem mtime), and low disk space. Complements self-audit by mechanically enforcing learning capture habits.
.claude/skills/affaan-m-delivery-gate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 81% | 0% |
A Stop hook that checks three things before Claude can finish a session, using only deterministic checks — file modification timestamps, disk usage, and regex patterns on the transcript text. No AI inference.
This is distinct from reasoning gates (like self-audit): delivery-gate checks machine-verifiable facts; self-audit checks output quality across four reasoning dimensions. Together they form defense in depth:
This is the same pattern as CI pipeline gates — automated, deterministic checks that verify machine-readable facts rather than trusting self-reported status.
| Check | Mechanism | On Hit | |-------|-----------|--------| | Rationalization patterns | Regex on transcript tail | Warning only (never blocks) | | Stale learning libraries | mtime on 5 configurable paths | Warning if some stale; Block if >=3 stale OR growth-log stale + complex task | | Disk space < 50GB | shutil.disk_usage | Warning | | Disk space < 15GB | shutil.disk_usage | Block (exit 2) |
Rationalization detection warns about patterns like "skip tests for now" and "pre-existing bug" — surface signals that thinking may have been cut short. It never blocks on its own, because regex heuristics can false-positive. The blocking conditions are: disk critical, >=3 learning libs stale, OR growth-log specifically stale (all require complex task >=3 edits).
Claude Code's built-in checks cover code quality (build → type → lint → test). But there's a different failure mode: the agent produces working code while the session hygiene was neglected — learning not captured, rationalized shortcuts, disk running out silently.
Over many sessions of "ship and forget," the human hasn't grown. This hook enforces the habit: complex task → must touch learning libraries.
bashcp quality-gate.py ~/.claude/scripts/
Add to ~/.claude/settings.json:
json{ "hooks": { "Stop": [{ "hooks": [{ "type": "command", "command": "python3 ~/.claude/scripts/quality-gate.py", "timeout": 5000 }] }] } }
Create these files in your project's memory directory. The hook checks if at least one was updated today:
memory/
├── growth-log/ # Daily learning entries (directory)
├── decisions/log.md # Decision log
├── output-index.md # Index of session outputs
├── ratings-tracker.md # Skill ratings over time
└── tooling_capabilities.md # Known tools inventoryCustomize the LIBS dict to match your own file structure.
Edit quality-gate.py:
| Variable | Default | Purpose | |----------|---------|---------| | RATIONALIZE | 4 patterns | Regex patterns for rationalization detection | | LIBS | 5 libraries | Files/dirs to check for today's updates | | COMPLEX_THRESHOLD | 3 | Edit/Write calls to classify as complex | | DISK_WARN_GB | 50 | Warn below this | | DISK_CRIT_GB | 15 | Block below this |
Simple session — allowed:
edit_count=1 (< 3, not complex) → exit 0Complex task, learning captured — allowed:
edit_count=5 (complex) → checks LIBS → growth-log updated today → exit 0Complex task, no learning — BLOCKED:
edit_count=4 (complex) → checks LIBS → all 5 stale → exit 2
stderr: "Blocked: complex task completed but no learning captured today."Low disk space — BLOCKED:
disk_free=12GB < 15GB critical → exit 2
stderr: "Blocked: disk space at 12GB (threshold: 15GB)."The hook enforces the habit of touching learning libraries, not the quality of what was recorded. If output-index.md is updated but growth-log is skipped, the hook passes (1 of 5 libraries touched). This is by design: mechanical gates check machine-verifiable facts. For content quality verification, pair with self-audit.
from __future__ import annotations)This code went through 4 rounds of automated code review (CodeRabbit + Greptile) with 9 real bugs found and fixed.
self-audit — Reasoning quality gate (completeness/consistency/groundedness/honesty)verification-loop — Code quality checks (build/type/lint/test)gateguard — PreToolUse safety gate| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→fail | 13,581 | 14,424 | +6% | 1 | 1 | 0% | 2,522 | 3,996 | +58% | 0 | 0 | — |
case-01 | fail→pass | 24,277 | 7,640 | -69% | 1 | 1 | 0% | 3,565 | 2,867 | -20% | 0 | 0 | — |
case-02 | fail→pass | 15,458 | 10,138 | -34% | 1 | 1 | 0% | 2,573 | 2,460 | -4% | 0 | 0 | — |
case-03 | fail→pass | 13,181 | 9,435 | -28% | 1 | 1 | 0% | 2,330 | 3,056 | +31% | 0 | 0 | — |
case-04 | fail→fail | 14,555 | 17,847 | +23% | 1 | 1 | 0% | 2,668 | 4,284 | +61% | 0 | 0 | — |
case-06 | fail→fail | 16,069 | 10,096 | -37% | 1 | 1 | 0% | 3,053 | 3,210 | +5% | 0 | 0 | — |
case-07 | fail→pass | 10,492 | 6,079 | -42% | 1 | 1 | 0% | 1,833 | 2,265 | +24% | 0 | 0 | — |
case-08 | fail→pass | 6,544 | 3,610 | -45% | 1 | 1 | 0% | 1,056 | 1,913 | +81% | 0 | 0 | — |
case-09 | fail→pass | 15,246 | 5,426 | -64% | 1 | 1 | 0% | 866 | 1,997 | +131% | 0 | 0 | — |
case-10 | fail→pass | 24,391 | 4,361 | -82% | 1 | 1 | 0% | 1,112 | 1,959 | +76% | 0 | 0 | — |
case-11 | fail→pass | 6,565 | 3,015 | -54% | 1 | 1 | 0% | 1,080 | 1,711 | +58% | 0 | 0 | — |
case-12 | pass→pass | 5,532 | 3,854 | -30% | 1 | 1 | 0% | 910 | 1,940 | +113% | 0 | 0 | — |
case-13 | fail→pass | 10,825 | 3,729 | -66% | 1 | 1 | 0% | 1,079 | 1,888 | +75% | 0 | 0 | — |
case-14 | pass→pass | 9,638 | 5,269 | -45% | 1 | 1 | 0% | 1,945 | 2,199 | +13% | 0 | 0 | — |
case-15 | pass→pass | 9,687 | 4,377 | -55% | 1 | 1 | 0% | 1,976 | 2,114 | +7% | 0 | 0 | — |
case-16 | fail→pass | 11,380 | 3,707 | -67% | 1 | 1 | 0% | 2,133 | 1,828 | -14% | 0 | 0 | — |
case-17 | fail→pass | 17,048 | 11,556 | -32% | 1 | 1 | 0% | 2,266 | 3,321 | +47% | 0 | 0 | — |
case-18 | fail→pass | 11,986 | 2,132 | -82% | 1 | 1 | 0% | 2,119 | 1,487 | -30% | 0 | 0 | — |
case-19 | fail→fail | 18,897 | 11,401 | -40% | 1 | 1 | 0% | 968 | 3,467 | +258% | 0 | 0 | — |
case-20 | pass→pass | 39,288 | 4,144 | -89% | 1 | 1 | 0% | 1,586 | 1,881 | +19% | 0 | 0 | — |
case-21 | pass→pass | 41,581 | 10,072 | -76% | 1 | 1 | 0% | 1,615 | 2,960 | +83% | 0 | 0 | — |
case-22 | fail→pass | 5,518 | 4,199 | -24% | 1 | 1 | 0% | 891 | 1,949 | +119% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/3/2026 | +60% |
Other measured skills in the registry, with their headline benchmark lift.