Install any skill in seconds. Free to start, no credit card required.
Get Started Free →A ship gate that runs before any production deploy: checks the silent failure modes that make a deploy 'succeed' while prod stays broken, then verifies the live revision instead of trusting deploy output.
.claude/skills/sickn33-pre-ship-gate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 37% | 0% |
Most bad deploys do not fail loudly. The pipeline goes green, the CLI prints "deployed", and the old or broken version is still what users hit. This skill is the gate you run right before a production deploy and right after, so an agent stops trusting deploy output and starts confirming what is actually live. It exists because "the deploy command exited 0" and "the new version is serving traffic" are two different facts, and agents routinely confuse them.
The gate has three phases. Do not skip to phase 3.
Walk the silent failure catalog. These are the modes that let a deploy "succeed" while production stays broken. For each one, confirm it or flag it. Do not assume.
The human or the deploy tooling runs the actual command. This skill does not execute the production deploy itself. It gates it.
Confirm the running system, not the deploy log.
bash# You intended to ship this commit INTENDED="$(git rev-parse --short HEAD)" # Ask the running service what it is actually serving LIVE="$(curl -fsS https://your-service.example.com/health | jq -r '.revision')" if [ "$INTENDED" = "$LIVE" ]; then echo "Live revision $LIVE matches intended $INTENDED: verified shipped." else echo "MISMATCH: intended $INTENDED but live is $LIVE. Do not report shipped." fi
markdownPRE-SHIP GATE, verdict: HOLD - Migrations: 1 pending (add_users_status_col): NOT yet applied to prod. BLOCK. - Feature flags: new_checkout flag is OFF in prod. Enabling required post-deploy. - Build assets: new bundle hash confirmed (a1b2c3 != previous 9f8e7d). OK. - Release pointer: deploy updates active symlink. OK. - Rollout: canary at 10%, manual promote required. NOTE. - Env/secrets: STRIPE_KEY present in prod. OK. Reason for HOLD: run migration add_users_status_col before cutover, or the new code will 500 on /orders.
Solution: The check is hitting a cached edge or the old pod. Verify the revision field in the response, not just the status code.
Solution: Sequence migrations before cutover, or gate the code path behind a flag until the migration lands.
Solution: Confirm the traffic pointer or promotion step, not just the upload step.
curl -fsS against a status endpoint and are illustrative. Replace the placeholder host and version field with your own before use.@codebase-audit-pre-push: clean and audit the code before it ever reaches a deploy.@dos-verify-done-claims: verify a "done" claim against git ground truth after the fact.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 19,055 | 5,377 | -72% | 1 | 1 | 0% | 1,870 | 2,509 | +34% | 0 | 0 | — |
case-02 | fail→pass | 10,350 | 6,877 | -34% | 1 | 1 | 0% | 1,736 | 2,693 | +55% | 0 | 0 | — |
case-03 | pass→pass | 11,408 | 5,867 | -49% | 1 | 1 | 0% | 2,158 | 2,662 | +23% | 0 | 0 | — |
case-04 | fail→fail | 6,955 | 5,325 | -23% | 1 | 1 | 0% | 1,261 | 2,682 | +113% | 0 | 0 | — |
case-05 | fail→fail | 8,537 | 6,282 | -26% | 1 | 1 | 0% | 1,631 | 2,691 | +65% | 0 | 0 | — |
case-06 | pass→fail | 10,575 | 5,103 | -52% | 1 | 1 | 0% | 1,769 | 2,424 | +37% | 0 | 0 | — |
case-07 | fail→pass | 8,570 | 5,637 | -34% | 1 | 1 | 0% | 1,563 | 2,637 | +69% | 0 | 0 | — |
case-08 | fail→pass | 11,184 | 4,217 | -62% | 1 | 1 | 0% | 2,026 | 2,290 | +13% | 0 | 0 | — |
case-09 | pass→pass | 7,956 | 4,276 | -46% | 1 | 1 | 0% | 1,325 | 2,320 | +75% | 0 | 0 | — |
case-10 | pass→pass | 6,499 | 3,718 | -43% | 1 | 1 | 0% | 1,178 | 2,187 | +86% | 0 | 0 | — |
case-11 | pass→fail | 4,774 | 4,737 | -1% | 1 | 1 | 0% | 1,039 | 2,376 | +129% | 0 | 0 | — |
case-12 | pass→pass | 17,987 | 8,152 | -55% | 1 | 1 | 0% | 2,785 | 2,988 | +7% | 0 | 0 | — |
case-13 | pass→pass | 9,883 | 7,961 | -19% | 1 | 1 | 0% | 1,911 | 2,920 | +53% | 0 | 0 | — |
case-14 | fail→pass | 9,385 | 3,556 | -62% | 1 | 1 | 0% | 1,601 | 2,101 | +31% | 0 | 0 | — |
case-15 | pass→pass | 8,727 | 6,454 | -26% | 1 | 1 | 0% | 1,522 | 2,691 | +77% | 0 | 0 | — |
case-16 | pass→pass | 6,996 | 5,181 | -26% | 1 | 1 | 0% | 1,270 | 2,426 | +91% | 0 | 0 | — |
case-17 | pass→pass | 8,399 | 7,379 | -12% | 1 | 1 | 0% | 1,461 | 2,785 | +91% | 0 | 0 | — |
case-18 | pass→pass | 8,388 | 5,055 | -40% | 1 | 1 | 0% | 1,481 | 2,459 | +66% | 0 | 0 | — |
case-19 | pass→pass | 14,149 | 6,121 | -57% | 1 | 1 | 0% | 2,211 | 2,586 | +17% | 0 | 0 | — |
case-20 | pass→pass | 3,364 | 3,395 | +1% | 1 | 1 | 0% | 563 | 2,167 | +285% | 0 | 0 | — |
case-21 | pass→pass | 8,160 | 8,888 | +9% | 1 | 1 | 0% | 1,372 | 3,058 | +123% | 0 | 0 | — |
case-22 | pass→pass | 7,594 | 7,525 | -1% | 1 | 1 | 0% | 1,476 | 2,924 | +98% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.