Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute Fly.io production deployment checklist with health checks, auto-scaling, monitoring, and rollback procedures. Trigger: "fly.io production", "fly.io go-live", "fly.io prod checklist".
.claude/skills/jeremylongshore-flyio-prod-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -6% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 88% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 35% | 0% |
Make go-live a signed evidence boundary. The gate must cover the actual workload: stateless Machines, region-bound volumes, Managed Postgres, release commands, health semantics, autostart, networking, observability, billing, and tested rollback.
Record app, organization, environment, domains, regions, process groups, data stores, objectives, on-call route, and accepted risks.
Use scoped expiring tokens, protect production environments, inspect secret names and digests without values, and prove rotation and revocation.
Validate image, architecture, ports, processes, signals, release command, health checks, strategy, unavailable capacity, and volume compatibility.
Confirm Machine redundancy where required, regional placement, dependency failure behavior, Managed Postgres or volume backup boundaries, restore tests, and reconciliation.
Prove logs, metrics, health, alerts, provider-status escalation, resource inventory, cost owner, and cleanup of obsolete billable resources.
Run the approved release path, observe acceptance gates, restore the prior image or config in a drill, and close only after actual fleet reconciliation.
Production deploy uses the narrowest app or organization token in a protected environment. Read-only monitoring should not share deploy authority. Remember that anyone with deploy access can deploy code that reads runtime App secrets.
Use Read and Grep to inspect application configuration, deployment evidence, provider documentation, fixtures, logs, schemas, and existing tests before proposing a change. Use Write or Edit only for an approved plan, configuration, implementation, test, or redacted receipt. Do not create, deploy, scale, restart, stop, suspend, destroy, rotate, revoke, expose, or migrate live Fly.io resources without explicit operator approval.
Return the target organization, app, environment, region set, Machine or database identifiers, source-contract fingerprint, evidence, unresolved risks, rollback state, and final decision without exposing tokens, secrets, connection strings, or customer data.
A web app with Managed Postgres passes only after the team proves two healthy Machines, scoped deploy and read-only tokens, a rolling deploy, database restore and reconnect, region-aware alerts, current cost ownership, and rollback to the previous image.
| Failure | Response | | --- | --- | | Required evidence is missing | Keep the gate open and assign the exact proof, owner, and deadline. | | Exception has no expiry | Reject it or add a compensating control, approver, review date, and rollback trigger. | | Rollback drill fails | Do not approve go-live until image, config, secret, and data recovery paths are corrected. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 19,806 | 10,464 | -47% | 1 | 1 | 0% | 2,309 | 2,174 | -6% | 0 | 0 | — |
case-05 | pass→pass | 13,232 | 12,370 | -7% | 1 | 1 | 0% | 1,302 | 2,449 | +88% | 0 | 0 | — |
case-01 | pass→pass | 16,797 | 13,816 | -18% | 1 | 1 | 0% | 1,990 | 2,686 | +35% | 0 | 0 | — |
case-02 | pass→pass | 14,422 | 14,837 | +3% | 1 | 1 | 0% | 1,571 | 2,891 | +84% | 0 | 0 | — |
case-03 | pass→pass | 17,565 | 15,576 | -11% | 1 | 1 | 0% | 1,828 | 2,878 | +57% | 0 | 0 | — |
case-06 | pass→pass | 16,019 | 15,234 | -5% | 1 | 1 | 0% | 1,721 | 2,792 | +62% | 0 | 0 | — |
case-07 | pass→pass | 19,882 | 16,380 | -18% | 1 | 1 | 0% | 2,356 | 2,987 | +27% | 0 | 0 | — |
case-08 | pass→pass | 20,994 | 18,196 | -13% | 1 | 1 | 0% | 2,723 | 3,310 | +22% | 0 | 0 | — |
case-09 | pass→pass | 16,128 | 13,861 | -14% | 1 | 1 | 0% | 1,620 | 2,663 | +64% | 0 | 0 | — |
case-10 | fail→fail | 14,349 | 16,755 | +17% | 1 | 1 | 0% | 1,553 | 3,087 | +99% | 0 | 0 | — |
case-11 | pass→pass | 13,992 | 13,826 | -1% | 1 | 1 | 0% | 1,973 | 2,551 | +29% | 0 | 0 | — |
case-12 | fail→pass | 20,541 | 7,947 | -61% | 1 | 1 | 0% | 2,638 | 1,683 | -36% | 0 | 0 | — |
case-13 | pass→pass | 12,501 | 8,015 | -36% | 1 | 1 | 0% | 1,356 | 1,702 | +26% | 0 | 0 | — |
case-14 | pass→pass | 18,387 | 11,508 | -37% | 1 | 1 | 0% | 1,830 | 2,242 | +23% | 0 | 0 | — |
case-15 | pass→pass | 20,407 | 13,640 | -33% | 1 | 1 | 0% | 2,442 | 2,667 | +9% | 0 | 0 | — |
case-16 | fail→pass | 11,843 | 7,348 | -38% | 1 | 1 | 0% | 1,254 | 1,640 | +31% | 0 | 0 | — |
case-17 | pass→pass | 11,493 | 7,917 | -31% | 1 | 1 | 0% | 1,012 | 1,632 | +61% | 0 | 0 | — |
case-18 | pass→pass | 11,397 | 11,495 | +1% | 1 | 1 | 0% | 1,087 | 2,435 | +124% | 0 | 0 | — |
case-19 | pass→pass | 20,254 | 23,419 | +16% | 1 | 1 | 0% | 2,758 | 4,831 | +75% | 0 | 0 | — |
case-20 | pass→pass | 20,579 | 23,304 | +13% | 1 | 1 | 0% | 2,744 | 4,497 | +64% | 0 | 0 | — |
case-21 | pass→pass | 14,349 | 12,214 | -15% | 1 | 1 | 0% | 1,545 | 2,561 | +66% | 0 | 0 | — |
case-22 | pass→pass | 7,759 | 8,041 | +4% | 1 | 1 | 0% | 442 | 1,677 | +279% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v2, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/30/2026 | +26% |
Other measured skills in the registry, with their headline benchmark lift.