Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when adding, retiring, or auditing feature flags. Triggers on "add a flag", "ship behind a flag", "rollout plan", "kill switch", "stale flags", "flag debt", "LaunchDarkly", "GrowthBook", "Statsig", "Unleash", "Flipt", or any progressive-delivery question. Ships flag debt scanner, rollout planner, and kill-switch auditor (all stdlib Python), 4 references on flag taxonomy + provider trade-offs + rollout strategies + lifecycle, plus a /flag-cleanup slash command.
.claude/skills/alirezarezvani-feature-flags-architect/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 108% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 113% | 0% |
End-to-end discipline for feature flags: classify them, ship them, ramp them, and retire them. Most teams treat flags as throwaway if-statements; this skill treats them as a controlled lifecycle with measurable debt.
ifrequest → design → ship → ramp → cleanup → archiveFlags that skip cleanup become debt: dead branches, stale defaults, untested code paths, unbounded blast radius. The three scripts in this skill enforce the lifecycle.
bash# 1. Audit the repo for flag debt python scripts/flag_debt_scanner.py --repo . --max-age-days 90 # 2. Plan a progressive rollout for a new flag python scripts/rollout_planner.py --population 100000 --target-percent 100 --duration-days 14 --strategy ring # 3. Verify every flag has a documented kill switch python scripts/kill_switch_audit.py --repo . --flag-doc docs/feature-flags.md
Different flag types have different lifespans and ownership. Misclassifying creates debt.
| Type | Purpose | Typical lifespan | Owner | Cleanup trigger | |---|---|---|---|---| | Release | Hide unfinished features in production | days–weeks | Eng | 100% rollout reached | | Experiment | A/B test variants | weeks | Product/Marketing | Test concluded; winner picked | | Operational | Circuit breakers, perf toggles, kill switches | months–years | Eng/SRE | Replaced by autoscaling/feature retirement | | Permission | Entitlements per user/account/plan | years (permanent) | Product | Plan/role removed |
Only Release and Experiment flags should be on a debt-scanner watchlist. Operational and Permission flags are by design long-lived. See references/flag_taxonomy.md for decision tree.
All three are stdlib-only. Run with --help.
flag_debt_scanner.pyFinds flags older than --max-age-days with low usage, suggesting candidates for cleanup.
bashpython scripts/flag_debt_scanner.py --repo . --max-age-days 90 --format text python scripts/flag_debt_scanner.py --repo . --max-age-days 60 --format json > debt.json
Detection heuristic:
--repo for code references matching common flag-call patterns:flag("..."), isFlagEnabled("..."), featureFlag("..."), getFlag("...")client.variation("...", ...), unleash.isEnabled("..."), growthbook.feature("...")git log --diff-filter=A -S <name>).--max-age-days ago AND used in ≤--min-uses places.Outputs flag name, age in days, file references, suggested action. JSON mode is CI-friendly.
rollout_planner.pyGenerates a phased rollout schedule from population size, target percent, duration, and strategy.
bashpython scripts/rollout_planner.py --population 100000 --target-percent 100 --duration-days 14 --strategy ring python scripts/rollout_planner.py --population 50000 --target-percent 25 --duration-days 7 --strategy linear python scripts/rollout_planner.py --population 1000000 --target-percent 100 --duration-days 30 --strategy log
Strategies:
ring: 1% → 5% → 25% → 50% → 100%, evenly spaced. Default for risky launches.linear: constant rate per day. Default for medium-risk.log: rapid early, slow tail. Default for low-risk launches with confidence.cohort: by named cohort (internal → beta → free → paid → all).Outputs a markdown table with date, percent, expected user count, abort criteria, and verification step per phase.
kill_switch_audit.pyCross-references code-discovered flags against documentation to verify each has a kill switch path written down.
bashpython scripts/kill_switch_audit.py --repo . --flag-doc docs/feature-flags.md python scripts/kill_switch_audit.py --repo . --flag-doc runbooks/flags.md --format json
What it checks:
--flag-docUse as a pre-merge gate before any new flag ships.
| Provider | Best for | Pricing model | Lock-in risk | OSS option | |---|---|---|---|---| | LaunchDarkly | Enterprise, complex targeting, audit/compliance | Per-MAU, expensive | High | No | | GrowthBook | Mid-market, A/B testing focused, OSS-friendly | Per-MAU + OSS | Low | Yes (self-host) | | Statsig | Growth/product teams, advanced experimentation | Free tier + per-MAU | Medium | No | | Unleash | OSS-first, self-hosted, dev-friendly | OSS + Enterprise | Low | Yes | | Flipt | Lightweight, k8s-native, simple needs | OSS-only | None | Yes | | DIY | <100 flags, no targeting, full control | None | None | N/A |
Decision rules:
references/provider_comparison.md for detail.1. Classify: which of the 4 flag types?
→ Release (most common for engineering work)
2. Run rollout_planner.py to design the ramp
3. Add flag entry to docs/feature-flags.md BEFORE writing code:
- name, owner, type, kill-switch trigger, dashboard URL
4. Write the code with the flag
5. Run kill_switch_audit.py — must pass before merge
6. Deploy at 0%; verify kill switch works
7. Execute rollout schedule; abort if abort criteria met
8. At 100% for 7+ days: remove flag, delete dead branch, archive doc entry1. Run flag_debt_scanner.py --repo . --max-age-days 90 > debt.md
2. For each flagged item:
a. Confirm it reached 100% (or was killed)
b. Find the issue/PR that introduced it; verify owner agrees to remove
c. Delete dead branches; remove flag config
d. Run kill_switch_audit.py — should now show one fewer flag
3. Update CHANGELOG: "Removed N stale flags"1. Estimate flag count (current + 12-month projection)
2. Required features:
- Targeting rules (user, account, geo, %)?
- A/B testing + stats?
- Audit log / SOC2?
- Self-hosting / data residency?
3. Pricing budget (MAU * cost-per-MAU)
4. See provider_comparison.md decision tree
5. Build a 30-day proof-of-concept before signing1. Identify the failure modes:
- Latency spike (which threshold?)
- Error rate spike (which threshold?)
- Business metric regression (which threshold?)
2. Wire each to an abort:
- Manual: dashboard link + on-call playbook
- Automated: alert threshold flips flag back to 0%
3. Test the kill switch in staging BEFORE production rollout
4. Document in flag-doc; pass kill_switch_audit.pyreferences/flag_taxonomy.md — 4 types, decision tree, ownership, lifespanreferences/provider_comparison.md — LaunchDarkly / GrowthBook / Statsig / Unleash / Flipt / DIY trade-offsreferences/rollout_strategies.md — ring / linear / log / cohort / geo, abort criteria, monitoringreferences/flag_lifecycle.md — request → design → ship → ramp → cleanup → archive/flag-cleanup — Run the full cleanup workflow on the current repo: scan for debt, generate a removal plan, audit kill switches.
assets/flag_request_template.md — fill-in form for new flag requests (name, owner, type, kill switch, rollout plan)if (FLAG_FOO) 50 places — should be a Permission flag with a runtime config, not a Release flagA team using this skill should achieve:
kill_switch_audit.py at merge timeflag_debt_scanner.py --max-age-days 90 returns ≤5 stale flags repo-wide| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | fail→pass | 10,925 | 3,308 | -70% | 1 | 1 | 0% | 2,118 | 3,084 | +46% | 0 | 0 | — |
case-01 | fail→fail | 12,913 | 8,324 | -36% | 1 | 1 | 0% | 2,409 | 2,820 | +17% | 0 | 0 | — |
case-02 | fail→fail | 21,085 | 23,334 | +11% | 1 | 1 | 0% | 3,500 | 6,320 | +81% | 0 | 0 | — |
case-03 | fail→pass | 14,663 | 15,778 | +8% | 1 | 1 | 0% | 2,581 | 5,357 | +108% | 0 | 0 | — |
case-04 | pass→pass | 31,641 | 12,409 | -61% | 1 | 1 | 0% | 2,942 | 4,675 | +59% | 0 | 0 | — |
case-05 | fail→pass | 16,618 | 4,993 | -70% | 1 | 1 | 0% | 3,155 | 3,461 | +10% | 0 | 0 | — |
case-06 | pass→pass | 12,549 | 14,790 | +18% | 1 | 1 | 0% | 2,166 | 5,089 | +135% | 0 | 0 | — |
case-07 | pass→pass | 11,817 | 4,779 | -60% | 1 | 1 | 0% | 1,861 | 3,278 | +76% | 0 | 0 | — |
case-09 | fail→fail | 11,265 | 11,847 | +5% | 1 | 1 | 0% | 2,068 | 4,605 | +123% | 0 | 0 | — |
case-10 | pass→pass | 13,223 | 11,061 | -16% | 1 | 1 | 0% | 2,213 | 4,426 | +100% | 0 | 0 | — |
case-11 | fail→pass | 9,566 | 5,120 | -46% | 1 | 1 | 0% | 1,562 | 3,394 | +117% | 0 | 0 | — |
case-12 | pass→pass | 9,213 | 8,466 | -8% | 1 | 1 | 0% | 1,547 | 3,725 | +141% | 0 | 0 | — |
case-13 | fail→pass | 12,974 | 13,966 | +8% | 1 | 1 | 0% | 2,408 | 5,131 | +113% | 0 | 0 | — |
case-14 | pass→pass | 19,576 | 18,125 | -7% | 1 | 1 | 0% | 3,484 | 5,736 | +65% | 0 | 0 | — |
case-15 | fail→fail | 6,243 | 6,999 | +12% | 1 | 1 | 0% | 1,162 | 3,732 | +221% | 0 | 0 | — |
case-16 | fail→pass | 8,355 | 4,194 | -50% | 1 | 1 | 0% | 1,890 | 3,372 | +78% | 0 | 0 | — |
case-17 | fail→pass | 7,606 | 1,920 | -75% | 1 | 1 | 0% | 1,344 | 2,702 | +101% | 0 | 0 | — |
case-18 | fail→pass | 12,681 | 12,546 | -1% | 1 | 1 | 0% | 2,384 | 4,915 | +106% | 0 | 0 | — |
case-19 | fail→pass | 4,749 | 3,892 | -18% | 1 | 1 | 0% | 792 | 3,154 | +298% | 0 | 0 | — |
case-20 | pass→pass | 5,967 | 7,569 | +27% | 1 | 1 | 0% | 1,391 | 4,096 | +194% | 0 | 0 | — |
case-21 | fail→pass | 9,114 | 9,790 | +7% | 1 | 1 | 0% | 1,957 | 4,466 | +128% | 0 | 0 | — |
case-22 | pass→pass | 8,139 | 7,474 | -8% | 1 | 1 | 0% | 1,869 | 4,216 | +126% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.