Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Guides experiment state transitions: launching, pausing, resuming, ending, shipping variants, archiving, resetting, and duplicating. Covers preconditions, implications for variant assignment and analysis, and the decision framework for when to use each action. TRIGGER when: user asks to launch, pause, resume, end, ship, archive, reset, or duplicate an experiment. DO NOT TRIGGER when: user is creat
.claude/skills/kunanonj-cursor-plugin-posthog-managing-experiment-lifecycle/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 329% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 2% | 0% |
This skill covers experiment state transitions — what each action does, when to use it, and how it affects variant assignment and analysis.
textdraft ──launch──▶ running ──end──▶ stopped ──archive──▶ archived │ ▲ │ pause resume ship_variant │ │ (also ends if running) ▼ │ paused (flag inactive, still "running" status) Any non-draft state ──reset──▶ draft
For each action, the two key questions:
experiment-launch)Transitions draft → running. Activates the feature flag and sets start_date.
start_dateNo request body needed.
experiment-pause)Deactivates the feature flag. Users fall back to the default experience (typically control).
/decide — no new exposure events recordedNo request body. Use experiment-resume to reactivate.
experiment-resume)Reactivates the feature flag after a pause. Users are re-bucketed deterministically into the same variants.
No request body.
experiment-end)Sets end_date and transitions to stopped. The feature flag is NOT modified.
end_dateOptional body: conclusion ("won", "lost", "inconclusive", "stopped_early", "invalid") and conclusion_comment.
Use this when you want to freeze results without changing what users see.
experiment-ship-variant)Rewrites the feature flag so the selected variant is served to 100% of users.
Always confirm with the user before shipping — this permanently rewrites the feature flag.
Required: variant_key (e.g. "test"). Optional: conclusion, conclusion_comment.
Returns 409 if an approval policy requires review before the flag change.
experiment-archive)Hides a stopped experiment from the default list view.
No request body. Can be restored by setting archived=false via experiment-update.
experiment-reset)Returns an experiment to draft state. Clears start_date, end_date, conclusion, and archived.
start_date is adjusted after re-launchNo request body.
experiment-duplicate)Creates a copy as a new draft with fresh dates and no results.
Important: always provide a unique feature_flag_key different from the original. If the same key is used, both experiments share a flag — changes to one affect both.
Optional: custom name (defaults to "Original Name (Copy)").
| Situation | Action | Tool | | -------------------------------------------------- | ------------------------ | ------------------------- | | Draft ready, flag implemented, metrics set | Launch | experiment-launch | | Clear winner, significant results | Ship the winning variant | experiment-ship-variant | | No significant difference after sufficient time | End as inconclusive | experiment-end | | Something wrong, need to stop exposure temporarily | Pause | experiment-pause | | Resume after pause | Resume | experiment-resume | | Experiment ended, ready to clean up | Archive | experiment-archive | | Need to start over with same config | Reset to draft | experiment-reset | | Want a similar experiment with a fresh start | Duplicate | experiment-duplicate |
All lifecycle actions require an experiment ID. If you don't have one, load the finding-experiments skill to resolve the user's reference (name, description, "latest", etc.) to a concrete ID before proceeding.
| Error message | Meaning | | --------------------------------------- | ------------------------------------ | | "Experiment has already been launched." | Can't launch a non-draft experiment | | "Experiment has not been launched yet." | Can't end/pause/ship a draft | | "Experiment has already ended." | Can't end/pause a stopped experiment | | "Experiment is already paused." | Use resume instead | | "Experiment is not paused." | It's already active | | "Experiment is already in draft state." | Nothing to reset | | "Experiment is already archived." | Already done |
When you get a 400, explain the situation to the user rather than retrying.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 7,157 | 4,664 | -35% | 1 | 1 | 0% | 546 | 2,342 | +329% | 0 | 0 | — |
case-02 | fail→fail | 10,826 | 8,006 | -26% | 1 | 1 | 0% | 1,572 | 1,754 | +12% | 0 | 0 | — |
case-03 | fail→pass | 9,804 | 4,856 | -50% | 1 | 1 | 0% | 1,652 | 2,139 | +29% | 0 | 0 | — |
case-12 | fail→pass | 8,945 | 2,013 | -77% | 1 | 1 | 0% | 1,611 | 1,712 | +6% | 0 | 0 | — |
case-04 | fail→pass | 8,488 | 5,605 | -34% | 1 | 1 | 0% | 1,534 | 2,288 | +49% | 0 | 0 | — |
case-05 | pass→pass | 7,809 | 5,011 | -36% | 1 | 1 | 0% | 1,160 | 2,040 | +76% | 0 | 0 | — |
case-06 | pass→pass | 8,592 | 3,219 | -63% | 1 | 1 | 0% | 1,241 | 1,910 | +54% | 0 | 0 | — |
case-07 | pass→pass | 5,789 | 3,159 | -45% | 1 | 1 | 0% | 974 | 1,872 | +92% | 0 | 0 | — |
case-08 | pass→pass | 7,262 | 3,717 | -49% | 1 | 1 | 0% | 1,232 | 2,032 | +65% | 0 | 0 | — |
case-09 | pass→pass | 11,649 | 4,524 | -61% | 1 | 1 | 0% | 1,678 | 2,121 | +26% | 0 | 0 | — |
case-10 | pass→pass | 10,214 | 3,843 | -62% | 1 | 1 | 0% | 1,629 | 2,028 | +24% | 0 | 0 | — |
case-11 | fail→pass | 11,256 | 3,484 | -69% | 1 | 1 | 0% | 1,908 | 1,937 | +2% | 0 | 0 | — |
case-13 | pass→pass | 4,935 | 2,560 | -48% | 1 | 1 | 0% | 791 | 1,819 | +130% | 0 | 0 | — |
case-14 | pass→pass | 4,600 | 2,488 | -46% | 1 | 1 | 0% | 645 | 1,755 | +172% | 0 | 0 | — |
case-15 | pass→pass | 6,665 | 2,487 | -63% | 1 | 1 | 0% | 1,076 | 1,731 | +61% | 0 | 0 | — |
case-16 | pass→pass | 12,241 | 3,706 | -70% | 1 | 1 | 0% | 1,777 | 1,993 | +12% | 0 | 0 | — |
case-17 | pass→pass | 9,612 | 1,917 | -80% | 1 | 1 | 0% | 1,661 | 1,721 | +4% | 0 | 0 | — |
case-18 | fail→pass | 4,536 | 1,992 | -56% | 1 | 1 | 0% | 703 | 1,665 | +137% | 0 | 0 | — |
case-19 | pass→pass | 7,357 | 2,691 | -63% | 1 | 1 | 0% | 1,093 | 1,831 | +68% | 0 | 0 | — |
case-20 | pass→pass | 4,116 | 4,619 | +12% | 1 | 1 | 0% | 605 | 2,131 | +252% | 0 | 0 | — |
case-21 | fail→pass | 8,727 | 6,160 | -29% | 1 | 1 | 0% | 1,567 | 2,468 | +57% | 0 | 0 | — |
case-22 | pass→pass | 11,116 | 15,804 | +42% | 1 | 1 | 0% | 1,911 | 3,912 | +105% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.