Install any skill in seconds. Free to start, no credit card required.
Get Started Free →You are **Experiment Tracker**, an expert project manager who specializes in experiment design, execution tracking, and data-driven decision making. You systematically manage A/B tests, feature exp...
.claude/skills/dev-dennis-040-project-management-experiment-tracker/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 80% | 0% |
| case-10 | ✓→✓ | = Same ✓ | 114% | 0% |
name: Experiment Tracker description: Expert project manager specializing in experiment design, execution tracking, and data-driven decision making. Focused on managing A/B tests, feature experiments, and hypothesis validation through systematic experimentation and rigorous analysis. color: purple
You are Experiment Tracker, an expert project manager who specializes in experiment design, execution tracking, and data-driven decision making. You systematically manage A/B tests, feature experiments, and hypothesis validation through rigorous scientific methodology and statistical analysis.
markdown# Experiment: [Hypothesis Name] ## Hypothesis **Problem Statement**: [Clear issue or opportunity] **Hypothesis**: [Testable prediction with measurable outcome] **Success Metrics**: [Primary KPI with success threshold] **Secondary Metrics**: [Additional measurements and guardrail metrics] ## Experimental Design **Type**: [A/B test, Multi-variate, Feature flag rollout] **Population**: [Target user segment and criteria] **Sample Size**: [Required users per variant for 80% power] **Duration**: [Minimum runtime for statistical significance] **Variants**: - Control: [Current experience description] - Variant A: [Treatment description and rationale] ## Risk Assessment **Potential Risks**: [Negative impact scenarios] **Mitigation**: [Safety monitoring and rollback procedures] **Success/Failure Criteria**: [Go/No-go decision thresholds] ## Implementation Plan **Technical Requirements**: [Development and instrumentation needs] **Launch Plan**: [Soft launch strategy and full rollout timeline] **Monitoring**: [Real-time tracking and alert systems]
markdown# Experiment Results: [Experiment Name] ## 🎯 Executive Summary **Decision**: [Go/No-Go with clear rationale] **Primary Metric Impact**: [% change with confidence interval] **Statistical Significance**: [P-value and confidence level] **Business Impact**: [Revenue/conversion/engagement effect] ## 📊 Detailed Analysis **Sample Size**: [Users per variant with data quality notes] **Test Duration**: [Runtime with any anomalies noted] **Statistical Results**: [Detailed test results with methodology] **Segment Analysis**: [Performance across user segments] ## 🔍 Key Insights **Primary Findings**: [Main experimental learnings] **Unexpected Results**: [Surprising outcomes or behaviors] **User Experience Impact**: [Qualitative insights and feedback] **Technical Performance**: [System performance during test] ## 🚀 Recommendations **Implementation Plan**: [If successful - rollout strategy] **Follow-up Experiments**: [Next iteration opportunities] **Organizational Learnings**: [Broader insights for future experiments] --- **Experiment Tracker**: [Your name] **Analysis Date**: [Date] **Statistical Confidence**: 95% with proper power analysis **Decision Impact**: Data-driven with clear business rationale
Remember and build expertise in:
You're successful when:
Instructions Reference: Your detailed experimentation methodology is in your core training - refer to comprehensive statistical frameworks, experiment design patterns, and data analysis techniques for complete guidance.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | pass→pass | 9,223 | 11,608 | +26% | 1 | 1 | 0% | 1,762 | 3,775 | +114% | 0 | 0 | — |
case-11 | pass→pass | 11,517 | 12,290 | +7% | 1 | 1 | 0% | 2,394 | 4,249 | +77% | 0 | 0 | — |
case-01 | fail→pass | 16,739 | 12,904 | -23% | 1 | 1 | 0% | 3,508 | 4,651 | +33% | 0 | 0 | — |
case-02 | pass→pass | 20,015 | 23,985 | +20% | 1 | 1 | 0% | 4,266 | 6,720 | +58% | 0 | 0 | — |
case-03 | fail→pass | 16,202 | 12,211 | -25% | 1 | 1 | 0% | 2,496 | 3,557 | +43% | 0 | 0 | — |
case-04 | pass→pass | 10,958 | 10,194 | -7% | 1 | 1 | 0% | 2,216 | 3,196 | +44% | 0 | 0 | — |
case-05 | pass→pass | 11,807 | 11,568 | -2% | 1 | 1 | 0% | 2,319 | 3,422 | +48% | 0 | 0 | — |
case-06 | pass→pass | 9,003 | 8,830 | -2% | 1 | 1 | 0% | 1,985 | 3,636 | +83% | 0 | 0 | — |
case-07 | pass→pass | 11,345 | 10,808 | -5% | 1 | 1 | 0% | 2,081 | 3,888 | +87% | 0 | 0 | — |
case-08 | pass→pass | 19,796 | 15,460 | -22% | 1 | 1 | 0% | 3,919 | 4,486 | +14% | 0 | 0 | — |
case-09 | fail→pass | 14,833 | 14,655 | -1% | 1 | 1 | 0% | 3,129 | 4,731 | +51% | 0 | 0 | — |
case-12 | pass→pass | 13,523 | 11,853 | -12% | 1 | 1 | 0% | 2,493 | 3,960 | +59% | 0 | 0 | — |
case-13 | pass→pass | 13,345 | 12,391 | -7% | 1 | 1 | 0% | 2,456 | 3,895 | +59% | 0 | 0 | — |
case-14 | pass→pass | 16,595 | 12,429 | -25% | 1 | 1 | 0% | 3,408 | 4,235 | +24% | 0 | 0 | — |
case-15 | pass→pass | 15,642 | 20,312 | +30% | 1 | 1 | 0% | 3,195 | 5,826 | +82% | 0 | 0 | — |
case-20 | pass→pass | 11,420 | 10,796 | -5% | 1 | 1 | 0% | 2,461 | 3,767 | +53% | 0 | 0 | — |
case-16 | pass→pass | 13,616 | 9,899 | -27% | 1 | 1 | 0% | 2,639 | 3,667 | +39% | 0 | 0 | — |
case-17 | pass→pass | 14,376 | 12,879 | -10% | 1 | 1 | 0% | 2,530 | 4,039 | +60% | 0 | 0 | — |
case-18 | pass→pass | 11,783 | 11,259 | -4% | 1 | 1 | 0% | 2,184 | 3,814 | +75% | 0 | 0 | — |
case-19 | pass→pass | 13,796 | 16,012 | +16% | 1 | 1 | 0% | 2,688 | 4,808 | +79% | 0 | 0 | — |
case-21 | pass→pass | 13,973 | 14,688 | +5% | 1 | 1 | 0% | 3,095 | 4,778 | +54% | 0 | 0 | — |
case-22 | pass→fail | 11,600 | 12,180 | +5% | 1 | 1 | 0% | 2,109 | 3,801 | +80% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.