Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Product analytics agent for KPI definition, dashboard setup, experiment design, and test result interpretation. Use when a product question needs numbers — e.g., defining activation/retention KPIs and a dashboard spec for a new feature, or sizing an A/B test and judging whether the result is significant enough to ship.
.claude/skills/alirezarezvani-cs-product-analyst/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 244% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -27% | 0% |
The cs-product-analyst agent turns product questions into measurable answers. It orchestrates the product-analytics and experiment-designer skills to define metric frameworks, compute retention/cohort/funnel metrics from raw CSV exports, size experiments before they run, and interpret results after they finish — separating statistical significance from practical business significance.
Use this agent instead of cs-product-manager when the work is quantitative: the PM agent decides what to build; this agent measures whether it worked.
Skill Locations:
../../product-team/skills/product-analytics/ (SKILL.md)../../product-team/skills/experiment-designer/ (SKILL.md)../../product-team/skills/product-analytics/scripts/metrics_calculator.pypython ../../product-team/skills/product-analytics/scripts/metrics_calculator.py retention events.csv (subcommands: retention, cohort, funnel)../../product-team/skills/experiment-designer/scripts/sample_size_calculator.pypython ../../product-team/skills/experiment-designer/scripts/sample_size_calculator.py --baseline-rate 0.12 --mde 0.02 --mde-type absolute --daily-samples 800Goal: Define the decision metric, supporting metrics, and guardrails for a feature before any analysis runs.
Steps:
Expected Output: A one-page metric spec with primary KPI, guardrails, and dashboard layout.
Goal: Quantify how users actually behave from raw event exports.
Steps:
metrics_calculator.py retention|cohort|funnel on the exportExpected Output: Retention curve / cohort matrix / funnel table with a written interpretation and one recommended action.
Goal: Size a test before launch; judge the result after.
Steps:
sample_size_calculator.py to get required n and runtime at current trafficExpected Output: Pre-registered test plan, then a decision memo with effect size, confidence, guardrail status, and recommendation.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | fail→fail | 10,039 | 13,182 | +31% | 1 | 1 | 0% | 1,651 | 3,159 | +91% | 0 | 0 | — |
case-03 | fail→pass | 14,469 | 15,972 | +10% | 1 | 1 | 0% | 2,722 | 3,918 | +44% | 0 | 0 | — |
case-14 | pass→pass | 11,784 | 6,841 | -42% | 1 | 1 | 0% | 2,006 | 2,167 | +8% | 0 | 0 | — |
case-02 | fail→fail | 11,159 | 13,931 | +25% | 1 | 1 | 0% | 2,128 | 3,189 | +50% | 0 | 0 | — |
case-01 | fail→pass | 7,123 | 29,372 | +312% | 1 | 1 | 0% | 1,493 | 5,132 | +244% | 0 | 0 | — |
case-04 | fail→pass | 11,867 | 3,039 | -74% | 1 | 1 | 0% | 2,422 | 1,481 | -39% | 0 | 0 | — |
case-05 | fail→pass | 33,636 | 3,679 | -89% | 1 | 1 | 0% | 2,369 | 1,696 | -28% | 0 | 0 | — |
case-06 | fail→pass | 9,869 | 2,981 | -70% | 1 | 1 | 0% | 2,057 | 1,504 | -27% | 0 | 0 | — |
case-07 | fail→pass | 11,523 | 4,111 | -64% | 1 | 1 | 0% | 2,610 | 1,902 | -27% | 0 | 0 | — |
case-08 | fail→fail | 13,735 | 10,995 | -20% | 1 | 1 | 0% | 2,265 | 2,977 | +31% | 0 | 0 | — |
case-09 | pass→pass | 13,359 | 9,053 | -32% | 1 | 1 | 0% | 2,310 | 2,706 | +17% | 0 | 0 | — |
case-10 | pass→pass | 8,976 | 7,346 | -18% | 1 | 1 | 0% | 1,756 | 2,266 | +29% | 0 | 0 | — |
case-11 | pass→pass | 11,538 | 12,143 | +5% | 1 | 1 | 0% | 2,093 | 2,714 | +30% | 0 | 0 | — |
case-12 | pass→pass | 13,269 | 12,798 | -4% | 1 | 1 | 0% | 2,127 | 3,209 | +51% | 0 | 0 | — |
case-13 | fail→pass | 12,188 | 8,085 | -34% | 1 | 1 | 0% | 2,140 | 2,447 | +14% | 0 | 0 | — |
case-15 | fail→fail | 12,017 | 18,204 | +51% | 1 | 1 | 0% | 2,236 | 4,034 | +80% | 0 | 0 | — |
case-16 | pass→pass | 14,186 | 11,600 | -18% | 1 | 1 | 0% | 2,056 | 2,973 | +45% | 0 | 0 | — |
case-17 | fail→pass | 13,866 | 3,146 | -77% | 1 | 1 | 0% | 2,206 | 1,613 | -27% | 0 | 0 | — |
case-18 | fail→pass | 6,604 | 2,006 | -70% | 1 | 1 | 0% | 967 | 1,099 | +14% | 0 | 0 | — |
case-19 | fail→pass | 13,201 | 11,841 | -10% | 1 | 1 | 0% | 2,509 | 3,022 | +20% | 0 | 0 | — |
case-21 | fail→fail | 12,187 | 10,475 | -14% | 1 | 1 | 0% | 2,173 | 2,719 | +25% | 0 | 0 | — |
case-22 | fail→fail | 25,542 | 17,923 | -30% | 1 | 1 | 0% | 4,352 | 4,604 | +6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.