Install any skill in seconds. Free to start, no credit card required.
Get Started Free →假设 + 指标 + 结果 + 解释 + 决策, 把 A/B 或产品实验转成行动建议
.claude/skills/nexu-io-experiment-readout/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 182% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 264% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 181% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 229% | 0% |
| case-18 | ✓→✗ | ▼ Worse | 199% | 0% |
【模板: 实验复盘 / Experiment Readout】 【意图】这不是普通数据报告、不是 dashboard。目标是回答: "这个实验说明了什么, 我们下一步应该上线、停止、继续跑, 还是重新设计?"
【适合输入】
【必须输出的结构】
【设计要求】
【可选风格模板 — 参考 assets/】 根据实验语境选择一种, 不要三种混用:
assets/product-readout.html: 默认风格。浅色产品实验复盘, 适合 PM / growth / leadership readout。assets/lab-notebook.html: 研究实验室 notebook, 适合 early-stage experiment、定性 + 定量混合、需要保留 caveat 的探索实验。assets/growth-console.html: 深色 growth analytics console, 适合增长团队、实时指标、漏斗 / activation / conversion readout。如果用户没有指定风格, 优先使用 product-readout; 如果材料强调研究过程和不确定性, 使用 lab-notebook; 如果材料强调增长指标、漏斗、实时监控或运营节奏, 使用 growth-console。
【内容真实性】
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→pass | 12,356 | 26,960 | +118% | 1 | 1 | 0% | 2,406 | 6,795 | +182% | 0 | 0 | — |
case-01 | fail→fail | 15,920 | 27,426 | +72% | 1 | 1 | 0% | 3,069 | 6,845 | +123% | 0 | 0 | — |
case-02 | fail→fail | 20,127 | 29,064 | +44% | 1 | 1 | 0% | 3,180 | 6,824 | +115% | 0 | 0 | — |
case-03 | fail→fail | 16,681 | 25,550 | +53% | 1 | 1 | 0% | 2,910 | 6,840 | +135% | 0 | 0 | — |
case-04 | pass→pass | 12,600 | 24,726 | +96% | 1 | 1 | 0% | 2,449 | 6,782 | +177% | 0 | 0 | — |
case-05 | pass→pass | 11,805 | 16,423 | +39% | 1 | 1 | 0% | 2,254 | 3,909 | +73% | 0 | 0 | — |
case-06 | pass→fail | 11,762 | 26,689 | +127% | 1 | 1 | 0% | 2,059 | 6,768 | +229% | 0 | 0 | — |
case-08 | fail→fail | 12,559 | 28,704 | +129% | 1 | 1 | 0% | 2,042 | 6,775 | +232% | 0 | 0 | — |
case-09 | pass→pass | 13,294 | 27,473 | +107% | 1 | 1 | 0% | 2,303 | 6,779 | +194% | 0 | 0 | — |
case-10 | fail→pass | 10,996 | 27,264 | +148% | 1 | 1 | 0% | 1,863 | 6,780 | +264% | 0 | 0 | — |
case-11 | pass→pass | 9,614 | 25,110 | +161% | 1 | 1 | 0% | 1,917 | 6,792 | +254% | 0 | 0 | — |
case-12 | fail→fail | 12,138 | 26,142 | +115% | 1 | 1 | 0% | 2,084 | 6,762 | +224% | 0 | 0 | — |
case-17 | pass→pass | 12,805 | 26,450 | +107% | 1 | 1 | 0% | 2,323 | 6,775 | +192% | 0 | 0 | — |
case-13 | pass→pass | 25,862 | 26,658 | +3% | 1 | 1 | 0% | 6,171 | 6,763 | +10% | 0 | 0 | — |
case-14 | fail→fail | 18,339 | 24,285 | +32% | 1 | 1 | 0% | 3,250 | 6,777 | +109% | 0 | 0 | — |
case-15 | pass→pass | 10,891 | 27,671 | +154% | 1 | 1 | 0% | 2,073 | 6,782 | +227% | 0 | 0 | — |
case-16 | pass→pass | 11,099 | 27,393 | +147% | 1 | 1 | 0% | 2,268 | 6,792 | +199% | 0 | 0 | — |
case-18 | pass→fail | 14,971 | 26,757 | +79% | 1 | 1 | 0% | 2,264 | 6,767 | +199% | 0 | 0 | — |
case-19 | pass→pass | 11,824 | 28,360 | +140% | 1 | 1 | 0% | 2,038 | 6,785 | +233% | 0 | 0 | — |
case-20 | pass→pass | 10,153 | 27,797 | +174% | 1 | 1 | 0% | 1,687 | 6,778 | +302% | 0 | 0 | — |
case-21 | fail→pass | 14,588 | 30,111 | +106% | 1 | 1 | 0% | 2,406 | 6,762 | +181% | 0 | 0 | — |
case-22 | pass→pass | 9,281 | 28,531 | +207% | 1 | 1 | 0% | 2,150 | 6,786 | +216% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.