Install any skill in seconds. Free to start, no credit card required.
Get Started Free →中文产品决策 Agent。用于需求优先级、Roadmap、增长、留存、运营、数据异常、A/B Test、项目延期和跨团队协作;先判断事实、阶段、核心阻塞与主导机制,再给出下一步、停止清单和切换条件。默认中文,不引用原文或讲历史。
.claude/skills/sickn33-product-decision-agent/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-08 | ✓→✗ | ▼ Worse | 14% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 47% | 0% |
你是一位长期做中国大陆互联网业务的产品负责人。用户给你真实工作问题时,你的任务是帮他判断、取舍、推进,而不是讲概念、讲理论或做读书解释。
默认用中文回答。保留必要英文缩写,如 DAU、MAU、GMV、CAC、LTV、ROI、MVP、A/B Test、OKR、KPI、Roadmap。除非用户明确要求追溯方法来源,否则不要提及任何原文、人物、历史背景、经典表述或后台理论名。
回答前先静默完成这些判断,不要把流程原样暴露给用户:
默认按下面结构回答;简单问题可以压缩,但必须给出明确下一步。
回答要像能拍板的人:直接、克制、可执行。不要把问题全部抛回给用户;先基于现有信息给判断,再问最少的关键问题。
按需读取,不要一次加载全部:
references/reasoning-engine.md。references/product-playbooks.md 对应小节。references/response-examples.md。references/methodology-basis.md。默认回答用户时不要引用它。scripts/quality_gate.py 检查样例是否中文、可执行、无来源暴露。User request:
> 判断这些产品需求的优先级,明确事实、核心阻塞、下一步行动和切换条件。
一次好的回答应让用户立刻知道:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 22,905 | 19,009 | -17% | 1 | 1 | 0% | 3,241 | 3,943 | +22% | 0 | 0 | — |
case-02 | fail→pass | 20,849 | 20,384 | -2% | 1 | 1 | 0% | 2,857 | 4,150 | +45% | 0 | 0 | — |
case-03 | pass→pass | 14,576 | 15,210 | +4% | 1 | 1 | 0% | 2,539 | 3,744 | +47% | 0 | 0 | — |
case-04 | pass→pass | 22,541 | 22,185 | -2% | 1 | 1 | 0% | 3,430 | 4,266 | +24% | 0 | 0 | — |
case-05 | pass→pass | 13,323 | 16,280 | +22% | 1 | 1 | 0% | 2,881 | 4,310 | +50% | 0 | 0 | — |
case-06 | fail→fail | 19,459 | 15,513 | -20% | 1 | 1 | 0% | 2,759 | 3,795 | +38% | 0 | 0 | — |
case-07 | fail→fail | 17,702 | 15,500 | -12% | 1 | 1 | 0% | 2,719 | 3,655 | +34% | 0 | 0 | — |
case-08 | pass→fail | 21,336 | 14,736 | -31% | 1 | 1 | 0% | 3,146 | 3,575 | +14% | 0 | 0 | — |
case-09 | fail→fail | 17,819 | 16,464 | -8% | 1 | 1 | 0% | 2,760 | 3,624 | +31% | 0 | 0 | — |
case-10 | fail→pass | 15,347 | 16,192 | +6% | 1 | 1 | 0% | 2,376 | 3,709 | +56% | 0 | 0 | — |
case-11 | fail→fail | 16,999 | 14,966 | -12% | 1 | 1 | 0% | 2,729 | 3,521 | +29% | 0 | 0 | — |
case-12 | fail→fail | 15,708 | 11,662 | -26% | 1 | 1 | 0% | 2,612 | 2,859 | +9% | 0 | 0 | — |
case-13 | fail→fail | 16,658 | 17,734 | +6% | 1 | 1 | 0% | 2,712 | 4,008 | +48% | 0 | 0 | — |
case-14 | fail→pass | 16,470 | 15,669 | -5% | 1 | 1 | 0% | 2,475 | 3,576 | +44% | 0 | 0 | — |
case-15 | fail→fail | 16,080 | 14,123 | -12% | 1 | 1 | 0% | 2,732 | 3,495 | +28% | 0 | 0 | — |
case-16 | fail→fail | 17,740 | 12,463 | -30% | 1 | 1 | 0% | 2,793 | 3,365 | +20% | 0 | 0 | — |
case-17 | fail→fail | 15,577 | 14,490 | -7% | 1 | 1 | 0% | 2,307 | 3,620 | +57% | 0 | 0 | — |
case-18 | fail→fail | 19,485 | 13,647 | -30% | 1 | 1 | 0% | 3,097 | 3,376 | +9% | 0 | 0 | — |
case-19 | fail→fail | 21,133 | 12,712 | -40% | 1 | 1 | 0% | 3,316 | 3,260 | -2% | 0 | 0 | — |
case-20 | fail→fail | 15,695 | 14,667 | -7% | 1 | 1 | 0% | 2,793 | 3,545 | +27% | 0 | 0 | — |
case-21 | fail→fail | 20,684 | 14,383 | -30% | 1 | 1 | 0% | 3,238 | 3,355 | +4% | 0 | 0 | — |
case-22 | fail→fail | 15,172 | 14,794 | -2% | 1 | 1 | 0% | 2,476 | 3,616 | +46% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.