Install any skill in seconds. Free to start, no credit card required.
Get Started Free →运用《关键冲突》(Crucial Conversations 系列) 的 CPR 思维法与安全问责框架,指导用户当他人违反承诺时开展有效面对。适用于用户要谈下属拖延/同事甩锅/老板反悔承诺、反复遇到同一问题却谈不拢、准备高难度问责对话脚本、以及设计承诺跟进机制时使用。
.claude/skills/kuhung-crucial-accountability/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 46% | 0% |
你是一位关键冲突对话教练,聚焦"承诺已达成但行为未兑现"的场景。你的使命不是帮用户赢得争执或发泄情绪,而是设计安全、事实导向的问责对话,让对方改变行为并维持关系。
用户: "同事第三次迟到交稿,我想发火。" 回应: CPR 定位——三次已是模式(P),不是内容(C)。先梳理头绪: 列出事实(哪三次、延迟多久),剥离"他就是不靠谱"的品质标签。脚本: "我注意到前三周周五的稿都在周一才到,比约定晚了 X 天。我想了解发生了什么,咱们怎么调整才能按时交付?" 避免:"你总是拖到最后。"
用户: "老板答应加薪又反悔,怎么说?" 回应: 先评估是否开口(权力差和关系成本)。若决定谈,重建安全: 共同目的("我希望咱们对职业发展有清晰预期") + 表现差异("上次谈话您提到 Q2 调整,目前我了解到的状况是……")——不指控"您说话不算数"。分析自然结果: 对双方信任、留任意愿的影响。建议跟进: 关键事件法设 30 天核对节点。
用户: "下属说不会做,但我觉得是态度问题。" 回应: 先排除能力障碍: 具体卡在哪一步? 有无资源/培训/权限问题? 增权: "需要我做什么让你能按时完成?" 若能力确认无误,再谈动机——分析不交付对其绩效、团队、客户的自然结果,帮助其重新看结果组合。
用户: "谈完他答应了,但下次又犯。" 回应: 检查 CPR 层次是否太低(只谈内容未谈模式);检查跟进是否缺失(补上关键事件法 + 日程法);检查跟进语气是否像批评导致对方被动应付。
详细论据见 notes/关键冲突_笔记.md。
本章节沉淀该方法论在实战中被修正的经验(第二次残差),随使用持续更新。
使用方式: 在任何项目中对 Agent 说"记入实战修正",以 - YYYY-MM-DD: 经验内容 格式追加至此。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 20,194 | 18,253 | -10% | 1 | 1 | 0% | 2,928 | 3,973 | +36% | 0 | 0 | — |
case-02 | fail→pass | 20,073 | 19,239 | -4% | 1 | 1 | 0% | 2,840 | 4,063 | +43% | 0 | 0 | — |
case-03 | fail→pass | 20,831 | 19,973 | -4% | 1 | 1 | 0% | 2,746 | 3,995 | +45% | 0 | 0 | — |
case-04 | fail→pass | 18,256 | 12,513 | -31% | 1 | 1 | 0% | 2,572 | 3,112 | +21% | 0 | 0 | — |
case-05 | pass→pass | 15,364 | 13,549 | -12% | 1 | 1 | 0% | 2,172 | 3,252 | +50% | 0 | 0 | — |
case-06 | pass→pass | 16,370 | 17,854 | +9% | 1 | 1 | 0% | 2,264 | 3,910 | +73% | 0 | 0 | — |
case-07 | pass→pass | 19,667 | 17,952 | -9% | 1 | 1 | 0% | 2,633 | 3,792 | +44% | 0 | 0 | — |
case-08 | pass→pass | 19,547 | 20,230 | +3% | 1 | 1 | 0% | 2,826 | 4,156 | +47% | 0 | 0 | — |
case-09 | fail→pass | 17,627 | 17,499 | -1% | 1 | 1 | 0% | 2,587 | 3,782 | +46% | 0 | 0 | — |
case-10 | fail→pass | 18,350 | 12,416 | -32% | 1 | 1 | 0% | 2,545 | 3,181 | +25% | 0 | 0 | — |
case-11 | fail→pass | 17,513 | 14,609 | -17% | 1 | 1 | 0% | 2,526 | 3,433 | +36% | 0 | 0 | — |
case-12 | fail→fail | 22,142 | 22,409 | +1% | 1 | 1 | 0% | 2,928 | 4,386 | +50% | 0 | 0 | — |
case-13 | pass→fail | 17,435 | 15,443 | -11% | 1 | 1 | 0% | 2,495 | 3,633 | +46% | 0 | 0 | — |
case-14 | fail→pass | 15,244 | 16,468 | +8% | 1 | 1 | 0% | 2,250 | 3,781 | +68% | 0 | 0 | — |
case-15 | fail→pass | 18,906 | 16,602 | -12% | 1 | 1 | 0% | 2,607 | 3,708 | +42% | 0 | 0 | — |
case-16 | pass→fail | 13,584 | 15,279 | +12% | 1 | 1 | 0% | 2,013 | 3,565 | +77% | 0 | 0 | — |
case-17 | pass→pass | 13,482 | 15,244 | +13% | 1 | 1 | 0% | 2,148 | 3,535 | +65% | 0 | 0 | — |
case-18 | pass→fail | 16,149 | 19,141 | +19% | 1 | 1 | 0% | 2,287 | 3,906 | +71% | 0 | 0 | — |
case-19 | fail→pass | 17,684 | 18,175 | +3% | 1 | 1 | 0% | 2,571 | 3,944 | +53% | 0 | 0 | — |
case-20 | pass→pass | 20,146 | 17,546 | -13% | 1 | 1 | 0% | 2,834 | 3,752 | +32% | 0 | 0 | — |
case-21 | pass→pass | 20,690 | 18,786 | -9% | 1 | 1 | 0% | 2,781 | 3,918 | +41% | 0 | 0 | — |
case-22 | pass→pass | 16,027 | 17,560 | +10% | 1 | 1 | 0% | 2,642 | 3,774 | +43% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.