Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Reviews a SKILL.md file and evaluates whether it is well-constructed, using best practices distilled from real-world skill analysis. Use when the user asks to check, review, or validate a skill they have built or are building. Triggers: '帮我检查这个skill', '我写的skill有没有问题', '我写的skill好不好', 'review my SKILL.md', 'validate my skill', '检查一下我的skill'. Two modes: quick (P0 only, ~30 seconds) or full (P0+P1+P2, comprehensive). NOT for executing skills, NOT for creating new skills from scratch.
.claude/skills/wtfitsme-design-skill-checker/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 113% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 88% | 0% |
根据对7个真实开源 skill 案例的纵横深挖,提炼出的检查标准,评估一个 skill 的构建质量。
如果用户没有直接提供 SKILL.md 内容,询问: > "请把你的 SKILL.md 内容粘贴过来,或者告诉我文件路径。"
如果给了路径,用 Read 工具读取。
默认:完整模式(P0 + P1 + P2 全面检查,12条)
如果用户明确说"快速检查""只看 P0""时间紧",则切换到快速模式(只检查 P0 核心3条,约30秒)。否则直接执行完整检查。
按所选模式,逐条检查并判定:✅ 通过 / ⚠️ 需改进 / ❌ 缺失
见"输出格式"节。
P0-1 · description 质量 检查 description 是否同时包含:
判定:
P0-2 · 职责单一性 检查这个 skill 是否只干一件能干净归类的事。
判定:
P0-3 · 失败分支 检查 skill 是否有失败处理机制:当依赖的工具/API/数据源/外部服务失败时,是否写明了处理方式。
判定:
P1-1 · 配置/规则的具体性 检查 skill 内部的配置、规则、约束是否具体可执行(非口号)。
判定:
P1-2 · 产出可判定性 检查是否有明确的验收标准,让产出的好坏可以被判断。
判定:
P1-3 · 正文长度 检查 SKILL.md 正文行数是否在合理范围。
判定:
P1-4 · 渐进式披露 检查是否合理使用了 references/ 目录分担细节。
判定:
P2-1 · 样例质量 检查是否有具体的好/坏输入输出样例。
判定:
正反例参考(用于校准判定标准):
❌ 没有样例的 description(P2-1 = ❌):
description: "Helps you write better code."✅ 有具体样例的 description(P2-1 = ✅):
description: "Reviews Rust code for ownership and borrow-checker issues.
Triggers: 'my Rust code has lifetime errors', 'borrow checker keeps rejecting this'.
NOT for: general code style, Python/JS code."同理,P2-1 = ✅ 的条件:skill 正文内有"输入示例 → 输出示例"对,或明确的正反例对比表(如 insight 的3段对话demo、radar 的"10分/1分选题"对比)。仅有"样例可以提高效果"类的说明,判定为 ⚠️。
P2-2 · allowed-tools 设置(工具执行型 skill 必查) 如果 skill 会调用外部工具/执行命令,检查 frontmatter 是否有 allowed-tools 限定。
判定:
P2-3 · 破坏性操作保护(工具执行型 skill 必查) 如果 skill 会执行写入/删除/发布等不可逆操作,检查是否有前置确认机制。
判定:
P2-4 · 方法论 skill 专项(方法论封装型必查) 如果 skill 封装的是一套方法论/框架,检查:
判定:
P2-5 · 迭代友好性 检查是否有 last-updated 字段、版本号、或 CHANGELOG。
判定:
## Skill 快速检查报告
**Skill 名称**:{name}
**检查模式**:快速(P0 核心项)
| # | 检查项 | 结果 | 问题/建议 |
|---|---|---|---|
| P0-1 | description 质量 | ✅/⚠️/❌ | {一句话} |
| P0-2 | 职责单一性 | ✅/⚠️/❌ | {一句话} |
| P0-3 | 失败分支 | ✅/⚠️/N/A | {一句话} |
**快速结论**:{✅ 核心没问题,可继续 / ⚠️ 有X处需注意 / ❌ 有X处关键缺陷,建议修复后再用}
> 需要完整检查?告诉我"完整检查"。## Skill 完整检查报告
**Skill 名称**:{name}
**检查模式**:完整(P0 + P1 + P2)
### P0 · 核心项
| # | 检查项 | 结果 | 问题/建议 |
|---|---|---|---|
| P0-1 | description 质量 | | |
| P0-2 | 职责单一性 | | |
| P0-3 | 失败分支 | | |
### P1 · 重要项
| # | 检查项 | 结果 | 问题/建议 |
|---|---|---|---|
| P1-1 | 配置/规则具体性 | | |
| P1-2 | 产出可判定性 | | |
| P1-3 | 正文长度 | | |
| P1-4 | 渐进式披露 | | |
### P2 · 加分项
| # | 检查项 | 结果 | 问题/建议 |
|---|---|---|---|
| P2-1 | 样例质量 | | |
| P2-2 | allowed-tools | | |
| P2-3 | 破坏性操作保护 | | |
| P2-4 | 方法论专项 | | |
| P2-5 | 迭代友好性 | | |
### 综合评分
| 层级 | 通过/总数 | 状态 |
|---|---|---|
| P0 核心 | X/3 | ✅/⚠️/❌ |
| P1 重要 | X/4 | ✅/⚠️/❌ |
| P2 加分 | X/Y(适用项)| ✅/⚠️/❌ |
**总体评级**:
- ✅ **可发布**:P0 全通过,P1 无❌
- ⚠️ **建议改进后发布**:P0 全通过,P1 有⚠️
- 🔧 **需修复**:P0 有⚠️或 P1 有❌
- ❌ **不建议使用**:P0 有❌
### 优先修复清单
{列出所有❌和重要⚠️,按 P0→P1→P2 排序,每条给具体改法}检查过程中,不要因为以下理由放水:
本 skill 的检查标准提炼自以下案例的纵深分析: wechat-topic-radar(失败分支/配置具体性)、typefully(allowed-tools/破坏性操作)、daymade deep-research(质量门/诚实失败)、Jeffallan(渐进式披露/validator)、market-insight(方法论封装/伦理边界)、trailofbits(反偷懒清单/前置检查分级)、ask-questions(When NOT to Use/职责单一)
以上项目均为开源项目,分析仅供学习参考,设计思路归原作者所有。
references/sample-report.md — 读完检查标准想看实际报告长什么样时加载。包含一个"有典型问题的 mini-SKILL.md"和对应的完整检查报告输出,可用于校准判定尺度。references/sample-report.md:真实 mini-SKILL.md 输入 + 完整检查报告输出样例criar-skill(葡萄牙语),改为通用表述| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,377 | 13,933 | -15% | 1 | 1 | 0% | 2,873 | 5,808 | +102% | 0 | 0 | — |
case-02 | fail→fail | 7,857 | 5,338 | -32% | 1 | 1 | 0% | 1,221 | 3,485 | +185% | 0 | 0 | — |
case-03 | fail→pass | 17,120 | 15,006 | -12% | 1 | 1 | 0% | 2,941 | 5,828 | +98% | 0 | 0 | — |
case-04 | pass→pass | 16,305 | 16,047 | -2% | 1 | 1 | 0% | 2,854 | 5,841 | +105% | 0 | 0 | — |
case-05 | pass→pass | 3,084 | 3,071 | -0% | 1 | 1 | 0% | 474 | 3,574 | +654% | 0 | 0 | — |
case-06 | pass→pass | 11,740 | 8,550 | -27% | 1 | 1 | 0% | 1,887 | 4,392 | +133% | 0 | 0 | — |
case-07 | fail→pass | 15,186 | 13,824 | -9% | 1 | 1 | 0% | 2,669 | 5,684 | +113% | 0 | 0 | — |
case-08 | pass→pass | 18,013 | 14,683 | -18% | 1 | 1 | 0% | 3,084 | 5,835 | +89% | 0 | 0 | — |
case-09 | pass→pass | 16,053 | 14,497 | -10% | 1 | 1 | 0% | 2,643 | 5,949 | +125% | 0 | 0 | — |
case-10 | pass→pass | 11,537 | 5,978 | -48% | 1 | 1 | 0% | 1,805 | 4,018 | +123% | 0 | 0 | — |
case-11 | fail→pass | 19,243 | 12,943 | -33% | 1 | 1 | 0% | 3,083 | 5,239 | +70% | 0 | 0 | — |
case-12 | pass→pass | 14,348 | 11,888 | -17% | 1 | 1 | 0% | 2,353 | 5,426 | +131% | 0 | 0 | — |
case-13 | pass→pass | 15,812 | 10,914 | -31% | 1 | 1 | 0% | 2,545 | 5,192 | +104% | 0 | 0 | — |
case-14 | fail→pass | 18,578 | 15,730 | -15% | 1 | 1 | 0% | 3,129 | 5,895 | +88% | 0 | 0 | — |
case-15 | pass→pass | 13,240 | 11,421 | -14% | 1 | 1 | 0% | 2,402 | 5,496 | +129% | 0 | 0 | — |
case-16 | pass→pass | 16,485 | 14,304 | -13% | 1 | 1 | 0% | 2,792 | 5,698 | +104% | 0 | 0 | — |
case-17 | pass→pass | 15,623 | 11,458 | -27% | 1 | 1 | 0% | 2,629 | 5,319 | +102% | 0 | 0 | — |
case-18 | fail→pass | 16,663 | 3,333 | -80% | 1 | 1 | 0% | 2,606 | 3,507 | +35% | 0 | 0 | — |
case-19 | fail→fail | 20,002 | 10,144 | -49% | 1 | 1 | 0% | 3,398 | 5,315 | +56% | 0 | 0 | — |
case-20 | pass→pass | 15,423 | 16,124 | +5% | 1 | 1 | 0% | 2,460 | 5,194 | +111% | 0 | 0 | — |
case-21 | fail→pass | 11,161 | 5,983 | -46% | 1 | 1 | 0% | 1,849 | 4,069 | +120% | 0 | 0 | — |
case-22 | pass→pass | 14,644 | 7,095 | -52% | 1 | 1 | 0% | 2,176 | 4,269 | +96% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.