Install any skill in seconds. Free to start, no credit card required.
Get Started Free →SOP: Run the external ARA rigor-reviewer (Seal Level 2, six-dimension semantic review) over ../ara/ and pass its level2_report.json to the user
.claude/skills/yogsoth-ai-ara-rigor-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -66% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -59% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -29% | 0% |
Key question: 这份 ARA 的认识论严谨度如何?逻辑弧在结构上闭合了吗?
先确认外部 rigor-reviewer skill 可 load。不可用则提示安装并停下。
Skill load rigor-reviewer,传 <artifact_dir> = ../ara/。它对 ARA 跑六维语义审查(全是要读懂 + 推理的语义检查,不是结构校验):
rigor-reviewer 在 artifact 根目录写 level2_report.json(每维 1–5 分 + strengths/weaknesses/suggestions + severity 排序 findings + overall grade + 给作者的问题)。
context-exploring 补打捞过程线。本 SOP 不自动循环。
> 注意:rigor-reviewer 的 D1–D6 是 ARA 自己的维度,与 DARE 的 D1–D5 评判 > 标准是两套东西,不要混。本 SOP 只透传 ARA 的报告,不施加 DARE 的 D1–D5。
ara/level2_report.json + 一句话总结(grade + 最该关注的 finding),交付用户。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-17 | fail→fail | 14,216 | 4,308 | -70% | 1 | 1 | 0% | 2,512 | 1,158 | -54% | 0 | 0 | — |
case-01 | fail→fail | 4,840 | 4,014 | -17% | 1 | 1 | 0% | 697 | 695 | -0% | 0 | 0 | — |
case-02 | fail→fail | 10,754 | 5,883 | -45% | 1 | 1 | 0% | 1,570 | 752 | -52% | 0 | 0 | — |
case-03 | fail→fail | 16,991 | 4,477 | -74% | 1 | 1 | 0% | 2,599 | 672 | -74% | 0 | 0 | — |
case-04 | fail→pass | 8,893 | 6,508 | -27% | 1 | 1 | 0% | 1,145 | 1,507 | +32% | 0 | 0 | — |
case-05 | fail→pass | 14,865 | 2,138 | -86% | 1 | 1 | 0% | 2,197 | 743 | -66% | 0 | 0 | — |
case-06 | pass→pass | 13,692 | 3,776 | -72% | 1 | 1 | 0% | 1,831 | 1,077 | -41% | 0 | 0 | — |
case-07 | fail→pass | 15,039 | 7,096 | -53% | 1 | 1 | 0% | 2,182 | 1,563 | -28% | 0 | 0 | — |
case-08 | fail→fail | 12,061 | 1,998 | -83% | 1 | 1 | 0% | 1,751 | 773 | -56% | 0 | 0 | — |
case-09 | fail→fail | 12,514 | 5,883 | -53% | 1 | 1 | 0% | 1,860 | 1,405 | -24% | 0 | 0 | — |
case-10 | pass→pass | 11,137 | 4,463 | -60% | 1 | 1 | 0% | 1,789 | 1,198 | -33% | 0 | 0 | — |
case-11 | fail→fail | 11,342 | 3,690 | -67% | 1 | 1 | 0% | 1,612 | 996 | -38% | 0 | 0 | — |
case-12 | pass→pass | 14,171 | 7,379 | -48% | 1 | 1 | 0% | 2,092 | 1,621 | -23% | 0 | 0 | — |
case-13 | pass→pass | 11,558 | 3,012 | -74% | 1 | 1 | 0% | 1,759 | 946 | -46% | 0 | 0 | — |
case-14 | pass→pass | 10,498 | 3,470 | -67% | 1 | 1 | 0% | 1,586 | 928 | -41% | 0 | 0 | — |
case-15 | fail→pass | 15,675 | 3,338 | -79% | 1 | 1 | 0% | 2,278 | 928 | -59% | 0 | 0 | — |
case-16 | fail→fail | 8,534 | 1,775 | -79% | 1 | 1 | 0% | 1,299 | 739 | -43% | 0 | 0 | — |
case-18 | fail→pass | 8,537 | 2,677 | -69% | 1 | 1 | 0% | 1,223 | 866 | -29% | 0 | 0 | — |
case-19 | fail→fail | 14,208 | 10,704 | -25% | 1 | 1 | 0% | 2,001 | 1,956 | -2% | 0 | 0 | — |
case-20 | fail→pass | 12,947 | 3,481 | -73% | 1 | 1 | 0% | 1,874 | 1,031 | -45% | 0 | 0 | — |
case-21 | fail→fail | 8,236 | 7,065 | -14% | 1 | 1 | 0% | 1,252 | 1,693 | +35% | 0 | 0 | — |
case-22 | fail→fail | 4,623 | 11,835 | +156% | 1 | 1 | 0% | 669 | 2,042 | +205% | 0 | 0 | — |
case-23 | fail→fail | 6,063 | 8,160 | +35% | 1 | 1 | 0% | 943 | 1,358 | +44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +26 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.