Install any skill in seconds. Free to start, no credit card required.
Get Started Free →SOP: Turn the feeding plan into the compiler's $ARGUMENTS and run the external ARA compiler once inline to produce ../ara/
.claude/skills/yogsoth-ai-ara-compile/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-21 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -68% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 169% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -28% | 0% |
Key question: 怎么把投喂计划喂给外部 compiler,一次抽出一份内部一致的 ARA?
先确认外部 compiler skill 可 load(ARA skills 已装:npx @ara-commons/ara-skills)。 若不可用,提示用户安装并停下,不要静默继续。
ARA 的 cross-layer binding(claim→proof→evidence、tree→claim)必须全局一致。 分批 compile 会各自从 C01 起撞 ID、断 tree,汇总等于重缝半成品 —— 正是 ARA 要消灭 的事。compiler 自带覆盖度循环(max 3 轮)+ 内建 Task 工具;真需要并行由它内部 自理,本 SOP 不越俎拆分。
$ARGUMENTS:--output ../ara/(与 context/ 平级,天然不会被下次 review 当 context 吃回去)。例: compiler context/2026-06-06-01-30-stage7-...md context/2026-06-05-...stage6...md \ context/figures/*.png \ --output ../ara/ \ 主干=stage7(报告线);其余为过程线/图片;大方向:<从 north-star-align 来的一段>
Skill load compiler,传上面的 $ARGUMENTS。compiler 跑 4 阶段(语义解构 → 认知映射 → src 层 → 探索图抽取)+ 覆盖度循环 + Seal Level 1。
若仍不过,把失败报告透传给用户,停。
<workspace>/ara/(logic/ src/ trace/ evidence/ PAPER.md),Level 1 已过。 交给 ara-rigor-review。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | fail→pass | 11,604 | 7,755 | -33% | 1 | 1 | 0% | 1,756 | 1,638 | -7% | 0 | 0 | — |
case-01 | fail→fail | 37,522 | 7,868 | -79% | 1 | 1 | 0% | 6,199 | 1,225 | -80% | 0 | 0 | — |
case-02 | fail→pass | 27,908 | 6,867 | -75% | 1 | 1 | 0% | 5,165 | 1,650 | -68% | 0 | 0 | — |
case-03 | fail→pass | 6,724 | 12,972 | +93% | 1 | 1 | 0% | 933 | 2,512 | +169% | 0 | 0 | — |
case-04 | pass→fail | 18,422 | 5,487 | -70% | 1 | 1 | 0% | 2,292 | 722 | -68% | 0 | 0 | — |
case-05 | pass→pass | 4,560 | 10,398 | +128% | 1 | 1 | 0% | 687 | 2,270 | +230% | 0 | 0 | — |
case-06 | pass→pass | 4,972 | 4,586 | -8% | 1 | 1 | 0% | 821 | 1,292 | +57% | 0 | 0 | — |
case-07 | pass→pass | 10,501 | 4,630 | -56% | 1 | 1 | 0% | 1,581 | 1,231 | -22% | 0 | 0 | — |
case-08 | fail→pass | 12,335 | 6,385 | -48% | 1 | 1 | 0% | 1,888 | 1,560 | -17% | 0 | 0 | — |
case-09 | fail→pass | 19,856 | 2,758 | -86% | 1 | 1 | 0% | 1,358 | 975 | -28% | 0 | 0 | — |
case-10 | fail→pass | 12,348 | 6,176 | -50% | 1 | 1 | 0% | 1,599 | 1,438 | -10% | 0 | 0 | — |
case-11 | pass→pass | 8,170 | 2,999 | -63% | 1 | 1 | 0% | 1,365 | 1,064 | -22% | 0 | 0 | — |
case-12 | fail→pass | 38,271 | 2,912 | -92% | 1 | 1 | 0% | 1,900 | 1,022 | -46% | 0 | 0 | — |
case-13 | fail→pass | 12,850 | 3,290 | -74% | 1 | 1 | 0% | 1,765 | 1,071 | -39% | 0 | 0 | — |
case-14 | pass→pass | 12,337 | 4,845 | -61% | 1 | 1 | 0% | 1,899 | 1,282 | -32% | 0 | 0 | — |
case-15 | fail→pass | 12,430 | 4,310 | -65% | 1 | 1 | 0% | 1,917 | 1,219 | -36% | 0 | 0 | — |
case-16 | fail→fail | 5,155 | 3,485 | -32% | 1 | 1 | 0% | 755 | 1,068 | +41% | 0 | 0 | — |
case-17 | pass→pass | 12,376 | 4,640 | -63% | 1 | 1 | 0% | 2,018 | 1,256 | -38% | 0 | 0 | — |
case-18 | fail→pass | 8,292 | 3,692 | -55% | 1 | 1 | 0% | 1,119 | 1,071 | -4% | 0 | 0 | — |
case-19 | fail→pass | 17,723 | 7,458 | -58% | 1 | 1 | 0% | 2,483 | 1,670 | -33% | 0 | 0 | — |
case-20 | fail→pass | 11,737 | 1,905 | -84% | 1 | 1 | 0% | 1,583 | 868 | -45% | 0 | 0 | — |
case-22 | fail→pass | 11,341 | 2,925 | -74% | 1 | 1 | 0% | 1,848 | 1,034 | -44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.