Install any skill in seconds. Free to start, no credit card required.
Get Started Free →SOP: Read context/INDEX.md and sort the whole directory into three ARA material types (report line, process line, images), locate the north-star file, and draft a feeding plan for the ARA compiler
.claude/skills/yogsoth-ai-context-exploring/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -9% | 0% |
Key question: 这批 context/ 里,对生成一份 ARA 最重要的素材分别是哪些,怎么分类喂给 compiler?
ARA 的 compiler 是逆向抽取器:它只抽源里已有的探索过程,绝不补写 (compiler Rule 9/14)。所以 ARA 的 exploration_tree 丰不丰满,不取决于 compiler, 取决于 context/ 有没有记下失败路径 / 被否方案 / 迭代 pivot。这把担子精确落在本 SOP:必须主动打捞过程线,而不是只挑末次成功 report。
context/INDEX.md。 它是总账:列 File / Phase / Topic / Checkpoints /Last Updated。正常 context 必有 INDEX;若缺失,停下并提示用户先维护 INDEX (不做无-INDEX 降级)。
(一条研究弧 = 一组连续推进同一 Topic 的文件)。
结论性产出、被确认的设计。这是 logic/claims.md + logic/problem.md 的源。
"我们本来想 X 但发现 Y 所以转向 Z"。这是 trace/exploration_tree.yaml 的源。 主动打捞:逐个 checkpoint 扫,凡是 dead_end / decision / pivot 痕迹都收。
evidence/figures +evidence/tables 的源。(正常 context 会有;没有就这类为空,不报错。)
*-north-star-*.md)。交给下游 north-star-align 深读。
markdown ## 投喂计划 ### 主干文件
### trace 素材
### 图片清单
### arc 范围
### 大方向
把投喂计划 + north-star 文件路径交给本 tactic 的下一步(north-star-align), 不落新文件。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 12,028 | 27,788 | +131% | 1 | 1 | 0% | 2,174 | 5,502 | +153% | 0 | 0 | — |
case-02 | fail→fail | 20,537 | 4,823 | -77% | 1 | 1 | 0% | 3,153 | 1,046 | -67% | 0 | 0 | — |
case-03 | fail→fail | 30,105 | 3,145 | -90% | 1 | 1 | 0% | 5,001 | 1,224 | -76% | 0 | 0 | — |
case-04 | fail→pass | 7,370 | 3,602 | -51% | 1 | 1 | 0% | 1,119 | 1,263 | +13% | 0 | 0 | — |
case-05 | fail→pass | 9,816 | 3,927 | -60% | 1 | 1 | 0% | 1,469 | 1,403 | -4% | 0 | 0 | — |
case-06 | fail→pass | 10,921 | 6,522 | -40% | 1 | 1 | 0% | 1,613 | 1,748 | +8% | 0 | 0 | — |
case-07 | fail→pass | 11,224 | 4,943 | -56% | 1 | 1 | 0% | 1,750 | 1,435 | -18% | 0 | 0 | — |
case-08 | pass→pass | 9,227 | 5,351 | -42% | 1 | 1 | 0% | 1,267 | 1,555 | +23% | 0 | 0 | — |
case-09 | pass→fail | 7,581 | 3,083 | -59% | 1 | 1 | 0% | 1,357 | 1,176 | -13% | 0 | 0 | — |
case-10 | fail→pass | 9,181 | 3,445 | -62% | 1 | 1 | 0% | 1,418 | 1,284 | -9% | 0 | 0 | — |
case-11 | fail→fail | 8,377 | 2,611 | -69% | 1 | 1 | 0% | 1,224 | 1,099 | -10% | 0 | 0 | — |
case-12 | fail→pass | 7,084 | 6,744 | -5% | 1 | 1 | 0% | 1,046 | 1,795 | +72% | 0 | 0 | — |
case-13 | pass→pass | 8,130 | 3,413 | -58% | 1 | 1 | 0% | 1,475 | 1,230 | -17% | 0 | 0 | — |
case-14 | fail→pass | 14,543 | 4,080 | -72% | 1 | 1 | 0% | 2,027 | 1,309 | -35% | 0 | 0 | — |
case-15 | fail→pass | 15,504 | 10,940 | -29% | 1 | 1 | 0% | 2,953 | 2,346 | -21% | 0 | 0 | — |
case-16 | fail→pass | 12,556 | 2,162 | -83% | 1 | 1 | 0% | 2,030 | 1,021 | -50% | 0 | 0 | — |
case-17 | pass→pass | 14,888 | 9,367 | -37% | 1 | 1 | 0% | 1,976 | 2,083 | +5% | 0 | 0 | — |
case-18 | fail→pass | 3,605 | 2,540 | -30% | 1 | 1 | 0% | 548 | 1,042 | +90% | 0 | 0 | — |
case-19 | fail→pass | 11,625 | 4,710 | -59% | 1 | 1 | 0% | 1,658 | 1,448 | -13% | 0 | 0 | — |
case-20 | fail→fail | 16,534 | 13,443 | -19% | 1 | 1 | 0% | 2,911 | 2,884 | -1% | 0 | 0 | — |
case-21 | fail→fail | 17,968 | 20,301 | +13% | 1 | 1 | 0% | 2,681 | 3,757 | +40% | 0 | 0 | — |
case-22 | fail→fail | 3,518 | 4,525 | +29% | 1 | 1 | 0% | 327 | 944 | +189% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.