Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Tactic: Review a context/ directory — sort material into ARA types, locate and align the north-star, and produce a feeding plan for the compiler
.claude/skills/yogsoth-ai-context-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -42% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -47% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 30% | 0% |
Key question: 这批 context 里,生成 ARA 最重要的素材是什么?大方向和用户对齐了吗?
Skill load context-exploring —— 读 INDEX → 整目录分三类素材(报告线/过程线/图片)→ 聚 arc 候选 → 定位 north-star 文件 → 草拟投喂计划。
Skill load north-star-align —— 深读 north-star → 提炼大方向 → 复用present-candidates/present-and-ask 与用户对齐 → 回填投喂计划的大方向 + arc 范围。
① 对齐过的大方向 ② 完整投喂计划(主干文件 / trace 素材 / 图片清单 / arc 范围 / 大方向),交给 compile-and-review。
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-exploring | SOP: Read context/INDEX.md and sort the whole directory into three ARA material types (report line, process line, images), locate the north-star file, and draft a feeding plan for the ARA compiler | | north-star-align | SOP: Deep-read the original north-star context, distill this ARA's overall direction, and align it with the user via the reused present-and-ask / present-candidates dialogue SOPs |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 19,304 | 9,794 | -49% | 1 | 1 | 0% | 2,398 | 1,916 | -20% | 0 | 0 | — |
case-04 | fail→pass | 11,576 | 4,751 | -59% | 1 | 1 | 0% | 1,605 | 1,085 | -32% | 0 | 0 | — |
case-01 | fail→fail | 5,676 | 2,321 | -59% | 1 | 1 | 0% | 838 | 682 | -19% | 0 | 0 | — |
case-02 | pass→fail | 5,743 | 5,819 | +1% | 1 | 1 | 0% | 1,046 | 1,176 | +12% | 0 | 0 | — |
case-05 | fail→pass | 15,282 | 6,773 | -56% | 1 | 1 | 0% | 2,240 | 1,298 | -42% | 0 | 0 | — |
case-06 | fail→pass | 9,640 | 2,832 | -71% | 1 | 1 | 0% | 1,416 | 753 | -47% | 0 | 0 | — |
case-07 | fail→pass | 4,273 | 3,761 | -12% | 1 | 1 | 0% | 724 | 944 | +30% | 0 | 0 | — |
case-08 | pass→pass | 14,656 | 3,749 | -74% | 1 | 1 | 0% | 2,143 | 924 | -57% | 0 | 0 | — |
case-09 | fail→pass | 9,773 | 3,748 | -62% | 1 | 1 | 0% | 1,452 | 979 | -33% | 0 | 0 | — |
case-10 | fail→pass | 9,932 | 3,913 | -61% | 1 | 1 | 0% | 1,460 | 930 | -36% | 0 | 0 | — |
case-11 | fail→pass | 11,758 | 2,793 | -76% | 1 | 1 | 0% | 1,698 | 790 | -53% | 0 | 0 | — |
case-12 | fail→pass | 12,048 | 4,394 | -64% | 1 | 1 | 0% | 1,935 | 998 | -48% | 0 | 0 | — |
case-13 | fail→pass | 9,726 | 5,922 | -39% | 1 | 1 | 0% | 1,542 | 1,261 | -18% | 0 | 0 | — |
case-14 | pass→pass | 13,367 | 8,026 | -40% | 1 | 1 | 0% | 1,844 | 1,621 | -12% | 0 | 0 | — |
case-15 | fail→pass | 9,984 | 4,571 | -54% | 1 | 1 | 0% | 1,447 | 1,145 | -21% | 0 | 0 | — |
case-16 | fail→pass | 12,466 | 3,260 | -74% | 1 | 1 | 0% | 1,908 | 846 | -56% | 0 | 0 | — |
case-17 | pass→pass | 11,421 | 2,900 | -75% | 1 | 1 | 0% | 1,767 | 858 | -51% | 0 | 0 | — |
case-18 | fail→pass | 12,648 | 4,118 | -67% | 1 | 1 | 0% | 1,862 | 995 | -47% | 0 | 0 | — |
case-19 | pass→pass | 12,807 | 6,979 | -46% | 1 | 1 | 0% | 1,857 | 1,550 | -17% | 0 | 0 | — |
case-20 | pass→pass | 36,044 | 16,137 | -55% | 1 | 1 | 0% | 6,140 | 2,795 | -54% | 0 | 0 | — |
case-21 | pass→pass | 5,596 | 16,669 | +198% | 1 | 1 | 0% | 943 | 3,181 | +237% | 0 | 0 | — |
case-22 | fail→fail | 2,214 | 4,102 | +85% | 1 | 1 | 0% | 318 | 892 | +181% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.