Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Campaign: Compile a context/ research record into an ARA (Agent-Native Research Artifact) and run a Level-2 epistemic review — no LaTeX, no narrative paper
.claude/skills/yogsoth-ai-ara-from-context/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -73% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 954% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -30% | 0% |
What this is: DARE 流水线最末端的"成文"环节。吃前面研究循环 (research ↔ experiment-execution 反复迭代)沉淀在 context/ 里的全部产物, 编译成一份 ARA(机器可执行的四层知识包),并做认识论审查。不写 LaTeX / 叙事论文 —— ARA 刻意反对 storytelling,要的是逻辑弧在结构上闭合。
Source of truth: 所有素材来自 context/。核心 = 末次 EE 的最终 report + 全程迭代轨迹 + 研究产出的图片。
Skill load context-review —— 回顾 context/,分三类素材,对齐大方向,产出投喂计划。
Skill load compile-and-review —— 一次 inline 跑外部 compiler 得 ../ara/,再跑 rigor-reviewer 得 level2_report.json。
运行需 ARA 的 compiler + rigor-reviewer skill 在位 (npx @ara-commons/ara-skills)。见本 repo README。
ara/(logic/ src/ trace/ evidence/ PAPER.md)+ ara/level2_report.json。
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | compile-and-review | Tactic: Compile the feeding plan into an ARA via the external compiler, then run Level-2 rigor review over it | | context-review | Tactic: Review a context/ directory — sort material into ARA types, locate and align the north-star, and produce a feeding plan for the compiler |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | ara-compile | SOP: Turn the feeding plan into the compiler's $ARGUMENTS and run the external ARA compiler once inline to produce ../ara/ | | ara-rigor-review | SOP: Run the external ARA rigor-reviewer (Seal Level 2, six-dimension semantic review) over ../ara/ and pass its level2_report.json to the user | | context-exploring | SOP: Read context/INDEX.md and sort the whole directory into three ARA material types (report line, process line, images), locate the north-star file, and draft a feeding plan for the ARA compiler | | north-star-align | SOP: Deep-read the original north-star context, distill this ARA's overall direction, and align it with the user via the reused present-and-ask / present-candidates dialogue SOPs |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,003 | 4,130 | -17% | 1 | 1 | 0% | 266 | 911 | +242% | 0 | 0 | — |
case-02 | fail→fail | 39,187 | 3,134 | -92% | 1 | 1 | 0% | 6,399 | 996 | -84% | 0 | 0 | — |
case-03 | fail→fail | 4,900 | 4,282 | -13% | 1 | 1 | 0% | 239 | 924 | +287% | 0 | 0 | — |
case-04 | fail→pass | 39,728 | 12,257 | -69% | 1 | 1 | 0% | 6,180 | 1,692 | -73% | 0 | 0 | — |
case-05 | fail→pass | 4,544 | 8,680 | +91% | 1 | 1 | 0% | 190 | 2,002 | +954% | 0 | 0 | — |
case-06 | fail→pass | 5,087 | 5,390 | +6% | 1 | 1 | 0% | 873 | 1,351 | +55% | 0 | 0 | — |
case-07 | fail→pass | 15,669 | 7,056 | -55% | 1 | 1 | 0% | 2,495 | 1,786 | -28% | 0 | 0 | — |
case-08 | pass→pass | 6,254 | 1,681 | -73% | 1 | 1 | 0% | 939 | 898 | -4% | 0 | 0 | — |
case-09 | pass→pass | 13,087 | 7,147 | -45% | 1 | 1 | 0% | 2,055 | 1,734 | -16% | 0 | 0 | — |
case-10 | fail→fail | 7,319 | 2,892 | -60% | 1 | 1 | 0% | 965 | 1,144 | +19% | 0 | 0 | — |
case-11 | fail→fail | 9,020 | 3,674 | -59% | 1 | 1 | 0% | 1,338 | 1,235 | -8% | 0 | 0 | — |
case-12 | pass→pass | 15,922 | 4,617 | -71% | 1 | 1 | 0% | 2,220 | 1,375 | -38% | 0 | 0 | — |
case-13 | fail→pass | 14,043 | 5,699 | -59% | 1 | 1 | 0% | 2,229 | 1,557 | -30% | 0 | 0 | — |
case-14 | fail→pass | 11,914 | 3,105 | -74% | 1 | 1 | 0% | 1,825 | 1,133 | -38% | 0 | 0 | — |
case-15 | fail→pass | 10,065 | 1,909 | -81% | 1 | 1 | 0% | 1,545 | 932 | -40% | 0 | 0 | — |
case-16 | fail→pass | 27,413 | 2,181 | -92% | 1 | 1 | 0% | 1,244 | 958 | -23% | 0 | 0 | — |
case-17 | fail→pass | 9,968 | 3,030 | -70% | 1 | 1 | 0% | 1,403 | 1,121 | -20% | 0 | 0 | — |
case-18 | fail→pass | 9,138 | 1,742 | -81% | 1 | 1 | 0% | 1,373 | 893 | -35% | 0 | 0 | — |
case-19 | fail→fail | 7,646 | 7,268 | -5% | 1 | 1 | 0% | 1,083 | 1,662 | +53% | 0 | 0 | — |
case-20 | pass→pass | 12,300 | 5,030 | -59% | 1 | 1 | 0% | 1,719 | 1,414 | -18% | 0 | 0 | — |
case-21 | fail→pass | 9,225 | 2,348 | -75% | 1 | 1 | 0% | 1,278 | 960 | -25% | 0 | 0 | — |
case-22 | fail→pass | 7,017 | 2,837 | -60% | 1 | 1 | 0% | 1,072 | 1,165 | +9% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.