Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build a 2+ level taxonomy (`outline/taxonomy.yml`) from a core paper set and scope constraints, with short descriptions per node. **Trigger**: taxonomy, taxonomy builder, 分类, 主题树, taxonomy.yml. **Use when**: survey/snapshot 的结构阶段(NO PROSE),已有 `papers/core_set.csv`,需要生成可映射且读者友好的主题结构。 **Skip if**: 已经有批准过且可映射的 taxonomy(不要无意义重构)。 **Network**: none. **Guardrail**: 避免泛化占位桶;保持 2+ 层且每节点有具体描述。
.claude/skills/willoscar-taxonomy-builder/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -49% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -28% | 0% |
Build outline/taxonomy.yml from papers/core_set.csv.
P0 compatibility note:
outline/taxonomy.yml, YAML list, >=2 levels, concrete descriptions).assets/domain_packs/*.yaml instead of Python prose.scripts/run.py stays a deterministic scaffold/helper: detect domain pack -> load pack when available -> otherwise fall back to the generic builder.references/overview.mdreferences/taxonomy_principles.mdreferences/domain_pack_<domain>.md and assets/domain_packs/<domain>.yamlreferences/archetypes_generic.mdreferences/examples_good.md and references/examples_bad.mdCurrent compatibility packs:
llm_agentsgen_imageembodied_airag_evaluationCreate outline/taxonomy.refined.ok only after reviewing a manually refined taxonomy. The marker is honored only while it is newer than the taxonomy, its upstream evidence, and the generator; stale markers are removed and the prior taxonomy is backed up before regeneration.
papers/core_set.csv (required)papers/papers_dedup.jsonlDECISIONS.md, GOAL.md, queries.mdoutline/taxonomy.ymlassets/taxonomy_schema.json: machine-readable shape for domain packs / output expectationsassets/domain_packs/*.yaml: compatibility domain packs for supported domainsUse scripts/run.py only for deterministic help:
GOAL.md / queries.md explicitly match its detection contractRefine the generated taxonomy before marking the unit DONE if:
Overview, Benchmarks, Open Problems, Misc)uv run python .codex/skills/taxonomy-builder/scripts/run.py --helpuv run python .codex/skills/taxonomy-builder/scripts/run.py --workspace <workspace>When running in compatibility mode, scripts/run.py currently reads:
papers/core_set.csv as the required corpus inputpapers/papers_dedup.jsonl when present for generic corpus signalsGOAL.md and queries.md as the authoritative domain-pack selection intent; corpus term co-occurrence cannot override themuv run python .codex/skills/taxonomy-builder/scripts/run.py --workspace <workspace>--workspace <dir>--top-k <int>--min-freq <int>--unit-id <id>--inputs <a;b;...>--outputs <a;b;...>--checkpoint <C*>uv run python .codex/skills/taxonomy-builder/scripts/run.py --workspace <workspace>GOAL.md, queries.md, and the pack detect rules before changing Python.outline/taxonomy.yml already contains a real non-placeholder taxonomy, the script intentionally returns without overwriting it.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,784 | 5,018 | +5% | 1 | 1 | 0% | 279 | 1,236 | +343% | 0 | 0 | — |
case-22 | pass→pass | 18,702 | 19,806 | +6% | 1 | 1 | 0% | 2,834 | 3,971 | +40% | 0 | 0 | — |
case-02 | fail→fail | 28,094 | 4,380 | -84% | 1 | 1 | 0% | 5,307 | 1,113 | -79% | 0 | 0 | — |
case-03 | fail→fail | 15,435 | 4,583 | -70% | 1 | 1 | 0% | 2,376 | 1,083 | -54% | 0 | 0 | — |
case-04 | fail→pass | 11,757 | 2,582 | -78% | 1 | 1 | 0% | 1,701 | 1,498 | -12% | 0 | 0 | — |
case-05 | pass→pass | 6,767 | 2,658 | -61% | 1 | 1 | 0% | 942 | 1,370 | +45% | 0 | 0 | — |
case-06 | fail→pass | 20,612 | 1,854 | -91% | 1 | 1 | 0% | 1,395 | 1,190 | -15% | 0 | 0 | — |
case-07 | fail→pass | 11,150 | 4,563 | -59% | 1 | 1 | 0% | 1,902 | 1,781 | -6% | 0 | 0 | — |
case-08 | pass→pass | 9,495 | 2,427 | -74% | 1 | 1 | 0% | 1,413 | 1,355 | -4% | 0 | 0 | — |
case-09 | fail→pass | 12,251 | 1,641 | -87% | 1 | 1 | 0% | 2,336 | 1,188 | -49% | 0 | 0 | — |
case-10 | fail→pass | 9,552 | 1,617 | -83% | 1 | 1 | 0% | 1,535 | 1,102 | -28% | 0 | 0 | — |
case-11 | pass→pass | 9,429 | 4,749 | -50% | 1 | 1 | 0% | 1,510 | 1,597 | +6% | 0 | 0 | — |
case-12 | pass→pass | 3,956 | 2,873 | -27% | 1 | 1 | 0% | 519 | 1,428 | +175% | 0 | 0 | — |
case-13 | pass→pass | 9,575 | 3,842 | -60% | 1 | 1 | 0% | 1,579 | 1,595 | +1% | 0 | 0 | — |
case-14 | fail→pass | 7,357 | 2,097 | -71% | 1 | 1 | 0% | 1,327 | 1,273 | -4% | 0 | 0 | — |
case-15 | fail→pass | 9,573 | 1,644 | -83% | 1 | 1 | 0% | 1,856 | 1,194 | -36% | 0 | 0 | — |
case-16 | fail→pass | 10,419 | 1,552 | -85% | 1 | 1 | 0% | 1,728 | 1,143 | -34% | 0 | 0 | — |
case-17 | fail→pass | 7,510 | 2,389 | -68% | 1 | 1 | 0% | 1,239 | 1,273 | +3% | 0 | 0 | — |
case-18 | fail→pass | 11,775 | 3,251 | -72% | 1 | 1 | 0% | 1,940 | 1,464 | -25% | 0 | 0 | — |
case-19 | fail→pass | 12,845 | 1,928 | -85% | 1 | 1 | 0% | 2,127 | 1,235 | -42% | 0 | 0 | — |
case-20 | pass→pass | 16,524 | 12,183 | -26% | 1 | 1 | 0% | 3,251 | 3,275 | +1% | 0 | 0 | — |
case-21 | pass→pass | 16,582 | 19,752 | +19% | 1 | 1 | 0% | 3,239 | 4,074 | +26% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 18 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.