Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate a complete, executable Research Spec from North Star + user input. Strategy-level skill that orchestrates questioning, outline, and spec writing.
.claude/skills/yogsoth-ai-writing-specs/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 4% | 0% |
You are generating a Research Spec — a document that is simultaneously human-readable and machine-executable. Another CC session will later read this spec and execute it step by step.
Before invoking this skill, the following MUST exist:
Read skills/research-catalog/SKILL.md in its entirety. Internalize:
Invoke these 3 SOPs sequentially. Each asks 2-3 focused questions:
scope-clarification — research boundaries, depth vs breadthcampaign-selection — which campaigns to include/emphasize/skipconstraint-elicitation — time budget, existing knowledge, hard constraintsWhen the outline includes an experiment-execution stage, default to SUGGESTING ara-from-context as a closing stage (compile the results into an ARA); the user may decline. Without experiment-execution, still offer it if the user wants the research packaged. ara-from-context is the 10th optional package — never a forced tail.
Synthesize the North Star, ResearchBrief, and user answers into a 5-10 line outline:
Stage 1: [campaign] ([strategies]) — [topic/focus]
Stage 2: [campaign] ([strategies]) — [topic/focus]
...
Stage N: experiment-execution (experiment-design) — [topic]
Stage N+1: ara-from-context — compile the research into an ARA (OPTIONAL closing stage; include only if the user wants the results packaged for agent reproduction; user may remove)Present this outline to the user. Wait for confirmation. User may adjust stages, reorder, add, or remove.
Expand the confirmed outline into a complete Research Spec. Follow this schema exactly:
# Research Spec: <Topic>
> Generated: YYYY-MM-DD
> North Star: <one sentence>
> Scope: <N> stages, estimated <M> sessions
> Source: de-anthropocentric-research-engineFor EACH stage, write ALL of these fields:
When the outline includes an ara-from-context stage, write its Execution Steps so the stage opens with an explicit user-decision gate, NOT an automatic transition:
"实验评估通过,是否进入 ARA 成文?(compile results into an ARA)" — proceed only on approval.
the spec's experiment-execution stage and the research-catalog to choose it at run time.
new spec field, and do not reuse the Backtrack Condition mechanism (that is for retreat, this is a forward go/no-go decision).
ara-from-context (context-review → compile-and-review), producingara/ + ara/level2_report.json. Note the external prerequisite: npx @ara-commons/ara-skills must be installed (compiler + rigor-reviewer).
Invoke spec-self-review SOP. This is MANDATORY and cannot be skipped.
Present the completed spec to the user for review. Wait for approval or change requests. If changes requested, revise and re-run self-review.
Save the spec to: docs/de-anthropocentric/specs/YYYY-MM-DD-<topic>-spec.md
Inform the user: "Spec complete. To execute, invoke executing-specs with the path to this spec file."
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | campaign-selection | Structured questioning SOP to determine which campaigns to include, emphasize, or skip. Used during spec generation. | | constraint-elicitation | Structured questioning SOP to identify practical constraints that shape the research spec. Used during spec generation. | | research-catalog | Capability menu for the research engine. Lists the 10 freely-composable research packages, what each does, when to reach for it, and a pointer to its full skill table. Read this after north-star crystallization to decide which packages to use — no fixed order. Also serves as the skill-index (capability map). | | scope-clarification | Structured questioning SOP to determine research boundaries, depth, and breadth. Used during spec generation. | | spec-self-review | Quality gate for Research Specs. Checks for placeholders, consistency, scope, ambiguity, context protocol, and quantification. Mandatory before user review. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 22,528 | 15,487 | -31% | 1 | 1 | 0% | 1,188 | 3,257 | +174% | 0 | 0 | — |
case-02 | fail→fail | 19,267 | 16,273 | -16% | 1 | 1 | 0% | 2,566 | 1,711 | -33% | 0 | 0 | — |
case-03 | fail→fail | 35,277 | 15,828 | -55% | 1 | 1 | 0% | 5,305 | 1,923 | -64% | 0 | 0 | — |
case-04 | fail→fail | 15,613 | 20,037 | +28% | 1 | 1 | 0% | 342 | 2,278 | +566% | 0 | 0 | — |
case-05 | fail→pass | 27,971 | 32,869 | +18% | 1 | 1 | 0% | 3,299 | 4,411 | +34% | 0 | 0 | — |
case-06 | fail→fail | 8,456 | 33,332 | +294% | 1 | 1 | 0% | 1,380 | 2,034 | +47% | 0 | 0 | — |
case-07 | fail→fail | 47,749 | 34,192 | -28% | 1 | 1 | 0% | 8,235 | 6,003 | -27% | 0 | 0 | — |
case-08 | fail→fail | 51,088 | 10,696 | -79% | 1 | 1 | 0% | 8,246 | 1,748 | -79% | 0 | 0 | — |
case-09 | fail→fail | 12,481 | 13,288 | +6% | 1 | 1 | 0% | 1,221 | 3,591 | +194% | 0 | 0 | — |
case-10 | fail→pass | 14,921 | 8,726 | -42% | 1 | 1 | 0% | 2,273 | 2,103 | -7% | 0 | 0 | — |
case-11 | fail→pass | 14,750 | 3,019 | -80% | 1 | 1 | 0% | 1,646 | 1,929 | +17% | 0 | 0 | — |
case-12 | fail→pass | 21,166 | 13,892 | -34% | 1 | 1 | 0% | 3,107 | 3,775 | +21% | 0 | 0 | — |
case-13 | pass→pass | 12,917 | 8,552 | -34% | 1 | 1 | 0% | 2,193 | 2,815 | +28% | 0 | 0 | — |
case-14 | pass→pass | 12,497 | 4,021 | -68% | 1 | 1 | 0% | 1,827 | 2,065 | +13% | 0 | 0 | — |
case-15 | fail→pass | 17,449 | 7,279 | -58% | 1 | 1 | 0% | 2,613 | 2,706 | +4% | 0 | 0 | — |
case-16 | pass→pass | 12,530 | 4,211 | -66% | 1 | 1 | 0% | 2,059 | 2,114 | +3% | 0 | 0 | — |
case-17 | fail→pass | 8,725 | 1,580 | -82% | 1 | 1 | 0% | 1,531 | 1,686 | +10% | 0 | 0 | — |
case-18 | pass→pass | 11,870 | 2,737 | -77% | 1 | 1 | 0% | 1,812 | 1,912 | +6% | 0 | 0 | — |
case-23 | fail→pass | 15,501 | 4,437 | -71% | 1 | 1 | 0% | 2,404 | 2,037 | -15% | 0 | 0 | — |
case-19 | fail→pass | 5,522 | 3,675 | -33% | 1 | 1 | 0% | 980 | 2,098 | +114% | 0 | 0 | — |
case-20 | fail→pass | 5,812 | 2,390 | -59% | 1 | 1 | 0% | 958 | 1,812 | +89% | 0 | 0 | — |
case-21 | fail→pass | 21,591 | 2,060 | -90% | 1 | 1 | 0% | 3,520 | 1,692 | -52% | 0 | 0 | — |
case-22 | fail→pass | 35,281 | 11,071 | -69% | 1 | 1 | 0% | 1,044 | 3,105 | +197% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.