Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Route an unbound research Goal to one Workflow or materialize its initial C0 kickoff block; use during initial binding, not for later checkpoint summaries or approval.
.claude/skills/willoscar-pipeline-router/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 3464% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 115% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -18% | 0% |
Routing is a commitment: select one Workflow from the desired Artifact and evidence method, record that choice, and expose the next human Decision.
GOAL.md, or the current user request when the Workspace is new.PIPELINE.lock.md, UNITS.csv, STATUS.md, and DECISIONS.md whenpresent.
docs/PIPELINE_TAXONOMY.md and candidate Pipeline front matter duringselection.
PIPELINE.lock.md for a newly bound Goal.DECISIONS.md.queries.md when the selected retrieval Workflow starts at C0.STATUS.md.Read the Goal and identify:
When a routing discriminator is missing, write the smallest grouped question set to DECISIONS.md and stop at that Decision.
Completion criterion: every fact that can change the Workflow choice is known or explicitly bounded in DECISIONS.md.
Load docs/PIPELINE_TAXONOMY.md, then inspect only the candidate Pipeline contracts. Choose from target Artifact and evidence method rather than keyword matching alone. Treat delivery profiles and course-report use cases as overlays inside their existing Workflow family.
Completion criterion: exactly one executable Pipeline is selected and its target Artifacts match the Goal.
Write PIPELINE.lock.md with:
textpipeline: <pipeline path> units_template: <template path from Pipeline front matter> locked_at: <YYYY-MM-DD>
Initialize UNITS.csv from that template when the Workspace is new. Preserve a valid existing lock; a route change is an explicit operator Decision followed by reinitialization or migration, never a silent rewrite.
Completion criterion: the lock, Unit template, and Workspace projection refer to the same Pipeline.
Materialize the C0 kickoff block and approval checkbox in DECISIONS.md. Seed queries.md from the Goal when retrieval is part of the selected Workflow. Later checkpoints use checkpoint-brief. Historical Workspaces whose saved Unit still invokes pipeline-router after C0 are delegated to checkpoint-brief with a deprecation warning; the router never approves them.
The deterministic helper may be used after selection:
bashuv run python .codex/skills/pipeline-router/scripts/run.py \ --workspace workspaces/<name> \ --checkpoint C0
Completion criterion: DECISIONS.md contains the C0 checkpoint and one clear approval or answer surface; retrieval Workflows have a non-empty C0 query seed.
Check that the locked Pipeline exists, its Unit template exists, and the active checkpoint is represented in both STATUS.md and DECISIONS.md.
Completion criterion: the Pipeline Runner can continue from Workspace files without making another routing inference.
docs/PIPELINE_TAXONOMY.md only while selecting or deliberately changinga Workflow.
pipelines/*.pipeline.md is the executioncontract and the taxonomy leaves context.
assets/pipeline-selection-form.md only when missing routing facts requirea human answer.
--help when C0 materialization needs debugging; thehelper records the initial route projection but does not choose one.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,235 | 4,058 | -22% | 1 | 1 | 0% | 346 | 1,032 | +198% | 0 | 0 | — |
case-02 | fail→pass | 3,756 | 34,583 | +821% | 1 | 1 | 0% | 198 | 7,056 | +3464% | 0 | 0 | — |
case-03 | fail→fail | 3,601 | 3,747 | +4% | 1 | 1 | 0% | 228 | 1,017 | +346% | 0 | 0 | — |
case-04 | fail→pass | 7,696 | 15,340 | +99% | 1 | 1 | 0% | 1,274 | 2,742 | +115% | 0 | 0 | — |
case-05 | fail→fail | 7,957 | 4,114 | -48% | 1 | 1 | 0% | 1,244 | 1,536 | +23% | 0 | 0 | — |
case-10 | fail→pass | 5,781 | 2,133 | -63% | 1 | 1 | 0% | 785 | 1,190 | +52% | 0 | 0 | — |
case-06 | fail→pass | 12,682 | 4,728 | -63% | 1 | 1 | 0% | 1,988 | 1,656 | -17% | 0 | 0 | — |
case-07 | fail→pass | 12,645 | 3,760 | -70% | 1 | 1 | 0% | 1,780 | 1,462 | -18% | 0 | 0 | — |
case-08 | pass→pass | 8,098 | 6,581 | -19% | 1 | 1 | 0% | 1,094 | 1,925 | +76% | 0 | 0 | — |
case-09 | fail→pass | 9,495 | 2,175 | -77% | 1 | 1 | 0% | 1,614 | 1,172 | -27% | 0 | 0 | — |
case-11 | fail→pass | 14,728 | 3,601 | -76% | 1 | 1 | 0% | 2,125 | 1,401 | -34% | 0 | 0 | — |
case-12 | fail→pass | 13,022 | 3,756 | -71% | 1 | 1 | 0% | 1,868 | 1,405 | -25% | 0 | 0 | — |
case-13 | pass→pass | 12,524 | 2,605 | -79% | 1 | 1 | 0% | 1,846 | 1,250 | -32% | 0 | 0 | — |
case-14 | fail→pass | 7,235 | 1,958 | -73% | 1 | 1 | 0% | 1,042 | 1,135 | +9% | 0 | 0 | — |
case-15 | fail→pass | 12,565 | 2,938 | -77% | 1 | 1 | 0% | 2,043 | 1,327 | -35% | 0 | 0 | — |
case-16 | fail→pass | 10,348 | 4,073 | -61% | 1 | 1 | 0% | 1,544 | 1,423 | -8% | 0 | 0 | — |
case-17 | fail→fail | 13,199 | 3,331 | -75% | 1 | 1 | 0% | 2,055 | 1,406 | -32% | 0 | 0 | — |
case-18 | pass→pass | 8,580 | 3,955 | -54% | 1 | 1 | 0% | 1,204 | 1,480 | +23% | 0 | 0 | — |
case-19 | fail→pass | 8,688 | 6,782 | -22% | 1 | 1 | 0% | 1,257 | 2,035 | +62% | 0 | 0 | — |
case-20 | fail→fail | 11,620 | 6,914 | -40% | 1 | 1 | 0% | 2,099 | 1,155 | -45% | 0 | 0 | — |
case-21 | fail→fail | 17,036 | 12,267 | -28% | 1 | 1 | 0% | 2,906 | 2,885 | -1% | 0 | 0 | — |
case-22 | fail→fail | 9,493 | 5,385 | -43% | 1 | 1 | 0% | 1,404 | 1,096 | -22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.