Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build or validate the source manifest for the source-tutorial pipeline from user-provided URLs/files. **Trigger**: source manifest, sources list, tutorial sources, url list, 资料清单, 教程来源. **Use when**: `source-tutorial` 的 C1,需要把多源输入落成统一的 `sources/manifest.yml`,并在内容不完整时显式阻塞。 **Skip if**: 已经有完整且经过确认的 `sources/manifest.yml`。 **Network**: none. **Guardrail**: 不要伪造来源;manifest 不完整时应返回 BLOCKED,而不是假装完成。
.claude/skills/willoscar-source-manifest/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -55% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -49% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -38% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -49% | 0% |
Goal: collect the source candidates for source-tutorial into one explicit manifest before ingest starts.
GOAL.mdDECISIONS.mdsources/manifest.ymlGOAL.md and any existing DECISIONS.md notes to understand what kinds of sources the tutorial is expected to use.sources/manifest.yml does not exist, scaffold it with a minimal example.source_idkindlocatorlabelkind: video, prefer transcript_locator unless the platform exposes public subtitleskindwebpagepdfmarkdownrepodocs_sitevideoFor video:
transcript_locatortranscript_locatoruv run python .codex/skills/source-manifest/scripts/run.py --workspace <workspace>--workspace <dir> (required)--unit-id <U###>--inputs <semicolon-separated>--outputs <semicolon-separated>--checkpoint <C#>uv run python .codex/skills/source-manifest/scripts/run.py --workspace <workspace>Cause:
Fix:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,803 | 4,780 | -0% | 1 | 1 | 0% | 243 | 691 | +184% | 0 | 0 | — |
case-02 | fail→fail | 9,297 | 4,436 | -52% | 1 | 1 | 0% | 509 | 639 | +26% | 0 | 0 | — |
case-03 | fail→fail | 4,569 | 4,811 | +5% | 1 | 1 | 0% | 204 | 607 | +198% | 0 | 0 | — |
case-04 | fail→pass | 11,457 | 5,772 | -50% | 1 | 1 | 0% | 1,962 | 1,420 | -28% | 0 | 0 | — |
case-05 | pass→pass | 8,445 | 2,131 | -75% | 1 | 1 | 0% | 1,480 | 789 | -47% | 0 | 0 | — |
case-06 | fail→pass | 11,871 | 2,180 | -82% | 1 | 1 | 0% | 1,758 | 785 | -55% | 0 | 0 | — |
case-07 | fail→pass | 10,035 | 2,496 | -75% | 1 | 1 | 0% | 1,559 | 788 | -49% | 0 | 0 | — |
case-08 | fail→pass | 10,785 | 3,304 | -69% | 1 | 1 | 0% | 1,572 | 981 | -38% | 0 | 0 | — |
case-09 | pass→pass | 12,430 | 3,113 | -75% | 1 | 1 | 0% | 1,768 | 857 | -52% | 0 | 0 | — |
case-10 | pass→pass | 15,079 | 3,609 | -76% | 1 | 1 | 0% | 2,306 | 1,000 | -57% | 0 | 0 | — |
case-11 | fail→pass | 11,720 | 2,801 | -76% | 1 | 1 | 0% | 1,709 | 870 | -49% | 0 | 0 | — |
case-12 | fail→pass | 8,315 | 2,139 | -74% | 1 | 1 | 0% | 1,296 | 772 | -40% | 0 | 0 | — |
case-13 | fail→pass | 7,148 | 2,859 | -60% | 1 | 1 | 0% | 1,216 | 903 | -26% | 0 | 0 | — |
case-14 | pass→pass | 12,666 | 2,065 | -84% | 1 | 1 | 0% | 2,026 | 774 | -62% | 0 | 0 | — |
case-15 | fail→pass | 11,862 | 3,405 | -71% | 1 | 1 | 0% | 1,631 | 1,046 | -36% | 0 | 0 | — |
case-16 | fail→pass | 8,072 | 2,546 | -68% | 1 | 1 | 0% | 1,341 | 776 | -42% | 0 | 0 | — |
case-17 | fail→pass | 10,584 | 1,649 | -84% | 1 | 1 | 0% | 1,539 | 635 | -59% | 0 | 0 | — |
case-18 | pass→pass | 21,791 | 2,646 | -88% | 1 | 1 | 0% | 3,359 | 869 | -74% | 0 | 0 | — |
case-19 | fail→pass | 5,252 | 1,618 | -69% | 1 | 1 | 0% | 738 | 655 | -11% | 0 | 0 | — |
case-20 | pass→pass | 2,592 | 2,925 | +13% | 1 | 1 | 0% | 388 | 866 | +123% | 0 | 0 | — |
case-21 | pass→pass | 3,383 | 3,686 | +9% | 1 | 1 | 0% | 525 | 969 | +85% | 0 | 0 | — |
case-22 | pass→pass | 2,545 | 2,491 | -2% | 1 | 1 | 0% | 363 | 813 | +124% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.