Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a tutorial-specific quality gate and write a PASS/FAIL report for the final tutorial deliverable. **Trigger**: tutorial self-loop, tutorial quality gate, tutorial pass/fail, 教程自循环, 教程质量门. **Use when**: `source-tutorial` 的 C3,已经有 `output/TUTORIAL.md`,想在交付前确认它满足教程合同而不是普通长文。 **Skip if**: 还没有 tutorial 正文。 **Network**: none. **Guardrail**: 报告缺口时不要发明内容;把失败清楚地路由回 tutorial 写作阶段。
.claude/skills/willoscar-tutorial-selfloop/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -51% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -56% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -69% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -62% | 0% |
Goal: converge the tutorial deliverable to a stable teaching-quality bar.
The gate should compare output/TUTORIAL.md against the intended module shape from outline/module_plan.yml, rather than acting like a generic prose checker.
output/TUTORIAL.mdoutline/module_plan.ymloutput/TUTORIAL_SELFLOOP_TODO.mduv run python .codex/skills/tutorial-selfloop/scripts/run.py --workspace <workspace>--workspace <dir> (required)--unit-id <U###>--inputs <semicolon-separated>--outputs <semicolon-separated>--checkpoint <C#>uv run python .codex/skills/tutorial-selfloop/scripts/run.py --workspace <workspace>| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 2,170 | 6,222 | +187% | 1 | 1 | 0% | 309 | 630 | +104% | 0 | 0 | — |
case-02 | fail→fail | 20,412 | 4,600 | -77% | 1 | 1 | 0% | 3,222 | 518 | -84% | 0 | 0 | — |
case-03 | fail→fail | 4,452 | 5,137 | +15% | 1 | 1 | 0% | 704 | 545 | -23% | 0 | 0 | — |
case-04 | fail→pass | 8,785 | 2,336 | -73% | 1 | 1 | 0% | 1,386 | 675 | -51% | 0 | 0 | — |
case-05 | pass→pass | 8,443 | 2,048 | -76% | 1 | 1 | 0% | 1,288 | 552 | -57% | 0 | 0 | — |
case-06 | fail→pass | 10,273 | 2,925 | -72% | 1 | 1 | 0% | 1,551 | 685 | -56% | 0 | 0 | — |
case-07 | fail→pass | 6,351 | 1,892 | -70% | 1 | 1 | 0% | 877 | 528 | -40% | 0 | 0 | — |
case-08 | fail→pass | 14,313 | 2,651 | -81% | 1 | 1 | 0% | 2,129 | 660 | -69% | 0 | 0 | — |
case-09 | fail→pass | 9,996 | 1,841 | -82% | 1 | 1 | 0% | 1,458 | 554 | -62% | 0 | 0 | — |
case-10 | fail→pass | 6,866 | 1,873 | -73% | 1 | 1 | 0% | 1,011 | 464 | -54% | 0 | 0 | — |
case-11 | pass→pass | 4,815 | 2,954 | -39% | 1 | 1 | 0% | 772 | 677 | -12% | 0 | 0 | — |
case-12 | fail→pass | 9,114 | 1,845 | -80% | 1 | 1 | 0% | 1,321 | 532 | -60% | 0 | 0 | — |
case-13 | pass→pass | 10,815 | 1,982 | -82% | 1 | 1 | 0% | 1,491 | 515 | -65% | 0 | 0 | — |
case-14 | pass→pass | 10,718 | 1,599 | -85% | 1 | 1 | 0% | 1,468 | 458 | -69% | 0 | 0 | — |
case-15 | fail→pass | 14,513 | 6,367 | -56% | 1 | 1 | 0% | 2,107 | 1,167 | -45% | 0 | 0 | — |
case-16 | pass→pass | 15,840 | 1,787 | -89% | 1 | 1 | 0% | 1,391 | 460 | -67% | 0 | 0 | — |
case-17 | fail→pass | 10,784 | 3,114 | -71% | 1 | 1 | 0% | 1,489 | 753 | -49% | 0 | 0 | — |
case-18 | fail→pass | 8,324 | 1,505 | -82% | 1 | 1 | 0% | 1,197 | 432 | -64% | 0 | 0 | — |
case-19 | fail→pass | 12,151 | 1,688 | -86% | 1 | 1 | 0% | 1,874 | 507 | -73% | 0 | 0 | — |
case-20 | pass→pass | 13,461 | 19,935 | +48% | 1 | 1 | 0% | 2,142 | 3,294 | +54% | 0 | 0 | — |
case-21 | pass→pass | 14,976 | 20,857 | +39% | 1 | 1 | 0% | 2,541 | 3,947 | +55% | 0 | 0 | — |
case-22 | pass→pass | 12,194 | 10,860 | -11% | 1 | 1 | 0% | 1,990 | 2,091 | +5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.