Install any skill in seconds. Free to start, no credit card required.
Get Started Free →维护中文毕业论文的 `codex_md/question_list.md`:把本轮问题、边界、优先级、协作方案和验收口径结构化,作为整条 thesis pipeline 的控制面。 **Trigger**: 毕业论文问题清单, thesis question list, 论文修改清单, 本轮目标, 结构问题梳理, review问题整理. **Use when**: 你已经有一批材料或上一轮 review 结果,需要明确这一轮到底修什么、不修什么,并给后续重构与编译复查提供统一入口。 **Skip if**: 当前只是在做一次性局部措辞修改,且没有形成新一轮结构/证据/编译问题。 **Network**: none. **Guardrail**: 不在这里写正文;不把问题单写成长篇散文;每条问题必须可执行、可验收。
.claude/skills/willoscar-thesis-question-list/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -54% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -38% | 0% |
这个 skill 负责维护毕业论文流程的控制面文档:
codex_md/question_list.md它不是备忘录,而是这一轮工作的单一入口。
codex_md/material_index.mdcodex_md/missing_info.mdcodex_md/*.mdclaude_md/review_checklist.mdcodex_md/question_list.mdAlways read:
references/overview.mdreferences/question_entry_schema.mdreferences/examples_good.mdreferences/examples_bad.mdMachine-readable contract:
assets/question_list_contract.jsonuv run python .codex/skills/thesis-question-list/scripts/run.py --workspace <workspace>| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,616 | 8,938 | +147% | 1 | 1 | 0% | 205 | 1,182 | +477% | 0 | 0 | — |
case-02 | fail→fail | 28,950 | 6,114 | -79% | 1 | 1 | 0% | 4,937 | 816 | -83% | 0 | 0 | — |
case-03 | fail→fail | 17,574 | 7,256 | -59% | 1 | 1 | 0% | 2,664 | 866 | -67% | 0 | 0 | — |
case-04 | fail→fail | 2,713 | 53,236 | +1862% | 1 | 1 | 0% | 355 | 597 | +68% | 0 | 0 | — |
case-05 | fail→fail | 4,738 | 5,531 | +17% | 1 | 1 | 0% | 653 | 643 | -2% | 0 | 0 | — |
case-06 | pass→pass | 16,014 | 9,797 | -39% | 1 | 1 | 0% | 2,889 | 2,040 | -29% | 0 | 0 | — |
case-07 | pass→pass | 8,973 | 4,726 | -47% | 1 | 1 | 0% | 1,285 | 1,228 | -4% | 0 | 0 | — |
case-08 | pass→pass | 5,881 | 7,476 | +27% | 1 | 1 | 0% | 786 | 1,522 | +94% | 0 | 0 | — |
case-09 | pass→pass | 6,706 | 3,012 | -55% | 1 | 1 | 0% | 971 | 871 | -10% | 0 | 0 | — |
case-10 | pass→pass | 6,233 | 4,068 | -35% | 1 | 1 | 0% | 1,010 | 1,057 | +5% | 0 | 0 | — |
case-11 | pass→pass | 8,964 | 3,225 | -64% | 1 | 1 | 0% | 1,410 | 910 | -35% | 0 | 0 | — |
case-12 | pass→pass | 6,136 | 3,755 | -39% | 1 | 1 | 0% | 880 | 986 | +12% | 0 | 0 | — |
case-13 | fail→pass | 9,302 | 3,487 | -63% | 1 | 1 | 0% | 1,361 | 978 | -28% | 0 | 0 | — |
case-14 | pass→pass | 8,202 | 8,945 | +9% | 1 | 1 | 0% | 1,245 | 1,786 | +43% | 0 | 0 | — |
case-15 | fail→pass | 4,729 | 1,629 | -66% | 1 | 1 | 0% | 812 | 634 | -22% | 0 | 0 | — |
case-16 | fail→pass | 8,516 | 2,038 | -76% | 1 | 1 | 0% | 1,513 | 689 | -54% | 0 | 0 | — |
case-17 | fail→pass | 8,052 | 2,079 | -74% | 1 | 1 | 0% | 1,292 | 692 | -46% | 0 | 0 | — |
case-18 | fail→pass | 9,881 | 2,851 | -71% | 1 | 1 | 0% | 1,361 | 849 | -38% | 0 | 0 | — |
case-19 | fail→pass | 10,472 | 1,985 | -81% | 1 | 1 | 0% | 1,554 | 694 | -55% | 0 | 0 | — |
case-20 | fail→pass | 12,588 | 3,828 | -70% | 1 | 1 | 0% | 1,665 | 1,037 | -38% | 0 | 0 | — |
case-21 | pass→pass | 5,919 | 1,414 | -76% | 1 | 1 | 0% | 811 | 593 | -27% | 0 | 0 | — |
case-22 | fail→pass | 10,694 | 2,828 | -74% | 1 | 1 | 0% | 1,543 | 833 | -46% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.