Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Curate reader-facing survey tables for the Appendix (clean layout + high information density), using only in-scope evidence and existing citation keys. **Trigger**: appendix tables, publishable tables, survey tables, reader tables, 附录表格, 可发表表格, 综述表格. **Use when**: you have C4 artifacts (evidence packs + anchor sheet + citations) and want tables that look like a real survey (not internal logs).
.claude/skills/willoscar-appendix-table-writer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-22 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 359% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 20% | 0% |
The pipeline can produce index tables that are useful for planning/debugging, but read like internal artifacts.
This skill writes publishable, reader-facing tables that can live in an Appendix:
Index tables remain in outline/tables_index.md and should not be copied verbatim into the paper.
outline/table_schema.md (table intent + evidence mapping)outline/tables_index.md (internal index; optional but recommended)outline/subsection_briefs.jsonloutline/evidence_drafts.jsonloutline/anchor_sheet.jsonlcitations/ref.bibGOAL.mdRead as needed:
references/table_cell_hygiene.md when Appendix table cells still copy raw paper self-narration or generic result wrappersMachine-readable assets:
assets/table_cell_hygiene.jsonoutline/tables_appendix.mdMission: choose tables a reader actually wants in a survey Appendix.
Do:
Avoid:
Mission: make the table look publishable in LaTeX.
Do:
<br> sparingly (0-1 per cell; never a list dump)Avoid:
Mission: prevent hallucinations.
Do:
Avoid:
outline/tables_appendix.md must:
course_paper, or >=2 for survey / deep**Appendix Table A1. Representative systems by method family and evaluation setting**#, ##, ###) inside the file (the merger adds an Appendix heading)TODO, TBD, FIXME, ASCII three-dot ellipsis, unicode ellipsis)[@BibKey] (keys must exist in citations/ref.bib)GOAL.md (scope) and outline/table_schema.md (what each table must answer).queries.md:draft_profile when present; course_paper uses one strong reader table by default, while survey / deep retain at least two.outline/tables_index.md as a shortlist source, but do not paste it verbatim.outline/subsection_briefs.jsonl, outline/evidence_drafts.jsonl, and outline/anchor_sheet.jsonl (no guessing).citations/ref.bib.If you are unsure what to build, start with these two:
1) Method/architecture map (representative works)
2) Evaluation protocol / benchmark map
Optional third (only if it stays clean): 3) Risk / threat-surface map
Bad (index table / internal notes):
planning / memory / tools / eval / safety (slash dump)<br> linesGood (survey table):
Also good (avoid intermediate-artifact tells):
-> inside cells; prefer natural phrasing (e.g., "interleaves reasoning traces with tool actions").If you cannot fill a row without guessing:
evidence-draft / anchor-sheet for that area.uv run python .codex/skills/appendix-table-writer/scripts/run.py --helpuv run python .codex/skills/appendix-table-writer/scripts/run.py --workspace <workspace>--workspace <workspace> (required)--unit-id <id> (optional; used only for runner bookkeeping)--inputs <a;b;c> (optional; ignored by the validator; kept for runner compatibility)--outputs <relpath> (optional; defaults to outline/tables_appendix.md)--checkpoint <C#> (optional; ignored by the validator)uv run python .codex/skills/appendix-table-writer/scripts/run.py --workspace workspaces/e2e-agent-survey-latex-verify-YYYYMMDD-HHMMSS
uv run python .codex/skills/appendix-table-writer/scripts/run.py --workspace <workspace> --outputs outline/tables_appendix.md
Notes:
outline/tables_appendix.md from the existing evidence artifacts and then validates the result.output/TABLES_APPENDIX_REPORT.md.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | fail→pass | 9,527 | 7,232 | -24% | 1 | 1 | 0% | 1,398 | 2,777 | +99% | 0 | 0 | — |
case-01 | fail→fail | 4,637 | 6,629 | +43% | 1 | 1 | 0% | 260 | 2,076 | +698% | 0 | 0 | — |
case-02 | fail→fail | 5,895 | 5,379 | -9% | 1 | 1 | 0% | 223 | 2,022 | +807% | 0 | 0 | — |
case-03 | fail→fail | 17,961 | 6,180 | -66% | 1 | 1 | 0% | 3,832 | 1,932 | -50% | 0 | 0 | — |
case-04 | fail→fail | 9,256 | 7,550 | -18% | 1 | 1 | 0% | 1,833 | 2,233 | +22% | 0 | 0 | — |
case-05 | pass→pass | 9,471 | 5,430 | -43% | 1 | 1 | 0% | 1,706 | 2,759 | +62% | 0 | 0 | — |
case-06 | fail→fail | 14,796 | 8,416 | -43% | 1 | 1 | 0% | 2,927 | 2,191 | -25% | 0 | 0 | — |
case-07 | fail→fail | 9,201 | 4,010 | -56% | 1 | 1 | 0% | 1,605 | 2,496 | +56% | 0 | 0 | — |
case-08 | fail→pass | 7,556 | 2,775 | -63% | 1 | 1 | 0% | 1,325 | 2,223 | +68% | 0 | 0 | — |
case-09 | fail→pass | 9,076 | 3,767 | -58% | 1 | 1 | 0% | 1,587 | 2,436 | +53% | 0 | 0 | — |
case-10 | fail→pass | 4,644 | 12,047 | +159% | 1 | 1 | 0% | 859 | 3,942 | +359% | 0 | 0 | — |
case-11 | pass→pass | 11,120 | 5,962 | -46% | 1 | 1 | 0% | 1,850 | 2,878 | +56% | 0 | 0 | — |
case-12 | fail→fail | 12,093 | 2,186 | -82% | 1 | 1 | 0% | 1,908 | 2,056 | +8% | 0 | 0 | — |
case-13 | fail→fail | 7,579 | 4,771 | -37% | 1 | 1 | 0% | 1,249 | 2,561 | +105% | 0 | 0 | — |
case-14 | fail→pass | 12,094 | 3,802 | -69% | 1 | 1 | 0% | 1,916 | 2,301 | +20% | 0 | 0 | — |
case-15 | pass→pass | 6,385 | 12,859 | +101% | 1 | 1 | 0% | 1,047 | 2,509 | +140% | 0 | 0 | — |
case-21 | pass→pass | 7,684 | 5,149 | -33% | 1 | 1 | 0% | 1,378 | 2,675 | +94% | 0 | 0 | — |
case-16 | fail→fail | 7,945 | 5,423 | -32% | 1 | 1 | 0% | 1,360 | 2,818 | +107% | 0 | 0 | — |
case-17 | fail→pass | 10,101 | 3,933 | -61% | 1 | 1 | 0% | 1,995 | 2,384 | +19% | 0 | 0 | — |
case-18 | fail→pass | 2,751 | 10,541 | +283% | 1 | 1 | 0% | 367 | 3,824 | +942% | 0 | 0 | — |
case-19 | pass→pass | 5,875 | 9,297 | +58% | 1 | 1 | 0% | 1,112 | 3,257 | +193% | 0 | 0 | — |
case-20 | pass→pass | 16,330 | 13,991 | -14% | 1 | 1 | 0% | 2,542 | 3,848 | +51% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.