Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when drafting, structuring or revising the manuscript or report in Stage 07 (Writing) — shaping the contribution into one story, writing the abstract and introduction, fixing prose that reads generic or templated, ordering sentences for clarity, or deciding what Figure 1 should show.
.claude/skills/tangxiangru-paper-writing/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-24 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-23 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-07 | ✓→✓ | = Same ✓ | 26% | 0% |
Stage 07 turns approved stage summaries into a manuscript. The failure mode is not bad grammar — it is a draft that reports what was done instead of arguing one claim. This skill is the correction for that.
reference.md in this directory is the long-form treatment: reviewer reading order, the seven sentence-level principles from Gopen and Swan, mathematical notation habits, figure design, and a pre-submission checklist. Read it when a specific problem below is the one you have.
Before drafting anything, write the paper's claim as one sentence:
If you cannot write it, the framing has not converged, and no amount of prose will fix that. Go back to the Stage 02 hypothesis manifest and the Stage 06 analysis and find out which claim the evidence actually supports. Writing around a claim the results do not support is the most expensive mistake available at this stage.
Every section then serves that one claim. Related work, experiments and discussion support it; they are not independent mini-papers.
Delete openings that carry no information: "With the rapid development of…", "In recent years, X has attracted increasing attention…". Start at sentence 1.
By the end of the introduction the reader must have:
Contribution bullets state results, not activities. "We conduct extensive experiments on three datasets" is an activity. "We show retrieval recovers 12 points of accuracy that long-context prompting loses when evidence is diffuse" is a result.
Symptoms and fixes, in order of how often they apply:
| Symptom | Fix | | --- | --- | | Hedging stacked on hedging ("may potentially suggest") | State the claim, or state the limitation. Not both in one clause. | | Vague quantifiers ("significantly better", "a variety of") | Replace with the number. If there is no number, say what you actually observed. | | Ambiguous pronouns ("this shows…") | Name the referent: "this gap shows…". | | Verb buried at the end of a long sentence | Move the verb early; readers hold the subject in memory until they get it. | | Terminology drifting across sections | Pick one term per concept and use it everywhere. Synonyms read as different concepts. | | Every paragraph the same length and shape | Vary it. Uniform paragraph shape is the strongest tell of generated prose. |
Roughly equal time on: the title and abstract; the introduction; the figures; and everything else combined. Reviewers usually read title → abstract → figures → introduction → results, and form a verdict before the method section. Front-load accordingly.
For citation handling use the citation-discipline skill. For venue-specific required sections use venue-checklist.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 25,017 | 21,420 | -14% | 1 | 1 | 0% | 3,259 | 3,635 | +12% | 0 | 0 | — |
case-06 | fail→fail | 23,513 | 16,504 | -30% | 1 | 1 | 0% | 2,922 | 2,871 | -2% | 0 | 0 | — |
case-07 | pass→pass | 14,808 | 12,156 | -18% | 1 | 1 | 0% | 1,592 | 2,003 | +26% | 0 | 0 | — |
case-24 | fail→pass | 14,800 | 8,828 | -40% | 1 | 1 | 0% | 1,644 | 2,203 | +34% | 0 | 0 | — |
case-02 | pass→pass | 16,135 | 17,887 | +11% | 1 | 1 | 0% | 2,524 | 4,077 | +62% | 0 | 0 | — |
case-03 | pass→pass | 18,809 | 15,413 | -18% | 1 | 1 | 0% | 3,106 | 3,343 | +8% | 0 | 0 | — |
case-04 | pass→pass | 19,402 | 23,942 | +23% | 1 | 1 | 0% | 3,084 | 3,806 | +23% | 0 | 0 | — |
case-05 | pass→pass | 10,040 | 9,699 | -3% | 1 | 1 | 0% | 1,565 | 2,540 | +62% | 0 | 0 | — |
case-08 | pass→pass | 12,423 | 15,339 | +23% | 1 | 1 | 0% | 2,022 | 2,505 | +24% | 0 | 0 | — |
case-09 | pass→pass | 7,514 | 7,254 | -3% | 1 | 1 | 0% | 1,221 | 2,128 | +74% | 0 | 0 | — |
case-10 | fail→fail | 12,863 | 10,458 | -19% | 1 | 1 | 0% | 1,259 | 1,700 | +35% | 0 | 0 | — |
case-11 | pass→pass | 9,317 | 10,509 | +13% | 1 | 1 | 0% | 1,402 | 1,772 | +26% | 0 | 0 | — |
case-12 | pass→pass | 7,446 | 8,259 | +11% | 1 | 1 | 0% | 1,365 | 2,264 | +66% | 0 | 0 | — |
case-13 | pass→pass | 9,862 | 13,215 | +34% | 1 | 1 | 0% | 1,692 | 2,280 | +35% | 0 | 0 | — |
case-14 | pass→pass | 11,664 | 12,263 | +5% | 1 | 1 | 0% | 1,680 | 1,855 | +10% | 0 | 0 | — |
case-15 | fail→pass | 15,947 | 11,668 | -27% | 1 | 1 | 0% | 2,545 | 2,963 | +16% | 0 | 0 | — |
case-16 | pass→pass | 14,661 | 13,854 | -6% | 1 | 1 | 0% | 1,640 | 2,230 | +36% | 0 | 0 | — |
case-17 | pass→pass | 17,512 | 14,370 | -18% | 1 | 1 | 0% | 2,064 | 2,375 | +15% | 0 | 0 | — |
case-18 | pass→pass | 16,434 | 15,043 | -8% | 1 | 1 | 0% | 2,738 | 3,340 | +22% | 0 | 0 | — |
case-19 | pass→pass | 18,640 | 13,722 | -26% | 1 | 1 | 0% | 2,136 | 2,694 | +26% | 0 | 0 | — |
case-20 | pass→pass | 6,293 | 6,166 | -2% | 1 | 1 | 0% | 1,279 | 2,087 | +63% | 0 | 0 | — |
case-21 | pass→pass | 12,154 | 10,356 | -15% | 1 | 1 | 0% | 1,108 | 1,675 | +51% | 0 | 0 | — |
case-22 | pass→pass | 8,020 | 10,926 | +36% | 1 | 1 | 0% | 1,261 | 1,826 | +45% | 0 | 0 | — |
case-23 | fail→pass | 13,492 | 5,310 | -61% | 1 | 1 | 0% | 1,371 | 1,607 | +17% | 0 | 0 | — |
case-25 | pass→pass | 17,318 | 8,587 | -50% | 1 | 1 | 0% | 1,944 | 2,136 | +10% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +16 percentage points is the difference between those two pass rates over the 25 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.