Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when revising an ACL paper for computational-linguistics house style, covering task-first framing, linguistic examples tied to quantitative error analysis, scoping language claims to tested languages, LLM-era claim discipline, anonymous self-reference, Limitations prose, and compressing into the 8-page or 4-page ACL format.
.claude/skills/brycewang-stanford-acl-writing-style/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 63% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 20% | 0% |
Use this on the manuscript itself. ACL reviewers are NLP specialists who read for whether the paper understands language as well as models; the style that survives them is concrete, example-grounded, and precisely scoped.
goes in, what comes out, why it is hard, and for whom.
resource, new analysis, or new finding — ACL reviews are calibrated per type.
on page one; abstract problem statements without an example read as vague at this venue.
"English only" — and if it is, say that too.
| Reflex phrasing | ACL-safe phrasing | |---|---| | "LLMs cannot do X" | "The five models tested fail X under these prompts" | | "Our method understands Y" | "Improves the Y benchmark by n points; error classes A, B shrink" | | "Works across languages" | "Evaluated on de/hi/sw/zh/ar; typological coverage discussed in §7" | | "Significantly better" | Reserve for tested significance; give the test and p-value or interval | | "State-of-the-art" | Scope to the exact setting, model scale, and date checked |
Reviewers increasingly ask whether a result is a property of the task, the model snapshot, or the prompt; write so each claim names which.
illustrated behavior occurs, in which slice, under which condition. Cherry-picked generations presented as evidence is a named reject pattern.
non-English examples; sloppy linguistics costs credibility with exactly the reviewers who like the paper's topic.
narratively ("the model gets confused").
"In our previous work." Keep it in place until camera-ready.
ARR bars relying on documents unavailable to them.
submission and added at camera-ready.
do not shrink a long paper into four pages; re-argue it.
appendix; keep one summary row of each in the body (see acl-supplementary).
contrast beat a page of citations (see acl-related-work).
diagrams restating the text are the first cut.
coverage, model dependence, evaluation validity. Specificity here is protected — ACL instructs reviewers not to penalize honest limitations.
representational harm. A boilerplate ethics paragraph is worse than none.
textweak: "We leverage powerful LLMs to achieve impressive gains." strong: "Reranking with a 7B model cuts negation-scope errors from 31% to 12% of sampled failures (Table 4)." weak: "Performance is good across all settings." strong: "Gains hold on 4 of 5 languages; Swahili degrades (-1.2 F1), which §7 traces to tokenizer fragmentation."
"the verifier," and "the LLM judge" in three sections reads as three systems to a tired reviewer.
"hallucination," "faithfulness," and "robustness" each have three incompatible literatures behind them.
thereafter ("XNLI dev-matched," not "the dev set").
tables, figures, and prose alike.
→ rewrite as design with rationale.
table supports plus pointers into it.
ACL reviewers treat conclusions as summaries under oath.
text[Style diagnosis] task-first / model-first / survey-ish / underspecified [First-page fix] <one concrete rewrite> [Overclaim list] <claim -> scoped version> [Example-evidence gaps] <anecdotes lacking counts> [Compression plan] <cut / move / merge>
Other measured skills in the registry, with their headline benchmark lift.