Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when writing or revising a CoRL paper's prose — leading with the embodied task and the learned component, calibrating claims to evaluation scale, writing the mandatory Limitations section as a scored asset, fitting the argument into 8 pages, and satisfying a dual reviewer audience of ML and robotics readers.
.claude/skills/brycewang-stanford-corl-writing-style/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 63% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 71% | 0% |
A CoRL paper is read by two audiences at once: reviewers fluent in learning methods who will probe the algorithmic claim, and reviewers fluent in robots who will probe the physical claim. Prose that serves only one of them loses the other's score. The style guidance here is about keeping both readers oriented inside 8 pages — with a mandatory Limitations section spending part of that budget (CoRL 2026 instructions, corl.org, read 2026-07-08).
By the end of page 1, both audiences should be able to answer four questions:
representation? data?)
A reliable abstract shape: task problem → why existing learning approaches fall short → the idea in one sentence → headline evidence with its scale attached ("across 8 manipulation tasks, 5 seeds, 50 evaluation episodes each, on a real UR5") → the takeaway for the field.
Robot-learning results are stochastic and setup-dependent; the writing must carry those qualifiers without drowning in them. Calibrate at the sentence level:
| Overclaimed | Calibrated | |---|---| | "Our policy solves kitchen manipulation" | "Our policy reaches 76% mean success on the 6-task kitchen suite" | | "Transfers seamlessly to the real world" | "Transfers with an 11-point average sim-to-real drop (Table 4)" | | "Generalizes to unseen objects" | "Maintains 61% success on 10 held-out objects (vs 78% on training objects)" | | "Runs in real time" | "Runs at 15 Hz on the onboard Orin" | | "Robust to disturbances" | "Recovers from 8 of 12 scripted pushes (protocol in §5.3)" |
The pattern: attach the number, the scale, and the pointer. This is also rebuttal insurance — precise claims are defensible in one page; vibes are not.
CoRL makes Limitations mandatory and counts it inside the page limit, which changes its rhetorical status: reviewers treat it as part of the argument, not boilerplate. A strong one:
depth camera returns holes"), matching what the supplementary video shows.
cross-robot claim"), which preempts the corresponding review objection.
Weak versions — generic ("more experiments needed"), disguised advertising ("limited only by compute"), or contradicted by the video — actively cost points with this reviewer pool.
text1 Introduction 1.00 pp the four-question contract 2 Related work 0.75 pp three-lane positioning (corl-related-work) 3 Method 2.00 pp one architecture figure; learned vs engineered boundary drawn explicitly 4 Experimental setup 1.25 pp tasks, robot/sim, data, baselines, protocol — the reproducibility spine lives HERE, not appendix 5 Results 2.25 pp claims in subsection headers; per-axis analysis 6 Limitations 0.50 pp mandatory; specific; video-consistent 7 Conclusion 0.25 pp one paragraph (references + appendix follow, uncounted)
Adjust the split, but defend two invariants: the setup section is generous (robotics readers judge rigor there), and Limitations is protected (it is mandatory and cutting it to reclaim space is not an option).
frequency. ML readers get the problem formalized; robotics readers learn what the policy actually commands.
proposals are scripted; the insertion policy is learned"). Blurring this line is the most common honesty complaint in reviews of systems-flavored papers.
annotations often communicates a robot result faster than any paragraph — but every figure claim needs its number in a table too; filmstrips are anecdotes.
CoRL paper rarely needs custom operator symbols, and every nonstandard symbol taxes half your audience.
actually established at that strength.
episodes so tables stand alone when skimmed.
either shows it or it isn't so.
MPC, SDF) is not universal across your two audiences.
text[ ] Page-1 contract: task, learned component, evidence scale, insight [ ] Every abstract claim → number + scale + section pointer [ ] Learned vs engineered boundary stated explicitly [ ] Setup section carries protocol detail (not deferred to appendix) [ ] Limitations: specific, video-consistent, claim-bounding [ ] Captions self-contained with seeds × episodes [ ] No demo adjectives; no uncalibrated robustness language [ ] Both audiences can follow §3 (interface first, standard notation)
Style norms are community culture; recalibrate against recent accepted papers in the newest PMLR volume (v305 for CoRL 2025) and the live author instructions at corl.org each cycle.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 27,398 | 25,575 | -7% | 1 | 1 | 0% | 3,824 | 5,025 | +31% | 0 | 0 | — |
case-02 | fail→pass | 16,975 | 14,105 | -17% | 1 | 1 | 0% | 1,745 | 2,852 | +63% | 0 | 0 | — |
case-03 | pass→pass | 18,207 | 16,339 | -10% | 1 | 1 | 0% | 2,069 | 3,327 | +61% | 0 | 0 | — |
case-04 | pass→pass | 10,273 | 5,429 | -47% | 1 | 1 | 0% | 1,607 | 2,216 | +38% | 0 | 0 | — |
case-05 | pass→pass | 27,854 | 28,352 | +2% | 1 | 1 | 0% | 4,777 | 5,824 | +22% | 0 | 0 | — |
case-06 | fail→pass | 19,106 | 17,460 | -9% | 1 | 1 | 0% | 2,661 | 3,349 | +26% | 0 | 0 | — |
case-07 | fail→pass | 23,434 | 21,385 | -9% | 1 | 1 | 0% | 2,820 | 3,903 | +38% | 0 | 0 | — |
case-08 | pass→pass | 14,661 | 15,196 | +4% | 1 | 1 | 0% | 2,089 | 2,935 | +40% | 0 | 0 | — |
case-09 | fail→pass | 18,501 | 19,535 | +6% | 1 | 1 | 0% | 2,060 | 3,516 | +71% | 0 | 0 | — |
case-10 | fail→pass | 22,618 | 18,409 | -19% | 1 | 1 | 0% | 2,256 | 3,398 | +51% | 0 | 0 | — |
case-11 | fail→pass | 15,442 | 7,451 | -52% | 1 | 1 | 0% | 1,696 | 2,662 | +57% | 0 | 0 | — |
case-12 | fail→fail | 11,519 | 14,465 | +26% | 1 | 1 | 0% | 1,733 | 2,919 | +68% | 0 | 0 | — |
case-13 | fail→pass | 15,118 | 11,498 | -24% | 1 | 1 | 0% | 1,425 | 2,888 | +103% | 0 | 0 | — |
case-14 | fail→fail | 17,780 | 14,314 | -19% | 1 | 1 | 0% | 2,051 | 2,743 | +34% | 0 | 0 | — |
case-15 | fail→pass | 22,039 | 20,569 | -7% | 1 | 1 | 0% | 1,942 | 3,325 | +71% | 0 | 0 | — |
case-16 | pass→pass | 13,273 | 19,922 | +50% | 1 | 1 | 0% | 1,967 | 3,263 | +66% | 0 | 0 | — |
case-17 | pass→pass | 18,910 | 16,560 | -12% | 1 | 1 | 0% | 2,005 | 3,135 | +56% | 0 | 0 | — |
case-18 | pass→pass | 19,636 | 8,780 | -55% | 1 | 1 | 0% | 1,931 | 2,924 | +51% | 0 | 0 | — |
case-19 | fail→fail | 14,422 | 12,050 | -16% | 1 | 1 | 0% | 1,576 | 2,535 | +61% | 0 | 0 | — |
case-20 | pass→pass | 18,293 | 17,791 | -3% | 1 | 1 | 0% | 2,482 | 3,667 | +48% | 0 | 0 | — |
case-21 | pass→pass | 17,901 | 9,920 | -45% | 1 | 1 | 0% | 1,771 | 2,641 | +49% | 0 | 0 | — |
case-22 | pass→pass | 15,806 | 15,477 | -2% | 1 | 1 | 0% | 2,050 | 3,114 | +52% | 0 | 0 | — |
case-23 | fail→pass | 19,794 | 17,949 | -9% | 1 | 1 | 0% | 1,989 | 3,076 | +55% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +43 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.