Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when deciding what belongs in a CVPR supplementary upload versus the 8-page body, covering the one-week-later supplement deadline, video and qualitative-result norms in computer vision, anonymous code packaging, the no-external-links rule, and keeping decision-critical evidence out of material reviewers may skip.
.claude/skills/brycewang-stanford-cvpr-supplementary/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 110% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 95% | 0% |
CVPR gives supplementary material its own deadline — November 20 in the 2026 cycle, one week after the November 13 paper deadline — and vision reviewers open supplements more than reviewers at most venues, because qualitative evidence lives there. That makes the supplement a designed deliverable, not an attic. Facts below are the 2026 cycle as read on 2026-07-08.
One rule sorts everything: anything a reviewer needs to judge the claims must be in the 8-page body. The supplement is legitimately for depth on things the body already establishes. A useful decision table:
| Content | Body or supplement? | Reasoning | |---|---|---| | Headline quantitative results | Body | The claim itself | | Core ablation (the one that justifies the design) | Body | Reviewers rank papers on it | | Full ablation grids, extra datasets | Supplement | Depth on an established point | | Method details needed to believe the numbers | Body (compressed) | "See supplement" for correctness reads as evasion | | Derivations, proofs, extra math | Supplement | Vision reviewers spot-check | | Qualitative result grids beyond page 8 | Supplement | Expected and widely read | | Failure cases | Both: one in body, catalog in supplement | Honesty signal reviewers reward | | Video results (tracking, generation, 3D, robotics) | Supplement | The only sanctioned channel — external video links are banned | | Code + configs | Supplement | See cvpr-artifact-evaluation |
For temporal or 3D work, the video is the result, and since external links are prohibited, the uploaded file is your one shot:
methods (anonymized — "Ours" vs "24]", never lab-recognizable branding).
README.txt) mapping each clip to the paper section it supports.to survive a strict cap and check the current page before rendering hours of footage.
The supplement deadline trailing the paper by a week is a packaging window. The submitted PDF is frozen; a supplement that introduces results the body never mentions effectively smuggles content past the page limit, and reviewers treat it that way. Legitimate week-two work: rendering video, cleaning code, formatting the extended tables the body already cites. Every body reference ("see Suppl. B") must resolve — a dangling pointer is the most common supplement bug at every CVF venue.
textsupplement.zip ├── suppl.pdf # A: extended results · B: ablation grids · C: failures │ # D: implementation detail · E: derivations ├── videos/ │ ├── README.txt # clip → claim map, one line each │ ├── 01_main_comparison.mp4 │ └── 02_failure_modes.mp4 └── code/ # scrubbed; see cvpr-artifact-evaluation └── REPRODUCE.md # table/figure → command
Use the same numbering scheme as the paper (Suppl. Table 7 continues Table 6), start the supplement PDF with a half-page table of contents, and repeat each research question above its ablation grid so sections stand alone.
Seen across CVF-venue submission cycles; all are findable in a one-hour audit:
after a late reorganization. Grep the body for Suppl and resolve each.
the body's Table 1 in reviews ("which Table 1?").
config dump says 100, because one was updated post-freeze. Generate both from the recipe ledger.
watermark, or a dataset figure carrying a lab logo.
method differences being illustrated; export lossless crops for the regions the text argues about.
which reviewers close after page 2. Curated depth beats exhaustive depth.
The supplement is reviewed double-blind under the same rules as the paper: no author names in code headers, no institutional dataset paths, no acknowledgements, no grant IDs, no watermark from a lab's internal tooling. And the external-link ban applies — a "project page" URL inside the supplement PDF is a policy violation, not a shortcut.
Design for the observed reading pattern, not the hoped-for one: the supplement gets opened (a) right after the body's qualitative section, looking for more examples and the failure cases, (b) during score justification, verifying one implementation detail, and (c) in discussion week, when another reviewer cites it. Each entry point should land in seconds — which is why the table of contents, paper-continuous numbering, and the clip-to-claim README are not cosmetics but the difference between a supplement that testifies for you and one that never takes the stand.
pointers so the paper survives an unopened supplement either way.
text[Supplement audit] ship-ready / issues [Placement violations] decision-critical content found in supplement: <list or none> [Video] clips indexed · failures included · metadata stripped [Pointer check] body references resolving: <n/m> [Anonymity] clean / leaks: <locations> [Size] archive size vs current cap (cap 待核实)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 30,040 | 21,003 | -30% | 1 | 1 | 0% | 5,484 | 4,407 | -20% | 0 | 0 | — |
case-02 | fail→pass | 31,711 | 23,972 | -24% | 1 | 1 | 0% | 4,463 | 4,159 | -7% | 0 | 0 | — |
case-03 | fail→pass | 21,323 | 11,072 | -48% | 1 | 1 | 0% | 2,754 | 3,372 | +22% | 0 | 0 | — |
case-04 | pass→pass | 17,102 | 17,556 | +3% | 1 | 1 | 0% | 1,929 | 3,354 | +74% | 0 | 0 | — |
case-05 | pass→pass | 18,083 | 14,746 | -18% | 1 | 1 | 0% | 2,144 | 3,142 | +47% | 0 | 0 | — |
case-06 | fail→pass | 13,865 | 9,466 | -32% | 1 | 1 | 0% | 1,486 | 3,121 | +110% | 0 | 0 | — |
case-07 | fail→fail | 18,217 | 21,131 | +16% | 1 | 1 | 0% | 1,980 | 3,740 | +89% | 0 | 0 | — |
case-08 | pass→pass | 11,926 | 22,861 | +92% | 1 | 1 | 0% | 2,006 | 3,513 | +75% | 0 | 0 | — |
case-09 | pass→pass | 13,260 | 15,076 | +14% | 1 | 1 | 0% | 1,950 | 2,943 | +51% | 0 | 0 | — |
case-10 | pass→pass | 20,247 | 16,246 | -20% | 1 | 1 | 0% | 2,401 | 3,228 | +34% | 0 | 0 | — |
case-17 | fail→pass | 19,927 | 28,666 | +44% | 1 | 1 | 0% | 2,348 | 4,581 | +95% | 0 | 0 | — |
case-11 | fail→pass | 16,743 | 18,871 | +13% | 1 | 1 | 0% | 1,731 | 3,710 | +114% | 0 | 0 | — |
case-12 | pass→pass | 15,253 | 9,058 | -41% | 1 | 1 | 0% | 2,237 | 2,987 | +34% | 0 | 0 | — |
case-13 | pass→pass | 18,819 | 18,914 | +1% | 1 | 1 | 0% | 2,036 | 3,552 | +74% | 0 | 0 | — |
case-14 | pass→pass | 15,872 | 8,903 | -44% | 1 | 1 | 0% | 1,837 | 2,188 | +19% | 0 | 0 | — |
case-15 | fail→pass | 14,669 | 11,614 | -21% | 1 | 1 | 0% | 1,476 | 2,517 | +71% | 0 | 0 | — |
case-16 | fail→fail | 19,111 | 14,955 | -22% | 1 | 1 | 0% | 1,956 | 3,142 | +61% | 0 | 0 | — |
case-18 | fail→pass | 24,449 | 18,806 | -23% | 1 | 1 | 0% | 3,078 | 4,275 | +39% | 0 | 0 | — |
case-19 | fail→pass | 18,941 | 7,803 | -59% | 1 | 1 | 0% | 1,735 | 2,713 | +56% | 0 | 0 | — |
case-20 | pass→pass | 13,725 | 9,520 | -31% | 1 | 1 | 0% | 1,421 | 3,062 | +115% | 0 | 0 | — |
case-21 | pass→pass | 9,947 | 16,556 | +66% | 1 | 1 | 0% | 1,681 | 2,993 | +78% | 0 | 0 | — |
case-22 | pass→pass | 22,035 | 21,717 | -1% | 1 | 1 | 0% | 2,308 | 4,169 | +81% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.