Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when revising an ACM MM (ACM Multimedia) paper for house style — putting the cross-modal contribution on the first page, framing media (figures, video, audio) as evidence rather than decoration, making the fusion the visible claim, and compressing the argument into a 6-8 page ACM sigconf body with references-only overflow.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 23% | 0% |
Use this to revise an ACM Multimedia draft for the venue's expectations. The reader is a busy reviewer scanning across sixteen thematic areas; the paper must announce what is multimedia about it before the model diagram.
Lead the abstract and first paragraph with: problem → why one modality is insufficient → the cross-modal method or system → media-grounded evidence → why it matters for multimedia. The multimedia contribution belongs on page one; a paper that opens with a single-modality benchmark reads as a CVPR or ACL paper that wandered in.
video that is evidence for a claim, an audio clip that a reader can check.
engagement result behind it.
| Media element | Weak (decoration) | Strong (evidence) | |---|---|---| | Teaser figure | A pretty system diagram | The moment where modalities disagree and the method wins | | Qualitative grid | Cherry-picked successes | Paired success/failure across modalities with captions that state the point | | Supplementary video | "See our results" | A clip tied to a specific claim, with the baseline shown alongside | | Audio sample | An unlabeled waveform | The case the vision-only baseline misses, annotated |
The 6–8 page body is short by ACM standards, and figures compete with text for space.
supplement; keep the body's argument self-contained without them.
grasp the cross-modal claim.
and which move to the supplement.
textPass 1 (contribution): Does page one name the cross-modal/systems contribution? Pass 2 (fusion): Is the mechanism the claim, and is it ablated later? Pass 3 (media): Does each figure/clip support a specific claim, with a self-contained caption? Pass 4 (scope): Is every "better/more engaging/higher quality" tied to a measured result? Pass 5 (budget): Does the body fit 6-8 sigconf pages with overflow holding references only?
Your reviewers may come from different thematic areas, so the title and abstract have to be legible to a vision person, an audio person, and a systems person at once.
image," "cross-modal") so area chairs assign the right reviewers.
papers may read little more.
not share.
Treat the 6–8 page limit as a budget you allocate before writing prose:
textp1 intro: contribution + why one modality fails + teaser figure p2 related work (tight) + problem setup p3-4 method: the fusion/alignment mechanism, one architecture figure p5-6 experiments: main table, the decisive ablation, failure cases, user-study summary p7-8 discussion + limitations; references spill onto the overflow pages (references only)
If a section will not fit, move detail to the supplement rather than shrinking the font or margins — template tampering is a desk-reject risk, and a cramped body reads worse than a clean one with a fuller supplement.
with the seam and ablating it.
multimedia question.
by measuring or softening.
captions stand alone.
references only, and misuse risks desk reject.
text[First-page verdict] multimedia contribution up front / buried [Fusion visibility] mechanism is the claim / hidden behind a backbone [Media evidence] each element earns its place / decorative elements: <list> [Scope] claims matched to evidence / overclaims: <list> [Page budget] fits 6-8 sigconf pages / over by <n> [Top three fixes] <ordered>
Other measured skills in the registry, with their headline benchmark lift.