Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Automated visual tuning: a vision or video model rates rendered variants in a loop. Render several labeled variants into one artifact, ask the model to rate them and suggest better values, render the suggestions, ask it to pick the best, repeat until good — the model is the eye, you run the loop.
.claude/skills/sickn33-lookdev-auto/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 25% | 0% |
Use whenever "looks/feels right" is the success criterion and there's no cheap numeric metric — animation easing/timing, zoom/camera feel, color grade, layout/spacing, design params, render/encoder settings, prompt params. Use the automated counterpart to lookdev when there's no human to sit the loop.
_Source: connerkward/lookdev-auto-skill (MIT)._
When the target is "does this LOOK/FEEL right" (not a number you can minimize), a vision model (image) or video-understanding model (motion/timing) can be the judge in a tight optimize loop. Worked reference: the screenstudio-alternative skill (iteration.py) (tuned zoom-animation feel via fal-ai/video-understanding).
small spread. Annotate each variant's params ON the artifact (burn the label in: "A · 2.2Hz · ζ0.5"). Images → a labeled grid/contact sheet. Video/motion → a labeled sequence (label card or burned-in overlay before/over each clip) so the model can compare temporally.
rubric (define what "good" means — and what "too much"/"too little" look like). Ask for per-variant ratings + concrete suggested new values as JSON: {"ratings":{"A":n,...},"best_so_far":"X","suggest":[[p1,p2],...]}.
model's suggestions (+ carry the current best) into one artifact; ask it to pick the single best. Usually converges in 2 rounds.
round is 1 upload + 1 inference, not 6. Montage/grid beats a loop of single calls.
"variant A used X" context to carry → fewer tokens, fewer mistakes.
ONLY JSON"; regex the first {...}.
the whole asset. Cheaper render, smaller upload, faster inference. Apply the found params to the full render once.
render + token cost. Wide-but-sparse round 1, narrow round 2.
fixed anchors each round — gives the model a reference scale and exposes when its "best" is worse than the safe default (catch a bad recommendation early).
(smooth, subtle settle, not bouncy, not sluggish). Don't ask "which do you like" — that lets it echo your framing. A held-out criterion keeps the judge honest (see verify-outputs-rule: the check must be independent of what you tuned).
of re-rendering it.
skip round 2.
spatial things (layout, color, crop); only reach for a true video model when the thing being judged is temporal (easing, timing, motion smoothness) — those are invisible in stills.
don't pay a model per step.
and let them pick; a model's "best" isn't their best. (This is why the screen-studio spring auto-tune was dropped — the model's pick didn't match the owner's eye.)
winner against the safe default yourself before committing.
spacing perceptible; near-identical variants get noise-rated.
User request:
> Use @lookdev-auto for this task: Automated visual tuning: a vision or video model rates rendered variants in a loop.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-09 | pass→pass | 13,155 | 8,623 | -34% | 1 | 1 | 0% | 2,294 | 2,785 | +21% | 0 | 0 | — |
case-10 | pass→pass | 17,013 | 12,190 | -28% | 1 | 1 | 0% | 2,673 | 3,288 | +23% | 0 | 0 | — |
case-01 | fail→pass | 37,093 | 16,995 | -54% | 1 | 1 | 0% | 1,500 | 4,101 | +173% | 0 | 0 | — |
case-02 | fail→fail | 18,699 | 17,134 | -8% | 1 | 1 | 0% | 3,365 | 4,368 | +30% | 0 | 0 | — |
case-03 | fail→pass | 19,920 | 14,084 | -29% | 1 | 1 | 0% | 3,653 | 3,821 | +5% | 0 | 0 | — |
case-04 | pass→pass | 13,657 | 6,100 | -55% | 1 | 1 | 0% | 2,414 | 2,196 | -9% | 0 | 0 | — |
case-05 | pass→pass | 12,770 | 5,613 | -56% | 1 | 1 | 0% | 1,930 | 2,196 | +14% | 0 | 0 | — |
case-06 | pass→pass | 11,526 | 5,702 | -51% | 1 | 1 | 0% | 2,074 | 2,092 | +1% | 0 | 0 | — |
case-07 | pass→pass | 12,972 | 8,078 | -38% | 1 | 1 | 0% | 2,510 | 2,712 | +8% | 0 | 0 | — |
case-08 | fail→pass | 13,584 | 7,157 | -47% | 1 | 1 | 0% | 2,313 | 2,550 | +10% | 0 | 0 | — |
case-11 | fail→pass | 15,070 | 7,289 | -52% | 1 | 1 | 0% | 2,427 | 2,380 | -2% | 0 | 0 | — |
case-12 | fail→pass | 10,943 | 5,278 | -52% | 1 | 1 | 0% | 1,724 | 2,154 | +25% | 0 | 0 | — |
case-13 | pass→pass | 7,778 | 3,559 | -54% | 1 | 1 | 0% | 1,430 | 1,925 | +35% | 0 | 0 | — |
case-14 | pass→pass | 13,879 | 9,195 | -34% | 1 | 1 | 0% | 2,113 | 2,691 | +27% | 0 | 0 | — |
case-15 | pass→pass | 12,565 | 8,156 | -35% | 1 | 1 | 0% | 2,231 | 2,610 | +17% | 0 | 0 | — |
case-16 | fail→pass | 15,279 | 6,989 | -54% | 1 | 1 | 0% | 2,557 | 2,485 | -3% | 0 | 0 | — |
case-17 | pass→pass | 14,874 | 9,005 | -39% | 1 | 1 | 0% | 2,477 | 2,729 | +10% | 0 | 0 | — |
case-18 | pass→pass | 10,792 | 4,807 | -55% | 1 | 1 | 0% | 1,655 | 2,043 | +23% | 0 | 0 | — |
case-19 | pass→pass | 15,407 | 7,289 | -53% | 1 | 1 | 0% | 2,529 | 2,501 | -1% | 0 | 0 | — |
case-20 | pass→pass | 18,015 | 12,625 | -30% | 1 | 1 | 0% | 2,957 | 3,316 | +12% | 0 | 0 | — |
case-21 | fail→pass | 11,543 | 6,744 | -42% | 1 | 1 | 0% | 2,272 | 2,542 | +12% | 0 | 0 | — |
case-22 | pass→pass | 13,542 | 8,877 | -34% | 1 | 1 | 0% | 2,236 | 2,685 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 21 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.