Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Assemble a two-host fake-podcast skit ad from a config — per-line lipsync clips hard-concatenated in script order, scaled/padded to 1080×1920, WHITE bottom-center captions (up to 5 words per cue, broken on sentence punctuation, word-wrapped to stay in-frame, held at least 0.9s) built from each line's OWN ElevenLabs char-level timestamps (offset by cumulative clip start, never Whisper), and closed on a Playwright/PIL brand end card composited from the real wordmark — never AI-rendered text. This
.claude/skills/gooseworks-ai-render-podcast-skit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 19% | 0% |
Assemble a two-host fake-podcast skit ad from a config: a skeptic and a believer at an absurd themed podcast desk do a snappy back-and-forth about the product (the set is deliberately unrelated — that is the joke). Each line is its own lipsync clip so the edit can cut on the dialogue beat (~1.8s avg); this capability is the FREE, deterministic assembly that concatenates those clips, renders the WHITE captions, and appends the brand end card.
scripts/config.example.json is the worked example (Ladder run-02 "Laundromat 2am", ~49s 1080×1920 9:16, ~22 lines); scripts/PIPELINE.md maps every config block to its source step and scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are separate capabilities: one ElevenLabs with-timestamps VO per line (one voice per host) via create-vo-elevenlabs; two photoreal base stills at the themed desk plus ~10 expression variants (mouths NEUTRAL/CLOSED, gpt-image-2 quality=high, not nano-banana) via create-image-gpt-image-fal; and one lipsync clip per (still, VO) pair via create-video-fal. Given the per-line clips + their VO timestamps + the brand wordmark SVG, render-podcast-skit walks the scenes in script order, builds the global caption timeline, renders the WHITE captions, hard-concats the clips, auto-appends the end card, and final-encodes crf28 → the master. Re-cuts reuse the existing VOs / stills / clips and cost $0.
needs no music (an optional low ambience is a taste call, off by default).
order (scale/pad to 1080×1920, re-encode) — no dissolves.
global words.json by offsetting each line's char-level word timings by the cumulative clip start, group into ≤5-word cues broken on sentence-final punctuation, and render WHITE #FFFFFF bottom-center captions (black outline), word-wrapped to stay in-frame and held ≥0.9s — PIL PNG overlays when the host ffmpeg lacks libass (common), else ASS. (Yellow 3-word karaoke was the old style, rejected in testing.) Whisper on the rendered clips mistimes; the VO timestamps are ground truth.
lockup is a deterministic HTML → PNG → 2.5s mp4 from the brand's real wordmark SVG (black bg, brand wordmark, CTA pill, URL), auto-appended after the last line. A diffusion model garbles a wordmark.
burn ASS via libass), append the end-card mp4, and final-encode -preset slow -crf 28 + aac 96k → a 1080×1920 h264+aac master (~6MB for ~28s; the old -crf 20 produced ~16MB). No paid calls, no keys.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,093 | 15,525 | -14% | 1 | 1 | 0% | 4,125 | 4,570 | +11% | 0 | 0 | — |
case-02 | fail→fail | 18,060 | 3,345 | -81% | 1 | 1 | 0% | 3,935 | 1,385 | -65% | 0 | 0 | — |
case-03 | fail→pass | 19,200 | 22,918 | +19% | 1 | 1 | 0% | 4,168 | 6,363 | +53% | 0 | 0 | — |
case-04 | pass→pass | 3,529 | 2,612 | -26% | 1 | 1 | 0% | 636 | 1,360 | +114% | 0 | 0 | — |
case-05 | pass→pass | 2,860 | 2,511 | -12% | 1 | 1 | 0% | 506 | 1,305 | +158% | 0 | 0 | — |
case-06 | pass→pass | 5,141 | 2,702 | -47% | 1 | 1 | 0% | 1,074 | 1,371 | +28% | 0 | 0 | — |
case-07 | fail→pass | 9,965 | 7,141 | -28% | 1 | 1 | 0% | 1,686 | 2,133 | +27% | 0 | 0 | — |
case-08 | fail→pass | 9,202 | 2,509 | -73% | 1 | 1 | 0% | 1,575 | 1,270 | -19% | 0 | 0 | — |
case-09 | pass→pass | 10,940 | 1,820 | -83% | 1 | 1 | 0% | 1,913 | 1,195 | -38% | 0 | 0 | — |
case-10 | pass→fail | 10,986 | 1,812 | -84% | 1 | 1 | 0% | 2,030 | 1,176 | -42% | 0 | 0 | — |
case-11 | fail→pass | 8,586 | 4,746 | -45% | 1 | 1 | 0% | 1,526 | 1,812 | +19% | 0 | 0 | — |
case-12 | pass→pass | 12,045 | 5,252 | -56% | 1 | 1 | 0% | 2,063 | 1,830 | -11% | 0 | 0 | — |
case-13 | pass→pass | 9,685 | 4,738 | -51% | 1 | 1 | 0% | 1,581 | 1,685 | +7% | 0 | 0 | — |
case-14 | pass→pass | 9,369 | 3,531 | -62% | 1 | 1 | 0% | 1,605 | 1,446 | -10% | 0 | 0 | — |
case-15 | pass→pass | 12,196 | 2,132 | -83% | 1 | 1 | 0% | 2,236 | 1,297 | -42% | 0 | 0 | — |
case-16 | fail→pass | 13,308 | 2,544 | -81% | 1 | 1 | 0% | 2,539 | 1,354 | -47% | 0 | 0 | — |
case-17 | fail→pass | 8,298 | 2,373 | -71% | 1 | 1 | 0% | 1,531 | 1,336 | -13% | 0 | 0 | — |
case-18 | fail→pass | 9,023 | 1,693 | -81% | 1 | 1 | 0% | 1,750 | 1,214 | -31% | 0 | 0 | — |
case-19 | fail→pass | 14,991 | 6,882 | -54% | 1 | 1 | 0% | 2,849 | 2,338 | -18% | 0 | 0 | — |
case-20 | fail→pass | 14,099 | 5,011 | -64% | 1 | 1 | 0% | 2,457 | 1,746 | -29% | 0 | 0 | — |
case-21 | fail→pass | 11,045 | 3,550 | -68% | 1 | 1 | 0% | 2,006 | 1,576 | -21% | 0 | 0 | — |
case-22 | pass→pass | 8,101 | 1,713 | -79% | 1 | 1 | 0% | 1,451 | 1,216 | -16% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.