Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Orchestrates story, script, screenplay, concept, product promo, and multi-shot idea work into finished video. Use first when the user asks to make a video from a story or script; asks what next in a story video project; or needs a decision spanning script splitting, image refs, voices or VO, video clips, render strategy, Timeline ordering, or final Timeline handoff. Routes execution to script-compose, image-compose, voice-compose, and video-compose before those skills' CLIs are used.
.claude/skills/utopai-research-story-to-video-workflow/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 154% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 22% | 0% |
generate_* here.Use this ladder unless the user skips, reorders, supplies refs, or asks for a rough direct render:
script-compose production script; existing screenplay -> capture/adapt.script-compose splits <=15s dialogue-aware shots and extracts characters, material variants, detailed locations/location variants, and speaker/VO needs.image-compose creates useful visual anchors: base/variant character sheets and detailed location/detail anchors.voice-compose creates reusable anchors for every speaker and VO/narrator.shot_id when sequence order is unambiguous, then hand off to Timeline.Plan ahead internally, but only ask the next meaningful user-facing choice; the Consent and gates ladder fixes when render path and dispatch become askable.
| Need | Load next | |---|---| | Script capture, rewrite, split, or analysis | script-compose | | Character, location, storyboard, starting frame, or visual anchor | image-compose | | Narration, dialogue read, character voice, or audio node | voice-compose | | Clip render, continuation, audio refs, storyboard animation, or video prompt | video-compose | | Scene/ref grouping or canvas layout frames | groups-compose |
Capability skills own CLI flags, node grammar, refs, and recovery hints. PROJECT_AGENT.md owns shared failure handling.
audio_result.data.text is source of truth only for approved final narration/line reads.video-compose includes spoken text verbatim and treats voice samples as timbre anchors.Follow the project PROJECT_AGENT.md § "Recommendation and choice shape". Recommend one concrete next step. Add a second option only when there is a real tradeoff.
Before recommending refs/video, inspect workflow.json when needed and summarize only:
If the story implies more than roughly 3 minutes, recommend narrowing scope before clip planning.
After shot notes, missing video-bound character/location/voice anchors are the default next step; include a rough-direct skip when speed matters. Once anchors/user refs/rough-direct are settled, offer only a short ref review or clip-plan confirmation if ambiguity remains.
Ask only after the script/shot plan is settled and anchors, usable refs, rough-direct, or a simple single-clip case make rendering real. If anchors are still missing, return to Planning checkpoint.
Use project choice shape:
RenderChoose render path.Straight to video (Recommended)description: Fastest path to motion.
Storyboard firstdescription: Generate storyboard images first for composition control.
For storyboard-first, load image-compose Pattern 6: one composite mosaic per clip/<=15s shot note, subtype storyboard.
Ask only after render path is picked and a multi-clip plan exists. Skip for one clip. Use project choice shape:
DispatchChoose clip dispatch.(Recommended).Hybriddescription: Chain within continuous scenes; render separate scenes independently.
Paralleldescription: Render all clips independently.
Sequentialdescription: Each clip continues from the previous one; boundaries default to a hard cut to a new angle (avoids the same-shot seam) — keep a boundary same-shot only for an unbroken oner.
Signals: continuous scene/state -> sequential (hard-cut handoffs between clips); a single unbroken action the viewer must read as ONE motion -> one ≤15s clip, else sequential with a same-shot handoff; separate scenes/time jumps/wardrobe changes/montage -> parallel; continuous clusters separated by hard cuts -> hybrid. Do not chain video refs across location, time, wardrobe/state, dream/reality, or montage breaks.
After terminal generate_*:
ok:false, follow project failure handling and do not advance the pipeline.ok:true, identify the landed node id from the result or canvas state.workflow.json if shots, refs, voices, clips, or reel order affect the next decision.Typical priority:
Timeline owns reel order. Numeric video_result.data.shot_id means a clip is in the reel. When all planned story clips are ready and order is unambiguous, assign shot_id = 1..N with one updateBatch before handoff:
node "$PAI_REPO_ROOT/server/cli/canvas_mutate.js" \
--op updateBatch \
--payload-json '{"updates":[{"id":"<video_1>","patch":{"shot_id":1}},{"id":"<video_2>","patch":{"shot_id":2}}]}'Do not use generate_video.js --shot-id for speculative/partial ordering. Assign after clips land. Local export uses reel_stitch.js only on explicit request. Then tell the user to open Timeline to inspect and preview.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→pass | 14,724 | 7,080 | -52% | 1 | 1 | 0% | 2,249 | 2,791 | +24% | 0 | 0 | — |
case-01 | fail→fail | 11,294 | 9,174 | -19% | 1 | 1 | 0% | 1,944 | 3,218 | +66% | 0 | 0 | — |
case-02 | fail→fail | 15,714 | 6,942 | -56% | 1 | 1 | 0% | 2,528 | 2,915 | +15% | 0 | 0 | — |
case-03 | fail→fail | 5,575 | 13,575 | +143% | 1 | 1 | 0% | 1,089 | 2,262 | +108% | 0 | 0 | — |
case-04 | pass→fail | 12,310 | 7,170 | -42% | 1 | 1 | 0% | 2,097 | 2,783 | +33% | 0 | 0 | — |
case-05 | pass→pass | 13,073 | 8,506 | -35% | 1 | 1 | 0% | 2,285 | 3,118 | +36% | 0 | 0 | — |
case-06 | pass→pass | 9,472 | 5,458 | -42% | 1 | 1 | 0% | 1,791 | 2,776 | +55% | 0 | 0 | — |
case-07 | fail→pass | 9,645 | 1,865 | -81% | 1 | 1 | 0% | 1,583 | 1,949 | +23% | 0 | 0 | — |
case-08 | fail→pass | 10,224 | 4,087 | -60% | 1 | 1 | 0% | 1,673 | 2,168 | +30% | 0 | 0 | — |
case-09 | fail→pass | 5,505 | 1,961 | -64% | 1 | 1 | 0% | 795 | 2,023 | +154% | 0 | 0 | — |
case-10 | fail→fail | 6,105 | 2,895 | -53% | 1 | 1 | 0% | 993 | 2,187 | +120% | 0 | 0 | — |
case-11 | fail→pass | 9,976 | 1,588 | -84% | 1 | 1 | 0% | 1,605 | 1,961 | +22% | 0 | 0 | — |
case-12 | fail→pass | 22,713 | 5,065 | -78% | 1 | 1 | 0% | 1,745 | 2,442 | +40% | 0 | 0 | — |
case-14 | pass→pass | 11,437 | 3,858 | -66% | 1 | 1 | 0% | 1,883 | 2,485 | +32% | 0 | 0 | — |
case-15 | fail→pass | 8,167 | 2,647 | -68% | 1 | 1 | 0% | 1,262 | 2,141 | +70% | 0 | 0 | — |
case-16 | pass→pass | 8,667 | 3,115 | -64% | 1 | 1 | 0% | 1,461 | 2,210 | +51% | 0 | 0 | — |
case-17 | fail→pass | 9,007 | 2,257 | -75% | 1 | 1 | 0% | 1,531 | 2,051 | +34% | 0 | 0 | — |
case-18 | fail→fail | 7,308 | 1,693 | -77% | 1 | 1 | 0% | 1,088 | 1,958 | +80% | 0 | 0 | — |
case-19 | pass→pass | 8,627 | 2,601 | -70% | 1 | 1 | 0% | 1,320 | 2,071 | +57% | 0 | 0 | — |
case-20 | fail→pass | 9,548 | 3,500 | -63% | 1 | 1 | 0% | 1,486 | 2,340 | +57% | 0 | 0 | — |
case-21 | pass→pass | 14,128 | 5,101 | -64% | 1 | 1 | 0% | 2,209 | 2,474 | +12% | 0 | 0 | — |
case-22 | fail→pass | 5,426 | 2,874 | -47% | 1 | 1 | 0% | 795 | 2,174 | +173% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.