Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generate_voice.js CLI. Use before calling generate_voice.js; when the user asks to give a character a voice, preview how a character sounds, create reusable timbre anchors for every speaking character or VO/narration, or create exact narration/VO/final line audio.
.claude/skills/utopai-research-voice-compose/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -28% | 0% |
Default to one short reusable timbre sample per speaking character and one VO/narrator sample when narration exists. video-compose keeps actual shot dialogue/VO in the video prompt. Treat audio_result.data.text as downstream speech only for approved final narration/line reads.
Follow PROJECT_AGENT.md for context/staging. This skill owns voice prompt + CLI shape.
Triggers: "give / design a voice for character]", "what does character] sound like", "voices for all the characters on the canvas".
image_result for the person; don't gate on subtype. Read data.local_path before prompt; layer name/role/description on top. node "$PAI_REPO_ROOT/server/cli/generate_voice.js" \ --text "<line>" \ --prompt "<voice design brief>" \ --source-node-id <character.id>
> [age bracket] [gender], [timbre], [register], [pace], [accent if relevant]. [optional emotional color].
✅ "Mid-50s man, gravelly baritone, measured pace, slight rasp from decades of smoking, weary but steady." ✅ "Young woman, bright mezzo, warm, quick and percussive. Slight Southern lilt." ❌ "Detective Morris's voice." — names the character, not the voice. The model needs sound qualities.
text: 1-3 sentence in-character sample (≤200 chars), not every script line.Triggers: narrator voice, voice-over, "a voice that says X" without character, narration track, or explicit final line-read audio.
--source-node-id: node "$PAI_REPO_ROOT/server/cli/generate_voice.js" \ --text "<the narration line>" \ --prompt "<voice design brief>"
--text; then data.text is source of truth.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 8,370 | 3,113 | -63% | 1 | 1 | 0% | 1,380 | 1,155 | -16% | 0 | 0 | — |
case-02 | fail→pass | 9,493 | 5,577 | -41% | 1 | 1 | 0% | 1,610 | 1,650 | +2% | 0 | 0 | — |
case-03 | fail→pass | 9,641 | 3,353 | -65% | 1 | 1 | 0% | 1,642 | 1,195 | -27% | 0 | 0 | — |
case-04 | fail→pass | 11,780 | 5,734 | -51% | 1 | 1 | 0% | 1,821 | 1,386 | -24% | 0 | 0 | — |
case-05 | fail→pass | 9,010 | 3,024 | -66% | 1 | 1 | 0% | 1,464 | 1,054 | -28% | 0 | 0 | — |
case-06 | fail→pass | 8,003 | 6,454 | -19% | 1 | 1 | 0% | 1,307 | 1,682 | +29% | 0 | 0 | — |
case-07 | pass→pass | 7,718 | 3,831 | -50% | 1 | 1 | 0% | 1,150 | 1,290 | +12% | 0 | 0 | — |
case-08 | fail→pass | 6,196 | 5,069 | -18% | 1 | 1 | 0% | 1,070 | 1,459 | +36% | 0 | 0 | — |
case-09 | fail→pass | 6,480 | 2,712 | -58% | 1 | 1 | 0% | 1,074 | 988 | -8% | 0 | 0 | — |
case-10 | pass→pass | 2,027 | 2,012 | -1% | 1 | 1 | 0% | 306 | 878 | +187% | 0 | 0 | — |
case-11 | pass→pass | 11,122 | 3,546 | -68% | 1 | 1 | 0% | 1,637 | 1,139 | -30% | 0 | 0 | — |
case-12 | pass→pass | 11,664 | 6,017 | -48% | 1 | 1 | 0% | 1,725 | 1,577 | -9% | 0 | 0 | — |
case-13 | pass→pass | 8,416 | 4,115 | -51% | 1 | 1 | 0% | 1,772 | 1,302 | -27% | 0 | 0 | — |
case-14 | fail→pass | 12,824 | 4,299 | -66% | 1 | 1 | 0% | 2,064 | 1,309 | -37% | 0 | 0 | — |
case-15 | fail→pass | 9,968 | 2,025 | -80% | 1 | 1 | 0% | 1,398 | 894 | -36% | 0 | 0 | — |
case-16 | fail→pass | 12,545 | 3,040 | -76% | 1 | 1 | 0% | 2,066 | 1,035 | -50% | 0 | 0 | — |
case-17 | fail→pass | 8,434 | 2,267 | -73% | 1 | 1 | 0% | 1,410 | 882 | -37% | 0 | 0 | — |
case-18 | pass→pass | 9,139 | 2,622 | -71% | 1 | 1 | 0% | 1,565 | 926 | -41% | 0 | 0 | — |
case-19 | fail→pass | 4,122 | 1,237 | -70% | 1 | 1 | 0% | 616 | 709 | +15% | 0 | 0 | — |
case-20 | pass→pass | 8,769 | 1,807 | -79% | 1 | 1 | 0% | 1,460 | 721 | -51% | 0 | 0 | — |
case-21 | fail→pass | 9,860 | 1,584 | -84% | 1 | 1 | 0% | 1,474 | 757 | -49% | 0 | 0 | — |
case-22 | pass→pass | 9,693 | 2,575 | -73% | 1 | 1 | 0% | 1,535 | 967 | -37% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +64 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.