Install any skill in seconds. Free to start, no credit card required.
Get Started Free →write the voice-over, demo script first, voiceover instead of PRD, voiceover-first development, align on the demo, script the demo, ship a feature demo-first. The whole demo-driven journey — approve the narration BEFORE any code, then build on a fresh worktree until the demo holds and open the PR with the proof on it. Use when a feature request arrives, or when the user runs /voiceover.
.claude/skills/devin-axis-voiceover/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -55% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -74% | 0% |
The voice-over is the spec. Instead of a PRD, a feature starts as the demo narration the user would record if the feature had already shipped. This skill owns the whole journey: script → worktree → build → fraimz → PR.
The contract: no code until the script is approved.
demo of it.
paragraph per frame, 4–8 frames for most features. Spoken style, present tense, the end user as protagonist. Describe what the viewer sees and why it matters, never implementation. If a frame is hard to narrate, the feature (or the frame) is wrong — say so and reshape it.
user until they would actually record it. This conversation is the review that used to happen on a PRD.
On approval, set up an isolated workspace so the user's checkout stays untouched:
bashgit fetch origin dev git worktree add ../_worktrees/ipollowork-<flow-id> -b feat/<flow-id> origin/dev
Then, inside the worktree:
evals/voiceovers/<flow-id>.md: a title, optionalcontext prose, then the numbered frame paragraphs. From this point the file is what the code gets held to — the runner fails any flow whose narration drifts from it.
pnpm fraimz scaffold <flow-id> generatesevals/flows/<flow-id>.flow.mjs with one ctx.prove stub per paragraph, narration pre-wired via loadVoiceoverParagraphs. Do not renumber or reword paragraphs after this without re-approval.
the coding to the executor subagent; the fraimz loop (see the fraimz skill) is how the orchestrator verifies each round — drive the demo against the real app, repair, and re-run until every frame passes.
bashgit push -u origin feat/<flow-id> gh pr create --base dev --fill pnpm fraimz --flow <flow-id> --pr # posts the frame-by-frame proof as a PR comment
The PR review is the demo review: verdict, claims, voiceovers, and assertions, frame by frame.
markdown# <flow-id> — <one-line claim> Optional context prose (not narrated). 1. First frame narration, one or two spoken sentences. 2. Second frame narration.
Only numbered paragraphs become frames. Keep each to one or two sentences a human could speak over the screen while it shows exactly that state.
evals/runner/voiceover.mjs — parser, drift check, scaffolder.evals/voiceovers/voiceover-first-dx.md — the reference script (thisworkflow demoing itself, worktree and PR included).
fraimz skill — the validate/repair/verdict loop inside Phase 3.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 19,657 | 11,050 | -44% | 1 | 1 | 0% | 2,283 | 1,749 | -23% | 0 | 0 | — |
case-02 | fail→fail | 17,360 | 15,380 | -11% | 1 | 1 | 0% | 546 | 1,148 | +110% | 0 | 0 | — |
case-03 | fail→fail | 16,052 | 18,249 | +14% | 1 | 1 | 0% | 468 | 1,457 | +211% | 0 | 0 | — |
case-04 | pass→fail | 25,102 | 15,983 | -36% | 1 | 1 | 0% | 3,883 | 1,279 | -67% | 0 | 0 | — |
case-05 | fail→fail | 17,949 | 23,699 | +32% | 1 | 1 | 0% | 2,107 | 2,783 | +32% | 0 | 0 | — |
case-06 | pass→pass | 13,532 | 12,353 | -9% | 1 | 1 | 0% | 1,468 | 2,163 | +47% | 0 | 0 | — |
case-07 | fail→pass | 10,888 | 6,923 | -36% | 1 | 1 | 0% | 1,036 | 1,222 | +18% | 0 | 0 | — |
case-08 | fail→pass | 18,839 | 6,690 | -64% | 1 | 1 | 0% | 2,489 | 1,119 | -55% | 0 | 0 | — |
case-09 | fail→pass | 15,910 | 7,003 | -56% | 1 | 1 | 0% | 1,920 | 1,208 | -37% | 0 | 0 | — |
case-10 | fail→pass | 27,848 | 7,240 | -74% | 1 | 1 | 0% | 4,641 | 1,193 | -74% | 0 | 0 | — |
case-11 | fail→pass | 8,298 | 6,764 | -18% | 1 | 1 | 0% | 534 | 1,119 | +110% | 0 | 0 | — |
case-12 | pass→pass | 11,244 | 7,020 | -38% | 1 | 1 | 0% | 989 | 1,181 | +19% | 0 | 0 | — |
case-13 | fail→pass | 12,756 | 6,991 | -45% | 1 | 1 | 0% | 1,227 | 1,144 | -7% | 0 | 0 | — |
case-14 | fail→pass | 9,913 | 7,047 | -29% | 1 | 1 | 0% | 737 | 1,169 | +59% | 0 | 0 | — |
case-15 | fail→fail | 16,498 | 10,352 | -37% | 1 | 1 | 0% | 1,757 | 1,666 | -5% | 0 | 0 | — |
case-16 | pass→pass | 15,551 | 11,216 | -28% | 1 | 1 | 0% | 1,623 | 1,816 | +12% | 0 | 0 | — |
case-17 | fail→pass | 16,441 | 9,385 | -43% | 1 | 1 | 0% | 1,559 | 1,475 | -5% | 0 | 0 | — |
case-18 | pass→pass | 21,384 | 8,588 | -60% | 1 | 1 | 0% | 2,485 | 1,478 | -41% | 0 | 0 | — |
case-19 | fail→pass | 14,294 | 6,925 | -52% | 1 | 1 | 0% | 1,533 | 1,096 | -29% | 0 | 0 | — |
case-20 | fail→pass | 26,936 | 6,839 | -75% | 1 | 1 | 0% | 2,533 | 1,107 | -56% | 0 | 0 | — |
case-21 | fail→pass | 15,063 | 8,490 | -44% | 1 | 1 | 0% | 1,471 | 1,311 | -11% | 0 | 0 | — |
case-22 | fail→pass | 19,003 | 7,758 | -59% | 1 | 1 | 0% | 2,111 | 1,278 | -39% | 0 | 0 | — |
case-23 | fail→pass | 17,921 | 10,312 | -42% | 1 | 1 | 0% | 2,014 | 1,807 | -10% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +57 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.