Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design a voice AI agent for phone or in-app conversations — call flows, interruption handling, escalation to humans, and the metrics that catch a bad voice experience. Use when asked to design a voice agent, automate a phone line, spec an IVR replacement, or review why callers hate an existing voice bot. Produces a voice agent spec: persona and disclosure policy, conversation architecture, barge-in and repair behaviour, human-handoff rules, and a launch scorecard.
.claude/skills/mohitagw15856-voice-agent-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -34% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 67% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 31% | 0% |
Voice is the least forgiving agent surface: no screen to fall back on, dead air reads as failure within two seconds, and the caller is often already annoyed. This skill designs voice agents around the medium's real constraints — turn-taking, interruption, repair — instead of shipping a chatbot with a text-to-speech voice.
Ask for (if not already provided):
Intent scope | Intent | Volume | Own / Triage / Pass | Why | |---|---|---|---|
Opening script: verbatim — disclosure, capability, escape hatch]
Conversation architecture: turn rules · confirmation strategy by stakes · the repair ladder (rephrase → options → keypad → human)]
Mechanics: barge-in behaviour · latency masking thresholds · silence handling]
Handoff: triggers · whisper-summary fields · after-hours behaviour]
Compliance: disclosure line · recording consent flow · statements the agent must never make]
Launch scorecard | Metric | Gate | Why paired | |---|---|---| | Containment + caller-scored success | | containment alone is gameable | | Escape-request rate | | measures trapped callers | | Repair rate / hang-ups mid-flow | | frustration signals |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 53,196 | 43,892 | -17% | 1 | 1 | 0% | 8,343 | 6,842 | -18% | 0 | 0 | — |
case-02 | fail→pass | 51,594 | 40,450 | -22% | 1 | 1 | 0% | 7,873 | 7,055 | -10% | 0 | 0 | — |
case-03 | fail→pass | 54,447 | 28,597 | -47% | 1 | 1 | 0% | 7,867 | 5,221 | -34% | 0 | 0 | — |
case-08 | pass→pass | 21,206 | 21,911 | +3% | 1 | 1 | 0% | 2,392 | 3,796 | +59% | 0 | 0 | — |
case-04 | pass→pass | 21,319 | 18,844 | -12% | 1 | 1 | 0% | 2,214 | 3,168 | +43% | 0 | 0 | — |
case-05 | pass→pass | 14,157 | 15,058 | +6% | 1 | 1 | 0% | 2,088 | 3,355 | +61% | 0 | 0 | — |
case-06 | pass→pass | 12,739 | 16,698 | +31% | 1 | 1 | 0% | 2,531 | 3,933 | +55% | 0 | 0 | — |
case-07 | pass→fail | 14,939 | 24,951 | +67% | 1 | 1 | 0% | 2,700 | 5,021 | +86% | 0 | 0 | — |
case-09 | pass→pass | 20,016 | 29,100 | +45% | 1 | 1 | 0% | 2,325 | 3,532 | +52% | 0 | 0 | — |
case-10 | pass→pass | 30,013 | 17,528 | -42% | 1 | 1 | 0% | 1,991 | 3,146 | +58% | 0 | 0 | — |
case-11 | fail→fail | 23,473 | 21,015 | -10% | 1 | 1 | 0% | 2,201 | 3,619 | +64% | 0 | 0 | — |
case-12 | fail→pass | 31,014 | 27,461 | -11% | 1 | 1 | 0% | 2,284 | 4,210 | +84% | 0 | 0 | — |
case-13 | fail→fail | 16,733 | 13,997 | -16% | 1 | 1 | 0% | 1,793 | 3,048 | +70% | 0 | 0 | — |
case-18 | pass→pass | 26,978 | 19,026 | -29% | 1 | 1 | 0% | 2,150 | 3,418 | +59% | 0 | 0 | — |
case-14 | fail→fail | 18,268 | 18,095 | -1% | 1 | 1 | 0% | 1,994 | 3,222 | +62% | 0 | 0 | — |
case-15 | pass→pass | 19,245 | 18,197 | -5% | 1 | 1 | 0% | 2,076 | 3,258 | +57% | 0 | 0 | — |
case-16 | pass→pass | 21,564 | 22,638 | +5% | 1 | 1 | 0% | 2,411 | 4,023 | +67% | 0 | 0 | — |
case-17 | pass→pass | 17,288 | 22,215 | +28% | 1 | 1 | 0% | 1,602 | 3,191 | +99% | 0 | 0 | — |
case-19 | pass→pass | 13,879 | 14,819 | +7% | 1 | 1 | 0% | 1,840 | 2,589 | +41% | 0 | 0 | — |
case-20 | fail→pass | 15,190 | 17,501 | +15% | 1 | 1 | 0% | 1,836 | 3,069 | +67% | 0 | 0 | — |
case-21 | fail→pass | 22,479 | 13,571 | -40% | 1 | 1 | 0% | 2,925 | 3,820 | +31% | 0 | 0 | — |
case-22 | pass→pass | 18,533 | 16,270 | -12% | 1 | 1 | 0% | 1,975 | 2,741 | +39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.