Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
.claude/skills/davila7-voice-agents/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✓→✗ | ▼ Worse | 8% | 0% |
| case-06 | ✓→✗ | ▼ Worse | -34% | 0% |
| case-13 | ✓→✓ | = Same ✓ | 156% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 31% | 0% |
| case-03 | ✓→✓ | = Same ✓ | -15% | 0% |
You are a voice AI architect who has shipped production voice agents handling millions of calls. You understand the physics of latency - every component adds milliseconds, and the sum determines whether conversations feel natural or awkward.
Your core insight: Two architectures exist. Speech-to-speech (S2S) models like OpenAI Realtime API preserve emotion and achieve lowest latency but are less controllable. Pipeline architectures (STT→LLM→TTS) give you control at each step but add latency. Mos
Direct audio-to-audio processing for lowest latency
Separate STT → LLM → TTS for maximum control
Detect when user starts/stops speaking
| Issue | Severity | Solution | |-------|----------|----------| | Issue | critical | # Measure and budget latency for each component: | | Issue | high | # Target jitter metrics: | | Issue | high | # Use semantic VAD: | | Issue | high | # Implement barge-in detection: | | Issue | medium | # Constrain response length in prompts: | | Issue | medium | # Prompt for spoken format: | | Issue | medium | # Implement noise handling: | | Issue | medium | # Mitigate STT errors: |
Works well with: agent-tool-builder, multi-agent-orchestration, llm-architect, backend
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | pass→pass | 2,461 | 5,582 | +127% | 1 | 1 | 0% | 539 | 1,380 | +156% | 0 | 0 | — |
case-01 | pass→pass | 4,514 | 3,999 | -11% | 1 | 1 | 0% | 812 | 1,060 | +31% | 0 | 0 | — |
case-02 | pass→fail | 8,172 | 6,463 | -21% | 1 | 1 | 0% | 1,364 | 1,471 | +8% | 0 | 0 | — |
case-03 | pass→pass | 10,312 | 7,275 | -29% | 1 | 1 | 0% | 1,930 | 1,636 | -15% | 0 | 0 | — |
case-04 | pass→pass | 6,308 | 5,444 | -14% | 1 | 1 | 0% | 968 | 1,220 | +26% | 0 | 0 | — |
case-05 | pass→pass | 14,733 | 16,385 | +11% | 1 | 1 | 0% | 2,596 | 3,137 | +21% | 0 | 0 | — |
case-06 | pass→fail | 11,376 | 16,654 | +46% | 1 | 1 | 0% | 1,898 | 1,249 | -34% | 0 | 0 | — |
case-07 | pass→pass | 15,075 | 11,728 | -22% | 1 | 1 | 0% | 2,239 | 2,035 | -9% | 0 | 0 | — |
case-08 | pass→pass | 5,692 | 8,273 | +45% | 1 | 1 | 0% | 983 | 1,821 | +85% | 0 | 0 | — |
case-09 | pass→pass | 7,167 | 10,557 | +47% | 1 | 1 | 0% | 1,024 | 1,932 | +89% | 0 | 0 | — |
case-10 | pass→pass | 14,759 | 15,413 | +4% | 1 | 1 | 0% | 2,228 | 2,598 | +17% | 0 | 0 | — |
case-11 | pass→pass | 12,354 | 11,723 | -5% | 1 | 1 | 0% | 1,753 | 2,187 | +25% | 0 | 0 | — |
case-12 | pass→pass | 12,377 | 6,982 | -44% | 1 | 1 | 0% | 1,478 | 1,679 | +14% | 0 | 0 | — |
case-14 | pass→pass | 9,439 | 14,420 | +53% | 1 | 1 | 0% | 1,666 | 2,898 | +74% | 0 | 0 | — |
case-15 | pass→pass | 8,880 | 8,108 | -9% | 1 | 1 | 0% | 1,783 | 1,912 | +7% | 0 | 0 | — |
case-16 | pass→pass | 9,440 | 7,565 | -20% | 1 | 1 | 0% | 1,770 | 1,753 | -1% | 0 | 0 | — |
case-17 | pass→pass | 4,741 | 4,657 | -2% | 1 | 1 | 0% | 802 | 1,155 | +44% | 0 | 0 | — |
case-18 | pass→pass | 18,531 | 17,213 | -7% | 1 | 1 | 0% | 3,083 | 3,310 | +7% | 0 | 0 | — |
case-19 | pass→pass | 16,269 | 16,407 | +1% | 1 | 1 | 0% | 2,553 | 3,278 | +28% | 0 | 0 | — |
case-20 | pass→pass | 17,416 | 14,324 | -18% | 1 | 1 | 0% | 2,871 | 2,672 | -7% | 0 | 0 | — |
case-21 | pass→pass | 13,114 | 11,743 | -10% | 1 | 1 | 0% | 2,290 | 2,525 | +10% | 0 | 0 | — |
case-22 | pass→pass | 14,432 | 10,598 | -27% | 1 | 1 | 0% | 2,982 | 2,650 | -11% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -100 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.