Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
.claude/skills/davila7-speech/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 41% | 0% |
Generate spoken audio for the current project (narration, product demo voiceover, IVR prompts, accessibility reads). OpenAI remains the default with gpt-4o-mini-tts-2025-12-15; Atlas Cloud is available only when the user explicitly selects it. Prefer the bundled CLIs for deterministic, reproducible runs.
scripts/text_to_speech.py for the default OpenAI path, or scripts/atlas_text_to_speech.py only when Atlas Cloud was selected (see references/cli.md).tmp/speech/ for intermediate files (for example JSONL batches); delete when done.output/speech/ when working in this repo.--out or --out-dir to control output paths; keep filenames stable and descriptive.Prefer uv for dependency management.
OpenAI backend package:
uv pip install openaiIf uv is unavailable:
python3 -m pip install openaiThe Atlas Cloud backend uses only the Python standard library.
OPENAI_API_KEY.ATLASCLOUD_API_KEY.If the selected provider key is missing, give the user these steps:
OPENAI_API_KEY or ATLASCLOUD_API_KEY as an environment variable in their system.If installation isn't possible in this environment, tell the user which dependency is missing and how to install it locally.
gpt-4o-mini-tts-2025-12-15 unless the user requests another model.cedar. If the user wants a brighter tone, prefer marin.instructions are supported for GPT-4o mini TTS models, but not for tts-1 or tts-1-hd.--rpm at 50.OPENAI_API_KEY before any live API call.xai/tts-v1, voice eve, language auto, and require ATLASCLOUD_API_KEY.openai package) for default OpenAI calls; the dedicated Atlas CLI uses its asynchronous HTTP contract.Reformat user direction into a short, labeled spec. Only make implicit details explicit; do not invent new requirements.
Quick clarification (augmentation vs invention):
Template (include only relevant lines):
Voice Affect: <overall character and texture of the voice>
Tone: <attitude, formality, warmth>
Pacing: <slow, steady, brisk>
Emotion: <key emotions to convey>
Pronunciation: <words to enunciate or emphasize>
Pauses: <where to add intentional pauses>
Emphasis: <key words or phrases to stress>
Delivery: <cadence or rhythm notes>Augmentation rules:
Input text: "Welcome to the demo. Today we'll show how it works."
Instructions:
Voice Affect: Warm and composed.
Tone: Friendly and confident.
Pacing: Steady and moderate.
Emphasis: Stress "demo" and "show".{"input":"Thank you for calling. Please hold.","voice":"cedar","response_format":"mp3","out":"hold.mp3"}
{"input":"For sales, press 1. For support, press 2.","voice":"marin","instructions":"Tone: Clear and neutral. Pacing: Slow.","response_format":"wav"}More principles: references/prompting.md. Copy/paste specs: references/sample-prompts.md.
Use these modules when the request is for a specific delivery style. They provide targeted defaults and templates.
references/narration.mdreferences/voiceover.mdreferences/ivr.mdreferences/accessibility.mdreferences/cli.mdreferences/audio-api.mdreferences/atlas-cloud.mdreferences/voice-directions.mdreferences/codex-network.mdreferences/cli.md: how to run speech generation/batches via scripts/text_to_speech.py (commands, flags, recipes).references/audio-api.md: API parameters, limits, voice list.references/atlas-cloud.md: optional Atlas Cloud model, CLI, polling, and download contract.references/voice-directions.md: instruction patterns and examples.references/prompting.md: instruction best practices (structure, constraints, iteration patterns).references/sample-prompts.md: copy/paste instruction recipes (examples only; no extra theory).references/narration.md: templates + defaults for narration and explainers.references/voiceover.md: templates + defaults for product demo voiceovers.references/ivr.md: templates + defaults for IVR/phone prompts.references/accessibility.md: templates + defaults for accessibility reads.references/codex-network.md: environment/sandbox/network-approval troubleshooting.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,992 | 21,200 | +254% | 1 | 1 | 0% | 964 | 2,592 | +169% | 0 | 0 | — |
case-20 | pass→pass | 6,416 | 12,597 | +96% | 1 | 1 | 0% | 886 | 3,415 | +285% | 0 | 0 | — |
case-02 | fail→fail | 7,032 | 8,970 | +28% | 1 | 1 | 0% | 1,280 | 2,458 | +92% | 0 | 0 | — |
case-03 | fail→fail | 9,099 | 7,977 | -12% | 1 | 1 | 0% | 1,389 | 2,483 | +79% | 0 | 0 | — |
case-04 | fail→fail | 6,200 | 10,750 | +73% | 1 | 1 | 0% | 1,004 | 2,476 | +147% | 0 | 0 | — |
case-05 | fail→fail | 8,866 | 7,531 | -15% | 1 | 1 | 0% | 1,422 | 2,287 | +61% | 0 | 0 | — |
case-06 | fail→fail | 5,739 | 6,000 | +5% | 1 | 1 | 0% | 828 | 2,220 | +168% | 0 | 0 | — |
case-07 | fail→fail | 6,337 | 6,566 | +4% | 1 | 1 | 0% | 930 | 2,341 | +152% | 0 | 0 | — |
case-08 | fail→pass | 9,751 | 5,411 | -45% | 1 | 1 | 0% | 1,633 | 2,746 | +68% | 0 | 0 | — |
case-09 | fail→pass | 11,414 | 1,897 | -83% | 1 | 1 | 0% | 1,838 | 2,207 | +20% | 0 | 0 | — |
case-10 | pass→pass | 13,636 | 6,134 | -55% | 1 | 1 | 0% | 2,124 | 2,939 | +38% | 0 | 0 | — |
case-11 | fail→pass | 9,528 | 2,116 | -78% | 1 | 1 | 0% | 1,514 | 2,256 | +49% | 0 | 0 | — |
case-12 | pass→pass | 8,036 | 4,334 | -46% | 1 | 1 | 0% | 1,476 | 2,672 | +81% | 0 | 0 | — |
case-13 | pass→fail | 7,216 | 8,291 | +15% | 1 | 1 | 0% | 1,063 | 2,413 | +127% | 0 | 0 | — |
case-14 | pass→pass | 11,288 | 3,527 | -69% | 1 | 1 | 0% | 1,651 | 2,367 | +43% | 0 | 0 | — |
case-15 | fail→pass | 10,563 | 4,753 | -55% | 1 | 1 | 0% | 1,581 | 2,694 | +70% | 0 | 0 | — |
case-16 | fail→pass | 12,812 | 5,012 | -61% | 1 | 1 | 0% | 1,999 | 2,820 | +41% | 0 | 0 | — |
case-17 | pass→pass | 9,542 | 5,676 | -41% | 1 | 1 | 0% | 1,535 | 2,844 | +85% | 0 | 0 | — |
case-18 | pass→pass | 11,569 | 5,178 | -55% | 1 | 1 | 0% | 1,739 | 2,741 | +58% | 0 | 0 | — |
case-19 | fail→pass | 11,959 | 1,736 | -85% | 1 | 1 | 0% | 1,935 | 2,154 | +11% | 0 | 0 | — |
case-21 | pass→pass | 8,407 | 3,295 | -61% | 1 | 1 | 0% | 1,212 | 2,475 | +104% | 0 | 0 | — |
case-22 | fail→pass | 9,877 | 7,697 | -22% | 1 | 1 | 0% | 1,548 | 3,183 | +106% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 14 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 14 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/12/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.