Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate real audio narratives from text content using Azure OpenAI's Realtime API.
.claude/skills/sickn33-podcast-generation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -14% | 0% |
Generate real audio narratives from text content using Azure OpenAI's Realtime API.
envAZURE_OPENAI_AUDIO_API_KEY=your_realtime_api_key AZURE_OPENAI_AUDIO_ENDPOINT=https://your-resource.cognitiveservices.azure.com AZURE_OPENAI_AUDIO_DEPLOYMENT=gpt-realtime-mini
Note: Endpoint should NOT include /openai/v1/ - just the base URL.
pythonfrom openai import AsyncOpenAI import base64 # Convert HTTPS endpoint to WebSocket URL ws_url = endpoint.replace("https://", "wss://") + "/openai/v1" client = AsyncOpenAI( websocket_base_url=ws_url, api_key=api_key ) audio_chunks = [] transcript_parts = [] async with client.realtime.connect(model="gpt-realtime-mini") as conn: # Configure for audio-only output await conn.session.update(session={ "output_modalities": ["audio"], "instructions": "You are a narrator. Speak naturally." }) # Send text to narrate await conn.conversation.item.create(item={ "type": "message", "role": "user", "content": [{"type": "input_text", "text": prompt}] }) await conn.response.create() # Collect streaming events async for event in conn: if event.type == "response.output_audio.delta": audio_chunks.append(base64.b64decode(event.delta)) elif event.type == "response.output_audio_transcript.delta": transcript_parts.append(event.delta) elif event.type == "response.done": break # Convert PCM to WAV (see scripts/pcm_to_wav.py) pcm_audio = b''.join(audio_chunks) wav_audio = pcm_to_wav(pcm_audio, sample_rate=24000)
javascript// Convert base64 WAV to playable blob const base64ToBlob = (base64, mimeType) => { const bytes = atob(base64); const arr = new Uint8Array(bytes.length); for (let i = 0; i < bytes.length; i++) arr[i] = bytes.charCodeAt(i); return new Blob([arr], { type: mimeType }); }; const audioBlob = base64ToBlob(response.audio_data, 'audio/wav'); const audioUrl = URL.createObjectURL(audioBlob); new Audio(audioUrl).play();
| Voice | Character | |-------|-----------| | alloy | Neutral | | echo | Warm | | fable | Expressive | | onyx | Deep | | nova | Friendly | | shimmer | Clear |
response.output_audio.delta - Base64 audio chunkresponse.output_audio_transcript.delta - Transcript textresponse.done - Generation completeerror - Handle with event.error.messageThis skill is applicable to execute the workflow or actions described in the overview.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 17,170 | 15,117 | -12% | 1 | 1 | 0% | 3,705 | 4,362 | +18% | 0 | 0 | — |
case-02 | fail→pass | 19,998 | 11,625 | -42% | 1 | 1 | 0% | 4,251 | 3,704 | -13% | 0 | 0 | — |
case-03 | fail→pass | 19,564 | 13,308 | -32% | 1 | 1 | 0% | 4,172 | 4,091 | -2% | 0 | 0 | — |
case-04 | fail→pass | 15,357 | 2,345 | -85% | 1 | 1 | 0% | 1,558 | 1,428 | -8% | 0 | 0 | — |
case-05 | fail→pass | 8,896 | 3,021 | -66% | 1 | 1 | 0% | 1,881 | 1,616 | -14% | 0 | 0 | — |
case-06 | pass→pass | 4,891 | 2,702 | -45% | 1 | 1 | 0% | 858 | 1,515 | +77% | 0 | 0 | — |
case-07 | pass→pass | 10,553 | 4,655 | -56% | 1 | 1 | 0% | 1,967 | 1,874 | -5% | 0 | 0 | — |
case-08 | pass→pass | 7,475 | 3,224 | -57% | 1 | 1 | 0% | 1,320 | 1,580 | +20% | 0 | 0 | — |
case-09 | fail→fail | 4,153 | 1,671 | -60% | 1 | 1 | 0% | 809 | 1,338 | +65% | 0 | 0 | — |
case-14 | pass→pass | 6,487 | 1,030 | -84% | 1 | 1 | 0% | 1,148 | 1,171 | +2% | 0 | 0 | — |
case-10 | fail→pass | 3,837 | 1,556 | -59% | 1 | 1 | 0% | 705 | 1,276 | +81% | 0 | 0 | — |
case-11 | pass→fail | 6,198 | 2,082 | -66% | 1 | 1 | 0% | 1,116 | 1,355 | +21% | 0 | 0 | — |
case-12 | pass→fail | 3,126 | 1,408 | -55% | 1 | 1 | 0% | 492 | 1,257 | +155% | 0 | 0 | — |
case-13 | fail→pass | 10,785 | 2,110 | -80% | 1 | 1 | 0% | 1,786 | 1,335 | -25% | 0 | 0 | — |
case-15 | pass→pass | 7,612 | 1,501 | -80% | 1 | 1 | 0% | 1,224 | 1,233 | +1% | 0 | 0 | — |
case-16 | fail→pass | 7,817 | 1,580 | -80% | 1 | 1 | 0% | 1,276 | 1,196 | -6% | 0 | 0 | — |
case-17 | pass→pass | 4,928 | 5,198 | +5% | 1 | 1 | 0% | 832 | 2,036 | +145% | 0 | 0 | — |
case-18 | pass→pass | 4,412 | 2,746 | -38% | 1 | 1 | 0% | 597 | 1,487 | +149% | 0 | 0 | — |
case-19 | pass→pass | 5,336 | 2,611 | -51% | 1 | 1 | 0% | 1,085 | 1,468 | +35% | 0 | 0 | — |
case-20 | pass→pass | 8,052 | 5,148 | -36% | 1 | 1 | 0% | 1,503 | 2,011 | +34% | 0 | 0 | — |
case-21 | pass→pass | 16,505 | 12,245 | -26% | 1 | 1 | 0% | 3,234 | 3,277 | +1% | 0 | 0 | — |
case-22 | pass→pass | 17,562 | 14,256 | -19% | 1 | 1 | 0% | 3,443 | 3,532 | +3% | 0 | 0 | — |
case-23 | pass→pass | 13,342 | 8,642 | -35% | 1 | 1 | 0% | 2,554 | 2,872 | +12% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +26 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.