Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate audiobooks, podcasts, or educational audio content on demand. User provides an idea or topic, Claude AI writes a script, and ElevenLabs converts it to high-quality audio. Supports multiple formats (audiobook, podcast, educational), custom lengths, and voice effects. Use when asked to create audio content, make a podcast, generate an audiobook, or produce educational audio. Returns MP3 audio file via MEDIA token.
.claude/skills/sundial-org-audio-gen/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 126% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 150% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 428% | 0% |
Generate high-quality audiobooks, podcasts, or educational audio content on demand using AI-written scripts and ElevenLabs text-to-speech.
Create an audiobook chapter:
User: "Create a 5-minute audiobook chapter about a dragon discovering friendship"Generate a podcast:
User: "Make a 10-minute podcast about the history of coffee"Produce educational content:
User: "Generate a 15-minute educational audio explaining how neural networks work"Style: Narrative storytelling with emotional depth
[whispers], [excited], [serious] for impactExample Structure:
[Opening hook - set the scene]
[long pause]
[Story development with character emotions]
[short pause] between sentences
[long pause] between paragraphs
[Climax with dramatic tension]
[long pause]
[Resolution and emotional closure]Style: Conversational and engaging
Example Structure:
**Intro:** "Welcome to [topic]. I'm excited to share..."
[short pause]
**Main Content:** "Let's start with... [topic 1]"
[long pause] between segments
**Outro:** "Thanks for listening! Remember..."Style: Clear explanations for learning
[excited] for important pointsExample Structure:
**Introduction:** What is [topic] and why it matters?
**Main Content:**
- Concept 1: Explanation + Example
- Concept 2: Explanation + Example
- Concept 3: Explanation + Example
**Summary:** Key takeaways and next stepsWord Count to Duration Conversion:
Pacing: Average conversational speed is ~75 words per minute
Practical Limits:
Parse the user's request for:
target_words = target_minutes × 75Example: 10 minutes = 10 × 75 = 750 words
Write the complete script following these rules:
Content Guidelines:
Formatting Rules:
[short pause] after sentences (use sparingly, not every sentence)[long pause] between paragraphs or major sections[whispers], [shouts], [excited], [serious], [sarcastic], [sings], [laughs]Show the script to the user and ask:
Here's the [format] script I've created (approximately [length] minutes):
[Display the script]
Would you like me to:
1. Generate the audio now
2. Make changes to the script
3. Adjust the length or toneIf user requests changes:
If user approves:
Format the script for TTS:
[effect] formatInvoke the TTS script:
IMPORTANT: The ELEVENLABS_API_KEY environment variable is already configured in the system. Simply invoke the TTS script directly.
bashuv run /home/clawdbot/clawdbot/skills/sag/scripts/tts.py \ -o /tmp/audio-gen-[timestamp]-[topic-slug].mp3 \ -m eleven_multilingual_v2 \ "[formatted_script]"
For long scripts, use heredoc:
bashuv run /home/clawdbot/clawdbot/skills/sag/scripts/tts.py \ -o /tmp/audio-gen-[timestamp]-[topic-slug].mp3 \ -m eleven_multilingual_v2 \ "$(cat <<'EOF' [formatted_script] EOF )"
Return the result:
MEDIA:/tmp/audio-gen-[timestamp]-[topic-slug].mp3
Your [format] is ready! [Brief description of content]. Duration: approximately [X] minutes.Available voice modulation effects (use sparingly for impact):
[whispers] - Soft, intimate delivery[shouts] - Loud, emphatic delivery[excited] - Enthusiastic, energetic tone[serious] - Grave, solemn tone[sarcastic] - Ironic, mocking tone[sings] - Musical, melodic delivery[laughs] - Amused, jovial tone[short pause] - Brief silence (~0.5s)[long pause] - Extended silence (~1-2s)Best Practices:
If the generated script exceeds target by >20%:
The script I generated is [X] words ([Y] minutes), which is longer than your target of [Z] minutes. Would you like me to:
1. Condense it to fit the target length
2. Split it into multiple parts
3. Keep it as isIf the generated script is under target by >20%:
The script is [X] words ([Y] minutes), shorter than your target. Would you like me to:
1. Expand it with more detail
2. Add additional examples or stories
3. Generate as isIf the TTS script fails:
I've created the script, but I'm unable to generate the audio right now. Here's your script:
[Display script]
Error: [specific error message]
You can:
1. Check that ELEVENLABS_API_KEY is configured
2. Use the script with your own text-to-speech tool
3. Try again in a moment
4. Ask me to troubleshoot the audio generationCommon TTS Issues:
For unrealistic requests (e.g., "100-hour audiobook"):
That length would require [X] words and take significant time to generate. I recommend:
- Breaking it into multiple episodes/chapters
- Targeting 5-30 minutes per audio file
- Creating a series instead of one long fileTTS Implementation:
~/.clawdbot/clawdbot/skills/sag/scripts/tts.pyFile Storage:
/tmp/audio-gen/audio-gen-[timestamp]-[topic-slug].mp3API Requirements:
Supported Models:
eleven_multilingual_v2 - Best quality (default)eleven_turbo_v2 - Faster generationeleven_turbo_v2_5 - Fastest generationeleven_multilingual_v1 - Legacy modelCost Estimate:
Generation Time:
User: Create a 5-minute audiobook chapter about a lonely robot learning to feel emotions
Agent: I'll create a 5-minute sci-fi audiobook chapter. Let me write the story...
[Generates ~400-word narrative with emotional arc]
Here's the story I've created about Unit-7, a maintenance robot who discovers the meaning of friendship through an unexpected encounter. Would you like me to generate the audio or make changes first?
User: Generate the audio
Agent: [Invokes SAG with formatted script]
MEDIA:/tmp/audio-gen-1234567890-lonely-robot.mp3
Your audiobook chapter is ready! The story follows Unit-7's emotional awakening through a touching encounter with a young engineer. Duration: approximately 5 minutes.User: Make a 10-minute podcast about sustainable living tips
Agent: [Generates podcast script with intro, 5 practical tips, outro]
Here's a 10-minute podcast with 5 actionable sustainable living tips. The tone is encouraging and practical. Would you like me to:
1. Generate audio now
2. Make it more urgent/climate-focused
3. Add more specific examples
4. Adjust the length
User: Add more specific examples for each tip
Agent: [Regenerates with detailed examples]
Better? I've added specific examples like calculating your carbon footprint, composting at home, and choosing sustainable brands. Ready for audio?
User: Perfect!
Agent: MEDIA:/tmp/audio-gen-1234567890-sustainable-living.mp3
Your sustainable living podcast is ready with 5 practical tips, each with concrete examples you can start using today! Duration: 10 minutes.User: Create a 15-minute educational audio explaining quantum computing for beginners
Agent: I'll create a beginner-friendly 15-minute explanation of quantum computing...
[Generates ~1,100 words with analogies, examples, and clear structure]
Script preview:
- Introduction: What is quantum computing? (2 min)
- Classical vs Quantum: The key differences (4 min)
- Qubits and Superposition: How it works (5 min)
- Real Applications: What it means for you (3 min)
- Conclusion: The future of computing (1 min)
Should I generate the audio?
User: Yes please
Agent: MEDIA:/tmp/audio-gen-1234567890-quantum-computing.mp3
Your educational podcast on quantum computing is ready! It uses everyday analogies to explain qubits, superposition, and real-world applications. Duration: approximately 15 minutes.Use the audio-gen skill when users request:
Do NOT use for:
Users can request specific voices:
User: Create a podcast with the Rachel voice
Agent: [Checks voice configuration for "Rachel" voice ID and uses it in SAG command]For multi-part content:
User: Create a 3-episode series about space exploration
Agent: I'll create Episode 1 first. Each episode will be 10 minutes. Let's start with "The Dawn of the Space Age"...
[After completion]
Episode 1 is ready! Would you like me to continue with Episode 2?Mix formats for unique styles:
User: Create an educational podcast that tells a story
Agent: [Generates content that combines storytelling narrative with educational explanations]Issue: Audio sounds robotic or unnatural Solution: Add more pauses and voice effects. Use contractions and conversational language.
Issue: Script doesn't match requested length Solution: Regenerate with explicit word count target. Check calculations (75 words/min).
Issue: Content is too technical or too simple Solution: Ask user for target audience. Adjust complexity accordingly.
Issue: SAG command fails Solution: Check ELEVENLABS_API_KEY is set. Verify SAG skill is installed and working.
Issue: User wants to edit the script manually Solution: Provide the plain text script. User can modify it and paste back for audio generation.
💡 Pro Tip: Always generate the script first and get user approval before creating audio. This saves time and API costs, and ensures the user gets exactly what they want.
Other measured skills in the registry, with their headline benchmark lift.