Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Convert written medical content into podcast or video scripts optimized for audio delivery. Transforms academic papers, reports, and educational materials into engaging spoken-word formats with pronunciation guides, timing markers, and audio-friendly structure.
.claude/skills/leoyeai-audio-script-writer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 174% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 143% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 893% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 581% | 0% |
Content transformation tool that converts written medical and scientific materials into professionally structured audio scripts suitable for podcasts, educational videos, audiobooks, and voiceover narration.
Key Capabilities:
✅ Use this skill when:
❌ Do NOT use when:
Integration:
abstract-summarizer (content condensation), lay-summary-gen (patient-friendly language)medical-translation (multi-language scripts), voice-cloning-tool (AI narration)Convert written text to conversational audio style:
pythonfrom scripts.audio_writer import AudioScriptWriter writer = AudioScriptWriter() # Transform written content script = writer.convert_to_audio( source_text=research_paper, format="podcast", # podcast, video, lecture, audiobook target_audience="medical_students", duration_minutes=15 ) print(script.spoken_text) # Converts: "The pathophysiology of diabetes mellitus involves..." # To: "So what exactly happens in diabetes? Well, it all starts when..."
Transformation Rules: | Written Style | Audio Style | Example | |---------------|-------------|---------| | "Furthermore" | "Plus" | Less formal transitions | | " et al." | "and their colleagues" | Expand abbreviations | | Numbers in text | Spoken numbers | "15%" → "15 percent" | | Long sentences | 15-20 word max | Break into digestible chunks | | Passive voice | Active voice | "was observed" → "we saw" | | Citations | Omit or footnote | "(Smith et al., 2024)" → reference tone] |
Create phonetic spelling for medical terms:
python# Generate pronunciation guide pronunciation = writer.create_pronunciation_guide( text=script, include_phonetic=True, include_syllables=True ) # Output: # "Hyperlipidemia: hi-per-lip-i-DEE-mee-uh" # "Metformin: met-FOR-min" # "Atherosclerosis: ath-er-oh-skleh-ROH-sis"
Guide Elements:
Calculate speaking duration and mark pacing cues:
python# Analyze timing timing = writer.calculate_timing( script=script, speaking_rate="conversational", # slow, conversational, fast include_pauses=True ) print(f"Estimated duration: {timing.duration_minutes} minutes") print(f"Word count: {timing.word_count}") print(f"Pace: {timing.words_per_minute} WPM")
Speaking Rates: | Style | WPM | Use Case | |-------|-----|----------| | Slow/Educational | 120-130 | Patient education, complex topics | | Conversational | 140-160 | Podcasts, general audience | | Fast/News | 170-190 | Time-constrained content | | Variable | Varies | Dynamic pacing with pauses |
Pacing Cues:
[BREATHE] - Brief pause for narrator
[PAUSE 2s] - Two-second pause for emphasis
[SLOW DOWN] - Reduce pace for key point
[SPEED UP] - Increase energy/excitement
[BEAT] - Dramatic pauseGenerate scripts for different audio formats:
python# Podcast episode podcast = writer.create_podcast_script( content=article, episode_format="interview", # solo, interview, panel include_intro_music=True, ad_breaks=[5, 12] # minutes ) # Educational video video = writer.create_video_script( content=lecture_slides, visual_cues=True, # Mark where visuals change b_roll_notes=True # Suggest supplemental footage )
Format Types: | Format | Characteristics | Best For | |--------|-----------------|----------| | Podcast | Conversational, segments, ads | Long-form content, interviews | | Video | Visual cues, B-roll notes | YouTube, educational platforms | | Lecture | Structured, Q&A breaks | Online courses, training | | Audiobook | Chapter markers, consistent tone | Textbooks, memoirs | | News | Tight, factual, quick | Research briefs, updates |
Scenario: Convert published study to 15-minute podcast episode.
bash# Convert paper to podcast script python scripts/main.py \ --input paper.pdf \ --format podcast \ --duration 15 \ --style conversational \ --include-intro-outro \ --output podcast_script.txt # Generate pronunciation guide python scripts/main.py \ --input podcast_script.txt \ --generate-pronunciation \ --output pronunciation_guide.txt
Structure:
[INTRO MUSIC 5s]
HOST: Welcome to Medical Research Today. I'm your host...
[BREATHE]
HOST: Today we're diving into a fascinating study about...
[PAUSE]
HOST: So what did the researchers find? Well...
[BREATHE]
HOST: Dr. Smith, one of the study authors, explains...
[SOUND BITE: Interview clip]
...
[OUTRO MUSIC]Scenario: Convert lecture notes to video script for online course.
python# Create lecture script lecture = writer.create_lecture_script( notes=lecture_content, duration=45, # minutes break_intervals=[15, 30], # minutes for student breaks interaction_points=True # "Pause and think..." prompts ) # Add visual cues script = writer.add_visual_cues( script=lecture, slide_transitions=True, animation_notes=True )
Lecture Elements:
Scenario: Create audio guide for diabetes management.
python# Patient-friendly script patient_script = writer.create_patient_script( medical_content=diabetes_guide, reading_level=6, # 6th grade empathetic_tone=True, key_points_highlighted=True ) # Slow, clear pacing patient_script.adjust_pacing( wpm=130, pause_after_sentences=1.5 # seconds )
Patient Script Features:
Scenario: Adapt live presentation to YouTube video format.
bash# Convert presentation script python scripts/main.py \ --input presentation_transcript.txt \ --format video \ --platform youtube \ --include-hooks true \ --engagement-cues true \ --output youtube_script.txt
YouTube Optimization:
From research paper to published podcast:
bash# Step 1: Extract and summarize content python scripts/main.py \ --input paper.pdf \ --extract-key-points \ --output key_points.txt # Step 2: Convert to audio script python scripts/main.py \ --input key_points.txt \ --format podcast \ --duration 20 \ --output raw_script.txt # Step 3: Add production elements python scripts/main.py \ --input raw_script.txt \ --add-music-cues \ --add-sound-effects \ --add-pacing-marks \ --output production_script.txt # Step 4: Generate pronunciation guide python scripts/main.py \ --input production_script.txt \ --generate-pronunciation \ --output pronunciations.txt # Step 5: Create timing breakdown python scripts/main.py \ --input production_script.txt \ --calculate-timing \ --output timing_breakdown.txt
Python API:
pythonfrom scripts.audio_writer import AudioScriptWriter from scripts.pronunciation import PronunciationGuide from scripts.timing import TimingCalculator # Initialize writer = AudioScriptWriter() pronouncer = PronunciationGuide() timing = TimingCalculator() # Read source material with open("research_article.txt", "r") as f: content = f.read() # Step 1: Convert to spoken format script = writer.convert_to_audio( text=content, format="podcast", target_duration=15, # minutes audience="general_medical" ) # Step 2: Add production elements script_with_cues = writer.add_production_cues( script=script, music_stings=True, transition_effects=True ) # Step 3: Generate pronunciation guide medical_terms = pronouncer.extract_terms(script_with_cues) pronunciation_guide = pronouncer.create_guide(medical_terms) # Step 4: Calculate timing timing_analysis = timing.calculate( script=script_with_cues, speaking_rate=150 # WPM ) # Export complete production package writer.export_production_package( script=script_with_cues, pronunciation=pronunciation_guide, timing=timing_analysis, output_dir="podcast_production/" )
Content Quality:
Audio Optimization:
Production Quality:
Before Recording:
Content Issues:
Audio Issues:
Production Issues:
Available in references/ directory:
audio_writing_best_practices.md - Broadcast writing guidelinesmedical_pronunciation_guide.md - Common terms phoneticspodcast_production_standards.md - Industry format standardsaccessibility_guidelines.md - Inclusive audio contentplatform_requirements.md - YouTube, Spotify, Apple specsvoice_care_tips.md - Narrator health and performanceLocated in scripts/ directory:
main.py - CLI interface for script conversionaudio_writer.py - Core text-to-audio transformationpronunciation.py - Medical terminology phoneticstiming.py - Duration calculation and pacingformat_templates.py - Podcast, video, lecture templatesvoice_direction.py - Narrator cues and directionaccessibility.py - Alternative format generation| Parameter | Type | Default | Required | Description | |-----------|------|---------|----------|-------------| | --input, -i | string | - | No | Input text file path | | --output, -o | string | - | No | Output JSON file path (default: stdout) | | --text | string | - | No | Direct text input (alternative to --input) | | --duration, -d | int | 5 | No | Target duration in minutes | | --pace, -p | string | normal | No | Speaking pace (slow, normal, fast) | | --style, -s | string | conversational | No | Script style (conversational, formal, educational) |
bash# Convert from file python scripts/main.py --input article.txt --duration 5 --output script.json # Direct text input python scripts/main.py --text "Medical research findings..." --duration 3 # From stdin cat article.txt | python scripts/main.py --duration 5 --style conversational # With specific style and pace python scripts/main.py --input paper.txt --style educational --pace slow
| Risk Indicator | Assessment | Level | |----------------|------------|-------| | Code Execution | Python script executed locally | Low | | Network Access | No external API calls | Low | | File System Access | Read input files, write output files | Low | | Instruction Tampering | Standard prompt guidelines | Low | | Data Exposure | Output saved only to specified location | Low |
bash# Python 3.7+ # No additional packages required (uses standard library)
🎙️ Pro Tip: The best audio scripts sound natural when spoken. Always read your script aloud before finalizing—if you stumble over a sentence, your narrator will too. Revise for the ear, not the eye.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | pass→pass | 15,239 | 14,654 | -4% | 1 | 1 | 0% | 2,705 | 6,584 | +143% | 0 | 0 | — |
case-22 | fail→fail | 19,811 | 22,893 | +16% | 1 | 1 | 0% | 2,869 | 8,117 | +183% | 0 | 0 | — |
case-05 | pass→pass | 6,179 | 4,341 | -30% | 1 | 1 | 0% | 486 | 4,824 | +893% | 0 | 0 | — |
case-01 | pass→pass | 3,939 | 3,088 | -22% | 1 | 1 | 0% | 680 | 4,632 | +581% | 0 | 0 | — |
case-02 | pass→pass | 8,165 | 8,279 | +1% | 1 | 1 | 0% | 1,501 | 5,739 | +282% | 0 | 0 | — |
case-03 | pass→pass | 4,776 | 5,069 | +6% | 1 | 1 | 0% | 831 | 5,028 | +505% | 0 | 0 | — |
case-04 | pass→pass | 3,961 | 3,519 | -11% | 1 | 1 | 0% | 655 | 4,747 | +625% | 0 | 0 | — |
case-19 | fail→fail | 19,819 | 19,674 | -1% | 1 | 1 | 0% | 3,267 | 7,021 | +115% | 0 | 0 | — |
case-06 | pass→pass | 2,673 | 3,163 | +18% | 1 | 1 | 0% | 486 | 4,690 | +865% | 0 | 0 | — |
case-07 | pass→pass | 2,202 | 3,564 | +62% | 1 | 1 | 0% | 321 | 4,711 | +1368% | 0 | 0 | — |
case-08 | pass→pass | 2,150 | 2,211 | +3% | 1 | 1 | 0% | 321 | 4,477 | +1295% | 0 | 0 | — |
case-09 | pass→pass | 2,552 | 3,308 | +30% | 1 | 1 | 0% | 402 | 4,620 | +1049% | 0 | 0 | — |
case-20 | fail→fail | 14,608 | 35,665 | +144% | 1 | 1 | 0% | 2,580 | 9,768 | +279% | 0 | 0 | — |
case-10 | pass→pass | 3,471 | 2,897 | -17% | 1 | 1 | 0% | 633 | 4,596 | +626% | 0 | 0 | — |
case-11 | pass→pass | 10,804 | 2,591 | -76% | 1 | 1 | 0% | 1,580 | 4,565 | +189% | 0 | 0 | — |
case-12 | pass→pass | 2,760 | 2,805 | +2% | 1 | 1 | 0% | 520 | 4,665 | +797% | 0 | 0 | — |
case-13 | pass→pass | 17,962 | 1,486 | -92% | 1 | 1 | 0% | 3,092 | 4,368 | +41% | 0 | 0 | — |
case-14 | pass→pass | 5,233 | 2,828 | -46% | 1 | 1 | 0% | 746 | 4,604 | +517% | 0 | 0 | — |
case-15 | fail→pass | 10,923 | 2,551 | -77% | 1 | 1 | 0% | 1,667 | 4,560 | +174% | 0 | 0 | — |
case-16 | pass→pass | 3,165 | 2,366 | -25% | 1 | 1 | 0% | 538 | 4,515 | +739% | 0 | 0 | — |
case-17 | fail→pass | 13,511 | 9,306 | -31% | 1 | 1 | 0% | 2,058 | 5,612 | +173% | 0 | 0 | — |
case-18 | pass→pass | 4,690 | 2,769 | -41% | 1 | 1 | 0% | 812 | 4,536 | +459% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.