Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribes audio/video files using ElevenLabs Scribe v2 API. Use when transcribing audio files, generating transcripts, or converting speech to text.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 235% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 86% | 0% |
<objective> Transcribe audio or video files using the ElevenLabs Speech-to-Text API (Scribe v2). Accepts a file path and optional parameters, reads the API key from the project's .env file, and returns a formatted transcription with speaker diarization and audio event tagging. </objective>
<quick_start> Via slash command: /elevenlabs-transcribe path/to/audio.mp3 /elevenlabs-transcribe path/to/audio.mp3 --output transcript.txt --num-speakers 3
Requirements:
ELEVENLABS_API_KEY in the project's .env fileuv installed (dependencies auto-install via PEP 723)</quick_start>
<prerequisites> Before transcribing, verify:
uv is available (dependency installation is automatic via inline script metadata — no venv or manual pip install needed).env file where Claude is running: ELEVENLABS_API_KEY=your-key-here
MUST stop if the API key is missing — inform the user to add it to their .env file. </prerequisites>
<process>
Step 1: Parse user input
Extract the audio file path and any options from $ARGUMENTS or the user's message. Supported options:
--output <path> or -o <path> — where to save the transcript--language <code> — ISO-639 language code (e.g., eng, spa, fra, deu, jpn, zho)--num-speakers <n> — max speakers in the audio (1-32)--keyterms "term1" "term2" — words/phrases to bias transcription towards--timestamps none|word|character — timestamp granularity--no-diarize — disable speaker identification--no-audio-events — disable audio event tagging--json — output full JSON responseStep 2: Validate the audio file
Confirm the file path exists. Expand ~ paths. The script handles validation automatically but check early for a clear error message.
Step 3: Check for API key
bashgrep -q "ELEVENLABS_API_KEY=" .env 2>/dev/null && echo "API key configured" || echo "API key missing"
If missing, tell the user to add ELEVENLABS_API_KEY= to their .env file and stop.
Step 4: Run transcription
Dependencies are installed automatically by uv via inline script metadata (PEP 723). No venv or manual pip install needed.
Basic transcription (diarize + audio events + auto language):
bashuv run ~/.claude/skills/elevenlabs-transcribe/scripts/transcribe.py "<audio_file_path>"
With output file and options:
bashuv run ~/.claude/skills/elevenlabs-transcribe/scripts/transcribe.py "<audio_file_path>" --output transcript.txt --language eng --num-speakers 3
With key terms for better accuracy:
bashuv run ~/.claude/skills/elevenlabs-transcribe/scripts/transcribe.py "<audio_file_path>" --keyterms "technical term" "product name"
Full JSON response:
bashuv run ~/.claude/skills/elevenlabs-transcribe/scripts/transcribe.py "<audio_file_path>" --json --output result.json
Step 5: Present results
Format the transcription output cleanly for the user. If diarization is enabled, group text by speaker. Highlight any audio events detected. Example output:
[Speaker 0]: Hello, how are you doing today?
[Speaker 1]: I'm doing great, thanks for asking! (laughter)</process>
<script_options> | Flag | Description | Default | |------|-------------|---------| | <file> | Path to audio/video file (required) | - | | --output <path>, -o | Save transcription to file | stdout | | --language <code> | ISO-639 code (eng, spa, fra, deu, jpn, zho) | auto-detect | | --num-speakers <n> | Max speakers in audio (1-32) | auto-detect | | --keyterms "t1" "t2" | Terms to bias transcription towards (max 100) | none | | --timestamps <level> | Granularity: none, word, character | word | | --no-diarize | Disable speaker identification | diarize enabled | | --no-audio-events | Disable audio event tagging | events enabled | | --json | Output full JSON response | formatted text | </script_options>
<supported_formats> All major audio and video formats: mp3, wav, mp4, m4a, ogg, flac, webm, aac, wma, mov, avi, mkv, and more. Maximum file size: 3GB. </supported_formats>
<api_details>
</api_details>
<error_handling> | Error | Resolution | |-------|------------| | ELEVENLABS_API_KEY not found | Add key to .env file in current directory | | uv: command not found | Install uv: curl -LsSf https://astral.sh/uv/install.sh pipe to sh | | File not found | Verify the file path and expand any ~ | | 422 Validation Error | Check file format/size, ensure model_id is valid | | 401 Unauthorized | API key is invalid or expired | </error_handling>
<success_criteria>
.env without exposure in chat--output specified, file written to requested path</success_criteria>
Other measured skills in the registry, with their headline benchmark lift.