Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribe audio files to text using a local Whisper ASR service. Use when: the user wants to transcribe an audio file to text, convert speech to text from an audio recording, extract text from voice recordings, or needs transcription of podcasts/interviews/voice notes. Only works for English language audio. The skill saves the transcription to a text file in the agent's local folder.
.claude/skills/valtterimelkko-transcribe-audio/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 13% | 0% |
Transcribe audio files to text using the local Whisper ASR service running at http://localhost:9000.
curl http://localhost:9000/health)To transcribe an audio file:
bashcurl -X POST http://localhost:9000/asr \ -F "audio_file=@/path/to/audio.wav" \ -F "language=en"
Save the transcription to the agent's local folder:
bashcurl -X POST http://localhost:9000/asr \ -F "audio_file=@/path/to/audio.wav" \ -F "language=en" \ -o transcription.txt
Or with a specific output path:
bashOUTPUT_FILE="${PWD}/$(basename /path/to/audio.wav .wav).txt" curl -X POST http://localhost:9000/asr \ -F "audio_file=@/path/to/audio.wav" \ -F "language=en" \ -o "$OUTPUT_FILE"
bash curl -s http://localhost:9000/health Expected response: {"status": "ok"}
bash curl -X POST http://localhost:9000/asr \ -F "audio_file=@<audio-file-path>" \ -F "language=en" \ -o <output-file-path>
bash cat <output-file-path>
| Format | Extension | MIME Type | |--------|-----------|-----------| | OGG Vorbis | .oga, .ogg | audio/ogg | | MP3 | .mp3 | audio/mpeg | | M4A | .m4a | audio/mp4 | | WAV | .wav | audio/wav | | WebM | .webm | audio/webm |
tiny.en modelCommon issues and solutions:
| Error | Cause | Solution | |-------|-------|----------| | Connection refused | Whisper service not running | Start the local Whisper service using your own deployment method (for example Docker Compose in your chosen whisper-service directory) | | Empty response | Audio file is silent or corrupted | Check the audio file can be played | | Garbage text | Audio is not in English | Only English is supported | | Slow transcription | High CPU load or concurrent requests | Wait for other transcriptions to complete |
bash# Set variables AUDIO_FILE="/path/to/recording.wav" OUTPUT_FILE="${PWD}/transcription.txt" # Verify service if ! curl -s http://localhost:9000/health | grep -q '"status": "ok"'; then echo "Error: Whisper service is not running" exit 1 fi # Transcribe curl -X POST http://localhost:9000/asr \ -F "audio_file=@${AUDIO_FILE}" \ -F "language=en" \ -o "$OUTPUT_FILE" echo "Transcription saved to: $OUTPUT_FILE" cat "$OUTPUT_FILE"
bashfor audio_file in *.wav; do output_file="${PWD}/$(basename "$audio_file" .wav).txt" echo "Transcribing: $audio_file -> $output_file" curl -X POST http://localhost:9000/asr \ -F "audio_file=@${audio_file}" \ -F "language=en" \ -o "$output_file" done
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,618 | 5,120 | -23% | 1 | 1 | 0% | 468 | 1,279 | +173% | 0 | 0 | — |
case-02 | fail→fail | 6,485 | 8,185 | +26% | 1 | 1 | 0% | 450 | 1,721 | +282% | 0 | 0 | — |
case-03 | fail→fail | 5,936 | 4,577 | -23% | 1 | 1 | 0% | 261 | 1,333 | +411% | 0 | 0 | — |
case-04 | fail→pass | 10,784 | 10,978 | +2% | 1 | 1 | 0% | 2,099 | 3,240 | +54% | 0 | 0 | — |
case-05 | fail→pass | 18,625 | 13,500 | -28% | 1 | 1 | 0% | 3,491 | 3,752 | +7% | 0 | 0 | — |
case-06 | fail→pass | 10,833 | 7,221 | -33% | 1 | 1 | 0% | 1,843 | 2,354 | +28% | 0 | 0 | — |
case-07 | fail→pass | 9,503 | 2,930 | -69% | 1 | 1 | 0% | 2,165 | 1,552 | -28% | 0 | 0 | — |
case-08 | fail→pass | 7,113 | 2,376 | -67% | 1 | 1 | 0% | 1,272 | 1,433 | +13% | 0 | 0 | — |
case-09 | pass→pass | 2,870 | 2,307 | -20% | 1 | 1 | 0% | 554 | 1,483 | +168% | 0 | 0 | — |
case-10 | pass→pass | 7,644 | 2,758 | -64% | 1 | 1 | 0% | 1,523 | 1,561 | +2% | 0 | 0 | — |
case-11 | fail→pass | 8,035 | 3,986 | -50% | 1 | 1 | 0% | 1,584 | 1,842 | +16% | 0 | 0 | — |
case-12 | pass→pass | 11,083 | 6,429 | -42% | 1 | 1 | 0% | 2,067 | 2,215 | +7% | 0 | 0 | — |
case-13 | fail→pass | 12,335 | 6,155 | -50% | 1 | 1 | 0% | 2,050 | 2,093 | +2% | 0 | 0 | — |
case-14 | fail→pass | 12,911 | 6,268 | -51% | 1 | 1 | 0% | 2,234 | 2,102 | -6% | 0 | 0 | — |
case-15 | pass→pass | 8,804 | 3,254 | -63% | 1 | 1 | 0% | 1,616 | 1,504 | -7% | 0 | 0 | — |
case-16 | fail→pass | 7,930 | 18,659 | +135% | 1 | 1 | 0% | 1,630 | 4,406 | +170% | 0 | 0 | — |
case-17 | pass→pass | 10,111 | 1,898 | -81% | 1 | 1 | 0% | 1,792 | 1,395 | -22% | 0 | 0 | — |
case-18 | fail→pass | 8,693 | 2,570 | -70% | 1 | 1 | 0% | 1,717 | 1,552 | -10% | 0 | 0 | — |
case-19 | fail→pass | 10,258 | 4,151 | -60% | 1 | 1 | 0% | 1,979 | 1,892 | -4% | 0 | 0 | — |
case-20 | pass→pass | 14,773 | 8,607 | -42% | 1 | 1 | 0% | 2,496 | 2,439 | -2% | 0 | 0 | — |
case-21 | fail→pass | 8,704 | 6,190 | -29% | 1 | 1 | 0% | 1,657 | 1,765 | +7% | 0 | 0 | — |
case-22 | fail→pass | 11,949 | 2,037 | -83% | 1 | 1 | 0% | 2,245 | 1,398 | -38% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.