Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribe local audio files or voice notes into markdown files using the legacy VoiceNote Bot workflow: local Whisper primary, OpenAI transcription fallback, and OpenRouter GPT-5 nano cleanup. Make sure to use this whenever the user wants audio files, voice recordings, dictations, interviews, or Telegram-style voice notes turned into markdown/text files without using Telegram, especially for batch folders or when they want the same cleanup behaviour as the old Telegram path.
.claude/skills/valtterimelkko-audio-transcription-workflow/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 19 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 184% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 86% | 0% |
Use this skill when the user wants to convert one or more local audio files into cleaned markdown transcripts using the same core workflow as a legacy voice-note transcription bot path, but without Telegram, webhooks, or queues.
For each input audio file, the bundled script:
gpt-4o-mini-transcribe by default) if Whisper failsopenai/gpt-5-nano unless cleanup is explicitly skipped or the API key is unavailabletranscription_manifest.json summary into the output directoryUse this skill instead of a basic one-shot transcription skill when the user wants any of the following:
/path/to/*.m4a--recursiveSupported formats:
.oga, .ogg, .mp3, .m4a, .wav, .webm, .mp4, .mpeg, .mpgaEach transcript is saved as a .md file with a metadata header like this:
yaml--- source_file: "note-01.oga" source_path: "/absolute/path/note-01.oga" generated_at: "2026-05-28T10:00:00+00:00" workflow: "local_whisper_primary_openai_fallback_openrouter_cleanup" language_requested: "auto" transcription_provider: "whisper" transcription_model: "local-whisper" cleanup_applied: true cleanup_provider: "openrouter" cleanup_model: "openai/gpt-5-nano" warnings: [] ---
The markdown body below the header is the transcript the user should read or reuse.
Use this script:
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py ...
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py \ /path/to/voice-note.oga \ --output-dir /path/to/output
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py \ /path/to/a.oga /path/to/b.m4a /path/to/c.wav \ --output-dir /path/to/output
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py \ /path/to/audio-folder \ --output-dir /path/to/output
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py \ /path/to/audio-folder \ --recursive \ --output-dir /path/to/output
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py \ '/path/to/audio/*.m4a' \ --output-dir /path/to/output
If the user explicitly wants no cleanup:
bashpython3 ./skills/audio-transcription-workflow/scripts/transcribe_audio_to_markdown.py \ /path/to/audio-folder \ --output-dir /path/to/output \ --skip-cleanup
Expected environment variables:
WHISPER_URL — optional, defaults to http://localhost:9000/asrOPENAI_API_KEY — optional but needed for transcription fallbackOPENAI_TRANSCRIPTION_MODEL — optional, defaults to gpt-4o-mini-transcribeOPENROUTER_API_KEY — optional but needed for legacy-style cleanupThe bundled script tries to work for low-context agents too.
OPENROUTER_API_KEY and OPENAI_API_KEY are resolved in this order: current process environment → a shell startup file such as ~/.bashrc via the shared credential loader → an optional legacy dotenv file pointed to by AUDIO_TRANSCRIPTION_LEGACY_ENV if you deliberately enable that fallback.WHISPER_URL stays simple: current process environment first, otherwise the script uses its built-in localhost default.That means cleanup should usually work even when an agent has not been explicitly told where the OpenRouter key lives.
./transcriptions near the current task context..md files if you need to summarise results back to the user.transcription_manifest.json was savedOPENAI_API_KEY is missing, that file fails.WHISPER_URL is intentionally not pulled from any legacy dotenv fallback, because those files may contain hostnames that are wrong for direct host-side agent use.-2, -3, and so on.After running the script, give a short operational summary such as:
/path/to/meeting-notes/ and turn them into markdown transcripts.".oga files, but without Telegram.".m4a files into markdown with cleanup."| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | fail→pass | 12,208 | 6,950 | -43% | 1 | 1 | 0% | 2,497 | 2,913 | +17% | 0 | 0 | — |
case-01 | fail→fail | 20,812 | 6,410 | -69% | 1 | 1 | 0% | 3,123 | 1,962 | -37% | 0 | 0 | — |
case-02 | fail→fail | 11,701 | 5,275 | -55% | 1 | 1 | 0% | 2,377 | 1,966 | -17% | 0 | 0 | — |
case-03 | fail→fail | 7,554 | 4,937 | -35% | 1 | 1 | 0% | 1,379 | 1,935 | +40% | 0 | 0 | — |
case-05 | fail→pass | 10,866 | 9,829 | -10% | 1 | 1 | 0% | 2,269 | 2,931 | +29% | 0 | 0 | — |
case-06 | fail→fail | 7,991 | 5,748 | -28% | 1 | 1 | 0% | 1,501 | 1,899 | +27% | 0 | 0 | — |
case-07 | pass→pass | 8,334 | 2,333 | -72% | 1 | 1 | 0% | 1,466 | 2,034 | +39% | 0 | 0 | — |
case-08 | fail→pass | 7,607 | 3,579 | -53% | 1 | 1 | 0% | 1,196 | 2,296 | +92% | 0 | 0 | — |
case-09 | fail→pass | 4,034 | 1,531 | -62% | 1 | 1 | 0% | 661 | 1,879 | +184% | 0 | 0 | — |
case-10 | fail→pass | 6,227 | 1,620 | -74% | 1 | 1 | 0% | 1,001 | 1,858 | +86% | 0 | 0 | — |
case-11 | fail→pass | 11,543 | 3,307 | -71% | 1 | 1 | 0% | 2,056 | 2,054 | -0% | 0 | 0 | — |
case-12 | fail→pass | 17,608 | 5,216 | -70% | 1 | 1 | 0% | 2,410 | 2,448 | +2% | 0 | 0 | — |
case-13 | pass→pass | 9,531 | 4,180 | -56% | 1 | 1 | 0% | 1,549 | 2,477 | +60% | 0 | 0 | — |
case-14 | fail→fail | 5,551 | 5,667 | +2% | 1 | 1 | 0% | 915 | 2,630 | +187% | 0 | 0 | — |
case-15 | pass→pass | 8,321 | 6,184 | -26% | 1 | 1 | 0% | 1,382 | 2,424 | +75% | 0 | 0 | — |
case-16 | pass→pass | 8,131 | 3,852 | -53% | 1 | 1 | 0% | 1,311 | 2,239 | +71% | 0 | 0 | — |
case-17 | fail→pass | 10,820 | 2,632 | -76% | 1 | 1 | 0% | 1,723 | 2,093 | +21% | 0 | 0 | — |
case-18 | fail→pass | 11,309 | 2,194 | -81% | 1 | 1 | 0% | 1,877 | 2,056 | +10% | 0 | 0 | — |
case-19 | fail→pass | 12,267 | 1,527 | -88% | 1 | 1 | 0% | 1,487 | 1,859 | +25% | 0 | 0 | — |
case-20 | pass→pass | 10,080 | 7,078 | -30% | 1 | 1 | 0% | 1,716 | 2,931 | +71% | 0 | 0 | — |
case-21 | pass→pass | 9,724 | 6,146 | -37% | 1 | 1 | 0% | 1,858 | 2,822 | +52% | 0 | 0 | — |
case-22 | pass→pass | 13,210 | 5,747 | -56% | 1 | 1 | 0% | 1,851 | 2,656 | +43% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.