Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
.claude/skills/davila7-transcribe/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -48% | 0% |
Transcribe audio using OpenAI, with optional speaker diarization when requested. Prefer the bundled CLI for deterministic, repeatable runs.
OPENAI_API_KEY is set. If missing, ask the user to set it locally (do not ask them to paste the key).transcribe_diarize.py CLI with sensible defaults (fast text transcription).output/transcribe/ when working in this repo.gpt-4o-mini-transcribe with --response-format text for fast transcription.--model gpt-4o-transcribe-diarize --response-format diarized_json.--chunking-strategy auto.gpt-4o-transcribe-diarize.output/transcribe/<job-id>/ for evaluation runs.--out-dir for multiple files to avoid overwriting.Prefer uv for dependency management.
uv pip install openaiIf uv is unavailable:
python3 -m pip install openaiOPENAI_API_KEY must be set for live API calls.bashexport CODEX_HOME="${CODEX_HOME:-$HOME/.codex}" export TRANSCRIBE_CLI="$CODEX_HOME/skills/transcribe/scripts/transcribe_diarize.py"
User-scoped skills install under $CODEX_HOME/skills (default: ~/.codex/skills).
Single file (fast text default):
python3 "$TRANSCRIBE_CLI" \
path/to/audio.wav \
--out transcript.txtDiarization with known speakers (up to 4):
python3 "$TRANSCRIBE_CLI" \
meeting.m4a \
--model gpt-4o-transcribe-diarize \
--known-speaker "Alice=refs/alice.wav" \
--known-speaker "Bob=refs/bob.wav" \
--response-format diarized_json \
--out-dir output/transcribe/meetingPlain text output (explicit):
python3 "$TRANSCRIBE_CLI" \
interview.mp3 \
--response-format text \
--out interview.txtreferences/api.md: supported formats, limits, response formats, and known-speaker notes.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-23 | pass→pass | 7,427 | 2,298 | -69% | 1 | 1 | 0% | 1,376 | 1,100 | -20% | 0 | 0 | — |
case-01 | fail→fail | 5,952 | 8,183 | +37% | 1 | 1 | 0% | 1,045 | 1,257 | +20% | 0 | 0 | — |
case-02 | fail→fail | 17,172 | 8,305 | -52% | 1 | 1 | 0% | 3,420 | 1,228 | -64% | 0 | 0 | — |
case-03 | fail→fail | 5,863 | 7,369 | +26% | 1 | 1 | 0% | 1,006 | 1,026 | +2% | 0 | 0 | — |
case-04 | pass→pass | 6,914 | 5,119 | -26% | 1 | 1 | 0% | 1,271 | 1,639 | +29% | 0 | 0 | — |
case-05 | fail→pass | 5,718 | 2,635 | -54% | 1 | 1 | 0% | 965 | 1,169 | +21% | 0 | 0 | — |
case-06 | fail→pass | 11,341 | 4,992 | -56% | 1 | 1 | 0% | 1,953 | 1,574 | -19% | 0 | 0 | — |
case-07 | pass→pass | 9,169 | 1,963 | -79% | 1 | 1 | 0% | 1,629 | 1,044 | -36% | 0 | 0 | — |
case-08 | fail→pass | 7,946 | 4,928 | -38% | 1 | 1 | 0% | 1,542 | 1,667 | +8% | 0 | 0 | — |
case-09 | fail→pass | 7,962 | 3,270 | -59% | 1 | 1 | 0% | 1,308 | 1,389 | +6% | 0 | 0 | — |
case-10 | fail→pass | 18,434 | 4,716 | -74% | 1 | 1 | 0% | 3,367 | 1,748 | -48% | 0 | 0 | — |
case-11 | pass→pass | 4,871 | 3,960 | -19% | 1 | 1 | 0% | 866 | 1,372 | +58% | 0 | 0 | — |
case-12 | pass→pass | 19,936 | 17,664 | -11% | 1 | 1 | 0% | 4,098 | 4,269 | +4% | 0 | 0 | — |
case-13 | pass→pass | 8,374 | 6,109 | -27% | 1 | 1 | 0% | 2,027 | 2,067 | +2% | 0 | 0 | — |
case-14 | pass→pass | 11,491 | 8,273 | -28% | 1 | 1 | 0% | 2,578 | 2,442 | -5% | 0 | 0 | — |
case-15 | fail→pass | 3,312 | 1,654 | -50% | 1 | 1 | 0% | 678 | 1,021 | +51% | 0 | 0 | — |
case-16 | fail→pass | 6,415 | 2,932 | -54% | 1 | 1 | 0% | 1,244 | 1,367 | +10% | 0 | 0 | — |
case-17 | fail→pass | 3,194 | 1,987 | -38% | 1 | 1 | 0% | 515 | 1,026 | +99% | 0 | 0 | — |
case-18 | fail→pass | 7,912 | 2,143 | -73% | 1 | 1 | 0% | 1,541 | 1,113 | -28% | 0 | 0 | — |
case-19 | fail→fail | 7,384 | 1,870 | -75% | 1 | 1 | 0% | 1,454 | 1,114 | -23% | 0 | 0 | — |
case-20 | fail→pass | 3,530 | 1,574 | -55% | 1 | 1 | 0% | 654 | 959 | +47% | 0 | 0 | — |
case-21 | fail→pass | 8,194 | 2,611 | -68% | 1 | 1 | 0% | 1,419 | 1,211 | -15% | 0 | 0 | — |
case-22 | pass→pass | 6,028 | 2,142 | -64% | 1 | 1 | 0% | 1,024 | 1,102 | +8% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.