Loading skill
Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribe audio and video files using the configured speech-to-text provider
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -64% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -69% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -56% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -57% | 0% |
Transcribe audio and video files using the configured speech-to-text provider. Supports multiple STT providers including OpenAI Whisper, Deepgram, and Google Gemini — the active provider is selected in Settings under Speech-to-Text (services.stt).
file_path (absolute path to a local audio or video file) to transcribe.services.stt) is shared between transcription and telephony call paths.When adding or modifying an STT provider, follow the onboarding checklist at assistant/docs/stt-provider-onboarding.md. That document covers the daemon catalog, config schema, adapter wiring, client catalog parity, and required tests.
Other measured skills in the registry, with their headline benchmark lift.