Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transcribe audio files using Google's Gemini API or Vertex AI
.claude/skills/sundial-org-gemini-stt/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 146% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -23% | 0% |
Transcribe audio files using Google's Gemini API or Vertex AI. Default model is gemini-2.0-flash-lite for fastest transcription.
bashgcloud auth application-default login gcloud config set project YOUR_PROJECT_ID
The script will automatically detect and use ADC when available.
Set GEMINI_API_KEY in environment (e.g., ~/.env or ~/.clawdbot/.env)
.ogg / .opus (Telegram voice messages).mp3.wav.m4abash# Auto-detect auth (tries ADC first, then GEMINI_API_KEY) python ~/.claude/skills/gemini-stt/transcribe.py /path/to/audio.ogg # Force Vertex AI python ~/.claude/skills/gemini-stt/transcribe.py /path/to/audio.ogg --vertex # With a specific model python ~/.claude/skills/gemini-stt/transcribe.py /path/to/audio.ogg --model gemini-2.5-pro # Vertex AI with specific project and region python ~/.claude/skills/gemini-stt/transcribe.py /path/to/audio.ogg --vertex --project my-project --region us-central1 # With Clawdbot media python ~/.claude/skills/gemini-stt/transcribe.py ~/.clawdbot/media/inbound/voice-message.ogg
| Option | Description | |--------|-------------| | <audio_file> | Path to the audio file (required) | | --model, -m | Gemini model to use (default: gemini-2.0-flash-lite) | | --vertex, -v | Force use of Vertex AI with ADC | | --project, -p | GCP project ID (for Vertex, defaults to gcloud config) | | --region, -r | GCP region (for Vertex, default: us-central1) |
Any Gemini model that supports audio input can be used. Recommended models:
| Model | Notes | |-------|-------| | gemini-2.0-flash-lite | Default. Fastest transcription speed. | | gemini-2.0-flash | Fast and cost-effective. | | gemini-2.5-flash-lite | Lightweight 2.5 model. | | gemini-2.5-flash | Balanced speed and quality. | | gemini-2.5-pro | Higher quality, slower. | | gemini-3-flash-preview | Latest flash model. | | gemini-3-pro-preview | Latest pro model, best quality. |
See Gemini API Models for the latest list.
For Clawdbot voice message handling:
bash# Transcribe incoming voice message TRANSCRIPT=$(python ~/.claude/skills/gemini-stt/transcribe.py "$AUDIO_PATH") echo "User said: $TRANSCRIPT"
The script exits with code 1 and prints to stderr on:
Other measured skills in the registry, with their headline benchmark lift.