Install any skill in seconds. Free to start, no credit card required.
Get Started Free →OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
.claude/skills/graniet-whisper/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 141% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 224% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 70% | 0% |
This skill is repo-local and stays inactive until explicitly activated.
When the original instructions refer to legacy tool names, use these Kheish mappings:
terminal => bashweb_extract => web_fetch, plus web_search when discovery is neededsearch_files => grep_search and glob_searchbrowser_* tools require a browser-capable surfaced tool or MCP; if none is available, use the closest available surface and say so explicitlyWhen the instructions mention local helper files, resolve them from ${KHEISH_SKILL_DIR}.
OpenAI's multilingual speech recognition model.
Use when:
Metrics:
Use alternatives instead:
bash# Requires Python 3.8-3.11 pip install -U openai-whisper # Requires ffmpeg # macOS: brew install ffmpeg # Ubuntu: sudo apt install ffmpeg # Windows: choco install ffmpeg
pythonimport whisper # Load model model = whisper.load_model("base") # Transcribe result = model.transcribe("audio.mp3") # Print text print(result["text"]) # Access segments for segment in result["segments"]: print(f"[{segment['start']:.2f}s - {segment['end']:.2f}s] {segment['text']}")
python# Available models models = ["tiny", "base", "small", "medium", "large", "turbo"] # Load specific model model = whisper.load_model("turbo") # Fastest, good quality
| Model | Parameters | English-only | Multilingual | Speed | VRAM | |-------|------------|--------------|--------------|-------|------| | tiny | 39M | ✓ | ✓ | ~32x | ~1 GB | | base | 74M | ✓ | ✓ | ~16x | ~1 GB | | small | 244M | ✓ | ✓ | ~6x | ~2 GB | | medium | 769M | ✓ | ✓ | ~2x | ~5 GB | | large | 1550M | ✗ | ✓ | 1x | ~10 GB | | turbo | 809M | ✗ | ✓ | ~8x | ~6 GB |
Recommendation: Use turbo for best speed/quality, base for prototyping
python# Auto-detect language result = model.transcribe("audio.mp3") # Specify language (faster) result = model.transcribe("audio.mp3", language="en") # Supported: en, es, fr, de, it, pt, ru, ja, ko, zh, and 89 more
python# Transcription (default) result = model.transcribe("audio.mp3", task="transcribe") # Translation to English result = model.transcribe("spanish.mp3", task="translate") # Input: Spanish audio → Output: English text
python# Improve accuracy with context result = model.transcribe( "audio.mp3", initial_prompt="This is a technical podcast about machine learning and AI." ) # Helps with: # - Technical terms # - Proper nouns # - Domain-specific vocabulary
python# Word-level timestamps result = model.transcribe("audio.mp3", word_timestamps=True) for segment in result["segments"]: for word in segment["words"]: print(f"{word['word']} ({word['start']:.2f}s - {word['end']:.2f}s)")
python# Retry with different temperatures if confidence low result = model.transcribe( "audio.mp3", temperature=(0.0, 0.2, 0.4, 0.6, 0.8, 1.0) )
bash# Basic transcription whisper audio.mp3 # Specify model whisper audio.mp3 --model turbo # Output formats whisper audio.mp3 --output_format txt # Plain text whisper audio.mp3 --output_format srt # Subtitles whisper audio.mp3 --output_format vtt # WebVTT whisper audio.mp3 --output_format json # JSON with timestamps # Language whisper audio.mp3 --language Spanish # Translation whisper spanish.mp3 --task translate
pythonimport os audio_files = ["file1.mp3", "file2.mp3", "file3.mp3"] for audio_file in audio_files: print(f"Transcribing {audio_file}...") result = model.transcribe(audio_file) # Save to file output_file = audio_file.replace(".mp3", ".txt") with open(output_file, "w") as f: f.write(result["text"])
python# For streaming audio, use faster-whisper # pip install faster-whisper from faster_whisper import WhisperModel model = WhisperModel("base", device="cuda", compute_type="float16") # Transcribe with streaming segments, info = model.transcribe("audio.mp3", beam_size=5) for segment in segments: print(f"[{segment.start:.2f}s -> {segment.end:.2f}s] {segment.text}")
pythonimport whisper # Automatically uses GPU if available model = whisper.load_model("turbo") # Force CPU model = whisper.load_model("turbo", device="cpu") # Force GPU model = whisper.load_model("turbo", device="cuda") # 10-20× faster on GPU
bash# Generate SRT subtitles whisper video.mp4 --output_format srt --language English # Output: video.srt
pythonfrom langchain.document_loaders import WhisperTranscriptionLoader loader = WhisperTranscriptionLoader(file_path="audio.mp3") docs = loader.load() # Use transcription in RAG from langchain_chroma import Chroma from langchain_openai import OpenAIEmbeddings vectorstore = Chroma.from_documents(docs, OpenAIEmbeddings())
bash# Use ffmpeg to extract audio ffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav # Then transcribe whisper audio.wav
| Model | Real-time factor (CPU) | Real-time factor (GPU) | |-------|------------------------|------------------------| | tiny | ~0.32 | ~0.01 | | base | ~0.16 | ~0.01 | | turbo | ~0.08 | ~0.01 | | large | ~1.0 | ~0.05 |
Real-time factor: 0.1 = 10× faster than real-time
Top-supported languages:
Full list: 99 languages total
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | pass→pass | 13,542 | 8,803 | -35% | 1 | 1 | 0% | 2,408 | 3,784 | +57% | 0 | 0 | — |
case-01 | fail→pass | 7,848 | 7,107 | -9% | 1 | 1 | 0% | 1,479 | 3,568 | +141% | 0 | 0 | — |
case-11 | pass→pass | 2,550 | 1,664 | -35% | 1 | 1 | 0% | 395 | 2,505 | +534% | 0 | 0 | — |
case-02 | pass→pass | 3,682 | 3,088 | -16% | 1 | 1 | 0% | 676 | 2,616 | +287% | 0 | 0 | — |
case-03 | pass→pass | 3,607 | 2,213 | -39% | 1 | 1 | 0% | 690 | 2,644 | +283% | 0 | 0 | — |
case-04 | fail→pass | 8,114 | 2,866 | -65% | 1 | 1 | 0% | 1,466 | 2,748 | +87% | 0 | 0 | — |
case-05 | pass→pass | 7,837 | 2,809 | -64% | 1 | 1 | 0% | 1,636 | 2,766 | +69% | 0 | 0 | — |
case-20 | fail→fail | 13,272 | 11,279 | -15% | 1 | 1 | 0% | 2,586 | 4,651 | +80% | 0 | 0 | — |
case-06 | fail→pass | 4,836 | 2,889 | -40% | 1 | 1 | 0% | 868 | 2,813 | +224% | 0 | 0 | — |
case-07 | pass→pass | 2,697 | 2,084 | -23% | 1 | 1 | 0% | 406 | 2,593 | +539% | 0 | 0 | — |
case-08 | pass→pass | 10,829 | 5,705 | -47% | 1 | 1 | 0% | 2,061 | 3,329 | +62% | 0 | 0 | — |
case-09 | fail→pass | 7,955 | 5,254 | -34% | 1 | 1 | 0% | 1,523 | 3,189 | +109% | 0 | 0 | — |
case-10 | pass→pass | 6,449 | 2,125 | -67% | 1 | 1 | 0% | 1,084 | 2,577 | +138% | 0 | 0 | — |
case-12 | pass→pass | 11,085 | 5,143 | -54% | 1 | 1 | 0% | 1,969 | 3,226 | +64% | 0 | 0 | — |
case-13 | pass→fail | 14,033 | 5,372 | -62% | 1 | 1 | 0% | 2,281 | 3,168 | +39% | 0 | 0 | — |
case-14 | pass→pass | 2,243 | 1,678 | -25% | 1 | 1 | 0% | 383 | 2,527 | +560% | 0 | 0 | — |
case-15 | pass→pass | 6,460 | 2,576 | -60% | 1 | 1 | 0% | 1,118 | 2,652 | +137% | 0 | 0 | — |
case-16 | pass→pass | 5,879 | 3,889 | -34% | 1 | 1 | 0% | 1,254 | 2,973 | +137% | 0 | 0 | — |
case-17 | pass→pass | 2,997 | 1,917 | -36% | 1 | 1 | 0% | 514 | 2,453 | +377% | 0 | 0 | — |
case-18 | pass→pass | 4,066 | 2,446 | -40% | 1 | 1 | 0% | 755 | 2,678 | +255% | 0 | 0 | — |
case-19 | pass→pass | 7,889 | 9,064 | +15% | 1 | 1 | 0% | 1,557 | 2,839 | +82% | 0 | 0 | — |
case-22 | fail→pass | 11,602 | 6,314 | -46% | 1 | 1 | 0% | 1,992 | 3,393 | +70% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.