Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast. Tier 2 (RSS+Groq Whisper) is the recommended approach -- fast, free, and most reliable. Also use when asked to "get transcript", "find episode", "summarize podcast", or "search podcast content". Do not use for general web scraping or non-podcast audio transcription.
.claude/skills/varnan-tech-podcast-transcript-fetcher/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 141% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 81% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 20% | 0% |
Fetch transcripts from 5 supported podcasts. Tier 2 (RSS+Groq Whisper) is the recommended approach -- fast, free, and the most reliable across all podcasts. Tier 1 free sources are best-effort (limited availability). Tier 3 Taddy API is the premium/commercial option.
bash# Get latest episode transcript (auto-detects best method) python scripts/get_transcript.py "Lenny's Podcast" --latest # Search by episode title or number python scripts/get_transcript.py 20vc --episode "Marc Andreessen" python scripts/get_transcript.py dwarkesh --episode 15 # Force specific method python scripts/get_transcript.py "cheeky pint" --latest --method whisper python scripts/get_transcript.py a16z --latest --method taddy # Save to file python scripts/get_transcript.py lennys --latest --output transcript.md # List all supported podcasts python scripts/get_transcript.py --list-podcasts
| Podcast | Tier 1 (best-effort) | Tier 2 RSS+Whisper RECOMMENDED] | Tier 3 Taddy (premium) | |---------|---------------------|-----------------------------------|------------------------| | Lenny's Podcast | GitHub archive (269 transcripts) | ✅ Substack RSS | ✅ Covered | | Dwarkesh Podcast | Website scrape + Substack PDF | ✅ Substack RSS | ✅ Covered | | Cheeky Pint | (none) | ✅ Transistor.fm RSS | ✅ Covered | | 20VC | Substack PDF | ✅ Libsyn RSS | ✅ Covered | | A16z Podcast | Website scrape | ✅ Simplecast RSS | ✅ Covered |
bash# Core (always required) pip install requests # Cloud transcription (recommended — fast, free tier) pip install groq export GROQ_API_KEY="your-key" # Get at https://console.groq.com # Local transcription (free, needs ~5GB RAM) pip install faster-whisper # Audio compression (for Groq's 25 MB limit — Windows: winget/scoop) # winget install ffmpeg or scoop install ffmpeg # Taddy API (commercial, optional) export TADDY_API_KEY="your-key" # Get at https://taddy.org
The script auto-selects the best method. Tier 2 is the default recommendation:
Tier 1 → Tier 2 (RECOMMENDED) → Tier 3
(best-effort) (Whisper) (Taddy API premium)Tier 1: Free direct sources (best-effort, limited availability)
ChatPRD/lennys-podcast-transcripts and searches by titleTier 2: RSS + Whisper transcription RECOMMENDED]
Tier 3: Taddy API (commercial/premium)
TADDY_API_KEY ($75/mo+)Once you have a transcript, pipe it to the agent for analysis:
markdownI have this transcript from [podcast]. Can you: 1. Summarize the key arguments 2. Extract 3 actionable insights 3. Identify any controversial claims 4. Compare with [other podcast] on the same topic
| Scenario | Command | |----------|---------| | Latest episode | get_transcript.py "Lenny's Podcast" --latest | | Specific episode by title | get_transcript.py 20vc --episode "Sam Altman" | | Episode by number | get_transcript.py dwarkesh --episode 42 | | Force Whisper transcription (Tier 2, recommended) | get_transcript.py a16z --latest --method whisper | | Force Taddy API (premium) | get_transcript.py lennys --latest --method taddy | | Save to Markdown | get_transcript.py cheeky-pint --latest --output episode.md | | JSON output | get_transcript.py dwarkesh --latest --json |
| Scenario | Command | |----------|---------| | Search all podcasts by keyword | get_transcript.py --search "Marc Andreessen" | | Search by guest name | get_transcript.py --guest "Sam Altman" | | Search within one podcast | get_transcript.py "Lenny's Podcast" --search "vibe coding" | | Batch-transcribe last N episodes | get_transcript.py "Dwarkesh Podcast" --last 5 | | Search + transcribe top matches | get_transcript.py --search "AI safety" --transcribe | | Pipeline with custom count | get_transcript.py --search "scaling laws" --transcribe --transcribe-count 5 | | Filtered search pipeline | get_transcript.py "A16z Podcast" --search "crypto" --transcribe |
Batch transcription saves to output/ with per-podcast subdirectories:
output/dwarkesh-podcast/Dwarkesh Podcast_2024-01-15_agi-is-still-30-years-away.md
output/20vc/20 Minutes VC (20VC)_2024-03-10_funding-round-analysis.mdEach file includes a YAML frontmatter header:
yaml--- podcast: Dwarkesh Podcast episode: AGI is still 30 years away date: 2024-01-15 url: https://... source: whisper ---
The registry at scripts/podcasts.json maps each podcast to its RSS feeds, transcript sources, and API endpoints. To add new podcasts:
json{ "id": "new-podcast", "name": "New Podcast", "rss": "https://example.com/feed.xml", "transcript_sources": { "primary": {"type": "website_scrape", "url": "https://example.com"} } }
| Problem | Solution | |---------|----------| | "No transcript found" | Tier 2 (RSS+Whisper) is the recommended approach. If auto mode fails, try --method whisper to force it. | | RSS fetch fails | RSS feeds may change; check scripts/podcasts.json for current URLs | | Audio download slow | Large MP3s can take minutes on slow connections | | Groq rate limited | Wait or switch to local faster-whisper | | Taddy not returning transcripts | Some episodes lack transcripts; try --method whisper | | Podcast not in registry | Add it to scripts/podcasts.json | | Unicode error on Windows | Fixed: script auto-reconfigures stdout to UTF-8; saved files use UTF-8 encoding | | Audio > 25 MB for Groq | Install ffmpeg: winget install ffmpeg (Windows) or brew install ffmpeg (macOS) |
| Podcast | Old Feed (broken) | Current Feed | |---------|------------------|--------------| | Cheeky Pint | feeds.transistor.fm/the-cheeky-pint (404) | feeds.transistor.fm/cheeky-pint-with-john-collison | | 20VC | feeds.simplecast.com/3GxrMqOd (404) | feeds.libsyn.com/61840/rss | | A16z | feeds.simplecast.com/0cJfpoz2 (404) | feeds.simplecast.com/JGE3yC0V |
GROQ_API_KEY in your env or .env fileOther measured skills in the registry, with their headline benchmark lift.