Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Install and use crv (claude-real-video) — a tool that lets any AI agent watch videos by extracting scene-aware keyframes, deduplicating them, and transcribing audio. Use when the user shares a video URL or file and wants it analyzed, summarized, or discussed.
.claude/skills/huangchihhungleo-claude-real-video-for-agents/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 80% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -17% | 0% |
crv (claude-real-video) is a CLI tool that extracts meaningful frames and transcripts from videos so AI agents can "see" and "read" them. It uses scene-change detection (not fixed-interval sampling), sliding-window deduplication, and optional Whisper transcription.
Key advantage: Same 58-second clip at fixed 1fps = 58 frames. crv keeps the 26 that actually differ, and --grid packs them into 3 contact sheets. Fewer tokens, nothing missed.
bash# macOS brew install ffmpeg # Linux sudo apt install ffmpeg # Windows winget install Gyan.FFmpeg
bash# Recommended: with audio transcription support pip install "claude-real-video[whisper]" # Core only (frames + dedup) pip install claude-real-video
The [whisper] extra never installs itself — without it there is no speech-to-text (videos that ship their own subtitles still get a transcript).
bashcrv --help ffmpeg -version
Run the bundled installer to symlink this skill into all detected agent platforms:
bashbash install-skill.sh
Or manually copy to your agent's skill directory:
bash# Claude Code cp -r skills/claude-real-video-for-agents ~/.claude/skills/ # Codex cp -r skills/claude-real-video-for-agents ~/.codex/skills/ # OpenCode cp -r skills/claude-real-video-for-agents ~/.opencode/skills/ # Gemini CLI cp -r skills/claude-real-video-for-agents ~/.gemini/skills/
bashcrv "https://www.youtube.com/watch?v=VIDEO_ID"
Output in crv-out/:
frames/ — deduplicated keyframestranscript.txt — plain-text transcriptMANIFEST.txt — summary for LLM consumptionbashcrv "https://youtu.be/VIDEO_ID" -o crv-out --grid --why "what the user wants to know"
--grid — tiles frames into 3x3 contact sheets (cuts image count ~9x)--why — focuses the analysis on a specific questionbashcrv lecture.mp4 -o out --lang en
bashcrv clip.mp4 --no-transcribe
bashcrv "https://..." --cookies cookies.txt crv "https://..." --cookies-from-browser chrome
bashcrv tutorial.mp4 --adaptive
bashcrv "https://youtu.be/..." --why "pricing strategy" --kb ~/notes
bashcrv video.mp4 --viewer # Opens viewer.html — video + keyframes + transcript, fully offline
When a user shares a video (URL or file path):
--grid and --why:bash crv "<url-or-path>" -o crv-out --grid --why "<user's question>" For long videos, cap frames: --max-frames 60
Use one output folder per video (e.g. -o crv-out/<slug>). A folder that already holds an analysis is refused; pass --overwrite to replace it.
MANIFEST.txt first — it summarizes the run (frame counts, frames dir) and includes the transcript. Frames are named in chronological order; per-segment transcript timings live in transcript.json when available (there are no per-frame timestamps).crv-out/grids/ (each is a 3x3 sequence of consecutive keyframes, chronological). Only read individual crv-out/frames/*.jpg when you need a close-up.transcript.json) where available.| Flag | Default | Description | |---|---|---| | source (positional) | — | Video URL or local file path | | -o, --out | crv-out | Output directory | | --overwrite | off | Replace a previous analysis living in the output directory (without this, a non-empty output dir is refused to avoid mixing videos) | | --scene | 0.30 | Scene-change sensitivity (0-1, lower = more frames) | | --fps-floor | 1.0 | Guarantee at least one frame every N seconds | | --max-frames | 150 | Hard cap on total frames | | --adaptive | off | Adaptive scene detection for slow-changing content | | --text-anchors | off | Force frames at subtitle-cue timestamps — needs a sidecar .srt/.vtt or embedded subtitle track (burned-in captions can't be detected) | | --lang | auto | Whisper language (en, zh, auto, etc.) | | --cookies | — | Netscape cookie file for login-gated sources | | --cookies-from-browser | — | Read cookies from browser (chrome, safari, firefox, edge) | | --no-transcribe | off | Skip audio transcription | | --viewer | off | Write a local viewer.html | | --whisper-model | base | Whisper model size (tiny, base, small, medium, large, turbo — turbo: near large-v2 accuracy, ~8x faster) | | --dedup-threshold | 8 | % of pixels that must change for a new frame (higher = fewer frames kept) | | --dedup-window | 4 | Compare against last N kept frames (1 = consecutive-only) | | --report | off | Keep dropped frames + write report.html | | --why | — | Viewing intent, e.g. --why "find the pricing strategy" — focuses the model's analysis | | --grid | off | Tile frames into 3x3 contact sheets | | --kb | — | Save as dated markdown note to knowledge-base folder | | --keep-audio | off | Save full soundtrack as audio.m4a (for Gemini, GPT-4o, etc.) |
pythonfrom claude_real_video import process result = process("https://youtu.be/...", "out", lang="en") print(result.frame_count, result.transcript_path)
crv-out/
├── MANIFEST.txt # Summary for the LLM
├── frames/ # Deduplicated keyframes
├── transcript.txt # Plain-text transcript
├── grids/ # 3x3 contact sheets (with --grid)
├── audio.m4a # Full soundtrack (with --keep-audio)
├── viewer.html # Local viewer (with --viewer)
├── report.html # Dedup report (with --report)
└── dropped/ # Dropped frames (with --report)--grid — it dramatically reduces token usage while preserving visual continuity.--why — it focuses the analysis on what the user actually cares about.--max-frames 60 for long videos (>10 min) to stay within context limits.--no-transcribe when the user only cares about visuals (thumbnails, UI, slides).--keep-audio when the user asks about music, tone, or sound effects.--adaptive for screencasts, tutorials, or slow-moving content.MANIFEST.txt before frames — it has the run summary and the transcript.transcript.json when it exists (e.g., "At 0:42, the presenter says..."); frames themselves carry order, not timestamps.--overwrite to replace it.--cookies option is for your own authorized access.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,204 | 5,754 | +80% | 1 | 1 | 0% | 493 | 2,530 | +413% | 0 | 0 | — |
case-02 | fail→fail | 4,889 | 6,911 | +41% | 1 | 1 | 0% | 787 | 2,516 | +220% | 0 | 0 | — |
case-03 | fail→fail | 6,225 | 5,056 | -19% | 1 | 1 | 0% | 1,300 | 2,486 | +91% | 0 | 0 | — |
case-04 | fail→pass | 15,231 | 6,239 | -59% | 1 | 1 | 0% | 2,799 | 3,256 | +16% | 0 | 0 | — |
case-05 | pass→pass | 5,001 | 4,504 | -10% | 1 | 1 | 0% | 808 | 2,450 | +203% | 0 | 0 | — |
case-06 | fail→pass | 9,482 | 2,227 | -77% | 1 | 1 | 0% | 1,822 | 2,421 | +33% | 0 | 0 | — |
case-07 | fail→pass | 11,162 | 1,716 | -85% | 1 | 1 | 0% | 2,056 | 2,395 | +16% | 0 | 0 | — |
case-08 | fail→pass | 7,234 | 1,938 | -73% | 1 | 1 | 0% | 1,382 | 2,494 | +80% | 0 | 0 | — |
case-09 | pass→pass | 8,750 | 3,021 | -65% | 1 | 1 | 0% | 1,689 | 2,358 | +40% | 0 | 0 | — |
case-10 | fail→pass | 17,158 | 1,774 | -90% | 1 | 1 | 0% | 2,946 | 2,440 | -17% | 0 | 0 | — |
case-11 | fail→pass | 6,564 | 2,047 | -69% | 1 | 1 | 0% | 1,214 | 2,578 | +112% | 0 | 0 | — |
case-12 | fail→pass | 12,478 | 1,749 | -86% | 1 | 1 | 0% | 2,334 | 2,395 | +3% | 0 | 0 | — |
case-13 | fail→pass | 13,361 | 1,617 | -88% | 1 | 1 | 0% | 2,403 | 2,370 | -1% | 0 | 0 | — |
case-14 | pass→pass | 8,297 | 1,112 | -87% | 1 | 1 | 0% | 1,579 | 2,344 | +48% | 0 | 0 | — |
case-15 | fail→pass | 12,402 | 2,156 | -83% | 1 | 1 | 0% | 2,137 | 2,390 | +12% | 0 | 0 | — |
case-16 | fail→pass | 8,823 | 3,085 | -65% | 1 | 1 | 0% | 1,584 | 2,703 | +71% | 0 | 0 | — |
case-17 | fail→fail | 2,424 | 2,428 | +0% | 1 | 1 | 0% | 446 | 2,453 | +450% | 0 | 0 | — |
case-18 | fail→pass | 9,319 | 3,011 | -68% | 1 | 1 | 0% | 2,187 | 2,731 | +25% | 0 | 0 | — |
case-19 | pass→pass | 7,654 | 2,702 | -65% | 1 | 1 | 0% | 1,411 | 2,647 | +88% | 0 | 0 | — |
case-20 | pass→pass | 8,390 | 5,018 | -40% | 1 | 1 | 0% | 1,764 | 3,194 | +81% | 0 | 0 | — |
case-21 | pass→fail | 11,632 | 4,846 | -58% | 1 | 1 | 0% | 2,384 | 3,172 | +33% | 0 | 0 | — |
case-22 | pass→pass | 12,240 | 7,301 | -40% | 1 | 1 | 0% | 2,321 | 3,596 | +55% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.