Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.
.claude/skills/video-production-buddy-video-understand/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -35% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 149% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -45% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -26% | 0% |
Understand video content locally using ffmpeg for frame extraction and Whisper for transcription. Fully offline, no API keys required.
ffmpeg + ffprobe (required): brew install ffmpegopenai-whisper (optional, for transcription): pip install openai-whisperbash# Scene detection + transcribe (default) python3 skills/video-understand/scripts/understand_video.py video.mp4 # Keyframe extraction python3 skills/video-understand/scripts/understand_video.py video.mp4 -m keyframe # Regular interval extraction python3 skills/video-understand/scripts/understand_video.py video.mp4 -m interval # Limit frames extracted python3 skills/video-understand/scripts/understand_video.py video.mp4 --max-frames 10 # Use a larger Whisper model python3 skills/video-understand/scripts/understand_video.py video.mp4 --whisper-model small # Frames only, skip transcription python3 skills/video-understand/scripts/understand_video.py video.mp4 --no-transcribe # Quiet mode (JSON only, no progress) python3 skills/video-understand/scripts/understand_video.py video.mp4 -q # Output to file python3 skills/video-understand/scripts/understand_video.py video.mp4 -o result.json
| Flag | Description | |------|-------------| | video | Input video file (positional, required) | | -m, --mode | Extraction mode: scene (default), keyframe, interval | | --max-frames | Maximum frames to keep (default: 20) | | --whisper-model | Whisper model size: tiny, base, small, medium, large (default: base) | | --no-transcribe | Skip audio transcription, extract frames only | | -o, --output | Write result JSON to file instead of stdout | | -q, --quiet | Suppress progress messages, output only JSON |
| Mode | How it works | Best for | |------|-------------|----------| | scene | Detects scene changes via ffmpeg select='gt(scene,0.3)' | Most videos, varied content | | keyframe | Extracts I-frames (codec keyframes) | Encoded video with natural keyframe placement | | interval | Evenly spaced frames based on duration and max-frames | Fixed sampling, predictable output |
If scene mode detects no scene changes, it automatically falls back to interval mode.
The script outputs JSON to stdout (or file with -o). See references/output-format.md for the full schema.
json{ "video": "video.mp4", "duration": 18.076, "resolution": {"width": 1224, "height": 1080}, "mode": "scene", "frames": [ {"path": "/abs/path/frame_0001.jpg", "timestamp": 0.0, "timestamp_formatted": "00:00"} ], "frame_count": 12, "transcript": [ {"start": 0.0, "end": 2.5, "text": "Hello and welcome..."} ], "text": "Full transcript...", "note": "Use the Read tool to view frame images for visual understanding." }
Use the Read tool on frame image paths to visually inspect extracted frames.
references/output-format.md -- Full JSON output schema documentationOther measured skills in the registry, with their headline benchmark lift.