Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Set up, repair, or verify the local AI short-video environment: video-use, FFmpeg, the Source Han Sans TW subtitle font, ElevenLabs credentials, and optional HyperFrames skills. Use whenever a user asks to install, configure, fix, reconnect, or check this editing environment. Do not use for Premiere or CapCut help, or to edit/transcribe media; hand those requests to the editing workflow after setup is verified.
.claude/skills/jaycheng1103-chatgpt-video-editing-setup/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 94% | 32 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 166% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-18 | ✓→✓ | = Same ✓ | -7% | 0% |
| case-20 | ✓→✓ | = Same ✓ | -52% | 0% |
Prepare or repair the environment without starting creative work. The outcome is an evidence-backed readiness report, not a transcript, upload, edit, preview, or render.
Read the setup runbook for commands and install choices. Read security and verification before handling credentials or declaring anything ready.
~/Developer/video-use and, only when requested, ~/Developer/hyperframes. Recognize normal checkouts and linked worktrees with git rev-parse, then capture each repository's exact origin, branch or detached state, commit, and status. Also check available runtime tools, the Source Han Sans TW subtitle font in the user's font directory, and the active agent's Skills location. Do not print secrets.
accepted origins are the exact official HTTPS URLs in the runbook. A missing or different origin and a dirty status are also hard stops: report the evidence, never rewrite the remote automatically, and do not pull, reset, overwrite, install dependencies, or register a Skill.
downloads, large downloads, or Skills-directory changes. Obtain explicit approval before any of them. Inspection and a no-cost local version check do not imply approval to mutate.
preflight immediately before dependency installation or Skill registration. Install/register the complete video-use repository so its helpers remain available. Install the Source Han Sans TW subtitle font only from the official Adobe Fonts repository, never overwriting an existing font file. Treat HyperFrames as optional unless the user specifically needs HTML, CSS, or GSAP animation; when requested, require Node.js 22 or newer.
protected ~/Developer/video-use/.env path. Never echo, log, or commit a credential.
files. Run HyperFrames repository, Node, lockfile, and Core Skills checks only if HyperFrames was explicitly approved and installed; otherwise report it as not requested. Do not upload media, call transcription, create an edit directory, or edit/render video.
command outcomes, remaining gaps, and the explicit next action. Never claim readiness without successful evidence.
Before any later first media upload to ElevenLabs, identify the source file, state that it will be uploaded for transcription with ElevenLabs Scribe v2, note that account quota or charges may apply, and obtain consent. Setup itself ends before that point.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | fail→fail | 14,943 | 20,813 | +39% | 1 | 1 | 0% | 1,664 | 1,366 | -18% | 0 | 0 | — |
case-12 | fail→fail | 18,379 | 21,043 | +14% | 1 | 1 | 0% | 2,146 | 1,327 | -38% | 0 | 0 | — |
case-13 | fail→fail | 9,990 | 29,081 | +191% | 1 | 1 | 0% | 833 | 1,925 | +131% | 0 | 0 | — |
case-01 | fail→fail | 22,483 | 24,807 | +10% | 1 | 1 | 0% | 3,520 | 1,210 | -66% | 0 | 0 | — |
case-02 | fail→fail | 25,508 | 18,422 | -28% | 1 | 1 | 0% | 736 | 1,280 | +74% | 0 | 0 | — |
case-03 | fail→fail | 22,374 | 19,361 | -13% | 1 | 1 | 0% | 562 | 1,206 | +115% | 0 | 0 | — |
case-04 | fail→fail | 12,589 | 21,752 | +73% | 1 | 1 | 0% | 1,357 | 1,993 | +47% | 0 | 0 | — |
case-05 | fail→fail | 17,645 | 25,827 | +46% | 1 | 1 | 0% | 2,432 | 2,028 | -17% | 0 | 0 | — |
case-07 | fail→fail | 10,215 | 30,297 | +197% | 1 | 1 | 0% | 1,986 | 2,037 | +3% | 0 | 0 | — |
case-08 | fail→fail | 17,455 | 18,801 | +8% | 1 | 1 | 0% | 2,297 | 2,407 | +5% | 0 | 0 | — |
case-09 | fail→fail | 13,858 | 15,882 | +15% | 1 | 1 | 0% | 577 | 1,257 | +118% | 0 | 0 | — |
case-10 | fail→fail | 11,076 | 30,344 | +174% | 1 | 1 | 0% | 1,080 | 2,582 | +139% | 0 | 0 | — |
case-11 | fail→fail | 10,141 | 22,484 | +122% | 1 | 1 | 0% | 847 | 1,300 | +53% | 0 | 0 | — |
case-14 | fail→pass | 9,122 | 31,542 | +246% | 1 | 1 | 0% | 1,459 | 3,888 | +166% | 0 | 0 | — |
case-15 | fail→fail | 7,723 | 20,112 | +160% | 1 | 1 | 0% | 443 | 1,705 | +285% | 0 | 0 | — |
case-16 | fail→pass | 23,290 | 11,018 | -53% | 1 | 1 | 0% | 2,412 | 1,879 | -22% | 0 | 0 | — |
case-17 | fail→fail | 12,041 | 33,456 | +178% | 1 | 1 | 0% | 1,082 | 2,141 | +98% | 0 | 0 | — |
case-18 | pass→pass | 15,534 | 11,835 | -24% | 1 | 1 | 0% | 2,021 | 1,884 | -7% | 0 | 0 | — |
case-19 | fail→fail | 10,725 | 9,457 | -12% | 1 | 1 | 0% | 1,682 | 1,310 | -22% | 0 | 0 | — |
case-20 | pass→pass | 25,239 | 9,465 | -62% | 1 | 1 | 0% | 3,017 | 1,446 | -52% | 0 | 0 | — |
case-21 | pass→pass | 13,418 | 14,729 | +10% | 1 | 1 | 0% | 1,519 | 1,883 | +24% | 0 | 0 | — |
case-22 | fail→pass | 15,191 | 10,949 | -28% | 1 | 1 | 0% | 2,177 | 1,648 | -24% | 0 | 0 | — |
case-23 | fail→fail | 11,900 | 21,261 | +79% | 1 | 1 | 0% | 1,310 | 1,342 | +2% | 0 | 0 | — |
case-24 | fail→fail | 18,417 | 19,632 | +7% | 1 | 1 | 0% | 2,220 | 1,341 | -40% | 0 | 0 | — |
case-25 | fail→fail | 13,778 | 21,027 | +53% | 1 | 1 | 0% | 1,371 | 1,775 | +29% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 8 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +12 percentage points is the difference between those two pass rates over the 8 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.