Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Dedicated YouTube thumbnail generator — interviews you for exactly the style elements you want (environment, text budget, extras, accent color), then renders high-contrast, vibrant, face-consistent thumbnails with Nano Banana Pro and verifies every frame before showing it. Use whenever you want to create, redo, or iterate thumbnail variants for a video — "make a thumbnail", "new version of B", "more realistic", "less text", "put the app on the screen", "another angle for the test". Renders into
.claude/skills/hassancs91-youtube-thumbnail/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 106% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 53% | 0% |
Turns a thumbnail concept into rendered, verified frames — with the style built from an element menu you pick from, not a fixed template. One number matters: CTR.
Division of labor: /packaging decides what to bet on (one fixed title, 3 distinct thumbnail levers, honesty checks). This skill decides what the frame looks like and renders it. Run standalone for iteration ("give me a new B"), or let packaging Stage 5 delegate here. Both inherit two non-negotiables from packaging:
brand.md. Never tone down tomatch the in-video brand.
Every render, regardless of selections: high contrast (deep darks OR clean brights, never flat) · vibrant saturated color with one accent doing the work · clear focal hierarchy (headline → hero object → face, three steps max) · big high-energy positive face from the media/library/faces/ kit · premium, clickable finish — crisp edges, no clutter · every text element perfectly spelled and razor sharp.
These came from a real before/after CTR comparison (the "Image 13" guide, 2026-08): the flat, low-contrast, no-hierarchy frame lost to the bold high-contrast hierarchical one ~4.4x on views. The axes are always-on; the decorations below are strictly opt-in.
Exact prompt language for every element lives in references/style-elements.md. Summary:
| Element | What it is | Cost | |---|---|---| | single-hook | ONE text element, huge (word or number) | the default; cleanest | | tiered-headline | small white italic context line + huge green grunge power line | +1 text block | | sticker | tilted yellow badge, black caps, red underline, arrow to the hero | +1; ≤2 short words or it garbles | | checkstrip | slim bottom strip, 3 green-check benefits | +3 micro-texts; busy fast | | brush-tag | brush-stroke tag at strip end (e.g. WITH CLAUDE) | +1 | | results-fan | fan/burst of result cards from the hero (pure imagery, no lettering) | busy-frame risk | | glow-burst | neon energy + particles radiating from the hero | cheap, usually worth it | | real-logo | pass the actual logo file as a --ref so it renders faithfully | extra ref | | real-app-screen | pass a real UI screenshot as a --ref for the on-screen app | extra ref; micro-text caveat |
Environments (pick one): dark-studio (neon-on-black, dramatic) · real-office (candid DSLR look, anchored to the real room visible in the face refs) · bright-graphic (saturated studio composite, the classic loud style).
Project dir, existing packaging/thumbs/ (what letters/versions exist), the locked title if packaging ran (thumbnail text must never repeat it), face kit present (no kit → stop and ask, see media/library/faces/README.md), GEMINI_API_KEY in .env.
Ask before generating — one AskUserQuestion round. Skip any axis the user already specified in their request; ask only what's still open.
single-hook (recommended default) / headline-plus-one (tieredheadline + at most ONE accent element) / full-kit (headline + sticker + checkstrip + tag — warn that it's the busy end; opt-in only).
prop/object, none.
Calibration note (real creator feedback, 2026-08): the default is MINIMAL text. The full-kit look was tried and rejected as "too much text"; high contrast + vibrant were explicitly kept. Never stack every element just because the guide contains them — each text block after the first must earn its slot.
Echo the exact hook text (and sticker/strip words if selected) before rendering. Check against the locked title: thumbnail text never repeats title words — the two combine into one message.
Assemble from references/style-elements.md blocks: labeled brief (Reference roles → Subject & Action → Composition & Hierarchy → Setting & Lighting → Text → Style), letter-by-letter spelling for every rendered word, positive framing with the one earned negation (sub-images carry no lettering/logos). Hardware rule: laptops/phones are "generic, plain dark bezel, no logos or lettering on the hardware" — otherwise you get a MacBook.
venv/Scripts/python tools/gen_thumbnail.py --prompt-file videos/<p>/packaging/thumbs/<X>.txt \
--out videos/<p>/packaging/thumbs/<X>.png --jpg --seed <N>--ref replaces that default — whenpassing a logo or app screenshot, re-include the face ref explicitly, and name each image's role in the prompt ("Image 1 = identity, Image 2 = the app").
real-app-screen source: pull a frame from the project's baked preview with ffmpeg, cropto the window, save to media/projects/<p>/ (reusable, committed path) — never screenshot by hand if the beat already exists in the video.
Read the image back: hook spelled right + legible · face reads as the creator · one dominant focal element · bright/saturated/positive expression · hands sane (count fingers, no phantom limbs) · no stray lettering in sub-images · hardware unbranded · honesty (screen/props match the real product) · 16:9, JPG <2MB. Fix-and-rerender before surfacing; flag anything borderline honestly (e.g. micro-text garble in a real-app screen).
letter; a restyle/refinement of the same lever bumps the number. One A/B/C test slot per letter family.
<X>.txt/.png, preserve the old one as <X>-v1.*.--seed to iterate composition-stable; change seed to explore. Change one thing at atime.
thumbs/_ABC-set.jpg) after every accepted render sothe set is always comparable at browse-wall scale.
| Symptom | Fix | |---|---| | Sticker/badge word garbled | ≤2 short words per badge, letters spelled out (L-I-F-E-T-I-M-E); re-render | | Checkstrip drops a checkmark | phrase as "each item a green circular checkmark immediately followed by…" | | MacBook bezel / brand on hardware | "generic slim laptop, plain dark bezel, no logos or lettering on the hardware" | | Real app screen: micro-text garbles | expected — invisible at browse size; if it matters, blur-text reroll or perspective-warp the real screenshot in post | | Face color cast (purple/green hand or skin) | name the light: "warm skin tones preserved; the screen adds only a faint cool glow" | | Clone/double renders a stranger | "BOTH men in this image are this exact man" + give the double a distinct material (wireframe, hologram) | | Office looks generic | anchor it: describe the real room visible in the face ref (wall color, plant) and say "match the real office behind him in Image 1" | | Word spelled right but flat/pasted | grunge texture + soft outer glow + hard drop shadow, restate "razor sharp" |
Model ids, API details, template history: .claude/skills/packaging/references/thumbnail-generation.md. Worked examples (proven prompts): videos/video-1/packaging/thumbs/*.txt — A2 (bright-graphic + full kit + real logo), B (dark-studio outcome scene), C (dark-studio clone), D2 (real-office + full kit), D3 (real-office + single-hook + real-app-screen).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 13,148 | 11,282 | -14% | 1 | 1 | 0% | 1,932 | 2,367 | +23% | 0 | 0 | — |
case-02 | fail→fail | 12,379 | 13,162 | +6% | 1 | 1 | 0% | 1,568 | 2,345 | +50% | 0 | 0 | — |
case-03 | fail→fail | 19,183 | 8,674 | -55% | 1 | 1 | 0% | 3,258 | 2,620 | -20% | 0 | 0 | — |
case-04 | fail→fail | 13,123 | 28,437 | +117% | 1 | 1 | 0% | 1,836 | 4,346 | +137% | 0 | 0 | — |
case-05 | fail→fail | 17,325 | 18,044 | +4% | 1 | 1 | 0% | 2,801 | 4,747 | +69% | 0 | 0 | — |
case-06 | fail→fail | 23,162 | 7,939 | -66% | 1 | 1 | 0% | 736 | 3,151 | +328% | 0 | 0 | — |
case-07 | pass→pass | 9,673 | 7,615 | -21% | 1 | 1 | 0% | 1,458 | 3,215 | +121% | 0 | 0 | — |
case-08 | fail→pass | 11,858 | 5,566 | -53% | 1 | 1 | 0% | 1,574 | 2,937 | +87% | 0 | 0 | — |
case-09 | fail→pass | 14,093 | 12,309 | -13% | 1 | 1 | 0% | 2,019 | 3,516 | +74% | 0 | 0 | — |
case-10 | fail→pass | 8,667 | 5,338 | -38% | 1 | 1 | 0% | 1,415 | 2,910 | +106% | 0 | 0 | — |
case-11 | pass→pass | 6,836 | 19,435 | +184% | 1 | 1 | 0% | 947 | 2,948 | +211% | 0 | 0 | — |
case-12 | fail→fail | 8,673 | 4,450 | -49% | 1 | 1 | 0% | 1,322 | 2,445 | +85% | 0 | 0 | — |
case-13 | fail→pass | 21,514 | 6,109 | -72% | 1 | 1 | 0% | 1,916 | 3,027 | +58% | 0 | 0 | — |
case-14 | fail→pass | 13,474 | 8,893 | -34% | 1 | 1 | 0% | 2,183 | 3,333 | +53% | 0 | 0 | — |
case-15 | fail→pass | 13,434 | 5,986 | -55% | 1 | 1 | 0% | 1,849 | 2,950 | +60% | 0 | 0 | — |
case-16 | fail→fail | 9,800 | 7,181 | -27% | 1 | 1 | 0% | 1,706 | 2,917 | +71% | 0 | 0 | — |
case-17 | pass→pass | 15,357 | 20,915 | +36% | 1 | 1 | 0% | 2,097 | 3,504 | +67% | 0 | 0 | — |
case-18 | fail→pass | 16,947 | 10,257 | -39% | 1 | 1 | 0% | 2,419 | 3,436 | +42% | 0 | 0 | — |
case-19 | fail→pass | 13,168 | 9,122 | -31% | 1 | 1 | 0% | 1,999 | 3,299 | +65% | 0 | 0 | — |
case-20 | fail→pass | 14,560 | 4,466 | -69% | 1 | 1 | 0% | 1,869 | 2,718 | +45% | 0 | 0 | — |
case-21 | fail→pass | 7,255 | 4,122 | -43% | 1 | 1 | 0% | 1,020 | 2,663 | +161% | 0 | 0 | — |
case-22 | fail→fail | 9,873 | 6,903 | -30% | 1 | 1 | 0% | 1,383 | 3,167 | +129% | 0 | 0 | — |
case-23 | pass→pass | 9,504 | 6,555 | -31% | 1 | 1 | 0% | 1,379 | 2,959 | +115% | 0 | 0 | — |
case-24 | pass→pass | 9,849 | 7,355 | -25% | 1 | 1 | 0% | 1,546 | 3,263 | +111% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 21 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +42 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.