Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when building, writing, refining, or structuring a Higgsfield AI prompt. Covers the MCSLA formula, prompt structure, narrative vs. timestamped formats, and how to write for both text-to-video and image-to-video workflows.
.claude/skills/osidemedia-higgsfield-prompt/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 603% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 500% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 671% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 458% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 663% | 0% |
Generated-checked block (scripts/build_index.py verifies anchors). Read the linked sections for full context — these lines are routing aids, not the rules themselves.
../../specs/model-specs.yaml →../higgsfield-seedance/SKILL.md § Official Prompt Architecture →Every high-performing Higgsfield prompt is built on five layers. Think of it as the cinematographer's checklist — fill in each layer and the model has everything it needs.
| Letter | Element | Description | Example | |--------|---------|-------------|---------| | M | Model | Which generation engine | "Use Kling 2.6" | | C | Camera | Named camera control | "FPV Drone shot weaving through the alley" | | S | Subject | Who/what + appearance | "A woman in a sand-colored suit, sharp eyes" | | L | Look | Style + color + lighting | "Cinematic, golden hour, anamorphic flare" | | A | Action | What happens in the scene | "She turns slowly, wind lifting her coat" |
Start from nothing — describe the entire scene from scratch. Best for: establishing scenes, abstract concepts, environments without a specific character.
[Subject + appearance].
[Environment — location, time, weather, atmosphere].
[Action — what happens and how].
[Camera — named control].
[Look — style + color grade].Example:
A lone astronaut stands on the surface of a red desert planet, helmet visor reflecting
twin moons rising on the horizon. Dust spirals slowly in the thin atmosphere.
She turns to face the camera, gloved hand raised in a slow salute.
Camera: slow Crane Up revealing the vast emptiness behind her.
Style: Cinematic, desaturated orange and deep blue, 2.35:1 anamorphic.Animate a provided still image. The image defines the starting frame. Best for: character consistency, product shots, portrait animation, storyboard bring-to-life.
[Reference the input image as the first frame].
[Describe what should move, change, or animate — not what is already visible].
[Camera — named control].
[Style/atmosphere cues].Example:
Starting from the provided image as the first frame.
The woman's hair lifts gently in the wind. She blinks slowly and turns her gaze
slightly to the left, a faint smile forming.
Camera: subtle Dolly In toward her face.
Style: Cinematic, warm afternoon light, shallow depth of field.Key rule for I2V: Do NOT re-describe what is already in the image. Only describe what should change or animate. Over-describing the static elements confuses the model. This applies equally to @ Image references in Seedance/Cinema Studio 3.0 — describe ONLY motion and camera movement, never what's already visible.
Write the scene as continuous action. No timestamps. Most natural for Higgsfield.
A detective pushes open the door to the rain-soaked rooftop, coat whipping in the wind.
She steps to the edge and looks down at the city below — a thousand lights blurring
through the downpour. Camera dollies slowly behind her, then cranes up to reveal the
skyline. Cinematic style, cold blue tones, 16:9.Only use when exact timing of separate actions matters — e.g., a transformation, a multi-phase action sequence, or a beat-synced music video. Maps to Cinema Studio 3.0's Custom multi-shot mode.
0–3s: Wide establishing shot. The fighter stands alone in the ring, chest heaving.
3–6s: Crash Zoom In on his face. Sweat on his brow, jaw clenched.
6–10s: 360 Orbit as he raises his fists. Crowd noise rises.> The #1 mistake in video prompting: over-describing appearance and under-describing > behavior. Give your subject something to DO. Give them an internal state that creates > visible behavior. A verb that describes motion or intention is more important than > adjectives.
Specificity beats generality:
Active verbs carry the scene:
Name the camera control: Higgsfield understands its own preset names. Always use them explicitly.
Lead with subject, end with style: Subject → Action → Camera → Style is the most reliable order.
Keep it under 200 words (short-form regime): Focused prompts outperform exhaustive ones. One clear intention > ten vague details. Regime exception (HARD RULE 8): block-scaffold production prompts — ../higgsfield-seedance/SKILL.md § Official Prompt Architecture — replace the word cap with structural lint; harvested production briefs run 218–2,059-word medians by register. The cap governs single-shot MCSLA prompts only.
Cinema Studio: Keep it under 512 characters: Cinema Studio has a hard 512-character limit on prompts (both 2.5 and 3.0).
See the Cinema Studio skill for full character budget details.
Before writing any prompt, answer these five questions. Vague prompts like "give me something cinematic" tell the AI nothing.
| Question | What to specify | |----------|----------------| | Who? | Subject + appearance (e.g. "a man in a leather jacket") | | Where? | Environment + atmosphere (e.g. "in a narrow aircraft galley, cold blue light") | | What's happening? | 1 primary action (e.g. "punches his opponent") | | Camera movement? | Named preset (e.g. "Handheld") or Cinema Studio Director Panel | | Mood/Genre? | Style + color grade, or Cinema Studio genre selection |
AI models can replicate real-life physics — but only so much at once. Asking for multiple complex actions in one clip overwhelms the model.
Rule: 1 primary action per clip, with 1–2 secondary actions max.
Break complex sequences into separate shots and stitch them in a video editor, or use Multi-Shot Manual mode to prompt each scene separately.
Fast Motion Trick: If fast motion keeps morphing or breaking, generate the scene in Slow Mo first, then speed it up in post (CapCut, Premiere, DaVinci). The model renders cleaner physics in slow motion.
Never leave a generic emotion in a prompt. "Sad" / "angry" / "surprised" / "scared" / "thoughtful" / "in love" — each is at least three or four distinct physical realizations, and the model renders a different version depending on which one your prompt invites. A prompt that says only "she looks surprised" produces a different shot every regeneration and degrades adherence across batches.
The rule: decompose the generic emotion into specific muscle movements, breath, eyes, and skin. If you can't decompose confidently, ask the user to choose a variant.
Clarification template — offer when the script or user supplies a generic emotion you cannot decompose without inventing detail:
> Which kind of surprise? > (a) Light positive — eyebrows lift, lips part softly, slow inhale > through the nose, no other movement. > (b) Shock — sharp inhale through the mouth, eyes widen, body > freezes in place, hand involuntarily lifts to chest. > (c) Disbelief — slow blink, head tilts a fraction, lips press > together, only one eyebrow lifts. > (d) Surprise-with-joy — eye light shifts (catchlight reads), > smile builds gradually, shoulders relax.
Same shape applies to any generic adjective — "tense" / "sad" / "angry" / "scared" / "thoughtful" / "in love" each decomposes into 3-5 distinct physical realizations. The decomposed prompt produces a performance; the generic prompt produces AI-video.
> Preset library alternative. For named micro-expression presets > that drop into a prompt without first-principles decomposition, > see ../higgsfield-soul/SKILL.md § Micro-Expressions. The catalog > covers most common emotional registers with locked physical > descriptors. Use the decompose-from-first-principles rule above > when no preset matches; use the preset library when one does.
Single-axis decomposition (above) names one register: angry / sad / surprised. Layered emotion names a composite state where two registers stack — anxious determination, tired tenderness, bitter amusement, cornered calculation. Production-team practice finds the model renders layered states better than single registers when the layering is described as one channel modulating another: the dominant state plus the underlying state plus the visible tell.
with shallow chest breath and a single hand at the side flexing open-closed (anxiety underneath).
(tenderness) with the body weight settled, slow blink interval (tiredness underneath).
with no smile crinkles at the eye corners (bitterness underneath).
Compose layered states by stacking decomposed physical realizations from the single-axis catalog. The dominant state goes in the face; the underlying state goes in breath, posture, and hand-state; the visible tell sits in the eyes.
For finer control, layer a tiny detail on top of an existing emotional cue: Roco is very upset, and his lower lip trembles. The base emotion gets the broad performance; the tiny detail gives the model a specific physical cue to render. Production-team discipline holds that the model renders the simple-emotion-plus- tiny-detail compound better than either an over-decomposed prompt or a too-generic one.
When a prompt involves Soul ID or any character who must stay consistent across shots, always split the output into two clearly labeled blocks:
Bad (mixed) — identity drifts:
A woman with sharp cheekbones and auburn hair in a blue trench coat runs through
a rain-soaked alley, her coat flapping, sharp cheekbones catching the neon light,
camera chasing her at full speed, her auburn hair streaming behind her.Good (separated) — identity stays locked:
Identity Block:
The Soul ID character — sharp cheekbones, auburn hair shoulder-length,
wearing a blue trench coat with silver buttons, lean athletic build.Motion Block:
She runs through a rain-soaked alley, coat flapping behind her.
Camera: Action Run — low behind, matching pace.
Neon reflections streak across wet concrete.
Style: Cinematic, cold blue shadows, warm neon accents. 16:9.When to apply this rule:
> Camera matches emotion, not just identity. The Motion Block describes WHAT > the character does and HOW the camera moves. The quality of the camera motion > — jittery handheld for anger, smooth handheld breathing for calm, static + slow > push for revelation — should track the focal character's emotional state. See > ../higgsfield-camera/SKILL.md § Camera-Emotion Sync for the 6-emotion movement > map and the emotional-arcs-within-a-shot pattern. For decomposing the underlying > generic emotion before picking a camera prescription, see § Generic-Emotion > Decomposition above.
Sub-skills can legitimately nominate different things for the same shot. higgsfield-camera § Camera-Emotion Sync nominates handheld-slow-low for sadness; higgsfield-prompt § Scene Archetype Router permits locked dolly-in for the Atmosphere archetype where mood-is-the-content. When two sub-skills nominate different camera moves (or motion presets, or style registers) for the same scene, resolve in this order:
When the resolution is non-obvious, surface it. Tell the user which sub-skill nominated what and why you picked one over the other — this is meta-correct behavior and lets the user override. Silent picking is the failure mode; transparent picking is the discipline.
| Mistake | Fix | |---------|-----| | Re-describing the image in I2V | Only describe what changes/moves | | Generic camera language | Use exact preset names | | No style specified | Always include visual style + color grade | | Too many actions in one shot | Split into separate generations and chain them | | Contradictory movements | Don't combine Dolly In + Dolly Out in same shot | | Prompt over 512 chars (Cinema Studio) | Cut text, reduce @ tags, use pronouns | | Describing impact before action | Just describe the action, let AI render the result | | Specific martial arts moves | Use general fighting energy instead of named moves | | Multiple @ Elements in action scenes | Use @ for static scenes, plain text for action | | Mixing identity + motion in one block | Separate into Identity Block + Motion Block (see above) | | Aspect ratio inside the prompt body | Set aspect in the Higgsfield UI / output-format header (per-model enum: e.g. Kling 3.0 accepts 16:9 / 9:16 / 1:1 only — check higgsfield model get <model> or MCP models_explore). Describe framing in plain language ("full body" / "chest-up" / "wide establishing") not numerical ratios. |
> Output ratio is an enum, not a free-form value — and anamorphic is a style register, not an output dimension. Output aspect ratio is a hard, enumerated platform spec — Kling 3.0 emits 16:9 / 9:16 / 1:1 and nothing else. "Anamorphic" is a cinematography register (anamorphic lens flares, letterboxed compositional read, >2:1 framing aesthetic) that the model can render within a 16:9 output. "16:9 anamorphic" written as a single phrase in the prompt body is incoherent — pick one. Output ratio belongs in the header (and must be one of the enum values for the chosen model — check higgsfield model get <model> or the MCP models_explore equivalent before assuming). Anamorphic style cues belong in the Look line ("anamorphic-style flares, letterboxed composition") as a style request, not as an output dimension.
> Negative constraints: For a comprehensive list of artifacts to avoid (floating limbs, > face warping, flickering textures, etc.) and the prompt phrasing to prevent them, see > ../shared/negative-constraints.md. Always check the relevant categories for your prompt type.
At a ~1.5% video / ~1% image acceptance bar, most misses are variance, not a broken prompt. Serial single-variable iteration is the right tool for a systematic miss — the prompt is genuinely wrong. Run it on a stochastic miss and you're "fixing" a prompt that was already right, burning credits to re-roll the same dice one at a time. So decide the fork before you touch the prompt:
wardrobe contaminates every time, the cut count is always wrong) → systematic. The prompt is wrong. Iterate it, one variable at a time (next section).
(performance flat on one roll, camera off on another, physics odd on a third) → stochastic. The prompt is right; the roll wasn't. Stop touching the prompt. Lock it, fire a batch, and cull.
You don't have to eyeball this. The ledger already classifies every reject as structural or stochastic, and ratio <project> prints a verdict per shot tag: iterate (structural-dominant), batch+sel (stochastic-dominant), mixed, or low-n. Below five logged rows (LOW_N_THRESHOLD) the split is noise — the ledger stays silent and you call it by eye. Read the verdict at the decision point; don't iterate against a batch+sel tag.
When the verdict is stochastic, the move is variance-harvesting: hold the same locked prompt constant, roll N at once (grid generation / Batch Size in Cinema Studio, DoP Lite for cheap rolls), and cull to the keeper. This is the opposite of the stylistic-fan-out exception in the next section — that varies N different looks; this rolls N identical attempts because the prompt is right and only the dice are the problem. They read alike and are economically distinct: fan-out explores, harvest exploits.
The cull rubric — how to pick the keeper from a batch. Batching is worthless without a disciplined select. Don't pick "the prettiest"; select against the falsifiable success criteria you locked before generating:
legibility, named physics anchors — any roll that fails one is out, however nice it looks. (These are your structural failure modes; a batch can't fix a structural miss, only dodge a stochastic one.)
performance, camera, or composition you were re-rolling for. Best one wins.
kept; log the culled rollsrejected with their real reject_reason so the denominator stays honest and the verdict keeps sharpening. A harvested batch is correctly logged as one prompt_hash, N rows, one keep + N−1 stochastic rejects.
all — stop harvesting and go iterate the prompt.
When a prompt is close-but-not-right and you're about to regenerate, change exactly one variable per attempt. Subject detail, composition, motion behavior, lighting, or style — pick the one that's wrong, change only that, regenerate.
Why it matters: if you change two variables and the result improves, you don't know which change drove the improvement. If the result regresses, you don't know which change broke it. Either way you've spent a generation and learned nothing about the prompt. Single-variable iteration gives every regeneration a clean cause-and-effect signal — you keep what works, drop what doesn't, and converge on the right prompt fast.
The exception: once the prompt is locked and you're varying purely for stylistic exploration (e.g., five lighting variants of an already-approved scene), batching changes is fine. The rule applies during refinement, not during fan-out. (Don't confuse this stylistic fan-out — N different looks — with variance-harvesting above, which rolls N identical locked prompts to beat a stochastic miss. Both batch; only one changes the prompt.)
Workflow:
way you expected?
If you find yourself wanting to "fix everything at once," stop and ask which fix matters most. That one goes in this regeneration; the rest wait their turn.
Iteration also accumulates clutter — old prompt edits that no longer apply, stale reference images attached from earlier shots, contradictory clauses layered atop one another, prompts that have grown so long the model loses the load-bearing pieces inside the noise. Four hygiene patterns from production practice:
the sticky-note prop but you've moved to character generation, delete the sticky-note block. Otherwise the model occasionally pulls the stale prop into the new generation.
accumulate by appending rather than replacing, the prompt collects contradictory clauses (tight close-up from the old version, wide establishing from the new one). Symptom: output degrades into model-confusion artifacts. Counter: when iterating, replace the relevant clause in place rather than appending a new one.
Polaroid was on the fridge in shot 1 but pulled off in shot 2), remove the now-stale reference from the prompt window so the model is not still trying to place it.
characters from accumulated detail, ask Claude (or whatever prompt- construction surface you use) to optimize / study the context / sanitize the prompt — consolidate redundant clauses, drop now-obsolete qualifiers, preserve the load-bearing structure. The sanitize pass keeps the prompt inside the cap and inside the model's effective attention window.
The Iteration Rule above assumes you can identify which one variable to change. When you can't — the prompt produces output that's vaguely off and you can't name why — run the 6-Pass Diagnostic Sequence to find it. Each pass isolates one variable, in order, and tests it before moving on.
The order is not arbitrary. Subject and action carry the heaviest token weight (early-prompt positioning); camera and style come next; audio and output controls sit at the periphery. Diagnosing in this order surfaces the highest- leverage problem first and stops you from chasing a style-pass fix when the real issue was the subject description three layers up.
| Pass | Variable | Question | |------|----------|----------| | 1 | Subject | Is the character / object / focal element described unambiguously? | | 2 | Action | Is the action concrete (physics-based) and singular for the shot? | | 3 | Camera | Is the camera move named (Director Panel preset or specific verb), not implied? | | 4 | Style | Is the look anchored (palette, grade, lens, lighting), not adjective-only? | | 5 | Audio | If audio is part of the output, is it described as a parallel track with concrete sounds? | | 6 | Output | Are aspect ratio, duration, and resolution set deliberately for the shot's needs? |
How to use it: start at Pass 1. If the result improves when you sharpen the subject, you've found your variable — return to the Iteration Rule loop and keep going. If Pass 1 doesn't move the result, lock the subject, advance to Pass 2, and so on. The sequence is a finder, not a refinement loop. Once you know which variable is wrong, the Iteration Rule takes over.
Don't run all six passes blindly. Six regenerations cost six credits. The sequence's value is the order — most prompt failures land on Pass 1 or Pass 2 because early-prompt tokens dominate. If you reach Pass 4 without moving the result, the prompt may need a structural rewrite, not iteration.
These best practices apply to Cinema Studio 3.0's generation engine (Business/Team plan) and complement the MCSLA formula above. They are not a replacement — use MCSLA as the primary framework, then apply these refinements.
> For the user-intent layer that sits above MCSLA — what working mode > you're in (Exploration / Continuation / Bridging / Repair) and how each > routes through Seedance's prompt modes — see > ../higgsfield-seedance/SKILL.md § Working Modes. The disambiguation > between working modes and prompt modes lives in the same file, > immediately above.
Tell the model WHAT you want and HOW it should FEEL, not every micro-detail. In the short-form regime, short prompts (30–100 words) consistently outperform long ones. (Block-scaffold production briefs are the other regime — structure replaces the cap there; see ../higgsfield-seedance/SKILL.md § Official Prompt Architecture.) The model is an AI director you collaborate with, not a render engine you command.
The Director's Formula maps directly to MCSLA:
| Director's Formula | MCSLA Layer | Priority | |-------------------|-------------|----------| | Subject | S (Subject) | First 20–30 words (early tokens carry heavy weight) | | Action | A (Action) | First 20–30 words | | Scene | — (Context) | Supporting detail | | Camera | C (Camera) | After subject + action | | Style | L (Look) | After camera | | Constraints | — (Guardrails) | End of prompt |
Key insight: Subject + Action should appear in the first 20–30 words of every prompt. Early tokens carry disproportionate weight in the generation engine.
Different genres perform best with different prompt lengths and lead elements:
| Genre | Lead With | Target Length | Example Lead | |-------|-----------|---------------|-------------| | Product / E-commerce | Subject | 30–50 words | "A matte-black wireless earbud case rotates slowly on a marble pedestal..." | | Lifestyle / Social | Action | 40–60 words | "She reaches for the coffee mug, steam curling upward..." | | Drama / Narrative | Scene | 60–100 words | "Rain hammers a narrow Tokyo alley at 2 AM, neon signs reflecting in puddles..." | | Music Video | Style | 50–80 words | "Anamorphic flares, crushed blacks, 16mm grain..." | | Landscape / Travel | Scene | 30–60 words | "Dawn breaks over a volcanic ridge, mist pouring through the caldera..." | | Commercial / Brand | Style | 40–70 words | "Clean white studio, soft even lighting, product hero moment..." | | Anime / Artistic | Style | 50–90 words | "Cel-shaded lines, saturated palette, Studio Ghibli cloud physics..." |
> A texture word in a Style lead ("16mm grain") is a look choice and belongs > there. The same word trailing a prompt as a bare quality plea softens the whole > frame instead — the distinction lives in ../shared/negative-constraints.md > § Whole-Frame Degradation.
Kill these words — they add zero information and waste tokens:
| Slop Word | Replace With | |-----------|-------------| | beautiful | (delete — describe the specific visual instead) | | stunning | (delete — describe what makes it striking) | | epic | large-scale, sweeping, towering | | amazing | (delete — show, don't tell) | | dynamic | fast-tracking, whip-pan, handheld | | energetic | sprinting, jumping, arms pumping | | cinematic camera movement | slow dolly push / crane up / tracking shot | | cool transition | match-cut / whip pan / smash cut | | cinematic / cinematic lighting | a named referent — a director ("Wes Anderson symmetry"), a lighting setup ("golden-hour backlight, long shadows"), or a lens spec ("anamorphic 2.39:1, lens flare from a practical light") | | high quality / high-res / 4K look | (delete — resolution is a render setting, not a prompt word) |
> Why the substitute matters, not just the deletion: generic adjectives are > high-frequency labels spread across a huge, diffuse slice of training data, so > they pull the output toward nothing in particular. A director name, a lighting > setup, or a lens spec samples a narrow, well-trained distribution and > actually moves the result. For the Seedance-specific treatment of this, see > ../higgsfield-seedance/SKILL.md § Prompt-Craft Laws → Name the thing.
Use concrete physics consequences instead of mood words. The model responds to observable, physical details:
fist connects, sweat flies off in slow motion, opponent's head snaps backdoor slams open, dust erupts from the frame, light floods the dark roomtires spin, gravel sprays backward, chassis drops as acceleration kicks inThe model cannot infer intensity from images alone. Use adverbs to guide interpretation:
slowly, dramatically, violently, gently, frantically, deliberately, cautiously, explosively
Example: "She turns slowly, eyes narrowing deliberately, then explosively lunges forward."
Every action prompt should follow this arc:
Example: "The fighter plants her feet, fists clenching (charge-up). She throws a spinning kick that connects with the sandbag (burst). The bag swings violently, chain rattling, sand dust puffing from the seams (aftermath)."
Cinema Studio 3.0's generation engine does not support negative prompt syntax. Do not write "no blur" or "avoid shaky camera." Instead, use positive constraints — describe what you WANT:
locked-off static camera, no movementsharp focus throughout, deep depth of fieldsubject in sharp focus, background falling into soft bokehbright, evenly lit, overcast daylightThe two depth-of-field spellings are not interchangeable — pick by which plane has to stay sharp (../shared/negative-constraints.md § Depth of field — two substitutes, two intents).
Describe audio separately in prompts. BGM, ambient SFX, and dialogue are handled as parallel tracks via dual-channel stereo generation:
A barista grinds coffee beans, pours steaming water over the filter.
Camera: tight close-up, slow dolly across the counter.
Style: warm tones, shallow depth of field.
Audio: the whir of the grinder, water bubbling through the filter,
ceramic mug placed on a wooden counter with a soft clink.
Soft jazz piano in the background, barely audible.Sound design descriptions like "the scratch of frosted glass, rustling plush fabric, gentle tapping on acrylic" directly influence the generated audio output.
Before writing a Seedance prompt, identify which archetype the scene fits. The archetype dictates camera behavior, spatial logic, and what changes across time. This is a planning layer on top of MCSLA — pick the archetype first, then fill in MCSLA.
| Archetype | Camera focus | Space dynamic | |-----------|-------------|---------------| | Pursuit | Distance closing/opening. Pursued ahead in frame, pursuer behind | Path narrows/opens | | Duel | Camera lower on dominant side; dominance MUST alternate | Fighters trade position | | Impact | Build-up slow → hit fast → aftermath slow | Point of contact = center |
Decision tree: Chase? → Pursuit. Two opponents trading advantage? → Duel. Single decisive contact moment? → Impact. None → default Duel.
Duel rule: neither side dominates more than one consecutive beat. If one fighter dominates the whole scene, describe it as a one-sided assault, not a duel.
| Archetype | What changes | Camera signature | |-----------|-------------|-----------------| | Journey | Position in space — road, flight, walking | Tracking, aerial, traveling alongside. Landscapes pass. | | Atmosphere | Nothing — mood IS the content. Rain on glass, empty street. | Minimal movement. Slow push-in or static hold. Micro-changes carry all drama. | | Reveal | Hidden → visible. Door opens, fog lifts, camera rounds corner. | Pan, crane, dolly reveal. Camera controls WHEN viewer sees the subject. |
Decision tree: Subject moves through space? → Journey. Something hidden becomes visible? → Reveal. Nothing changes, mood is the content? → Atmosphere. None → default Atmosphere.
| Archetype | Power dynamic | Camera signature | |-----------|--------------|-----------------| | Confrontation | Shifting — both push. Dominance trades per exchange. | Tight OTS, camera crosses axis on power shift. | | Interrogation | Asymmetric — one extracts, one resists. | Low-angle on questioner, push-in on silence. | | Negotiation | Balanced — both need something. | Symmetrical framing, matching shot sizes. |
Decision tree: Both pushing, dominance trading? → Confrontation. One extracting, one resisting? → Interrogation. Both need something, balanced? → Negotiation. None → default Confrontation.
Dialogue word limit: ~25–30 spoken words fit into 15 seconds. If the user provides more, keep the line where dominance flips (the power-shift exchange), 1 line before (setup), 1 line after (reaction). Convert the rest to physical behavior.
These are hard rendering constraints of the Seedance 2.0 engine — violating them causes broken output regardless of prompt quality.
> For the eight named substrate channels that micro-expressions > decompose into, see ../../vocab.md § Emotion as Visible Behavior — > Channels.
Every cut must change both shot size AND camera character. The scale runs extreme wide → wide → medium → MCU → close-up → ECU. Camera character: Handheld | Static | Stabilized tracking | Crane | Aerial — never repeat across a cut.
Bad (same camera character): MS handheld → CU handheld Good (both change): MS handheld → ECU static-locked
Inserts are sub-second (0.3–0.5s) dramatic punctuation at any shot size. Rules:
Never describe characters by age in Seedance prompts. Trigger words to avoid: boy, girl, child, kid, young, teen, little. Seedance age inference is unreliable and drifts across shots.
Scenes start already in progress unless the user explicitly says "starts with…" or "ends with…". Don't waste the first 2 seconds on setup beats.
> Full Seedance director reference including bilingual EN+ZH JSON output format is dropped in the project docs/ folder as Seedance 2 Skill.md — use it when you need the standalone director-mode prompt with scene-archetype routing and age-blind rules baked in.
higgsfield-soul — Character consistency, Soul ID, micro-expressionshiggsfield-camera — All named camera controlshiggsfield-style — Visual styles, color grades, lightinghiggsfield-models — Model selectionhiggsfield-troubleshoot — Fix failing generationstemplates/ — Annotated genre-specific prompt templates| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 23,919 | 13,171 | -45% | 1 | 1 | 0% | 1,705 | 11,994 | +603% | 0 | 0 | — |
case-02 | fail→pass | 15,789 | 21,642 | +37% | 1 | 1 | 0% | 2,089 | 12,532 | +500% | 0 | 0 | — |
case-03 | fail→fail | 15,107 | 14,803 | -2% | 1 | 1 | 0% | 1,844 | 11,987 | +550% | 0 | 0 | — |
case-04 | pass→pass | 12,356 | 13,277 | +7% | 1 | 1 | 0% | 1,522 | 11,743 | +672% | 0 | 0 | — |
case-05 | fail→fail | 11,178 | 10,955 | -2% | 1 | 1 | 0% | 1,512 | 11,512 | +661% | 0 | 0 | — |
case-06 | fail→pass | 11,349 | 8,034 | -29% | 1 | 1 | 0% | 1,456 | 11,221 | +671% | 0 | 0 | — |
case-07 | pass→pass | 11,572 | 12,610 | +9% | 1 | 1 | 0% | 1,529 | 11,897 | +678% | 0 | 0 | — |
case-08 | fail→pass | 38,057 | 10,345 | -73% | 1 | 1 | 0% | 2,044 | 11,411 | +458% | 0 | 0 | — |
case-09 | pass→pass | 15,718 | 9,990 | -36% | 1 | 1 | 0% | 2,004 | 11,238 | +461% | 0 | 0 | — |
case-10 | fail→pass | 13,485 | 13,927 | +3% | 1 | 1 | 0% | 1,547 | 11,807 | +663% | 0 | 0 | — |
case-11 | fail→pass | 22,666 | 11,989 | -47% | 1 | 1 | 0% | 2,195 | 11,920 | +443% | 0 | 0 | — |
case-12 | fail→pass | 16,822 | 12,373 | -26% | 1 | 1 | 0% | 1,850 | 11,916 | +544% | 0 | 0 | — |
case-13 | fail→pass | 14,767 | 9,098 | -38% | 1 | 1 | 0% | 1,968 | 11,429 | +481% | 0 | 0 | — |
case-14 | fail→pass | 12,621 | 9,505 | -25% | 1 | 1 | 0% | 1,537 | 11,327 | +637% | 0 | 0 | — |
case-15 | fail→pass | 18,765 | 18,530 | -1% | 1 | 1 | 0% | 2,294 | 12,841 | +460% | 0 | 0 | — |
case-16 | fail→pass | 10,240 | 8,836 | -14% | 1 | 1 | 0% | 1,354 | 11,482 | +748% | 0 | 0 | — |
case-17 | fail→fail | 14,584 | 11,769 | -19% | 1 | 1 | 0% | 1,665 | 11,395 | +584% | 0 | 0 | — |
case-18 | pass→pass | 13,717 | 9,922 | -28% | 1 | 1 | 0% | 1,585 | 11,365 | +617% | 0 | 0 | — |
case-19 | fail→fail | 15,017 | 35,393 | +136% | 1 | 1 | 0% | 1,561 | 12,302 | +688% | 0 | 0 | — |
case-20 | pass→pass | 19,840 | 19,666 | -1% | 1 | 1 | 0% | 2,691 | 12,773 | +375% | 0 | 0 | — |
case-21 | pass→pass | 24,407 | 13,424 | -45% | 1 | 1 | 0% | 2,206 | 12,190 | +453% | 0 | 0 | — |
case-22 | pass→pass | 18,500 | 16,793 | -9% | 1 | 1 | 0% | 3,075 | 12,854 | +318% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.