Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user asks about audio in Higgsfield videos, needs to add dialogue or lip-sync, wants sound effects or ambient sound in generated video, asks about music or BGM in output, or is using any audio-capable model (Kling 3.0, Seedance 1.5 Pro, Seedance 2.0, Veo 3/3.1, Grok Imagine Video). Also use when the user's prompt would benefit from audio direction but they haven't mentioned it. Also use when the user wants standalone audio — a soundtrack, ambience bed, multi-speaker scene audio (See
.claude/skills/osidemedia-higgsfield-audio/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 421% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 357% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 435% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 453% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 572% | 0% |
Routing aids — read the linked sections for the full rules.
@Audio1 is a conditioning INPUT — beat sync, the [AUDIO: Xs] script block, and the first-15s extraction trap →seed_audio, standalone) = whole-scene audio in ONE pass — multi-speaker dialogue + music + SFX + ambience mixed →seed_audio, qwen_audio_tts (NEW — Qwen 3.0 TTS Flash, expressive instructions + cloned voices), text2speech_v2 (5 engines incl. cozy_voice), plus 3 game-pipeline-only tools — distinct from in-video joint audio →| Model | Audio type | Dialogue | SFX | Ambient | BGM | Lip-sync | |-------|-----------|----------|-----|---------|-----|----------| | Kling 3.0 / Omni | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Multi-language | | Seedance 2.0 | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Multi-language | | Seedance 1.5 Pro | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Best lip-sync | | Veo 3 / 3.1 | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ English best | | Grok Imagine Video | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ | | All other models | ❌ | — | — | — | — | — |
"Native joint" means audio and video are generated simultaneously in one pass — not layered on after. This produces natural synchronization without post-production.
Models without native audio: add audio in post with Lipsync Studio or external tools.
Every audio-capable prompt should consider four layers. You don't need all four in every prompt, but knowing which to include gives the model clear direction.
Put dialogue in quotes. Be explicit about who speaks, their tone, and language.
She says: "We need to leave. Now."
He whispers: "Not yet."Best practices:
She speaks in Cantonese: "走啦"Taiwanese Mandarin, Shanghainese), Japanese, Korean, Spanish, Indonesian
Describe SFX at the point they happen. Tie them to visible actions.
The glass shatters on the floor — sharp crack, then settling tinkle.
Footsteps on wet concrete — splashing, rhythmic.
A door slams shut — heavy metal, echoing.Best practices:
Set the acoustic environment. This is the continuous sound bed.
Ambient: quiet café murmur, espresso machine, rain against windows.
Ambient: forest at night — crickets, distant owl, gentle wind through leaves.
Ambient: busy intersection — traffic, horns, construction in the distance.Best practices:
Don't name songs or artists (content filter). Describe the musical texture.
BGM: slow piano, minor key, melancholic.
BGM: tense orchestral build — low strings, rising.
BGM: lo-fi hip-hop beat, warm vinyl crackle, relaxed.Best practices:
NO BGM is a spec, no music is a preference[DEMO — Joey cinema-director-v3, 2026-08-16] [UNPROVEN HERE] When a piece must carry no score, the phrase matters. no music reads as a weak stylistic preference and loses to the model's strong prior that generated video wants a bed under it. NO BGM reads as a production term — a hard spec — and is the form to write. Expand it once on first use so the abbreviation is unambiguous, then let it carry.
Lead positive, then negate. Name what the audio is before naming what it is not — diegetic sources tied to surfaces and materials, plus room tone. A suppression clause with nothing positive in it leaves the model to decide what "silence" sounds like, and it decides in favour of a pad.
Audio: diegetic sound only — footsteps on wet stone, fabric shift, breath, room tone.
NO BGM — no background music of any kind. No score, no soundtrack, no instrumental,
no underscore, no ambient musical pad, no drone, no tone bed. Nothing musical at any point.Promote it to the top on a scene that must land silent. Audio instructions carry more weight early; by the time the model reaches a closing audio block it has already decided what the piece sounds like. State NO BGM in the header alongside shot count and cut policy, then restate it as the closing audio clause.
> Enumerate with care — this cuts against the house rule on negation. > ../shared/negative-constraints.md and the repo's staging-reference doctrine both > hold that naming a thing under a negation ships the token anyway and can prime > the very output you are refusing. The source's long list (score, soundtrack, > instrumental, underscore, ambient pad, drone, tone bed, swell, sting, humming, > whistling, lyrics) is its answer to a real failure — a bare negation lets the model > supply an "ambient texture" and consider the instruction honoured — but it is one > practitioner's fix and is not measured here. Default to the short form above; > reach for the full enumeration only when a short form has already failed on the > shot in front of you, and expect the enumeration itself to carry some priming risk.
Attached-track lock. When an audio or video track is attached it is the sole and complete audio source, and it owns all internal timing — never impose per-beat timing on a lipsync take:
AUDIO: the attached clip @Video 1 is the sole and complete audio source for this
sequence. Generate no additional audio of any kind — no room tone, no foley, no
ambience, no breath, no added dialogue, NO BGM.The unheard-track technique lets bodies perform to music that is not in the mix — useful when the score is added in post: state that the track is inaudible, then describe the performance against it ("singing roughly in time to the unheard 87 BPM beat"), plus the diegetic layer that IS heard.
Add audio cues naturally within your prompt or as a dedicated block at the end.
A woman walks into a quiet library. Her heels click on the marble floor — each step
echoing. She whispers to the librarian: "Do you have the Collected Letters?"
Distant page turns. A clock ticks somewhere above.[Scene description — visual content, action, camera]
Audio:
Dialogue: She says "We leave at dawn." He replies: "I'll be ready."
SFX: coffee cup set down, chair scraping back
Ambient: early morning kitchen — birds outside, kettle just boiled
BGM: none — silence emphasizes the tensionLip-sync is the most failure-prone audio feature. Follow these rules strictly:
> Expressive facial acting around the words — forced smiles, leaking fear, > mixed emotions during a spoken line — is driven separately by FACS Action Unit > codes per beat. Let lip-sync shape the phonemes; schedule the brow/eye/cheek > AUs for the performance. See ../higgsfield-facs/SKILL.md § Dialogue & > Monologue Facial Acting.
locked-off static camera or slow Dolly In onlynodding, turning head, looking aroundcompete with the lip engine and cause desync
the generative audio engine to override your dialogue
Multi-person lip-sync matching is an unresolved limitation across all models. The production workaround:
Field-observed word budgets for reliable lip-sync in a ~15s in-video Seedance dialogue clip — not official limits, and not the same as how many words the model can voice. The acoustic budget ≠ reliable-sync budget: the model will happily speak more words than it can keep synced to the mouth.
| Language | Reliable-sync budget (~15s clip) | Notes | |----------|----------------------------------|-------| | English | ~16–20 words (5–10 per line) | Strongest Western language | | Mandarin | — | Strongest sync overall | | Russian | ~10–15 words | Weak — budget conservatively | | Japanese / Korean | Under-tested | No reliable field numbers yet |
Cross-language sizing unit: "one short sentence ≈ one breath." Write dialogue in breath-sized sentences and count breaths, not seconds.
On surfaces that accept a spoken-voice reference, an attached rights-cleared voice recording drives lip-sync directly — the model syncs the mouth to your recording instead of synthesizing a voice first. This is the most reliable field-reported path for non-English dialogue (it sidesteps the weak-language sync budgets above). Rights-sensitive: only use recordings you have clear rights to — cloned or scraped voices are out.
@Audio1)The most under-used Seedance 2.0 capability: an uploaded audio file is a conditioning input, not just an output track. The model spec lists audio as a reference media role alongside image / video, and generate_audio (native sound output) is documented as independent of the audio reference medias — i.e. the uploaded file conditions the generation, and whether the clip also gets generated sound is a separate switch.
This means @Audio1 has two distinct jobs, and you pick one per shot:
| Use | What @Audio1 does | Prompt discipline | |-----|---------------------|-------------------| | Audio-as-output | Plays the uploaded track unmodified as the clip's soundtrack | Timestamp-anchor it (plays exactly as uploaded from 0s to end) and remove all ambient/SFX/music tokens so the engine doesn't override it (see § Seedance 2.0 below) | | Audio-as-driver (beat sync) | Drives the visuals — cut timing, camera acceleration, action pace, energy peaks | Write the audio→visual mapping explicitly (below). The clip can still get generated sound, or set generate_audio false for visuals-only. | | Audio-as-performance | The character on screen performs the track — hums, sings, plays along, moves to it — hitting the actual notes | Scope it to the one property you want (§ Scope an audio reference, below). Unscoped, it also lends its voice. |
> Why it works (author's model — empirical, not in the official spec): the > temporal branch that reasons about motion and pacing reads the sound's > structure — beat positions, dynamic contour, timbral texture, song-structure > sections — and maps it to visual rhythm. Treat the mechanism as a working > model; treat the capability (audio reference role) as confirmed.
Every image reference in this stack is scoped in both directions: what it locks, and what must be read past ("ignore the sheet's grey background", "not its camera vantage"). Audio references have had no equivalent vocabulary, and they need one for the same reason — a sound file carries several properties at once, and an unscoped reference lends all of them.
The case that shows it DEMO — Higgsfield "AI Love Stories" tutorial, 2026-08]: a character had to hum a specific tune on camera. Without a reference the model invented a different melody every take — the reported result of the no-reference control was a performance that was off-key with no rhythm. With a voice memo of the tune attached, it hit the notes on the first take.
But the memo was somebody else's voice, and the character has his own. So the reference was scoped in prose:
@Audio1 is the reference for the HUMMED LINE only. Take from it ONLY the
melody: the exact notes, pitches and intervals, the tempo, the phrasing and
the rhythmic pause — note for note, beat for beat, nothing improvised. Do NOT
copy the voice, timbre or vocal identity heard in the recording — he hums in
his OWN natural speaking voice, the same voice he speaks his lines with. All
spoken dialogue is performed as scripted below, not taken from any audio.The pattern, generalised — three parts, and the third is the one that gets skipped:
intervals; tempo and phrasing; the rhythmic pause. Not "the song".
Sound files carry a performer as well as a performance.
speaking voice, the same one he speaks his lines with". A reference that is only told what not to do leaves the model to pick, and it will pick the reference.
The same three-part shape works for other audio properties:
| Riding | Excluded | Sourced instead from | |---|---|---| | melody, tempo, phrasing | voice, timbre, identity | the character's own speaking voice | | rhythm and accent pattern | instrumentation, key | the scene's own diegetic sound | | emotional contour, dynamics | the words | the scripted dialogue |
Scope note. Whether a reference track becomes the spoken output differs by model line — see § Audio by Model. The scoping vocabulary above is about which property transfers, and is written to be read alongside whatever that line does with an attached track, not instead of it.
Rights: the memo above was recorded by the person for this purpose. Same constraint as the voice-reference lip-sync path — use recordings you have clear rights to.
UNPROVEN HERE] — one production, one property (melody). The mechanism is already confirmed (audio is a reference media role); what is unproven here is how far prose scoping steers it. Cheap to test on any tune you own.
Upload an MP3 as @Audio1, then map audio characteristics to visual elements. The minimum is three sentences, each handling one thing — rhythm source / which visual responds / how energy maps to the arc:
Use @Audio1 as the rhythmic foundation. Sync camera transitions to the beat
positions. Visual energy builds with the audio crescendo and peaks at the drop.You can assign different visual elements to different audio characteristics — mixing audio-to-visual the way you'd mix a track:
@Audio1 drives the visual rhythm. Camera cuts land on the downbeats. Subject
movement accelerates into the build, holds at the peak, releases on the drop.
Colour temperature shifts warmer with the crescendo.Camera ← beat position. Movement ← dynamic contour. Colour ← overall energy arc.
It stacks with other references — character from @Image1, camera style from @Video1, rhythm from @Audio1, processed together:
@Image1 as character reference. Follow @Video1 camera-movement style. @Audio1 as
rhythmic foundation — sync all camera transitions to the beat positions.
Character movement should pulse with the music.The one constraint: @Video1 camera style and @Audio1 rhythm have to be temporally compatible. A slow continuous dolly pulled from a video reference fighting an EDM track sends the temporal branch conflicting instructions — same failure class as mixing reference images of clashing styles. Pick references that can coexist. (Sibling of ../higgsfield-seedance/SKILL.md § Reference Roles → Load-Bearing Rule: references stay in their lanes.)
[EMPIRICAL — MiniMax H3 skill corpus, re-derived; cross-model editing craft] Beat sync governs what happens inside a clip; these three laws govern the timeline the clips land on:
in post — never per-clip audio stitched end to end. A join in the music is audible before a join in the picture is visible.
drop. Never hard-cut inside a sung vowel unless the incoming shot is an ECU whose mouth shape continues that vowel: lip continuity is an edit constraint, not only a prompt constraint.
match grade exactly. One unified fine-grain pass plus one LUT across the whole timeline, applied as a deliberate finishing step, hides the inter-clip color variance that would otherwise read as a continuity error.
[AUDIO: Xs] script block — dialogue + SFX + lip-sync from text aloneNo microphone, no recording. A timestamped script inside the prompt text generates voices, SFX, and lip-sync. Quoted text → speech with automatic lip-sync; physical descriptions → sound effects. Each marker is a timestamp in the clip:
[AUDIO: 0s] heavy footsteps on concrete, echoing in a corridor
[AUDIO: 2s] door bursting open, impact bang
[AUDIO: 3s] character says "Nobody move"
[AUDIO: 5s] tense silence, distant traffic
[AUDIO: 7s] character says "Put it down. Slowly."
[AUDIO: 9s] object placed on table, soft thudThe model generates the voice first, then maps facial movement to the waveform — so lip-sync quality is mostly set by how precisely you wrote the dialogue. Exact quoted text outperforms paraphrase. It works across languages (write the line in Spanish/Japanese/French → speech with phoneme-level lip-sync in that language).
This obeys the same physical rules as § Lip-Sync Rules above: a strong @Image1 character reference gives a consistent mouth structure to animate, and close-up framing beats wide (a small face has too few pixels to sync). Keep individual dialogue beats inside the 3–8s accuracy window.
It combines with beat sync in one generation — uploaded music as the rhythmic foundation, the script block as foreground dialogue/SFX, cuts synced to the beat:
@Audio1 as background music. Sync camera transitions to the beats.
[AUDIO: 0s] music from @Audio1 begins
[AUDIO: 3s] character says "This changes everything"
[AUDIO: 5s] sharp breath — beat drop hits simultaneously
[AUDIO: 8s] character says "Let's go"The audio reference limit is 15s, and the model takes the first 15s of whatever you upload. Drop in a full 3-minute track and you almost always feed it the intro — low energy, often ambient, no rhythmic drive. Nothing for the temporal branch to map.
The right 15s follow a build → drop arc: rising tension into a peak. That dynamic gradient is what becomes visual energy structure. A segment with uniform energy gives the model beats to detect but no arc — output is rhythmically synced but dramatically flat.
Where the window lives:
Extract exactly that segment before uploading. MP3 at ≥256kbps — lower bitrate degrades beat detection. Don't upload the full track and hope; pick the window, cut it, upload that. (Flipping the workflow — audio in first, visuals built around it — changes the output at a structural level, not subtly.)
Audio Speaker Attribution Format (V3/O3):
[Speaker: Character Name] "dialogue" in a [warm/confident/excited] [male/female] voice with [accent].
Add [sound: footsteps / rain / door closing] when [action].
Background ambient: [environment description].[EMPIRICAL — third-party, China Ark lane, 2026-08]:per-clip duration is enforced at 1.8–15.2s, and the total across all attached clips is also capped at ≤15.2s — three individually-legal 6s clips get rejected (captured 400 errors). Measured on the China Ark lane by a third party; Higgsfield's own proxy enforcement is unverified — if a multi-clip attach fails, this total cap is the first suspect.
"Audio @Audio1 plays exactly as uploaded from 0s to end. Do not modify."Then remove all ambient/SFX/music tokens to prevent the generative engine from overriding.
@Audio1 is also a visual driver — beat sync, the [AUDIO: Xs] script block,and the first-15s extraction trap are all in § Audio as a Conditioning Input above.
> Diegetic-only convention for the prompt body — a > prompt-authoring discipline that sits on top of Seedance 2.0's > audio capability. BGM is a valid audio layer (see § The Four > Audio Layers above) — that's what Seedance can generate. The > diegetic-only convention is what you should write in the > prompt body: only sounds that physically exist in the scene > (footsteps on wet pavement, fabric whip on motion, breath, room > tone, weather, weapon fire, crowd reaction, stage haze) rather > than naming songs, lyrics, or score cues. If music is intended > for the final cut, layer it in post rather than in the prompt > body. > > Two reasons the discipline matters even though BGM is > supported: (i) score descriptors ("dramatic strings", > "orchestral swell") underdetermine the generated audio and > routinely produce generic music beds at odds with the scene; > (ii) the timestamp-anchoring + remove-all-music-tokens > pattern in the bullets above already enforces this discipline > when an MP3 audio reference is uploaded — the diegetic-only > convention generalizes that pattern to all Seedance prompts > whether or not an audio reference is attached.
"This must be it," he murmured.tires screeching loudly| Problem | Cause | Fix | |---------|-------|-----| | Lip-sync completely off | Audio > 8s, or head motion tokens present | Trim to 5s, remove nodding/turning tokens | | Model replaces uploaded audio | Ambient/music tokens in prompt invite generative override | Add timestamp anchoring phrase, remove all ambient/music tokens | | Dialogue missing entirely | Non-MP3 format used (Seedance 2.0) | Convert to MP3 128-320kbps | | SFX drowns out dialogue | Too many SFX cues competing | Reduce to 1-2 SFX per shot, prioritize dialogue | | Audio sounds robotic | Flat emotional cues | Add emotional direction: "says warmly", "whispers with urgency" | | Background music too loud | BGM description too prominent in prompt | Move BGM to end of prompt, reduce detail, or say "subtle BGM" |
Not every prompt needs audio direction. Skip audio cues when:
> Negative constraints: For audio-specific artifacts (lip-sync desync, background music > overriding dialogue, SFX drowning dialogue) and their prevention phrases, see > ../shared/negative-constraints.md — Temporal/Consistency Artifacts section.
Cinema Studio 3.0 introduces native audio-video joint generation — a fundamental shift from models that treat audio as a post-processing step.
Audio is generated simultaneously with video via a unified multimodal architecture. This means:
Always describe audio as a separate section in your prompts. The generation engine handles three parallel audio tracks:
A chef slices vegetables rapidly on a wooden cutting board.
Camera: tight close-up tracking the knife.
Style: warm kitchen lighting, shallow depth of field.
Audio: rhythmic chopping on wood, oil sizzling in a nearby pan,
soft clinking of ceramic bowls. Light acoustic guitar BGM.| Parameter | Limit | |-----------|-------| | Accepted formats | MP3, WAV | | Max audio clips | 3 per generation | | Combined duration | ≤15s total | | Single file size | <15MB | | MP3 bitrate | 128–320 kbps |
Available but experimental in Cinema Studio 3.0:
Control speaking style, accent, and language by referencing a video with the desired voice:
Voiceover tone references @Video1. The narrator describes the product
in a warm, conversational tone. "This changes everything."
Audio @Audio1 plays exactly as uploaded from 0s to end.
Do not modify or replace the audio content.Dialects written directly in the prompt work — the model understands regional speech patterns. Write dialogue in the target dialect for authentic delivery.
When uploading reference audio that must play unmodified:
Audio @Audio1 plays exactly as uploaded from 0s to end.
Do not modify or replace the audio content.Then remove all ambient/SFX/music tokens from the prompt to prevent the generation engine from overriding the uploaded audio with generated sound.
Describe specific foley, not generic moods:
Wrong: nice ambient sounds, pleasant background noise
Right: the scratch of frosted glass, rustling of plush fabric, gentle tapping on acrylic, popping of bubble wrap, wooden floor creaking under bare feet
Specific sound descriptions directly influence the generated audio output. The more precise the foley description, the more accurate the result.
Separate from in-video joint audio above: Seed Audio 1.0 (ByteDance, released 2026-06-23 at the FORCE conference) is a standalone one-pass whole-scene audio generator. One generation produces multi-speaker dialogue + music + SFX + ambience, already mixed — a radio-drama scene, not a single voice track. Use it to build a soundtrack for footage you'll assemble in post, or scene audio that has no video at all.
| You need | Use | Why | |----------|-----|-----| | A whole scene's soundtrack: several speakers + music + SFX + ambience, mixed in one pass | Seed Audio 1.0 (seed_audio) | One-pass scene audio; script-style prompt drives the whole mix | | One clean voice track (narration, single-speaker VO) | text2speech_v2 (pick an engine) | Single-voice TTS — simpler, engine-selectable | | Sound baked into the generated video, synced to on-screen action and lips | Seedance generate_audio (in-video) | Native joint generation — audio and visuals in the same pass (see § Audio as a Conditioning Input) |
Model id seed_audio (output_type audio). Parameters:
| Param | Range / options | Default | |-------|-----------------|---------| | format | wav / mp3 / pcm / ogg_opus | wav | | sample_rate | 8000–48000 Hz | 24000 | | speech_rate | −50..100 | 0 | | loudness_rate | −50..100 | 0 | | pitch_rate | −12..+12 | 0 | | voice_type + voice_id | preset \| element — must travel together | none |
Media roles: image_references + audio_references. Per the fal schema: up to 3 reference audio clips (each ≤30s, ≤10MB) XOR one image reference — image and audio refs cannot combine. Reference audio inputs in the prompt by load order: @Audio1, @Audio2, @Audio3.
Everything in this subsection is community-converged practice, not spec — treat as a starting point, not a guarantee. Write the prompt as a radio-drama script:
[Scene: busy coffee shop, morning]Host (warm, upbeat): "…"[sound: espresso machine, soft jazz fades in]format: wav) when the audio is headed for postCompact worked example:
[Scene: rain-soaked night market, closing time]
Vendor (tired, warm): "Last skewers — half price, take them."
Girl (excited): "Two! No — three!"
[sound: rain drumming on tarp canopy, a scooter passing in the distance]
Vendor (chuckling): "Three it is. Careful, they're hot."
[sound: coins dropped on a metal tray, charcoal hiss]
Music: a lonely muted trumpet fades in under the rain, wistful but hopeful.The live standalone-audio catalog, reconciled against the models_explore snapshot of 2026-08-01 (../../specs/models_explore_snapshot_audio_2026-08-01.json; generated table: ../../specs/AUDIO-MODEL-SPECS.md, machine twin ../../specs/audio-model-specs.json — regenerate with python3 scripts/sync_specs.py --type audio). The Audio tab's UI tools — Voiceover (text → speech), Change Voice (swap a voice in any video), Translation (translate speech in any video) — sit on top of these models:
| Model id | Name | What it does | Availability | |----------|------|--------------|--------------| | seed_audio | Seed Audio 1.0 (ByteDance) | One-pass whole-scene audio: dialogue + music + SFX + ambience (§ above) | General | | qwen_audio_tts | Qwen Audio 3.0 TTS Flash (Alibaba) | Expressive TTS: natural-language instruction for emotion/dialect/speed, preset or cloned reference-element voices, 13 language hints | General (NEW 2026-08-01) | | text2speech_v2 | Text to Speech V2 | Single-voice TTS; engine via variant: elevenlabs, minimax, seed_speech, vibe_voice, cozy_voice (NEW); preset or reference-element voices (voice_type + voice_id) | General | | sonilo_music | Sonilo Music (FAL) | Text-to-music with controllable duration | Game pipeline only | | mirelo_text_to_audio | Mirelo Text to Audio (FAL) | Text-to-audio SFX with controllable duration | Game pipeline only | | inworld_text_to_speech | Inworld TTS (FAL) | Preset-voice TTS, ~110 voices across en/zh/ja/ko/es/fr/de/ru/… | Game pipeline only |
Engine picks within text2speech_v2: seed_speech when the deliverable is multilingual voiceover/narration; elevenlabs (Eleven v3) when fine emotional/tone control matters; vibe_voice for long-form narration. These are standalone audio generators — distinct from the native joint audio baked into Kling 3.0 / Seedance 2.0 / Veo during video generation. (Catalog reflects the 2026-08-01 snapshot; verify live before quoting pricing or availability.)
A post-generation alternative to prompting audio at all: upload the finished clip to Supercomputer and ask for an analyzed voice-over (e.g. "Analyze the video and create a voiceover for it in the style of wildlife documentaries") — the agent analyzes the footage, writes a script, and offers voices to pick from. Shown working in Higgsfield's Seedance-4K tutorial; useful when the visuals are already locked and only narration is missing.
higgsfield-seedance-vfx — Footage transforms whose payoff is a camera move synced to a spoken line (crash-zoom / push-in), or preserving the source talk track through a transform (SFX and source dialogue only); see ../higgsfield-seedance-vfx/references/dialogue-timing.mdhiggsfield-models — Which models support native audiohiggsfield-troubleshoot — Audio failure diagnosishiggsfield-cinema — Cinema Studio audio workflow with Kling 3.0higgsfield-vibe-motion — Motion graphics with audio (different from AI-generated audio)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 22,311 | 11,529 | -48% | 1 | 1 | 0% | 1,742 | 10,997 | +531% | 0 | 0 | — |
case-02 | fail→pass | 16,650 | 11,427 | -31% | 1 | 1 | 0% | 2,142 | 11,157 | +421% | 0 | 0 | — |
case-03 | fail→pass | 17,931 | 14,822 | -17% | 1 | 1 | 0% | 2,556 | 11,687 | +357% | 0 | 0 | — |
case-04 | fail→pass | 31,805 | 11,673 | -63% | 1 | 1 | 0% | 2,123 | 11,356 | +435% | 0 | 0 | — |
case-05 | pass→pass | 15,372 | 20,835 | +36% | 1 | 1 | 0% | 2,260 | 11,300 | +400% | 0 | 0 | — |
case-06 | fail→pass | 14,450 | 9,696 | -33% | 1 | 1 | 0% | 1,996 | 11,045 | +453% | 0 | 0 | — |
case-07 | fail→pass | 12,247 | 9,694 | -21% | 1 | 1 | 0% | 1,612 | 10,831 | +572% | 0 | 0 | — |
case-08 | pass→pass | 24,989 | 37,897 | +52% | 1 | 1 | 0% | 1,736 | 10,944 | +530% | 0 | 0 | — |
case-09 | pass→pass | 13,516 | 10,866 | -20% | 1 | 1 | 0% | 1,890 | 10,943 | +479% | 0 | 0 | — |
case-10 | fail→pass | 11,878 | 12,270 | +3% | 1 | 1 | 0% | 1,535 | 10,345 | +574% | 0 | 0 | — |
case-11 | fail→pass | 16,402 | 8,041 | -51% | 1 | 1 | 0% | 2,341 | 10,637 | +354% | 0 | 0 | — |
case-12 | fail→pass | 13,390 | 7,013 | -48% | 1 | 1 | 0% | 2,046 | 10,339 | +405% | 0 | 0 | — |
case-13 | fail→pass | 20,317 | 15,997 | -21% | 1 | 1 | 0% | 2,657 | 11,639 | +338% | 0 | 0 | — |
case-14 | fail→pass | 25,734 | 12,876 | -50% | 1 | 1 | 0% | 3,274 | 11,254 | +244% | 0 | 0 | — |
case-15 | fail→pass | 16,162 | 11,646 | -28% | 1 | 1 | 0% | 2,344 | 11,150 | +376% | 0 | 0 | — |
case-16 | fail→pass | 40,195 | 11,154 | -72% | 1 | 1 | 0% | 1,265 | 10,937 | +765% | 0 | 0 | — |
case-17 | pass→fail | 18,222 | 6,628 | -64% | 1 | 1 | 0% | 2,227 | 10,346 | +365% | 0 | 0 | — |
case-18 | fail→pass | 15,166 | 6,177 | -59% | 1 | 1 | 0% | 2,174 | 10,373 | +377% | 0 | 0 | — |
case-19 | fail→pass | 18,861 | 9,391 | -50% | 1 | 1 | 0% | 2,888 | 10,906 | +278% | 0 | 0 | — |
case-20 | fail→pass | 17,497 | 11,578 | -34% | 1 | 1 | 0% | 2,284 | 11,246 | +392% | 0 | 0 | — |
case-21 | fail→pass | 16,583 | 14,956 | -10% | 1 | 1 | 0% | 2,669 | 11,593 | +334% | 0 | 0 | — |
case-22 | pass→pass | 14,197 | 11,887 | -16% | 1 | 1 | 0% | 2,064 | 11,116 | +439% | 0 | 0 | — |
case-23 | pass→pass | 5,124 | 6,697 | +31% | 1 | 1 | 0% | 716 | 10,315 | +1341% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +65 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.