| name | create-cartoon-music-video |
| description | Produce a paid-social ad in the "cartoon music-video" format — every visual is a hand-crafted animated style (felt + foam-core, amigurumi yarn, claymation, paper-cut, painted, watercolor, etc.), every cut lands on a song bar, and a locked character carries the narrative across ~12–17 shots. The song IS the script. Reference run — the Coinbase "Bet on anything" prediction-markets debut. |
create-cartoon-music-video
Merged from: skills/molecules/create-cartoon-music-video/SKILL.md + skills/molecules/music-video-ad/SKILL.md (the advisory flagged the latter as a literal duplicate; both share the Coinbase "Bet on anything" lineage). The cartoon-specialized rules sit on top of the generic music-video-ad recipe; the generic format is preserved in ## Variants near the end.
Originally validated in content-goose at clients/coinbase/video-01-music-video-debut/. Adapt paths to your project.
Purpose
Encode the style + recipe for crafted-animation music-video ads where:
- The song carries the script — lyrics + vocals replace any VO or narrator.
- Every visual is generated via image-to-video AI in a single locked hand-crafted aesthetic (felt, yarn, clay, paper, painted, etc.).
- A locked character anchors 60%+ of shots so the cuts read as one continuous animation.
- Cuts land on the song's bar grid — hard cuts, beat-locked, no soft transitions.
- Word-burst captions burn in mid-frame with bold-italic accents on hook words.
This molecule wraps video-orchestrator/orchestrator (sequencing) AND molecules/music-video-ad (general music-video format). It adds the crafted-aesthetic tooling stack + the prompting rules + the content-filter fallback that we learned the hard way on the Coinbase debut.
If the format is "song-is-the-script BUT live-action talking head + b-roll", use molecules/music-video-ad directly. If the format is "narrator VO + cartoon scenes", use molecules/create-narrated-comic-strip-video. This molecule is specifically for AI-generated crafted/animated music-video ads.
When to use
ALL of these are true:
- The song fits in ~30–60s (12–18 bars at moderate tempo).
- The brand benefits from a warm/crafted aesthetic — fintech, consumer tech, lifestyle, education, fashion.
- A single locked character can carry the story (gender, age, look approved before production).
- You have credits for Higgsfield (~$30–60 per run for keyframes + animations).
- The CTA can fit on a static brand end-card.
Skip when:
- Talking-head VO format is wanted → use
molecules/music-video-ad generic or create-narrated-comic-strip-video.
- Format needs > 25 distinct shots (the locked-character premise breaks down at high shot count).
- Brand colour requires the entire visual world to be one solid colour (consider motion-graphics molecule instead).
- Compliance-sensitive product surface that needs explicit on-screen disclaimers in the visual world (better: keep disclaimers on the end-card; this molecule can't reliably burn small legible disclaimer text in the AI-generated scenes).
Style fingerprint
Non-negotiable for this molecule. Change one and you're making a different kind of video.
Visual
- Single hand-crafted aesthetic applied to every shot. Pick ONE: felt + foam-core, amigurumi yarn, claymation, paper-cut, watercolor, painted-illustration. The locked style descriptor threads into every keyframe prompt verbatim.
- Locked character appears in ≥ 60% of shots. Anchor portrait approved via human gate before any per-scene keyframe generates.
- Tableaux over animation. Each shot is a static composition with ONE intentional micro-action (the character does something specific). Camera barely moves.
- No readable text in AI-generated scenes. Brand wordmarks, product labels, captions all happen in post (PIL overlay + ffmpeg subtitle burn). The model garbles text.
- No human-scale reference objects. Pencils, rulers, hands-into-frame are a Higgsfield failure mode when prompted with "miniature scale." Add the negative explicitly to every prompt.
- 9:16 vertical at 1080×1920 native.
Audio
- Original song carries everything. No separate VO. Hook + verse + tag live in the lyrics.
- Word-level timestamps mandatory. Either from ElevenLabs Music
--with-timestamps (preferred, free, native) or Whisper post-transcript (fallback, expect ~20% drift on sung repetitions).
- Song mixed UNDER the end-card — the brand reveal plays with music still audible, with a 0.5s afade-out at the master's tail.
Structure
- Length: 30–60s, derived from the song duration after trim.
- Bar grid drives cuts. librosa beat-track → bars.json. Each shot occupies one bar (~3.0–3.5s at moderate tempo).
- End-card holds 3.0s with music playing underneath. Master ends with the music — no trailing silence.
- Hard cuts on the beat. No in-edit xfade transitions (FFmpeg fade/opacity-blend). The song's beat carries continuity. The Coinbase v5 in-edit-xfade test was rejected; this rule still stands.
- AI-interpolation transitions ARE allowed — these are different from in-edit xfades. Seedance
start_image + end_image interpolation generates a continuous-motion bridge clip between two scenes (optionally with a pivot keyframe like an ECU macro or altitude bridge). This is a video clip, not a soft-fade. It preserves cartoon-style frame integrity throughout — no opacity blending. Opt in by emitting transition_pairs.json at State 3. Cartoon-specific transition palette:
- Stylized morphs — yarn unraveling, paint dripping, ink bleeding, paper folding, clay reshaping. Seedance start+end excels at these (each frame is mid-transformation, no jump-cut). Use when the metaphor of transformation matters to the story (e.g. the Nike yarn→photoreal morph end card).
- Continuous camera moves — push-in to macro detail (a fiber, a stitch, a single brushstroke), then pull back. Sells the "hand-made miniature" world by drawing the viewer INTO the texture before pulling them back out. May not require a separate pivot keyframe if the scale change is large enough on its own.
- Whip-pan / slide transitions for split-screen → split-screen (e.g. 3-face triptych → 3-door triptych). These CAN use FFmpeg
xfade=transition=slideleft because they preserve the panel structure across the cut — no opacity-blend happens, just a horizontal slide. This is the one exception to the "no FFmpeg xfade" rule, and only because slideleft doesn't blend opacity.
- AVOID photoreal-style camera moves on non-photoreal scenes (e.g. a drone-shot ascent over a yarn city should still LOOK yarn-textured throughout — don't let the model introduce photoreal sheen mid-flight). The destination IS allowed to be photoreal (Nike yarn→photoreal morph end card), but transitional clips between two yarn scenes must hold the yarn aesthetic.
- Reference exemplar: the Nike Jordan "The Box" ad (
nike/ads/video-01-the-box/). 8 transition clips, 3 transition keyframes, ~75 credits added. Made the ad feel cinematic vs choppy. See atoms/planning/plan-scene-transitions/SKILL.md and atoms/video-generation/create-transition-{keyframe,clip}/SKILL.md.
State machine
Inherits from music-video-ad (which inherits from video-orchestrator/orchestrator). Adds the crafted-aesthetic-specific config at each state.
0. PREFLIGHT
1. BRAINSTORM (concept paragraph → 3–5 angles, biased toward crafted-style)
↓ HUMAN GATE: angle pick
2. DESIGN BRIEF (style descriptor, character descriptor, tempo target)
2.5 LYRIC LOCK (strip section markers; maintain lyrics-locked + lyrics-plain)
↓ HUMAN GATE: lyrics
2.6 MUSIC GENERATE (ElevenLabs Music plan-mode with timestamps)
↓ HUMAN GATE: song
3. SCENE PLAN (one bar per shot, character-action verb per shot)
3.5 CHARACTER ANCHOR (single nano_banana_2 image, verbatim descriptor)
↓ HUMAN GATE: anchor
3.5b ANGLES + LOCK (4 angles, drift check, write LOCKED.json)
5. KEYFRAMES ×N (nano_banana_2, anchor as medias, no scale-reference negatives)
5.5 ANIMATIONS ×N (Seedance 2.0, character-action-verb-leading prompts)
↓ Auto-fallback to Veo 3.1 Lite on ip_detected / nsfw
7. STITCH (ffmpeg concat with logo bug overlay; song mux + afade)
7.4 CAPTION SYNC (sync-captions-to-music, 130pt, no pill, accent bold-italic)
7.5 SCROLL TEST (hook-strength + synthetic-persona on first cut)
8. REVIEW (+ review-music-video-beat-sync + compare-with-reference-video)
9. POLISH (Whisper-test ship gate)
↓ HUMAN GATE: master review
10. DELIVER (upload-ad-sample → Goose Ads library)
Tool stack
| Stage | Default | Fallback / Alternative |
|---|
| Song generation | ElevenLabs Music v1, plan-mode, --with-timestamps | User-supplied + Groq Whisper word-level transcript |
| Beat detection | librosa beat_track 4/4 | librosa with --bpm-override if tempo is half-time-ambiguous |
| Character anchor | Higgsfield nano_banana_2, 9:16 | gpt_image_2 if photorealism wanted (rare for cartoon) |
| Character angles | nano_banana_2 with anchor as medias[role=image] | Soul ID training (~$15) if anchor-ref drift > 15% |
| Keyframes | nano_banana_2 with anchor + style descriptor | Re-roll with second-keyframe ref for cross-scene continuity |
| Animations | Seedance 2.0 std mode, 4s clips, character-action prompts | Veo 3.1 Lite on ip_detected / nsfw filter triggers |
| Stitch | ffmpeg concat + ffmpeg overlay (logo bug) | (xfade chain explored in v5 — REJECTED; hard cuts feel better in this format) |
| Captions | sync-captions-to-music + custom ASS via rebuild_captions.py | klap-karaoke style if user wants karaoke highlight |
| End-card | PIL composite (cairosvg + ImageDraw) + ffmpeg | hyperframe HTML/GSAP if motion is needed on the end-card |
| Publish | upload-ad-sample (skills/legacy/upload-ad-sample) | Operator-uploaded externally |
Hard rules (learned the hard way — break these and the cascade returns)
-
Strip […] section markers from lyrics before any music-gen call. Otherwise the song literally sings "CTA, two bars" or "Verse 1, four bars". Apply via re.sub(r'\[[^\]]+\]', '', text).
-
Maintain TWO lyric files. lyrics-locked.md (human-readable: frontmatter + bar grid table + section markers) for review surfaces. lyrics-plain.md (just lyric text, one line per song line, no table/headers) for downstream tools — sync-captions-to-music parses this one.
-
Lead animation prompts with the character verb, not the camera move. "The felt protagonist scrolls his phone with his thumb..." — not "Slow push-in toward the felt protagonist...". Camera direction is secondary.
-
Always pass "no rulers, no pencils, no measurement tools, no hands, no human-scale reference objects" as negative direction to both image (nano_banana_2) and video (seedance / veo) prompts when the scene is a "miniature" / "macro" composition. Higgsfield interprets "miniature" by adding scale-reference props.
-
Never put readable brand or product text in an AI-generated scene. The wordmark + subhead + product labels happen in post via PIL/ffmpeg overlay. Higgsfield garbles brand text into typos (PELETON, etc.).
-
Auto-fallback Seedance → Veo 3.1 Lite on ip_detected or nsfw status. Seedance's content filter false-positives on Oscar trophies, "intimate"/"leans toward" wording, and certain interior settings. Veo runs clean with the same start_image — re-issue immediately, don't surface a human gate for these.
-
End-card holds with music underneath. Mux the song over the WHOLE video including the end-card. Apply afade=t=out:st=<master_end - 0.5>:d=0.5 at the very tail. Video ends with music — no trailing silence.
-
Hard cuts on the beat. Do not xfade. Tested in v5 — softer transitions read worse in music-video format. Let the song carry continuity. Spend the visual budget on character motion + palette uniformity instead.
-
Captions stop at end-card start. sync-captions-to-music should drop or clip any chunk whose start >= caption_end_s where caption_end_s = body_end_s. The brand frame needs to breathe.
-
Logo bug always-on in the body, off on the end-card. Bottom-left, ~20% of frame width, white wordmark, 56px margin. Build via + PIL recolor → ffmpeg overlay. The end-card's big wordmark makes the bug redundant there.
Inputs
- Concept paragraph that fits the format (e.g. "an explainer rap for prediction markets on Coinbase").
- Brand name + Coinbase Blue / brand color hex + wordmark SVG.
- Locked-style aesthetic: felt+foam-core / amigurumi / clay / paper / etc.
- Locked-character descriptor (gender, age, build, skin, hair, wardrobe, key accessory).
- Climax line target + accent-word list.
- End-card subhead copy.
- Optional: reference video (e.g. Gum-of-Gods amigurumi for benchmark) → enables
compare-with-reference-video at State 8.
- Optional: user-supplied song MP3 (otherwise generate via ElevenLabs Music).
- Output folder (default
<brand>-ads/video-NN-<slug>/).
Workflow
This molecule is a thin wrapper. Its job is to seed each state with the crafted-aesthetic defaults + the hard rules.
Pre-flight setup
Write <video_folder>/.style-preferences.md carrying:
- Locked style descriptor block
- Target tempo range
- Caption style preset (
music-video, 130pt, mid-frame, no pill, accent bold-italic)
- Music provider (
create-music-elevenlabs default; with-timestamps required)
- Logo bug spec (path to SVG, target width, margin)
- End-card spec (background colour, subhead text, hold duration 3.0s, afade 0.5s)
Stand up storyboard.html shell — the single review surface — with placeholder sections for every state.
State 1 — Brainstorm
Bias toward concepts that:
- Have a song-able hook
- Work with a single locked craft style
- Feature one protagonist (not crowds, not multiple characters)
- Don't require on-screen product UI (Higgsfield can't render UI cleanly)
Output idea-brief.md with 3–5 angles, each with: hook, ICP, tonal axis, music spec (BPM + genre + vocal), visual style description, locked-character spec, climax line, accent words, end-card subhead, ~12-bar sample lyric.
Storyboard injection: full idea brief + 4-angle comparison table.
State 2 — Design brief
Force these required fields:
tempo_bpm target
genre_tags (specific instruments)
vocalist description (gender, style, ad-lib direction)
climax_bar + climax_line
accent_words array
style_descriptor (verbatim, threads into every image prompt)
character_descriptor (verbatim, threads into every character-bearing prompt)
end_card_subhead copy + brand color
- Compliance notes (any disclaimers required)
State 2.5 — Lyric lock
- Polish lyrics. Syllable check vs target BPM (10–14 syllables/bar at 124 BPM, 12 at 86, 10 at 74).
- Strip
[…] section markers — write a clean lyrics-plain.md alongside the human-readable lyrics-locked.md.
- HUMAN GATE on the text. Cheap iteration.
State 2.6 — Music generate
- Build
composition_plan.json from lyrics-plain.md (NEVER from lyrics-locked which has section markers).
- Run
atoms/music/create-music-elevenlabs/scripts/compose_detailed.sh --plan composition_plan.json --with-timestamps audio/.
- Normalize timestamps →
audio/word-timestamps.json.
- librosa beat-track →
audio/beats.json + audio/bars.json.
- HUMAN GATE on the song. Most important gate in the molecule — every downstream cost depends on it.
If user supplies the song:
- Drop at
audio/<name>.mp3.
- Run Groq Whisper with
timestamp_granularities[]=word for word-level transcript.
- Same librosa step.
- Expect ~20% drift on sung repetitions; plan for hand-patching the chorus captions later.
State 3 — Scene plan
For each bar:
- Bar number + time range
- Lyric line(s)
- Tableau description
- Character presence (target ≥ 60% of bars)
- ONE character-action verb (the thing the character does in this bar)
Output scene-plan.md. Storyboard injection: per-bar shot cards with placeholder mockups.
State 3.5 — Character lock
Anchor (HUMAN GATE): single nano_banana_2 call with the verbatim character descriptor block at 9:16. Surface in storyboard's character panel via Read. Operator approves or re-rolls.
Angles: 4 parallel nano_banana_2 calls with anchor as medias[role=image]. Drift check: hair shape, jersey/wardrobe color, key accessory placement, skin tone. Write LOCKED.json with method: anchor-ref. Soul ID escalation only if drift > 15%.
State 5 — Keyframes
For each scene:
- Character-bearing:
prompt = <verbatim style> + <verbatim character> + <tableau> + "no rulers/pencils/hands"; medias = [anchor].
- Style-only:
prompt = <verbatim style> + <tableau> + "no rulers/pencils/hands"; no medias.
- For continuity-critical scenes (same room reused across bars): pass BOTH anchor AND the prior scene's keyframe as
medias. Nano Banana 2 honors multiple references.
Fire in parallel. Spot-check 2–3 keyframes for character + style consistency before animating any.
State 5.5 — Animations
For each keyframe, mcp__higgsfield__generate_video call:
model: seedance_2_0
duration: 4
aspect_ratio: 9:16
medias: [{value: <keyframe_job_id>, role: start_image}]
prompt: <character-action verb leads> + <secondary background motion> + "Hand-articulated stop-motion craft feel preserved"
Auto-retry on Veo 3.1 Lite when Seedance returns:
status: ip_detected (usually Oscar shapes, branded objects)
status: nsfw (false positives on intimate / interior scenes)
Veo with the same start_image + lightly reworded prompt almost always renders. Add "no hands, no measurement tools" if Veo introduces hand intrusions.
State 7 — Stitch
scripts/stitch.py driver. For each body bar:
- Trim to bar duration (from
bars.json).
- Overlay the logo bug bottom-left (cairosvg-generated white wordmark PNG).
- ffmpeg concat all body bars + end-card (3.0s + 0.3s if using xfade, else exactly 3.0s).
- Mux full song over the silent concat. Apply
afade=t=out:st=<master_end - 0.5>:d=0.5 at the tail.
Hard cuts, no xfade. Tested and rejected.
State 7.4 — Caption sync
sync-captions-to-music script reads lyrics-plain.md + word-timestamps.json, produces chunks. Custom ASS via scripts/rebuild_captions.py:
- Font: New York 130pt
- Bold: yes
- Accent words:
\b1\i1 + \fs160 inline
- BorderStyle: 1 (outline+shadow), Outline: 8px, Shadow: 0 (no pill)
- Alignment: 5 (mid-frame center)
- Caption window: drop or clip any chunk with
start >= body_end_s so the brand reveal breathes
Burn over master-stitch → edits/master-final.mp4.
State 7.5 / 8 — Review
review-video-hook-strength (3s muted + sound test)
review-video-synthetic-persona (ICP-embodied watch)
review-music-video-beat-sync (caption onsets, cuts-on-bar, hook-on-the-one)
compare-with-reference-video if a reference path was set in implementation-brief.md
State 9 — Polish
Whisper-test ship gate: run Whisper on the final master and diff against lyrics-plain.md. If any locked lyric mistranscribes due to caption-burn opacity, surface as P0.
State 10 — Deliver
skills/legacy/upload-ad-sample three-call flow:
- POST
/api/ads-library/samples/upload-url → presigned PUT URL
- PUT video bytes to S3
- POST
/api/ads-library/samples to register
Repeat for thumbnail (climax frame is the default). Set is_published: true if operator confirms publish; is_featured: true if the sample should pin to homepage.
Output
Same as music-video-ad plus:
<video_folder>/HOW_TO.md — recipe for re-running this format
<video_folder>/LEARNINGS.md — what was learned this run + suggested skill changes
<video_folder>/.style-preferences.md — this molecule's contribution
Quality Checks
- Song has word-level timestamps (ElevenLabs OR Whisper).
lyrics-plain.md exists and has zero […] markers, zero markdown tables.
- Character locked via anchor-ref (or Soul ID) with drift < 15%.
- Every body bar has a keyframe + animation with character-action-verb prompt.
- Master ends WITH the music (no trailing silence; afade-out at tail).
- Logo bug visible in every body bar; suppressed on end-card.
- Captions stop at body_end (no caption burn-in over the brand frame).
- Storyboard.html is current — every state's section populated.
Failure Modes
- Song generated with section markers in the lyrics → "CTA, two bars" gets sung. Pre-flight strip is non-negotiable.
- First animation pass had camera-led prompts → motion invisible, ~$25 + 45 min wasted. Always pre-flight one animation prompt before firing all N.
- Seedance content filter stalls a wave → auto-fallback to Veo solves it without surfacing a gate.
- Whisper undercounts repeated hook lines → hand-patch chorus captions in chunks.json; document on the next run that ElevenLabs-native timestamps are the better path.
- End-card runs in silence → user perceives the video as ending early. Always mux song over the end-card with afade tail.
- Master uploaded then needs swap → legacy upload-ad-sample doesn't have a PATCH endpoint. Plan to delete+recreate or contact the admin. Don't auto-publish until the operator confirms v-final.
Reference run
coinbase/video-01-music-video-debut/ — the Coinbase "Bet on anything" prediction-markets debut. 54.9s, 1080×1920, felt+foam-core aesthetic, locked felt protagonist in Coinbase Blue jersey "1". 16 body bars + 3.0s end-card. Shipped + featured on Goose Ads (sample id a00f0d6d-8691-401d-9f9c-f35e956024d9).
See:
coinbase/video-01-music-video-debut/HOW_TO.md — the canonical step-by-step
coinbase/video-01-music-video-debut/LEARNINGS.md — what failed and what to change in the skills
coinbase/video-01-music-video-debut/storyboard.html — the single-review-surface
Composed Atoms
atoms/music/create-music-elevenlabs — original song with word-level timestamps (state 2.6)
atoms/video-generation/create-video-seedance — image-to-video, std mode (state 5.5 default)
atoms/video-generation/create-video-veo3 — fallback on ip_detected / nsfw (state 5.5 fallback)
atoms/image-generation/create-image-nano-banana-fal — character anchor + per-bar keyframes
atoms/assembly/stitch-videos-ffmpeg — concat + logo bug overlay + song mux + afade (state 7)
atoms/captions/burn-in-captions — mid-frame serif captions, accent bold-italic (state 7.4)
atoms/review/compare-with-reference-video — optional reference-anchored review (state 8)
To produce another cartoon-music-video for a different brand: change concept, lyrics, style descriptor, character descriptor, brand color, end-card copy. Everything else in the recipe stays the same.
Decision Rules
- One locked style descriptor (felt + foam-core / amigurumi yarn / claymation / paper-cut / etc.) threads into every keyframe prompt verbatim — never mix styles within one ad.
- Locked character anchors ≥ 60% of shots. Use anchor-ref (or Soul ID if drift > 15%) before any per-scene keyframe generates.
- Hard cuts on the beat — no in-edit ffmpeg xfade (rejected on Coinbase v5). AI-interpolation transitions via Seedance
start_image + end_image are a different mechanism and ARE allowed.
- Lead every animation prompt with the character verb, not the camera move ("The felt protagonist scrolls his phone..." NOT "Slow push-in toward...").
- Strip
[…] section markers from lyrics before any music-gen call — otherwise the song literally sings the section names.
- Auto-fallback Seedance → Veo 3.1 Lite on
ip_detected or nsfw content-filter triggers; never surface a human gate for these.
- End-card holds with music underneath; apply
afade=t=out:st=<end-0.5>:d=0.5 so the master ends with the music — no trailing silence.
Variants
This molecule is the cartoon-specialized form of the more general music-video-ad format. The
generic music-video-ad recipe (live-action or mixed-aesthetic music videos, song-is-the-script,
captions on the vocal, no locked aesthetic constraint) is preserved here so the merge does not
lose the broader contract.
Generic music-video-ad — when to use
Use the generic music-video-ad form when ALL of these are true:
- The product hook fits inside ~12–16 bars of lyrics (≈ 30–45 s).
- The ICP responds to rhythm/music-driven content (lo-fi explainer rap, R&B, dance-pop hook).
- The format benefits from a single locked visual style applied across many short tableaux (3–4 s each) rather than long animated shots.
- A recurring character or product can be locked once and re-rendered across all scenes.
- You're willing to pay for a song generation + word-timestamps (ElevenLabs Music API or Suno) before any visual credits.
Skip the generic form for VO-driven narrator ads (use create-narrated-comic-strip-video), anything > 60 s, or iMessage / motion-graphic / talking-head formats.
Generic music-video-ad — style fingerprint deltas vs cartoon
- Visual: generic form allows any single locked style (including live-action craft); cartoon variant restricts to AI-generated hand-crafted aesthetic.
- Caption rhythm: generic uses mid-frame serif with mixed regular + bold-italic; cartoon variant adds word-burst with 1–3 word chunks.
- Logo bug: cartoon variant requires persistent bottom-left logo bug in every body bar (off on end-card); generic form leaves logo placement to the brand.
- Transitions: generic form allows opt-in xfade for non-music-video product-explainer cuts (
--transition xfade); cartoon variant is hard cuts only.
Generic music-video-ad — reference
The generic reference is reference-videos/gum-of-gods-music-video.mp4 (third-party): 41 s vertical, crocheted amigurumi tableaux, single locked character, mid-tempo (~85 BPM) explainer-rap, mid-frame serif captions. Format read on the generic format: hook (problem) → mechanism → solution → aspirational payoff → CTA, brand reveal at the climax bar.
What this molecule does NOT cover
- Multi-character ads (locked-character premise assumes one protagonist).
- Live-action footage (use a different ad molecule).
- Long-form content > 60s (the locked-style premise gets repetitive past 20 shots).
- Cuts that require precise lip-sync to vocals (use
create-lipsync-hedra or similar instead).
- A/B variant generation (use
ugc-ad/create-hook-variant-pack to derive variants from a finished master).