| name | create-podcast-skit-ad |
| description | Produce a fabricated two-host podcast-skit video ad from scratch — two AI "hosts" (a skeptic + a believer) bantering at a themed-set desk with broadcast mics and over-ear headphones, tight bust framing, ~20-24 snappy back-and-forth cuts, ElevenLabs VO per line, Nano Banana base stills + expression variants, Veed Fabric lipsync per (still,line) pair, a brand end card, ffmpeg stitch, and burned ASS karaoke captions. The from-scratch base for the Ramp / Dutch / Ladder / Mindtrip format. Photoreal by default; hand off to `restyle-podcast-skit-ad` for felt/claymation/anime medium swaps and to `create-shortform-cuts-from-podcast-skit-ad` for 20s cutdowns. Use when the user asks for a "podcast ad", "fake podcast skit", "two-host banter ad", "Ramp-style podcast ad", or names that format. ~$12 + ~1-2h with review gates.
|
| tags | ["video","ads","podcast-skit","two-host","lipsync","character-lock"] |
create-podcast-skit-ad
Purpose
The production line for the fabricated two-host podcast-skit ad — the format
behind Ramp run-01, Dutch run-03, Ladder run-02, Mindtrip run-02. Two AI hosts
sit at a podcast desk on a themed set that is deliberately unrelated to the
product (a 24hr laundromat at 2am, a vet clinic, a butter-knife aisle) — the
mismatch is the joke. A skeptic and a believer trade ~20-24 short lines; the
believer keeps trying to explain the product plainly, the skeptic keeps tripping
over how plain it sounds. The brand surfaces mid-skit and resolves on an end card.
This molecule builds one from scratch. It is the counterpart to:
restyle-podcast-skit-ad — re-skin a finished skit into a new medium (felt/clay/anime).
create-shortform-cuts-from-podcast-skit-ad — cut a finished skit into 20s standalone Reels.
Canonical recipe mined from clients/ladder/ad-runs/run-02-podcast-skit/HOW_TO_MAKE_THIS_VIDEO.md (itself the Ramp v1 recipe generalized).
When to use
- "Make a podcast ad" / "fake podcast skit" / "two-host banter ad" / "Ramp-style podcast ad."
- A mid-funnel paid-social spot where two talking heads sell the product through deadpan banter, not a demo or UGC selfie.
Do NOT use for: a real podcast clip → ad (use podcast-ad/create-ffmpeg-motion-podcast-ad or podcast-clip-animated-ad), single-host UGC (use a ugc-ad diary/founder skill), or restyling/cutting an already-finished skit (use the two siblings above).
Composed Atoms
voiceover/format-script-to-style — apply the intonation legend (..., —, ?, CAPS) to each line before TTS.
voiceover/create-voiceover-elevenlabs — per-line VO via the with-timestamps endpoint (timestamps drive caption sync). Two voices: HER + HIM.
product-images/create-product-images-nanobanana — the 2 base stills (one per host) in the themed set, mouth neutral/closed.
product-images/edit-image-nano-banana — 10+ expression/pose variants, img2img anchored on each host's base still as the SOLE reference.
lipsync/create-lipsync-veed-fal — Veed Fabric 1.0 lipsync per (still, line-audio) pair.
end-cards/create-end-card-pil (or create-end-card-html) — brand end card with the real wordmark + CTA.
assembly/stitch-videos-ffmpeg — concat the per-line clips + end card.
captions/add-captions-burn — ASS karaoke, 3-word phrase cap, bottom-center, brand color.
review/watch + audio-editing/mix-master — QC the master and balance VO/bed.
Inputs
brand (required), product (required), brand_dir / run folder.
themed_set (required) — the absurd, product-unrelated location (e.g. "24hr laundromat at 2am"). The mismatch is the comedic engine.
angle (default: skeptic-vs-believer) — the two-host dynamic.
hosts — HER + HIM names + ElevenLabs voice_id + settings. Default cast: Brittney kPzsL2i3teMYv0FxEYQ6 + Brad T4x5CtnhOiichhcqFzgg, eleven_multilingual_v2. Swap Brittney → Brielle 6u6JbqK... (stability 0.30, style 0.35) if her read sounds AI-mechanical on a long monologue.
brand_wordmark (required) — path to the real logo/wordmark SVG/PNG for the end card. Never AI-render brand text.
caption_color (default brand accent, e.g. Ladder lime #FFEB3B), target_duration_sec (default ~45s, but lock it to the rendered VO length, not a guess).
cta — end-card CTA copy + URL.
Workflow
- Idea-brief lock. Skeptic-vs-believer two-host skit on the
themed_set. State brand guardrails explicitly (e.g. Ladder: no transformation claims; no competitor names — see (no-competitor-bashing-in-ads learning)). Human gate: approve the brief before any paid gen.
- Script with intonation marks →
working/script.json. 20-24 lines, ≤10 words each, alternating HER/HIM. Use the legend: ... trailing pause, — em-dash, ? rising, CAPS emphasis. No acronyms in VO (ElevenLabs spells them letter-by-letter — see (elevenlabs-acronyms learning)). Human gate: approve script.
- Render VO (
create-voiceover-elevenlabs, with-timestamps). One MP3 + one timestamp JSON per line + a manifest. Measure the real total VO length now and set the timeline to it — don't lock duration first ((render-vo-before-locking-timeline learning), (align-concat-to-words-json learning)).
- Two base stills (
create-product-images-nanobanana, 9:16): tight bust framing, broadcast mic foreground, over-ear headphones, mouth NEUTRAL/CLOSED (open mouth breaks lipsync). Generate HER first, then HIM using HER's still as a background reference so the set stays identical. iPhone-shot daylight register, not cinematic tungsten ((iphone-shot-not-cinematic learning), (iphone-shot-portrait-prompt learning)).
- Expression variants (
edit-image-nano-banana, img2img): ≥10 per cast, each anchored on that host's base still as the SOLE reference; the prompt describes ONLY the expression change ("curious questioning look, eyes wide open") so the background stays 100% consistent. Avoid "narrowed/squinting eyes" — trips NSFW reject; use "curious questioning look."
- End card (
create-end-card-pil): brand-color background, real wordmark (composited, never AI-rendered — (no-production-labels-in-keyframes learning)), CTA pill + URL → 2.5s static mp4.
- Lipsync (
create-lipsync-veed-fal, 720p, concurrency ~4): one clip per (chosen still, line audio) pair. Veed Fabric re-encodes audio with hiss — re-mux the original ElevenLabs MP3 back over each clip's video ((lipsync-audio-swap learning)).
Decision Rules
- Themed set must be unrelated to the product — that mismatch is the joke. A fitness app at a laundromat; a fintech at the DMV.
- Photoreal is the default cast (Ramp v1), NOT felt/clay — only go stylized if the user explicitly says so (
(ramp-format-default-v1-not-v2-felt learning)). For a medium swap on a finished cut, hand off to restyle-podcast-skit-ad.
- Base stills: mouth closed/neutral, always. Open-mouth stills fail Veed lipsync.
- Lock the timeline to rendered VO, never the reverse. Lines run faster than planned.
- No competitor names; no transformation/outcome claims unless the brand explicitly allows them.
- Brand text is composited, never AI-rendered (end card, any on-screen wordmark).
- Hand off shortform cutdowns to
create-shortform-cuts-from-podcast-skit-ad; don't linear-slice here.
Output
finals/<brand>-podcast-skit-master-v1.mp4 — 9:16, ~40-55s, captioned, mixed.
working/script.json, per-line VO MP3s + timestamps, base stills + variants, per-line lipsync clips, end-card mp4.
- Registered in the control-plane (
asset-manifest.json master active_master, render-outputs.json, history/versions.json) when run inside a VCP project.
Quality Checks
- Two distinct, consistent hosts; the set is identical across every cut (background never drifts).
- Lipsync reads natural on both hosts; no audio hiss (original MP3 re-muxed over Veed output).
- Captions cover every spoken word, 3-word cap, synced to VO timestamps; no two-layer stacking.
- Master duration matches the summed VO + end card (no
-shortest truncation; pad audio instead).
/watch pass: Whisper transcribes every line correctly; no NSFW artifacts, no open-mouth freeze.
Failure Modes
- Veed Fabric audio hiss → keep its video, re-mux the source ElevenLabs MP3 (mono, AAC 256k, loudnorm).
- Nano Banana NSFW on "narrowed eyes" → reword to "curious questioning look, eyes wide open."
- Open-mouth base still → lipsync fails; regenerate the still mouth-closed.
- fal.ai balance exhausted mid-batch ("User is locked. Reason: Exhausted balance.") → top up, retry only the failed line indices.
- Host reads AI-mechanical on a long line → swap Brittney → Brielle (stability 0.30, style 0.35).
- Concat truncates / desyncs on mixed fps or SAR → filter-concat with explicit
fps=N,setsar=1 per input.