Skip to main content

create-imessage-video-ad

Produce a 9:16 social-native ad that recreates an iMessage conversation reveal — bubbles pop in over time, composer types char-by-char, real Apple iMessage SFX hit on every send/receive, music bed underneath, brand end card with FREEPACK-style CTA. One continuous Playwright recording (no scene-cuts/glitches) in plain-chat or iPhone-framed (Dynamic Island over a background) variants, assembled to a master MP4 plus 9:16 / 1×1 export variants.

Jump to install

Source facts

Repository
criptogus/agent-evolve-network
Last source activity
July 4, 2026 at 17:42
Detected SKILL.md language
English
Stars
289
Forks
2

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
21 files

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
create-imessage-video-ad
description
Produce a 9:16 social-native ad that recreates an iMessage conversation reveal — bubbles pop in over time, composer types char-by-char, real Apple iMessage SFX hit on every send/receive, music bed underneath, brand end card with FREEPACK-style CTA. One continuous Playwright recording (no scene-cuts/glitches) in plain-chat or iPhone-framed (Dynamic Island over a background) variants, assembled to a master MP4 plus 9:16 / 1×1 export variants.
# create-imessage-video-ad > Built from these proven iMessage builds: `the Clinikally build` (newest), `the Catchback Cards build`, and `the Peloton build`. Adapt paths to your project. ## Purpose Use this skill when an ad concept calls for an authentic iMessage thread reveal: someone screenshots a product/result, sends it to a friend, the friend reacts and asks what app/service it is, and the conversation surfaces the brand + a CTA code. Produces a 9:16 social-native ad with real Apple iMessage SFX, optional music bed, and brand end card — one continuous Playwright recording assembled into a master MP4 plus 9:16 / 1×1 variants. ## Reference builds Three reference builds — copy the closest one as your starting point: - `the Clinikally build` — **NEWEST; default starting point.** iPhone-frame variant (Variant B: framed phone + Dynamic Island over a brand-relevant background photo), dark theme, **native 1080×1920 recording**, full composer drives on every sent bubble, NB2-generated hook/background/end-card-base assets, reference-styled end card. Friend-asks-friend inverse angle. 21.2s. - `the Catchback Cards build` — **abstract-glow end card**, gray peer avatar, product-flex hook (screenshot of a graded card with a price). 20.6s. Playful brand voice. Plain-chat variant. - `the Peloton build` — **photo-background end card** with real Peloton wordmark, indigo peer avatar, result-flex hook (Strava finish-line summary rendered locally as HTML→PNG). 19.0s. Grounded brand voice. Plain-chat variant. ## When to use - Reaction/discovery ads where the punchline is the *recipient's* curiosity ("what app is that?", "wait this is real?") - Promo-code reveals (FREEPACK / FIRSTPACK / WELCOME10) — the conversational delivery feels far less ad-like than a hard CTA card alone - Any time the brief mentions "fake DM", "screenshot of a chat", "iMessage style", "Tyler/peer" If the chat is just one beat (a single screenshot), reach for the atom `create-imessage-mockup` directly. Use *this* molecule only when you need a **multi-bubble timeline animated to video**. ## Composed Atoms - `create-imessage-mockup` — renders the iMessage HTML (plain or `with-iphone-frame` mode); this skill drives it via Playwright - `stitch-videos-ffmpeg` — final concat + crossfade - `scripts/fal_helpers.py` — (optional) NB2 (`fal-ai/nano-banana-2`) generation of photographic hook / background / end-card-base assets, Clinikally-style - (optional) ElevenLabs music generation OR a music asset already in the repo (search `**/music*.mp3`) ## Inputs 1. **`brief`** (required) — what the ad is selling, the CTA code, the peer's name (e.g. "Tyler"), the screenshot subject (product image path or AI-generated frame) 2. **`output_folder`** — usually `<brand>/ads/video-NN-imessage-<slug>/` 3. **`render_variant`** — `plain` (full-bleed chat, the catchback/peloton builds) or `iphone-frame` (framed phone + Dynamic Island over a background photo — the Clinikally build; **default for new builds**, it reads more native in-feed). `iphone-frame` additionally needs a `background_asset` (brand-relevant photo; NB2-generate if none exists). 4. **`peer_persona`** — name + 2-letter monogram/initials + avatar bg color (hex) — i.e. the non-self entry in the thread's `participants[]` 5. **`screenshot_asset`** — path to the image that gets sent in the first attachment bubble 6. **`bubbles`** — list of `{from: me|peer, text: string, delivered?: bool, typing_before?: bool}` 7. **`composer_drives`** — the composer text for **every** sent bubble (`{bubble_id, text, dur_sec}`). The typed text must equal the sent text — see Critical knowledge #13. 8. **`end_card`** — `{wordmark, code, tagline}` (+ optionally a `reference_image` for the reference-styled variant) — used to render the brand slate The master `thread.json` schema is identical to the `create-imessage-mockup` atom (typed `messages[]` — `text` / `typing` / `attachment` with `src`/`title`/`subtitle` — plus `participants[]`, `theme`, `header`); this skill adds **no new data structure**. The only new artifact is the **timeline** (see Workflow step 3). ## Critical knowledge — read before producing your first ad These are mistakes that have already been made and fixed in the canonical builds. Do not re-introduce them. ### 0. Use the real brand wordmark on the end card, never styled text CSS approximations of brand wordmarks look amateur even when typeface and kerning are close. The Peloton build caught this on first preview: a hand-styled "peloton" lowercase + manual dot got rejected immediately in favor of the real Wikimedia SVG. Drop the official SVG at `assets/<brand>-logo.svg` (sources: Wikimedia Commons, `brandfetch.com/<brand>.com`, brand press kit) and the bundled `render-end-card.template.js` injects it via the `<!--{{BRAND_LOGO_SVG}}-->` placeholder. Your only job is to color the paths. See `references/end-card-recipe.md` for the full end-card playbook (CTA pill color, logo contrast tricks, photo-background composition). ### 1. Use the real Apple iMessage SFX, not generic notification sounds **The "iMessage feel" is 70% the SFX.** Generic notification sounds break the illusion immediately. The Apple sounds are CC0 on BigSoundBank: - **Send (whoosh-up):** `https://bigsoundbank.com/UPLOAD/mp3/1313.mp3` — Apple "Message Sent" (~0.5s) - **Receive (tritone):** `https://bigsoundbank.com/UPLOAD/mp3/1111.mp3` — Apple "Note" (raw is 6s; trim to **1.4s with a 400ms fade-out**) Pre-bundled copies live at `assets/sfx/imessage-send.mp3` and `assets/sfx/imessage-receive.mp3` in this skill. **Reuse those** instead of re-downloading. Both are loudness-normalized to **-9 LUFS, peak 0dB** so they cut through the music bed without further tweaking. **LFS gotcha:** these files are git-lfs-tracked. If you `cp` them into a fresh ad folder and ffmpeg later reports `Invalid data found when processing input`, you got a 130-byte pointer file instead of the audio. Recover by running `git lfs pull --include="assets/sfx/*"` from the repo root, OR re-fetch from BigSoundBank with the recipe in #1 above. ### 2. NEVER play SFX on the typing indicator The typing-dots bubble appearing is a **silent state change** — iOS does not chime when someone starts typing. Adding a soft "receive" cue on the typing-pop sounds wrong even when quiet. Only fire `receive` SFX when the *actual text bubble* lands (i.e. on `typing-swap`). This was caught on review of the canonical build and explicitly removed. ### 3. Record the entire chat as ONE continuous Playwright session Do **not** record scene-by-scene and concat. Every page reload causes a micro-flicker at the cut, and "scrolling" has to be faked by *removing older bubbles* between scenes — which looks janky because it is janky. Architecture: - Render the full thread HTML with all messages present but each marked `popState: 'pending'` (which the CSS turns into `data-pending="1"` → `display:none`). - A single embedded driver script walks a `TIMELINE` array on `requestAnimationFrame`, removing `data-pending` and adding `pop-now` at the right moment for each event. - Auto-scroll the body via custom easing whenever a new bubble lands; do not rely on `scrollIntoView` (inconsistent in headless). - Compute the SFX cue list **deterministically from the timeline** (do not capture cues via `performance.now()` inside the page — there's a race with `setContent`'s load event that drops cues). `scripts/record-master.template.js` implements this end-to-end. Adapt the `TIMELINE` array and reuse the rest verbatim. ### 4. Pop-pending must remove rows from layout, not just hide them If `pop-pending` only sets `opacity: 0`, the row still occupies its full height — so the conversation pre-allocates all 15 bubbles' worth of space at frame 0, the auto-scroll has nothing to scroll to, and bubbles "pop into pre-allocated holes" instead of *arriving*. Use `display: none` (via `data-pending="1"`) so the conversation grows naturally as messages arrive. The atom's chat.css now has this rule baked in: ```css .row[data-pending="1"], .delivered-caption[data-pending="1"] { display: none !important; } ``` ### 5. End card stays static — no Ken-Burns, no zoom-pan The brand slate must **land hard**. A drifting end card reads as filler and undercuts the punch of the FREEPACK reveal. Render the end-card PNG once and ffmpeg it to a fixed-frame MP4: ```bash ffmpeg -y -loop 1 -i end-card.png -t 3.5 -r 30 \ -vf "scale=720:1280,format=yuv420p" \ -c:v libx264 -pix_fmt yuv420p -movflags +faststart end-card.mp4 ``` Crossfade from chat → end card with a **300ms** `xfade` in the stitch step (longer feels mushy, shorter feels like a hard cut). ### 6. ffmpeg amix divides by N inputs by default If you `amix` the music + the silent base + 11 SFX, every input gets divided by 13 unless you pass `normalize=0`. The audio will sound mysteriously quiet. Always: ``` amix=inputs=N:duration=first:dropout_transition=0:normalize=0 ``` Then take the mix through `volume=2.5,alimiter=limit=0.95` so it peaks at 0dB without clipping. ### 7. Music bed: lofi/hip-hop, 60Hz highpass, ~-10dB Keep the bed unobtrusive — `volume=0.30` plus `highpass=f=60` to clear room for the SFX. Fade out 1.5s before the end so the FREEPACK reveal can land in (relative) silence. If the bed is shorter than the ad, loop it in the mix with `aloop=loop=-1:size=2147483647,atrim=duration=<total>` (the Clinikally stitch does this) instead of hunting for a longer track. Any short lofi/hip-hop instrumental at `<your-project>/audio/music-bed.mp3` works (lab example: `coca-cola/ad-runs/video-01-museum-painting/audio/variants/A_hiphop/music.mp3`); generate one via ElevenLabs music if the repo has none. ### 8. Hook asset — HTML-mimic for app UIs, NB2 for photographic shots Route by what the hook *is*: - **App-UI screenshots** (Strava, Slack, a cancellation email, a statement): write an HTML file that mimics the real app's UI and render it to PNG with Playwright (~5 lines of script). AI-generating fake app UIs is unreliable — garbled chrome reads as slop. - **Photographic / lifestyle shots** (a throwback beach photo, a selfie-ish flex, the framed variant's background): generate with **Nano Banana 2** (`fal-ai/nano-banana-2/text-to-image`, via `scripts/fal_helpers.py`) — the Clinikally build generated its Goa throwback, beach background, and end-card base this way. Write each generator as a small `working/gen_<asset>.py` and save a `<asset>.meta.json` sidecar (prompt + model) next to the output for provenance. Two rules: 1. **Mimic the actual UI patterns, not just the vibe.** Abstract gradients + silhouettes read as "AI slop" (the first Peloton Strava attempt). Specific UI conventions (Strava-orange brand strip, polyline route, "Public · 2h ago" timestamp, real-feeling stats grid) read as real (the second attempt). 2. **Match the app's brand colors and typography exactly.** Get them from the app's marketing site or screenshots. See `examples/strava-card.example.html` for the canonical pattern. Render with: ```js const { chromium } = require('playwright'); const browser = await chromium.launch(); const ctx = await browser.newContext({ viewport: { width: 720, height: 880 }, deviceScaleFactor: 2 }); const page = await ctx.newPage(); await page.setContent(fs.readFileSync('strava-card.html', 'utf8'), { waitUntil: 'load' }); await page.waitForTimeout(150); await page.screenshot({ path: 'strava-card.png' }); ``` ### 9. Pacing curve — not flat 700ms gaps Real iMessage banter has rhythm: - One-word reactions ("bro no way" → "is that on an app"): **250–450ms** apart - Sentence reveals (after composer typing): **600–900ms** before the next event - Beat of silence (~600ms) before the final "bet" so it lands The canonical timeline in `record-master.template.js` is a good starting point — adjust event times rather than redesigning from scratch. ### 10. Variant B — iPhone-frame mode (the Clinikally pattern; default for new builds) Render with the atom's framed mode — `renderHTML(thread, { mode: 'with-iphone-frame' })` — so the chat plays inside a phone (Dynamic Island + status bar) over a brand-relevant background photo. Three gotchas the Clinikally build hit: - **The atom's framed branch hard-codes light mode.** Inject dark manually after render: `html.replace('<html>', '<html class="theme-dark">').replace('<body class="framed">', '<body class="framed theme-dark">')`. - **Layout math:** the framed UI is laid out at **514×914 logical** and scaled to the 1080×1920 output via `html { zoom: 2.10 }` (514 × 2.10 ≈ 1080). Use **fixed-px heights, not vh**, in every override. - **Scoped style overrides** (see the Clinikally `styleB()`): background photo as a data-URI on `body.framed`, compact `conv-header` (min-height 66px, 42px avatar), `.conversation` with `overflow: hidden; justify-content: flex-end`, attachment cards capped at `max-width: 58%`. In framed mode **the scroller is `.conversation`, not the document** — the driver must scroll that element. And the framed keyboard's `.input` starts empty: rebuild it on the first composer keystroke (inject `.composer-text` + caret + send button — the driver's `composerSpan()` does this). ### 11. Record at the NATIVE output resolution Playwright's `recordVideo` **cannot upscale** — set `viewport == recordVideo.size == the output size (1080×1920, deviceScaleFactor 1)` and let the `zoom` carry the layout to full size. The old 720×1280-then-lanczos-upscale flow ships a soft master; native recording is visibly crisper, especially bubble text. ### 12. Trim the paint offset so SFX cues align exactly `recordVideo` starts rolling at `newContext()`, but the page doesn't paint the iMessage shell until ~500ms later — so MP4 t=0 ≠ TIMELINE t=0 and every SFX cue lands late. Measure the delta (`Date.now()` at context creation → `__driverReady` after `setContent`), then trim it off the head: `ffmpeg -ss <offset> -i <recording> -t <TOTAL_DURATION> …`. After the trim, the cue list's `t` values line up with the video exactly. ### 13. EVERY sent bubble gets a full composer drive The composer-typed text must **equal** the sent bubble's text — a partial composer preview that "sends" a longer bubble reads as fake on a second watch. Pace at ~12–15 chars/sec with ±30% per-character jitter (`perChar * (0.7 + Math.random() * 0.6)`) so the typing feels like thumbs, not a metronome. Pair each send: `composer` (dur ≈ chars/14) → `pop` + `composer-clear` at the same `t`. ## Concept catalog Most iMessage ads fit one of these six angles. Use them as a brainstorming menu, not a checklist. Strongest hooks across all six: a specific number, a small act of self-trust, or a physically novel product mechanic. | Angle | Hook (the attachment) | Reveal | |---|---|---| | **Result-as-screenshot** | a number that brags by itself — race time, app summary, dashboard, FTP score | "X mins a day. that's it." | | **Setup flex** | photo of your space — tiny apartment, race-kit corner, home gym in a closet | "this is the whole gym" | | **Cancellation moment** | confirmation receipt — gym cancellation email, subscription cancelled page | "$X → $Y, do the math" | | **Feature-as-punchline** | short video clip of the feature in motion — rotating screen, transforming product | the mechanic *is* the brand | | **Friend-asks-friend (inverse)** | the peer initiates with the wow — "how are you doing this 😭" | you reply with the brand | | **Receipt-as-hook** | mundane financial document — credit-card statement, App Store receipt | small act of self-trust | Pick the strongest angle for the brand voice *before* writing copy. Most "the script is fine but feels off" feedback comes from picking the wrong angle, not bad copy. Catching that here is 30 seconds; catching it after recording is 30 minutes. ## End-card variants — which to pick The molecule ships two templates. Pick by brand voice, not aesthetic preference. See `references/end-card-recipe.md` for the full playbook. | Template | When to pick | |---|---| | `end-card.template.html` (abstract glow + product) | Brand voice is playful or product-led. Hook is a *thing* (CatchBack). | | `end-card-photo-bg.template.html` (athlete/persona photo) | Brand voice is grounded or lifestyle-led. Hook is a *person* (Peloton). | | Reference-styled (NB2 base + HTML overlay) | A specific end-card look is requested or referenced. NB2-generate the base image from the reference (Clinikally: `working/gen_endcard_v2.py` from `source/pellops-endcard-reference.jpeg`), then overlay the real wordmark + CTA in `end-card.html`. Text/logo stay HTML-crisp; only the backdrop is generated. | ## Workflow ### State 0 — Brainstorm angles (skip only if direct port of a reference)
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub