- name
- genvid-collage
- description
- Create a documentary-style paper-collage promo video (nature-documentary parody, "field notes", scrapbook cutouts, rubber stamps, binocular HUD) for a product, with every visual asset generated by an image/video model, a narrated voiceover timed to phrase onsets, licence-clean music, and a rendered vertical MP4. Use when asked for a viral-style Reels/Shorts/TikTok promo that "doesn't look like AI slop", a scrapbook or cutout-animation explainer, a "field notes" / "observation" style ad, or a video "like that famous one" built from generated characters, props, and short clips.
# GenVid Collage
Produce the finished MP4, not a storyboard. The film observes a relatable "species" (the customer's pain) like a wildlife documentary, then reveals the product as the fix. Every asset is generated for this film; every product claim is verified.
## Mandatory startup
1. Confirm three capabilities. Stop and ask for any that is missing; never substitute silently:
- **Narration:** a TTS provider plus how to reach its key (env var or secret). If the user has a router such as Speko, it may proxy several providers through one key. Never print or commit a key.
- **Image generation** with reference-image character consistency (for example Google Flow / Nano Banana, Gemini, or Grok Imagine), reachable through a signed-in browser the agent can drive.
- **Video generation** for 1–3 short clips with an image as the first frame (for example Veo in Flow). These usually cost credits. Get the user's go-ahead for the spend, and approve each generation once. Never retry a paid generation in a loop.
2. Load the installed `remotion-best-practices` skill in full (layout, audio, transitions, FFmpeg rules) and say that you are using it. If it is unavailable, say so and follow [references/production-pipeline.md](references/production-pipeline.md).
3. Read every reference before you write a line of the film:
- [references/style-bible.md](references/style-bible.md): the visual grammar, timings, and motion rules.
- [references/story-and-script.md](references/story-and-script.md): the documentary arc and the voiceover format.
- [references/asset-generation.md](references/asset-generation.md): the prompts, consistency method, downloading, and cutouts.
- [references/audio-and-timing.md](references/audio-and-timing.md): narration, verification, phrase onsets, and the music rules.
- [references/qa-checklist.md](references/qa-checklist.md): what to check before you call it done.
## Workflow
### 1. Brief and truth
Establish the product, the audience, the platform (default 1080×1920 at 30 fps, 80–120 s), the narration language, the CTA, and the "species" (the target user's daily pain). Verify every product claim against the codebase or the live product: channels, languages, setup time, trial terms, and what happens when the product doesn't know an answer. Never invent metrics, ratings, download counts, or testimonials. Jokes about the *problem* are fine; numbers about the *product* must be real.
### 2. Script first
Write the narration with the arc and line format in [references/story-and-script.md](references/story-and-script.md). Use one line per scene, 4–16 s each, with emotion tags where the model supports them. Then write the scene list with one focal visual per scene, plus the on-screen words that will land on spoken words.
### 3. Generate the voice before the visuals
The voiceover drives the timing:
1. Synthesize one file per line.
2. Transcribe each file back to text with any STT model to catch dropped words, garbled names, or tags read aloud. Regenerate the bad lines.
3. Measure the phrase onsets with `scripts/phrase-onsets.sh`.
These onsets are the beat sheet. Visual events are scheduled to them, never to guesses.
### 4. Generate the assets
Follow [references/asset-generation.md](references/asset-generation.md):
1. Make one hero character first, then every pose with that image as the reference.
2. Shoot cutouts on a plain light-grey seamless background.
3. Make props 1:1, and make backgrounds full-frame 9:16 with an empty middle.
4. Make 1–3 Veo clips from start frames.
5. Download at 2K with real clicks, one at a time.
6. Lift the subjects with `scripts/lift.swift` (macOS Vision), or another matting tool on other hosts.
7. Review everything on a contact sheet (`scripts/contact-sheet.sh`) before choosing.
### 5. Build the film
1. Copy `assets/template/` into the Remotion project, fix the import paths, and adapt it. You get:
- `kit.tsx`: paper, cutouts on twos, tape, notes, pen arrows, stamps, bubbles, binoculars, badge and confetti.
- `Film.tsx`: the scene table, the voiceover per scene, the music with a silence beat, and the transitions.
- `scenes.example.tsx`: two worked scenes.
2. Give every scene a `VO_AT` lead-in. Write beats as `at(sec)`, using the measured onset seconds.
3. Show real product UI in the solution half: embed the product's real components or faithful rebuilds inside a "pinned" scaled frame. Never invent a mock-up of the product.
### 6. Audio and render
1. Mix the music under the narration (about 0.2 gain), cut it to silence on the turning-point stamp, and bring it back on the reveal.
2. Render, then normalize the loudness to −14 LUFS with `scripts/finalize-audio.sh`.
3. If you share the film to a phone, also make a preview copy under 30 MB.
### 7. QA, then report
1. Render stills at each scene's key beat and build a strip.
2. Work through [references/qa-checklist.md](references/qa-checklist.md) and fix what it finds.
3. Re-render, then sample the final file every 6 s.
4. Report the path, duration, voice provider/model/voice, the generators and credits used, the music licence, what was verified, and any known imperfections.
Log every generated asset and track licence next to the project's other licence records.
## Completion standard
Finish only when all of these are true:
- The MP4 exists and has been sampled across its whole timeline.
- The narration was verified line by line by transcription.
- No product claim is unverified.
- The loudness is around −14 LUFS.
- Every temporary file you made (test stills, bundles, scratch audio) has been deleted.
GitHubで見る