Use whenever the user asks to tell, adapt, or animate any story as a stick-puppet story, puppet-theatre video, rod-puppet tale, or strictly stick-puppet animation. Builds a reusable GPT Image 2 asset library, then animates scenes with Google Gemini Omni Flash reference-to-video and native narration. No lip-sync is required or requested.
Installation
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
Use whenever the user asks to tell, adapt, or animate any story as a stick-puppet story, puppet-theatre video, rod-puppet tale, or strictly stick-puppet animation. Builds a reusable GPT Image 2 asset library, then animates scenes with Google Gemini Omni Flash reference-to-video and native narration. No lip-sync is required or requested.
Turn any narrative into a strictly stick-puppet production. First use GPT Image 2 to create a reusable visual asset pack—character puppets, prop puppets, modular scenery, and a stage/style reference. Then animate scene clips with Google Gemini Omni Flash via FAL reference-to-video, using the same references repeatedly to preserve identity. Omni supplies synchronized narration, music, and puppet-stage foley.
The governing rule is literal: every acting character, creature, vehicle, important prop, and moving story element must visibly be a physical puppet mounted on a stick or rod. Scenery may be a flat theatrical backdrop, but it must remain recognizably handmade stage scenery. Never drift into ordinary animation, live action, claymation, glossy 3D, or paper-collage characters without rods.
There is zero lip-sync requirement. Puppet mouths should normally be fixed. Communication comes from rod movement, tilt, bounce, pose, staging, narration, and sound—not mouth animation.
Operating model:
story → adaptation/beat map → reusable GPT Image 2 asset pack → per-clip GPT Image 2 scene keyframes → asset approval → Omni reference clips → dense QC/retry → assembly
When Omni drifts into hands/operators, lip-sync, volumetric characters, or unstable stage geometry, load references/omni-medium-hardening.md for the proven asset-level fixes, dominant-keyframe prompting pattern, 0.5-second QC procedure, and defect-specific corrective routing.
When to Use
Use automatically for:
“tell this as a stick-puppet story”
rod-puppet, stick-puppet, tabletop puppet-theatre, or handmade puppet-stage video
stories, myths, jokes, histories, explainers, ads, biographies, or abstract ideas requested in this exact medium
recurring series that should reuse the same cast and stage assets
Do not use for:
marionettes with strings unless the user explicitly asks to mix puppet types
hand puppets, sock puppets, shadow puppets, claymation, or generic cut-out animation
talking-head lip-sync
a single static illustration with no video request
Required Runtime and Routing
Asset creation
Use the configured GPT Image 2 image-generation tool (image_generate). Create assets before video generation. Use image editing with the accepted reference whenever a correction or variation must preserve a character.
Do not substitute another image model silently. If GPT Image 2 is unavailable, report the blocker.
Animation and audio routing
Default route — Google Gemini Omni Flash:
Unless the user explicitly asks for Seedance, use the dedicated FAL Omni MCP tools:
mcp__fal_omni__upload_asset
mcp__fal_omni__omni_reference_to_video — default and preferred
mcp__fal_omni__omni_image_to_video — only when exactly one reference is genuinely sufficient
mcp__fal_omni__omni_edit_video — localized visual repairs only
Google Gemini Omni Flash constraints:
aspect: 16:9 or 9:16
duration: 3–10 seconds per clip
native synchronized audio is generated with the video
the edit route cannot repair or replace speech
Use reference-to-video whenever reusable identity assets exist. Omni Flash is the normal path because it is good enough for this workflow and keeps production simpler.
Explicit Seedance override:
Use Seedance only when the user explicitly says to use Seedance for this production or comparison. Do not infer Seedance from availability, prior experiments, or the existence of Seedance tools.
For the Seedance branch:
load seedance-2, seedance-api-upload, and seedance-api-generate
upload every local asset through seedance_upload_asset
represent every supplied asset in elements; do not default to first_frame, last_frame, reference_images, reference_video, or reference_audio
use the exact GPT Image 2 scene keyframe as the dominant clip reference and include only active identity/prop references needed by that clip
preserve the same dense 0.5-second QC gates
expect Seedance native narration to mispronounce uncommon names or terms; validate with ASR
if Seedance visuals pass but narration does not, prefer overlaying the already approved narration track rather than rerolling otherwise-good video, provided the user did not explicitly require fully native Seedance audio
Never switch between Omni and Seedance silently. Report which route produced the final visuals and whether narration was native or overlaid.
Defaults
Setting
Default
Aspect
16:9 for stories; 9:16 for social requests
Clip length
8–10 seconds
Total length
determined by story; do not crush the plot into 30 seconds automatically
Narrator
warm, expressive storybook narrator
Dialogue
narrated indirectly unless exact dialogue materially improves the story
mostly locked frontal proscenium view with restrained pushes/pans
Retry ceiling
one paid corrective video attempt before asking
Ask only for missing choices that materially change the story, format, or spend. If the user supplies a story and says “make it,” infer sensible defaults and proceed to the approval checkpoint.
Strict Visual Grammar
What counts as a stick puppet
A valid puppet has all of these:
A flat or shallow-relief handmade body silhouette.
At least one clearly visible support/control rod extending downward or to the side.
Rigid or hinge-like movement consistent with a physical puppet.
A coherent material treatment: painted card, thin plywood, felt-faced board, or similar craft material.
Practical contact with a miniature/tabletop stage rather than free-floating cinematic animation.
For humanoids, one central support rod is sufficient; optional thin arm rods may appear. For creatures, vehicles, clouds, fire, magic, crowds, and objects that move, include visible rods too. Rod visibility is a feature, not a defect.
Forbidden drift
Every image and video prompt must explicitly avoid:
humans holding puppets or visible puppeteers
realistic living actors replacing puppet characters
invisible rods, floating characters, or fully articulated 3D characters
lip movement, mouth flaps, phoneme animation, or talking-head framing
marionette strings, hand-inside puppets, sock puppets, and shadow silhouettes
uncontrolled extra characters, duplicated limbs, or unexplained text
Hands may appear only if the user explicitly wants the puppeteers revealed. Otherwise frame below or around the control area so rods remain visible but operators do not.
Performance language without lip-sync
Use:
lean forward/back for intention
tilt for curiosity or doubt
quick vertical bounce for speech emphasis
slow sag for sadness or defeat
side-to-side tremble for fear
quarter-turn or silhouette swap for reactions
crossing paths, hiding behind flats, entering through wings
prop exchange, stage wipe, curtain, rotating backdrop, or scenery slide
Keep mouths fixed. If a generated mouth visibly speaks, the clip fails QC even if narration is correct.
Reusable GPT Image 2 Asset System
Create a project library, not disposable scene illustrations.
Character master — one clean full-body puppet per recurring character, neutral pose, fixed mouth, clearly visible rod, uncluttered contrasting background.
Character variants — only variants the story needs: profile/three-quarter orientation, one or two pose silhouettes, costume change, damaged/transformed state. Variants must be edited from the accepted master.
Prop board — reusable important objects as separate rod-mounted puppet props.
Scenery modules — one separate full-frame backdrop plate per location. Do not use a multi-location triptych as a video reference; it invites the model to remix architecture.
Scale/contact sheet — accepted cast together on the same stage, used to lock relative size and palette.
Scene keyframes — one GPT Image 2 composite tableau per clip, derived from the accepted masters and the exact location plate. This is the primary structural reference for Omni and must already contain the intended stage geometry, character placement, visible rods, lower masking rail, and fixed-mouth design.
Asset rules
One asset should have one job. Do not bake a character, prop, and background into every master reference. Scene keyframes are the deliberate exception: they are derived composites used to lock one clip's composition.
Keep the same silhouette, face marks, costume colors, rod position, material, and scale across variants.
Design characters with no functional mouth. Prefer a completely blank lower face or one tiny sealed painted dash integrated into the artwork. Never generate lips, teeth, an open mouth, a jaw seam, or a hinged lower face. If moustaches are used, they must be a single rigid painted shape, not a separate fuzzy layer.
Make all characters unmistakably planar 2D cardstock silhouettes: uniform edge thickness under 3 mm, no rounded anatomy, no volumetric noses, no realistic skin shading, no articulated jaw, no sculpted hands, and no depth beyond stacked paper details.
Add a permanent waist-high decorative lower stage masking rail to every stage and scene keyframe. Rods emerge through narrow slots behind this rail; operators and hands remain physically impossible to see. Do not show open space beneath the rods.
Use editing from the accepted master for variants; do not regenerate recurring characters from prose alone.
Generate each location as its own full-frame plate with fixed proscenium, floor line, curtain wings, masking rail, camera height, and lighting. Never use a scenery triptych directly in Omni.
Create every scene keyframe with GPT Image 2 using the accepted stage, active character masters, prop masters, and one location plate as image references. Reject keyframes that alter stage geometry or make a character volumetric.
Avoid transparent backgrounds unless the generation route reliably preserves them. A plain high-contrast studio background is acceptable for masters.
Save the original generated file locally before uploading it.
Give each asset a stable ID and filename such as char-fox-master.png, prop-key.png, set-forest-night.png, and keyframe-03.png.
GPT Image 2 master prompt pattern
Create a reusable production asset for a strictly physical stick-puppet theatre.
ASSET: [single character/prop/stage module].
CONSTRUCTION: flat handmade [painted card/thin plywood/felt-faced board], shallow relief, slightly imperfect cut edges, practical paint texture.
ROD: one clearly visible [wooden/black] control stick attached to [location], extending fully beyond the puppet; it must not disappear or be cropped.
IDENTITY: [stable silhouette, face marks, costume, palette, scale].
POSE: [neutral or named production pose]. Mouth fixed and closed; no lip-sync shape.
VIEW: [front/profile/three-quarter], orthographic production-reference feel, full object visible, generous margin.
BACKGROUND: simple contrasting studio backdrop, no scenery unless this is a scenery asset.
STRICTLY A PUPPET ASSET: no living actor, no puppeteer, no hand, no strings, no sock puppet, no clay, no glossy 3D, no free-floating animation, no text, no extra object.
Asset QC
Reject or repair an asset when:
the support rod is missing, ambiguous, detached, or cropped
the character looks alive rather than constructed
the mouth is open in a speech shape
recurring identity markers changed
the silhouette will be unreadable at phone size
the background contains accidental narrative elements
Completion criterion: every recurring entity has an accepted reusable master, required variants are derived from it, and the cast scale sheet is coherent.
Narrative Adaptation
Preserve the story’s causality, emotional turns, and ending. Simplify staging—not meaning.
For each beat, define:
narration or exact dialogue
which puppet assets appear
entrance, performance gesture, and exit
backdrop/prop changes
opening tableau and final hold
sound: narration, music, and practical puppet/stage foley
duration and transition
Prefer one dramatic action per clip. If a sentence needs several locations or reversals, split it. Target roughly 18–23 narrated words per 10 seconds, fewer for emotional pauses or technical language.
Dialogue does not imply lip-sync. When a character “speaks,” hold the fixed mouth and use a small rod bounce or tilt while the narrator/voice performs the line. Explicitly state this in every relevant clip prompt.
Completion criterion: the workspace exists and brief.md records the source story, requested format, and any non-negotiable facts.
2. Write the adaptation and asset manifest
Create production-plan.md from the linked template. Include the complete beat map and a deduplicated asset manifest. Reuse the same asset IDs across scenes.
Completion criterion: every story beat maps to a clip, every clip names its assets, and every recurring asset is generated once then referenced—not reinvented.
3. Present one approval checkpoint
Before paid generation, show:
concise adapted script/beat list
cast and reusable asset list
visual/material direction
aspect, duration estimate, selected video provider, and number of generated clips
approximate base video cost and one-retry ceiling, noting that current FAL prices may change
Ask for one approval covering story adaptation and production scope. If the user already explicitly instructed immediate production with sufficient detail, treat that as approval.
Completion criterion: approval is explicit or contained in the execution request.
4. Generate the reusable asset pack with GPT Image 2
Generate the stage/style bible first, then character masters, then variants by editing the accepted masters, then individual props and individual full-frame location plates, then the cast scale sheet. Finally generate one scene keyframe per clip with GPT Image 2, using the accepted masters as image references. The scene keyframe—not a loose style board—must lock the exact proscenium geometry, masking rail, backdrop, cast placement, and opening tableau that Omni should preserve.
Do not begin paid Omni generation until the masters and every scene keyframe pass asset QC. Inspect keyframes at full resolution for blank/sealed mouths, planar edges, unchanged costume/face marks, visible rods entering masked slots, impossible hand visibility, and location geometry matching the stage bible.
Completion criterion: all assets in the manifest and all clip keyframes exist locally and pass rod, identity, mouthless-face, planar-construction, operator-exclusion, framing, and structural-consistency checks.
5. Upload and index references
Upload accepted local assets using mcp__fal_omni__upload_asset. Record local path, public URL, role, and asset ID in uploads.json.
Keep reference order stable within a clip:
<IMAGE_REF_0> — stage/style bible
<IMAGE_REF_1> onward — speaking/primary characters in story importance order
then props
then scene backdrop
The prompt must bind every supplied image explicitly. Never pass mystery references.
Completion criterion: each asset needed by the next clip has a valid FAL URL and an unambiguous <IMAGE_REF_N> binding.
6. Write reference-to-video prompts
Use this structure:
FORMAT: [16:9/9:16], [duration] seconds, locked-off orthographic recording of a strictly planar physical tabletop stick-puppet theatre.
REFERENCE BINDINGS: <IMAGE_REF_0> is the exact scene keyframe and the primary structural truth. Preserve its proscenium, curtain wings, floor line, lower masking rail, lighting, character placement, scale, and negative space exactly. Remaining references are identity backups only: <IMAGE_REF_1> is [puppet], [...].
OPENING TABLEAU: Begin as a nearly exact copy of <IMAGE_REF_0>. Do not redesign, add depth, or reveal anything beneath the masking rail.
PUPPET CONSTRUCTION LOCK: Every character is a flat 2D painted-card silhouette with uniform paper-thin edges, rigid body, fixed limbs, no rounded anatomy, no realistic skin shading, and no 3D volume. Each has a visible bamboo rod entering a narrow slot behind the lower masking rail. Rods move; hands and operators never exist in frame.
FACE LOCK: Characters have no functional mouths. Preserve the blank lower face or tiny sealed painted dash exactly. No lips, teeth, tongue, jaw seam, mouth opening, facial speech motion, cheek motion, or lip-sync at any moment. Narration is fully off-screen.
PUPPET ACTION: [chronological rigid whole-puppet translations, tilts, and single-piece bounces only]. No articulated facial performance. Every moving narrative element remains visibly mounted on a control stick.
NARRATION: An off-screen [fixed narrator specification] says exactly: “[script]” No character emits sound.
AUDIO: clean close-mic off-screen narration; subtle stage creaks, card/wood taps, curtain or scenery-slide foley; restrained music under speech; no extra voices.
CAMERA: locked frontal orthographic proscenium view. No orbit, perspective change, rack focus, camera push, shallow depth of field, or view below the masking rail.
ENDING HOLD: [simple stable tableau].
STRICT AVOID: hands, fingers, arms, operators, puppeteers, human skin, visible control area, lips, mouth opening, lip-sync, jaw movement, 3D or rounded characters, realistic people, volumetric bodies, missing rods, floating figures, camera movement, changing architecture, moving floor line, altered curtains, extra scenery, marionette strings, hand puppets, sock puppets, claymation, glossy CGI, anime, text, subtitles, or extra characters.
Use reference-to-video even for scenery-heavy clips when character identity matters. A clip should usually use only the references it needs; too many references dilute control.
Completion criterion: every clip prompt identifies all references, exact narration, fixed-mouth rule, rod-visible action, and final hold.
7. Generate clips and QC immediately
Generate clips in story order with mcp__fal_omni__omni_reference_to_video. Put the approved scene keyframe first as <IMAGE_REF_0> and treat it as the structural source of truth; include only the active character/prop masters needed to reinforce identity. Download each result to clips/clip-NN.mp4 and inspect it before paying for the next dependency.
Check with ffprobe for duration, dimensions, codecs, and audio. Extract frames at 0.5-second intervals across the entire clip, plus tighter consecutive frames during any apparent speech gesture or fast movement. A three-frame contact sheet is insufficient: hands and lip-sync can appear for only a few frames.
Compare sampled frames directly against the scene keyframe and character masters. Reject a clip for any of these—even briefly:
any hand, finger, wrist, arm, operator, or human skin enters frame
a blank/sealed mouth opens, changes shape, gains lips/teeth, or exhibits facial speech motion
a flat puppet gains rounded anatomy, realistic skin shading, volumetric depth, articulated joints, or 3D facial geometry
proscenium, curtain wings, floor line, masking rail, camera angle, or backdrop architecture changes materially
a rod disappears, disconnects, turns into an arm, or reveals open operator space beneath it
narration is wrong/missing or a character voice is added
Repair policy:
localized visual defect with correct audio → try Omni video edit only when the defect can be removed without changing voice; instruction must name the exact timestamp range and end Keep everything else the same.
any hand/operator appearance → regenerate with a higher masking rail and less vertical puppet motion; do not merely ask to “hide the hand”
any mouth/lip-sync motion → regenerate from a mouthless scene keyframe; remove dialogue punctuation and state that all speech is off-screen narration with no character voice
3D drift → regenerate with a locked orthographic camera, planar silhouette wording, simpler whole-puppet translations, and the scene keyframe as dominant first reference
structural drift → reduce references, remove camera motion, and regenerate from the exact full-frame keyframe
identity drift → reduce reference count and foreground the affected puppet master
Paid correction count follows the user's approved policy. When unlimited corrections are authorized, that means correct until the acceptance gates pass, not blind unlimited rerolls: diagnose and change the controlling asset or prompt each time.
Completion criterion: each accepted clip passes technical, hand-free, planar-stick-puppet, identity, mouthless/no-lip-sync, structural-continuity, narrative, and audio checks at 0.5-second sampling.
8. Assemble and verify
Normalize accepted clips to one size/frame rate/audio format, concatenate with FFmpeg, and output H.264/AAC MP4. Avoid decorative crossfades that make rods or puppets ghost together; hard cuts, curtain wipes, scenery slides, and brief stage blackouts suit the medium.
Verify the final file with ffprobe, a contact sheet, and a complete listen/watch-through. Check every boundary for black-frame glitches, audio cuts, identity jumps, mouth movement, and puppet-medium drift.
Completion criterion: final.mp4 plays end to end, audio remains synchronized, every story beat is present, and every acting/moving entity remains a visible stick puppet.
Continuity Strategy
Do not feed previous video frames back merely to preserve identity; use the reusable GPT Image 2 masters as the stable source of truth. A prior final frame may be added only when a specific spatial handoff matters and reference capacity permits it.
Prefer theatrical resets:
curtain close/open
backdrop slides laterally
foreground wing wipes frame
stage lights dim and rise
a rod-mounted title/icon crosses as a wipe
visible scenery swap during a held narration beat
These transitions preserve the handmade language while preventing accumulated generative artifacts.
Narration and Voice
Pin the same narrator description in every prompt: language/accent, age range, vocal character, energy, pace, and recording quality. Any provider may vary voice or pronunciation across separately generated clips; do not promise perfect identity.
Prefer native provider narration when it passes per-clip ASR and listening checks. If several visually accepted clips mispronounce names or domain terms, a separately verified narration overlay is an allowed corrective workflow: preserve the provider-generated visuals, replace the flattened audio deliberately, verify sync and total duration, and disclose the provenance split. Do not casually layer competing speech or call an overlaid track “native audio.”
Proven Production Patterns
For multi-location stories, control comes from a layered reference hierarchy rather than one compact board:
one fixed stage/style bible with a permanent lower masking rail
one master image per recurring mouthless planar character
rod-mounted prop masters
one separate full-frame location plate per setting
one GPT Image 2 scene keyframe per clip, derived from the accepted masters
Put the scene keyframe first in Omni reference mode and include only the active identity masters after it. Never use a multi-location scenery board as the structural video reference; it encourages architectural remixing.
Inspect every generated clip at 0.5-second intervals across its full duration, with tighter consecutive-frame review around suspicious motion. Start/middle/end sampling can miss brief hands, lip-sync, and 3D drift. Check narration separately with per-clip ASR, then transcribe the assembled film again because correct source clips do not prove boundary-safe assembly.
See references/omni-medium-hardening.md for defect signatures, corrective routing, and the hardened dominant-keyframe workflow. Load references/provider-qc-and-recovery.md for provider-neutral dense QC plus Seedance elements, full-bleed recovery, timeout reconciliation, pronunciation checks, and verified narration-overlay assembly. See references/distant-lamp-production-pattern.md only for the earlier baseline recipe; when it conflicts with the hardening references or this SKILL.md, follow the hardened workflow.
Common Pitfalls
Calling any flat character a stick puppet. A visible attached control rod and physical rod-driven motion are mandatory.
Generating scene illustrations instead of reusable assets. Build masters and modular scenery first; derive variants from accepted masters.
Using Omni text-to-video despite having assets. This workflow is reference-first. Use reference-to-video.
Asking for lip-sync out of habit. Explicitly forbid mouth movement in images and every video clip.
Hiding the rods for cleanliness. Rods prove the medium. Keep them visible while operators remain out of frame.
Overloading references. Supply only the stage, active cast, key prop, and backdrop needed for that clip.
Letting scenery become cinematic reality. It must remain a theatrical flat/module with practical craft texture.
Using every puppet type interchangeably. No strings, hands-inside, socks, or shadow-only characters unless the user changes the brief.
Blind rerolls. Diagnose rod loss, mouth motion, identity drift, narration, or staging separately and change the prompt.
Assuming exact narrator consistency. Repeat the voice specification, but disclose that separate Omni clips can vary.
Skipping the full listen-through. Correct visuals do not prove complete narration or clean boundaries.
Creating moving non-puppet effects. If a cloud, flame, moon, vehicle, sign, or magical object acts, mount it visibly on a rod too.
Verification Checklist
GPT Image 2 created the reusable asset pack before video generation
Every recurring character has one accepted master and stable variants
Important props and moving effects are visibly rod-mounted
Scenery remains handmade theatrical scenery
Every active clip uses the explicitly requested reference-video provider; Gemini Omni is the default, while Seedance requests use uploaded elements
Every reference is explicitly bound in the prompt or represented as an elements entry
Every actor remains recognizably constructed and attached to a visible rod
No mouths animate; no lip-sync was requested or introduced
No puppeteers, hands, strings, sock puppets, clay, living actors, or glossy CGI appear
Narration is exact enough for the approved adaptation and audible over music/foley
Paid retries stayed within the approved ceiling
Final MP4 passed technical, visual-medium, identity, story, and audio review
Reusable assets, prompts, clips, and final path are reported to the user