| name | short-form-video-planning |
| description | Produces a short-form video plan for TikTok, Instagram Reels, or YouTube
Shorts with hook formula, single-concept structure, caption strategy, audio
selection criteria, and loop optimization. Use when the user asks to plan
a short video under 60 seconds, create a Reel or TikTok concept, or
structure a vertical video for social media.
Do NOT use for long-form video scripts over 60 seconds (use
video-script-writing), YouTube channel strategy (use youtube-video-strategy),
or video shot lists (use video-storyboard).
|
| license | Apache-2.0 |
| metadata | {"author":"foundry-skills","version":"1.0.0","tags":"video-production planning template","category":"design-creative","subcategory":"video-audio","depends":"","disclaimer":"none","difficulty":"intermediate"} |
Short Form Video Planning
When to Use
Use this skill when the user's request matches any of the following scenarios:
- The user wants to plan, script, or structure a single video intended for TikTok, Instagram Reels, or YouTube Shorts, where the finished video will be 60 seconds or under
- The user has a concept or topic and needs help turning it into a video that performs -- they want watch time, shares, saves, or follower conversions, not just a recorded talking head
- The user is planning a vertical video (9:16) for mobile-first viewing and needs guidance on hook design, beat structure, on-screen text placement, and caption strategy
- The user asks specifically about how to make a short video loop, increase completion rate, or survive the "first 3 seconds" of a scroll feed
- The user is adapting existing content (a blog post, a long-form video, a tweet thread) into a short-form format and needs help compressing it without losing impact
- The user is planning a series of short videos and needs each episode planned as a standalone unit
- The user wants to create a trend-responsive video and needs the trend structure mapped before content is assigned to it
Do NOT use this skill when:
- The user needs a full script with dialogue, scene descriptions, and shot-by-shot coverage for videos longer than 60 seconds -- use
video-script-writing instead
- The user needs a YouTube channel growth strategy, upload cadence, or metadata optimization beyond a single video caption -- use
youtube-video-strategy instead
- The user needs a detailed shot list with camera angles, lens choices, lighting setups, and editing sequences -- use
video-storyboard instead
- The user is planning a live stream, webinar, or event that happens to be recorded -- those formats follow different engagement mechanics entirely
- The user needs a paid advertising video (pre-roll, bumper ad, or sponsored content) -- ad video structure follows different retention curves and CTA placement rules than organic short-form
- The user is planning a podcast or audio-first content that will have a static visual or looping image over it -- this is not a short-form video in the algorithmic sense
- The user wants analytics or performance reporting on videos already published -- that is a separate audit and optimization workflow
Process
Step 1: Clarify the Core Inputs Before Planning Anything
Never begin planning without confirming these five parameters. Guessing on any of them produces a plan that cannot be executed.
- Topic and single core message: The video must be expressible in one declarative sentence of 10 words or fewer. If the user's concept takes two sentences to explain, the concept is not ready for short-form. Help them cut it before proceeding. Example of a ready concept: "Mouth breathing ruins your sleep quality." Example of an unready concept: "The relationship between breathing patterns, sleep architecture, and how modern diets contribute to airway restriction."
- Target platform: TikTok, Instagram Reels, YouTube Shorts, or cross-platform. Platform choice affects safe zones for on-screen text (platform UI overlaps differ), caption character limits, and audio discovery mechanics. TikTok has the most aggressive audio-driven discovery; YouTube Shorts discovery relies more on title and thumbnail for suggested video placement.
- Content type: Choose one -- educational, transformation, storytelling, tutorial, opinion/hot take, or entertainment/humor. The content type determines the beat structure in Step 3.
- Target audience: Not a demographic but a psychographic state. "25-35-year-old women" is not useful. "Someone who feels overwhelmed by meal prep but wants to cook healthier" is the usable version -- it contains the pain, desire, and situation that the hook must speak to.
- Duration target: 7-15 seconds (micro), 15-30 seconds (short), 30-45 seconds (standard), or 46-60 seconds (long-short). Shorter videos have higher completion rates but lower information density. The algorithm weights completion rate heavily -- a 10-second video watched fully beats a 60-second video watched 40% of the way through. Default to the shortest duration that can fully deliver the concept.
Step 2: Design the Hook -- The First 1.5 to 3 Seconds
The hook is not an introduction. It is a retention device. Research on short-form platforms consistently shows that 30-50% of viewers who will ultimately abandon a video do so within the first 3 seconds. The hook must interrupt the scroll and create an information gap or an emotional spike before the viewer consciously decides to keep watching.
- Pattern interrupt hook: Open with something visually or tonally unexpected. A chef saying "I hate cooking" before a recipe video. A personal finance creator holding cash and saying "I burned this on purpose." The incongruity creates cognitive dissonance that drives the viewer to resolve it by watching.
- Direct address hook: Speak to a specific situation the target viewer is in right now. "If you have been waking up tired every morning" lands on the exact person this video is for. The more specific the address, the higher the retention from the intended audience. Broad addresses ("Hey everyone") dilute this effect.
- Curiosity gap hook: Withhold the most interesting information and make it clear that the video contains it. "The reason your plants keep dying has nothing to do with how much you water them." The viewer must watch to close the gap. The curiosity gap works best when the withheld information is counterintuitive or contradicts a common assumption.
- Visual hook: The first frame itself must stop the scroll before any audio plays. (On TikTok, autoplay with sound is common; on Instagram Reels, many users view with sound off initially.) Use an unexpected visual -- a surprising result already visible, an unusual setting, or a before/after split that raises the question "how did that happen?"
- Stakes hook: Open by immediately stating what is at risk. "Most people making this mistake will never get promoted." Stakes hooks work well for professional, financial, and health content where the audience has something to lose.
- Duration constraint: The hook must be deliverable in 1.5-3 seconds. For spoken audio, this is approximately 5-10 words at natural speaking pace. For text overlay, this is one short phrase that can be read in under 2 seconds. Do not plan a hook longer than this -- if it takes 4+ seconds to deliver the hook, it is not a hook, it is a cold open, and cold opens kill short-form retention.
- Hook and thumbnail alignment: For YouTube Shorts, the thumbnail and title function as a pre-hook before the video starts. Plan the hook text to be consistent with but not identical to the title.
Step 3: Build the Beat Structure for the Chosen Content Type
Each content type has a proven beat sequence. Use the correct sequence -- do not invent a structure. The structure exists because it matches how attention behaves on short-form platforms: high at the start, drops around the 5-7 second mark, recovers if there is a pattern interrupt or escalation, and drops again at the end unless there is a payoff.
- Educational format (15-45 seconds): Hook (0-2s) -- Unexpected claim or counterintuitive fact (2-6s) -- Evidence or mechanism explanation (6-25s) -- Practical application (25-35s) -- Result statement that loops (35-45s). The key is that the "evidence" beat must validate the hook's claim without redundancy.
- Transformation format (20-45 seconds): Hook (0-2s) -- Before state: establish the problem clearly (2-8s) -- Process: show the change happening (8-30s) -- After reveal: the result that makes the process feel worth it (30-40s) -- Loop frame (40-45s). The transformation format lives or dies on how relatable the "before" state is. If the viewer does not recognize themselves in the before state, they disengage.
- Storytelling format (20-60 seconds): Hook (0-2s) -- Scene setting: who, where, what was happening (2-10s) -- Inciting moment: the thing that changed (10-20s) -- Tension or consequence (20-40s) -- Resolution or punchline (40-55s) -- Loop statement (55-60s). Storytelling is the hardest format to compress because setup context competes with brevity. Every second of setup must earn its place by making the payoff more impactful.
- Tutorial format (30-60 seconds): Hook showing the finished result (0-2s) -- "Here's how" bridge (2-4s) -- Step 1 (4-15s) -- Step 2 (15-30s) -- Step 3 if needed (30-45s) -- Final result reveal that matches the opening hook (45-60s). Tutorials require clear step numbering on-screen ("Step 1," "Step 2") because viewers often re-watch tutorials for reference -- the text overlays make re-watching useful.
- Opinion/Hot Take format (20-45 seconds): Hook -- a bold claim that will provoke agreement or disagreement (0-2s) -- Why most people believe the opposite (2-8s) -- The evidence or argument (8-30s) -- Restatement of the original claim with added nuance (30-40s) -- Engagement bait ending ("Tell me if you agree") (40-45s). Hot takes generate comment engagement, which the algorithm interprets as a positive signal. Plan for the controversy -- it is a feature, not a risk.
- Entertainment/Humor format (7-30 seconds): Pattern interrupt visual (0-1s) -- Setup (1-8s) -- Subversion or punchline (8-15s) -- Reaction beat if applicable (15-20s) -- Loop that makes the punchline funnier on second watch (20-30s). Comedy is the hardest format to plan because timing cannot be taught in a planning document -- but the beat positions for setup and punchline can be designated, and the length of the setup can be constrained to prevent over-explanation.
- Retention arc principle: Every beat transition is a micro-hook. At each transition point, plan a small payoff or new information that gives the viewer a reason to stay. A common mistake is building to one payoff at the end -- space the dopamine hits throughout.
Step 4: Plan the Audio Layer
Audio on short-form platforms serves two functions: it drives in-video retention (keeping the viewer watching) and it drives platform discovery (trending sounds are surfaced to more users). These two functions sometimes conflict -- a trending sound may not be the right tone for the content.
- Original narration: The best choice for educational, tutorial, and opinion content. Narration creates an authoritative voice that viewers associate with expertise. Pacing should be 130-160 words per minute for educational content and 160-200 words per minute for entertainment or hype content. Dead air (silence lasting more than 0.5 seconds) drops retention -- if you must pause for dramatic effect, limit it to a single beat of 0.5-1 second maximum.
- Trending audio: Trending sounds on TikTok peak at 2-4 weeks after they begin trending, then become associated with "dated" content and are algorithmically deprioritized. Use trending audio when the sound's mood, energy, or lyrical content genuinely fits the message -- forced trend usage looks inauthentic and reduces trust. When using trending audio, the visual story must carry the message because the audio will often be music or a sound effect, not narration.
- Music-only audio: Use for transformation, beauty, fashion, and lifestyle content where the visual is the primary communication channel. Choose music with a tempo that matches the cut rhythm. 120-128 BPM (typical dance/pop tempo) suits fast cuts every 0.5-1 second. 70-90 BPM suits slower, more deliberate content. Lyrics should be instrumental or have minimal, non-distracting words -- lyrics compete with on-screen text for cognitive attention.
- Hybrid audio (narration + music): The background music should sit at -12 to -18 dB below the narration volume. If a viewer can hear the music while focused on the voice, it is too loud. The music provides emotional texture; the voice provides information.
- Sound effects as audio texture: For tutorials and process videos, the ambient sounds of the activity (chopping, typing, mixing) function as audio texture that signals authenticity. Plan whether to use ambient sound or replace it with music -- using both simultaneously requires careful level mixing.
- Voice-off audio (no speaking, just captions): Some creators run videos with no voice for accessibility and silent-viewing audiences. If planning this format, every beat must be communicated through text overlays and visuals. Increase text overlay dwell time to minimum 2 seconds per screen (reading time for a short phrase). Rhythm must be carried by music and cut timing.
Step 5: Write the Caption Strategy
The caption is not a description of the video. It is a second engagement surface that extends the video's reach and drives specific algorithmic signals (saves, profile visits, comments).
- First visible line (the "above fold" line): On TikTok and Instagram, only the first 1-2 lines of the caption are visible before "more" -- this line must stand alone as compelling text. It should either reinforce the hook or add a dimension the video does not cover. Optimal length: 60-100 characters. Do not start with "In this video" -- that is a YouTube habit that wastes the above-fold space.
- Body text: 1-3 sentences that add context, a related tip, a surprising statistic, or a personal anecdote. The body text rewards viewers who engage enough to expand the caption -- it deepens the relationship and increases time spent on the profile.
- Call to action: One specific action per video. "Save this for later" drives saves, which is a strong algorithmic signal (it tells the algorithm this content has reference value). "Follow for part 2" drives follow-through. "Drop a [specific emoji] if you've experienced this" drives comments. Never use multiple CTAs in one caption -- it creates decision paralysis and reduces the rate of any single action.
- Hashtag strategy: 3-5 hashtags is the evidence-backed range. Using 20-30 hashtags is a legacy strategy from early Instagram that no longer drives meaningful discovery on any platform. Use: 1 broad category hashtag (500M+ posts), 1-2 niche hashtags (under 10M posts), 1 audience-specific hashtag (the community the viewer identifies with), and 1 content-specific hashtag (the specific topic, not the format). Never use #viral, #fyp, or #trending -- these are signal-diluted to the point of uselessness.
- Platform caption differences: TikTok captions have a 2,200-character limit but discovery weighting favors shorter captions. Instagram Reels captions have the same limit but benefit more from longer body text (Instagram's algorithm weighs caption engagement more heavily). YouTube Shorts captions are the title plus description -- the title functions as the above-fold line and must include the searchable keyword.
Step 6: Engineer the Loop
A looped video -- one where the viewer watches it a second time immediately -- generates a second play count and increases average watch time, which is one of the most heavily weighted algorithmic signals on all three platforms. A well-engineered loop can double the effective watch time of a 15-second video.
- Visual loop: The last frame should compositionally mirror or directly connect to the first frame. Same camera angle, similar lighting, similar subject framing. This creates a subliminal sense of continuity that makes the repeat feel seamless rather than jarring.
- Audio loop: The last spoken word or sound should resolve in a way that makes the first word feel like it naturally follows. Ending mid-thought ("And that's why you should always...") draws the viewer back to the beginning to hear it again -- but only if the beginning answers the dangling clause. Plan this explicitly.
- Semantic loop: The final piece of information should recontextualize the hook. When the viewer hears the hook again after watching the full video, it lands differently because they now know the answer. "Your dull knife is trying to hurt you" is more impactful the second time because the viewer now knows exactly how it does so.
- Result-as-loop: For tutorials and transformations, ending on the finished result -- which looks similar to the "promise" visual in the hook -- creates a loop where the viewer watches to verify they understand the process, then watches again to confirm the result.
- Loop pitfalls: Hard cuts to black, "Thanks for watching," and "Don't forget to like and subscribe" break the loop. These endings signal that the video is over and the viewer should leave -- never plan these endings.
Step 7: Specify the Visual and Production Plan
The visual plan does not need to reach storyboard detail (use video-storyboard for that) but must give enough direction to shoot or produce the video without ambiguity.
- Aspect ratio and safe zones: All short-form vertical video is 9:16 (1080x1920 pixels). The bottom 20% and top 10% of the frame are UI overlay zones on all three platforms -- text overlays, stickers, and critical visual information should never be placed in these zones. Place text in the middle 70% of the frame, favoring the center and upper-center areas.
- Subject placement: For talking-head content, place the face in the upper third of the frame, leaving the lower portion for text overlays and decorative elements. For tutorial/product content, center the subject/object in the middle third. For transformation content (before/after), left-right split or sequential cuts are both valid -- plan which you will use.
- Cut rhythm: The cut rhythm should match the audio. For music-driven content, cut on the beat or half-beat. For narrated content, cut at natural pause points in the narration -- cutting mid-sentence disorients the viewer. A general guide: educational content cuts every 3-8 seconds, entertainment content cuts every 0.5-3 seconds, storytelling can sustain 5-15 second shots if the visual is moving or expressive.
- On-screen text rules: Text must reinforce the audio, not duplicate it verbatim. If the narration says "A sharp knife controls where the blade goes," the text overlay should say "Sharp = you control it" -- the compressed version. Font size minimum is 48px at 1080px width (roughly 4.5% of frame height). Use one typeface with high contrast against the background (white text on a dark background or dark text with a white stroke).
- Lighting and framing guidance: Even, front-facing light avoids shadows on the face that read as low-production quality on small mobile screens. For product or object content, a simple plain background prevents visual distraction. The camera should be at eye level for talking-head content -- filming from below creates an unflattering angle that registers subconsciously as low authority.
- B-roll planning: For narrated educational content, plan one visual per main beat to illustrate the concept while the narration continues. B-roll should be filmed in 9:16 or cropped from landscape -- never plan letterboxed 16:9 inserts in a vertical video, as the black bars register as production errors to the viewer.
Step 8: Final Plan Validation
Before delivering the plan, run it through this checklist:
- Does the hook deliver its full impact in under 3 seconds? If reading it aloud takes longer, it is too long.
- Is there exactly one core concept? If the plan could generate a "part 1" and "part 2" naturally, the concept has not been compressed enough for this duration.
- Does the audio fill at least 80% of the video duration? Count the silence and compare.
- Is the caption's first line under 100 characters and self-contained?
- Does the final beat connect back to the first beat (visual, audio, or semantic loop)?
- Are all text overlays placed in the middle 70% of the frame?
- Is the hashtag count 3-5?
- Does the plan avoid any ending that signals the video is over (black screen, outro music, "thanks for watching")?
Output Format
## Short-Form Video Plan: [Video Title -- the working title, max 8 words]
**Platform:** [TikTok | Instagram Reels | YouTube Shorts | Cross-platform]
**Duration:** [7-15s | 15-30s | 30-45s | 46-60s -- specify exact target seconds]
**Content Type:** [educational | transformation | storytelling | tutorial | opinion | entertainment]
**Aspect Ratio:** 9:16 vertical (1080 x 1920px)
**Core Message:** [Single declarative sentence of 10 words or fewer]
---
### Hook (0:00 -- 0:02/0:03)
**Hook Type:** [pattern interrupt | direct address | curiosity gap | visual | stakes]
**Spoken Audio:** "[Exact words, or 'none' if visual-only]"
**On-Screen Text:** "[Exact text overlay, placement: center / upper-center / lower-center]"
**First Frame Visual:** [Describe what the viewer sees in frame 1 -- composition, subject, key visual element]
**Why This Hook Works:** [One sentence explaining the specific retention mechanic this hook uses]
---
### Beat Structure
| Timestamp | Beat Name | Duration | Spoken Audio | Visual | On-Screen Text | Retention Mechanic |
|-----------|-----------|----------|--------------|--------|----------------|-------------------|
| 0:00-0:02 | Hook | 2s | [audio] | [visual] | [text] | [pattern interrupt / curiosity / stakes] |
| 0:02-0:07 | [Beat 2] | Xs | [audio] | [visual] | [text] | [micro-hook / evidence / escalation] |
| 0:07-0:XX | [Beat 3] | Xs | [audio] | [visual] | [text] | [payoff / method / proof] |
| 0:XX-0:XX | [Beat 4] | Xs | [audio] | [visual] | [text] | [application / contrast / reveal] |
| 0:XX-0:XX | Loop Frame | 2-3s | [audio or silence] | [visual matching frame 1] | [text or none] | [loop trigger] |
---
### Audio Direction
**Audio Type:** [original narration | trending sound | music only | narration + music | voice-off + captions]
**Pacing:** [words per minute -- 130-160 for educational, 160-200 for entertainment]
**Music/Sound Description:** [specific tempo BPM range, genre, energy level, or "none"]
**Narration Notes:** [tone, authority level, emotional register -- calm/urgent/conversational/authoritative]
**Silence / Dead Air:** [where silence is intentional and how long -- maximum 0.5-1 second]
---
### Caption
**Above-Fold Line (max 100 characters):**
[First line text -- must work as a standalone statement]
**Body Text:**
[1-3 sentences of context, supporting fact, or related insight not covered in the video]
**Call to Action:**
[Single specific CTA -- save / follow / comment / share -- with exact phrasing]
**Hashtags (3-5):**
#[broad-category] #[niche-topic-1] #[niche-topic-2] #[audience-community] #[content-specific]
---
### Loop Engineering
**Visual Loop:** [How the last frame connects compositionally to the first frame]
**Audio Loop:** [How the ending audio leads naturally into the opening audio on repeat]
**Semantic Loop:** [How knowing the ending changes how the viewer hears the hook on re-watch]
**Loop Type:** [seamless / result-reveal / open-loop / semantic reframe]
---
### Visual and Production Notes
**Camera Setup:** [Position, angle, distance from subject]
**Lighting:** [Natural / ring light / softbox / ambient -- color temperature and direction]
**Subject Placement:** [Frame thirds placement -- face upper-third / object center / etc.]
**Safe Zone:** [Confirm all text and subject matter is in the middle 70% of frame]
**Cut Rhythm:** [Cut frequency and trigger -- every Xs / on beat / at narration pause points]
**B-Roll Needed:** [List each B-roll clip with description and duration]
**Text Graphics Needed:** [Count and exact text for each overlay]
**Props / Setup Items:** [List any physical items needed for filming]
**Estimated Takes:** [Realistic count based on content complexity]
Rules
-
One concept per video, enforced strictly. If the user's idea contains two distinct claims, two different methods, or requires two separate setups to explain, split it into two videos before planning. A video that tries to teach two things teaches neither -- viewers cannot retain two takeaways from a 30-second video, and the algorithm cannot categorize the content accurately for distribution.
-
The hook must be complete and impactful in 3 seconds or fewer. Measure this literally: speak the hook aloud at natural speaking pace and time it. If it takes more than 3 seconds, it is not a hook -- cut it. Common mistake: starting with a question that takes 5 seconds to ask. Cut the question to its sharpest form.
-
Audio must cover at least 80% of the video duration. Calculate: for a 30-second video, audio must run for at least 24 seconds. The only exception is a deliberate 1-second silence used as a dramatic device -- plan it explicitly and limit it to one instance per video. Unplanned silence reads as a production error and drops retention immediately.
-
Never plan content for landscape (16:9) or square (1:1) framing in a short-form video plan. Vertical 9:16 is not a preference -- it is the technical specification of every short-form platform. Horizontal content played in a short-form feed is displayed with pillar boxing (vertical black bars), which signals low-quality production to the viewer and reduces watch time.
-
On-screen text must never occupy the bottom 20% or top 10% of the frame. TikTok's UI places the like/comment/share column on the right side and username/caption overlay at the bottom. Instagram Reels has a similar UI. YouTube Shorts places the engagement buttons on the right. Text placed in these UI zones will be covered by interface elements on some devices, making the content unreadable for a portion of viewers.
-
Caption hashtags: 3-5 tags, no exceptions. Using more than 5 hashtags is not a neutral choice -- it actively dilutes distribution signals on TikTok and Instagram. Using zero hashtags removes the video from hashtag-based discovery feeds. Never include #fyp, #viral, or #trending -- these tags are used by billions of posts and provide zero discovery differentiation. One broad, two niche, one audience, one topic-specific is the proven formula.
-
Loop optimization is not optional. Every plan must specify all three loop dimensions: visual, audio, and semantic. A video plan without a loop design is an incomplete plan. The loop is a compounding return -- a 15-second video watched 3 times generates 45 seconds of watch time. The algorithm does not differentiate between one watch of 45 seconds and three watches of 15 seconds.
-
Edge Cases
Multi-Part Series Planning
When the user's concept is too large for a single video and a series is the right approach, each episode must be a complete standalone video that also functions as part of the series.
- The hook of every episode must work for a viewer who has never seen any other episode. Never design a hook that requires context from a previous video -- "As I mentioned in part 1" is a retention killer for new viewers and is the single most common series mistake.
- Add "Part [N] of [Total]" as a text overlay in the first 2 seconds, placed in the upper-center safe zone. This text should be small and secondary -- it signals to returning viewers that they are in the right place without dominating the hook for new viewers.
- End each episode with a loop that also functions as a cliffhanger or open question -- the loop satisfies the repeat-viewer behavior while the unanswered element drives series continuation. Example: a tutorial series episode ends on the finished step, which is also the raw starting material for the next step.
- Number the series concept before planning individual episodes: what is the complete arc? How many episodes? What does each episode deliver independently? What does the series deliver collectively? The answer to these questions shapes the beat structure of each episode.
Visual-Only Videos (No Narration, Music Only)
Some formats -- aesthetic content, fashion, design, food, nature -- communicate primarily through visuals and music with no spoken audio.
- Every narrative beat that would normally be carried by voice must instead be carried by a text overlay or a visual action. Map each beat in the structure table to its text-based communication equivalent.
- Text overlays must have a minimum dwell time of 2.0 seconds (this is the minimum reading time for a short phrase at comfortable pace for the average adult). For longer text, add 0.5 seconds per additional 5 words.
- Cut rhythm becomes the primary pacing tool -- plan cuts on the music beat (every half-beat for fast content, every full measure for slow content). Specify the BPM of the chosen music and calculate cut frequency: at 120 BPM with cuts on every beat, that is one cut per 0.5 seconds.
- The absence of voice means the video loses the authority and connection that a human voice provides. Compensate by ensuring the visual quality is exceptional -- this format has a higher visual quality threshold than talking-head content.
- On-screen text must carry the hook function. Plan the first text overlay to appear within the first 1.5 seconds, before the viewer has had time to scroll.
Platform-Exclusive Content (Single-Platform, Not Cross-Post)
When the user specifies a single platform rather than cross-platform, the plan must adapt to that platform's specific mechanics.
- TikTok-only: Audio discovery is the primary distribution mechanism. Trending sound selection matters more here than on any other platform. Comment engagement drives algorithm boost more on TikTok than saves do -- plan a comment-bait hook ("drop a [emoji] if this is you"). TikTok's UI places the username and caption in the lower-left and the engagement column on the right -- plan text overlays to avoid these zones specifically. TikTok also rewards watch-time loops most aggressively.
- Instagram Reels-only: Reels are surfaced in the Explore page, in the Reels tab, and in the main feed. Feed placement is important -- the caption above-fold line matters more here because feed viewers see the caption before engaging with the video. Instagram's algorithm weights saves and shares more heavily than comments relative to TikTok. Caption length can be more substantial (up to 300 characters above fold equivalent). Cover frame matters -- plan a specific cover frame that communicates the video's value as a static image.
- YouTube Shorts-only: Search-driven discovery means the title (which functions as the caption's above-fold line) must include a searchable keyword phrase. Shorts are surfaced in the Shorts feed, but also appear in regular YouTube search results -- this is a discovery advantage no other platform offers. Subscribers matter more here -- the plan should include a subscribe CTA, which is more natural on YouTube than on TikTok or Instagram. Chapter markers are not available for Shorts but the description can include keyword-rich context that aids search indexing.
Repurposing Existing Long-Form Content
When the user has a blog post, podcast episode, long YouTube video, or presentation and wants to adapt it to short-form, the process is compression, not extraction.
- Do not simply cut a 60-second clip from a longer video and call it a Reel -- this rarely works because long-form content is structured for a viewer who is already engaged, not for a viewer who needs to be earned. The hook, beats, and loop must be purpose-built for short-form.
- Extract the single most surprising, counterintuitive, or emotionally resonant moment from the long-form content. That moment is the basis for the short-form core message -- not the conclusion, not the introduction, not the summary.
- The long-form hook ("Today we're going to talk about...") must be replaced entirely with a short-form hook designed for the scroll feed.
- If the long-form content is audio (podcast), plan a video-first version where the audio clip is supported by visual proof, b-roll, or text overlays -- a podcast clip as-is cannot carry a short-form video because it was not designed for visual engagement.
Response Videos (Duet or Stitch Format)
Response videos (TikTok Duet, TikTok Stitch, or Instagram "responding to a comment" overlay) have a different structure because they begin with borrowed content.
- The borrowed clip (Stitch) typically runs 1-5 seconds and functions as the context setup -- but plan the response hook as if the viewer did not see the original video. The response must work for someone who watched the Stitch clip for the first time with no prior knowledge.
- In Stitch format, the borrowed clip ends and the creator's response begins -- plan the transition point explicitly. The first 1-2 seconds of the response portion is the response hook, which must be stronger than a typical hook because the viewer's attention has just been borrowed from another creator and must be re-captured.
- In Duet format, both videos play simultaneously side-by-side -- the original creator's video is on the right, and the responder's video is on the left. Plan the left-side content to react, comment, or demonstrate in sync with the right-side video. The Duet format reduces the available text overlay space because only half the screen is available.
- Always plan the response to be self-contained and valuable even without knowing the original creator's context. The response video will be surfaced to people who have never seen the original.
Promotional or Sponsored Content
When the video is paid or organic promotion for a product, service, or brand, the planning rules shift significantly.
- Platform disclosure requirements: TikTok, Instagram, and YouTube all require disclosure of paid partnerships. Plan for the disclosure label -- on TikTok and Instagram this is a platform-native label that appears at the top of the video; on YouTube Shorts it is disclosed in the description. The disclosure must not be the hook or the above-fold caption -- it is a required label, not a content beat.
- The 70/30 rule applies with strict enforcement: the first 70% of the video by timestamp must deliver genuine value with no product mention. The product is introduced in the final 30% as the natural solution to the problem established in the first 70%. Violating this rule -- opening with the brand name or product shot -- triggers the viewer's ad-avoidance response within 1-2 seconds and tanks completion rates.
- Authenticity of voice matters more in short-form than in any other ad format. Plan the script in the creator's natural voice and phrasing -- not in the brand's official language. Short-form audiences are expert at detecting "brand speak" inserted into a creator's content, and they disengage when they detect it.
- Plan a CTA that drives a specific action (link in bio, use a code, follow the brand account) -- not a vague brand awareness message. Short-form promotional content that does not drive a measurable action is difficult to justify from a campaign ROI perspective.
Reactive / Trend-Based Content with a Short Production Window
When the user needs to move fast because a trend is active now and will be irrelevant in 7-10 days:
- Prioritize the simplest possible production setup that still meets the minimum video quality threshold. A trending-format video filmed in 20 minutes on a phone with natural light will outperform a beautifully produced video that misses the trend window.
- Identify the trend's core mechanics first -- what structure, pacing, sound, and format define the trend? Most trends have 2-3 structural elements (a specific sound, a specific transition type, a specific text format). Map the user's message onto the minimum viable version of these elements.
- Plan for a single setup, single camera position, single take approach wherever possible. Complex production logistics kill trend responsiveness.
- Include a trend shelf-life estimate in the production notes: "This trend appears to be in its peak/rising/declining phase. If production takes more than [X days], consider whether the sound or format will still be algorithmically boosted at time of posting."
Example
Input: "Plan a 30-second TikTok and cross-post Instagram Reel about why most people's morning routines are backwards -- they start with their phone, which kills their focus for the rest of the day. Target audience: knowledge workers and remote workers who feel scattered and unproductive. Educational format."
Short-Form Video Plan: Your Morning Routine Is Backwards
Platform: TikTok (cross-post to Instagram Reels)
Duration: 30 seconds
Content Type: Educational
Aspect Ratio: 9:16 vertical (1080 x 1920px)
Core Message: Checking your phone first thing in the morning destroys deep focus for the entire day.
Hook (0:00 -- 0:02)
Hook Type: Curiosity gap + direct address hybrid
Spoken Audio: "The first thing you do every morning is ruining your focus."
On-Screen Text: "Your morning routine is backwards." (center frame, large, white text with dark stroke)
First Frame Visual: Extreme close-up of a phone screen lighting up with notifications -- no face yet, just the phone and the glow. The frame is slightly dim, morning-light color temperature.
Why This Hook Works: The statement makes a direct accusation about something the viewer does every single morning, creating both recognition ("that is me") and urgency ("I need to know why"). The curiosity gap -- "ruining your focus" without explaining how -- forces the viewer to stay to get the explanation.
Beat Structure
| Timestamp | Beat Name | Duration | Spoken Audio | Visual | On-Screen Text | Retention Mechanic |
|---|
| 0:00-0:02 | Hook | 2s | "The first thing you do every morning is ruining your focus." | ECU of phone screen lighting up with notifications, morning glow | "Your morning routine is backwards." | Direct accusation + curiosity gap |
| 0:02-0:08 | The Problem Mechanism | 6s | "When you check your phone first, your brain enters reactive mode -- you start responding to other people's priorities before you have set your own." | Creator on camera, eye level, plain background -- slight lean forward to signal importance | "Reactive mode = scattered all day" | Counterintuitive mechanism that explains a felt experience |
| 0:08-0:16 | The Evidence | 8s | "Research on cognitive priming shows that the first task you do in the morning sets your attentional state for the next 2 to 3 hours. Phone checking is the highest-distraction task possible. You are priming your brain for distraction." | B-roll: overhead shot of a hand reaching for a phone, then the same hand stopping and pulling back -- a deliberate visual contrast | "First task = 2-3 hrs of cognitive tone" | Specific number + scientific framing creates authority |
| 0:16-0:23 | The Fix | 7s | "The fix takes 10 minutes. Before you touch your phone: drink water, write the one thing you need to accomplish today, and do 5 minutes of any physical movement. Now you have set the attentional tone, not your inbox." | Creator demonstrating in real-time: picking up a glass of water, writing in a small notebook, doing 5 quick jumping jacks -- all in the same shot, fast cuts | "1. Water 2. One priority 3. 5 min movement" | Actionable simplicity -- 3 steps, concrete and achievable |
| 0:23-0:28 | Result Statement | 5s | "Your focus does not get destroyed before you have even started. You go into your first task already in the mode you want to be in." | Creator at desk, calm and focused, typing with purpose -- contrast to the distracted phone-checking at the start | "Focus before inbox." | Vivid contrast between the before state (implied) and the after state (shown) |
| 0:28-0:30 | Loop Frame | 2s | [Sound of a single phone notification -- one clear chime, then silence] | ECU of phone screen lighting up again -- the same shot as the opening frame, same notification glow |
Audio Direction
Audio Type: Original narration + minimal ambient sound
Pacing: 145 words per minute (calm, educational, authoritative -- not rushed)
Music/Sound Description: No background music during narration. The ambient sounds of the demonstration (water pouring, pen on paper, light movement) provide audio texture. The single phone chime at the end is the only "produced" sound effect.
Narration Notes: Measured and direct -- the tone of someone sharing a useful finding with a colleague, not lecturing. No upspeak, no vocal fry, no filler words ("um," "like," "you know"). The word "ruining" in the hook should be delivered with slight emphasis -- it is the emotional spike word.
Silence / Dead Air: One intentional 0.5-second pause between "your own" (end of Beat 2) and "Research on cognitive priming" (start of Beat 3) -- this pause makes the upcoming stat feel deliberate and important. No other silence.
Caption
Above-Fold Line (94 characters):
Stop checking your phone the second you wake up. Here's the science of what it costs you.
Body Text:
Cognitive priming is real -- the first 10-15 minutes of your mental state each morning compounds into your focus capacity for hours. Phone checking is reactive by design. You are training your brain to expect interruption before you have done a single thing for yourself.
Three swaps that take under 10 minutes: water first, write your one priority, move your body. That is it.
Call to Action:
Save this and try the three swaps tomorrow morning -- then tell me what happened.
Hashtags (5):
#productivity #morningroutine #deepwork #remotework #focushacks
Hashtag logic: #productivity (broad, 2B+ posts), #morningroutine (niche, 180M posts), #deepwork (niche, 12M posts), #remotework (audience-specific community, 450M posts), #focushacks (content-specific, 8M posts)
Loop Engineering
Visual Loop: The final frame is an identical shot to the opening frame -- the same extreme close-up of a phone screen lighting up with notifications, same color temperature, same composition. The viewer arriving back at the beginning sees the exact image that started the video, which now carries full meaning.
Audio Loop: The closing sound is a single phone notification chime, followed by 0.5 seconds of silence. This silence flows directly into the opening line "The first thing you do every morning is ruining your focus" -- the chime is the implied cause, and the opening accusation is the implied effect. The audio sequence feels continuous on repeat.
Semantic Loop: On first watch, "The first thing you do every morning is ruining your focus" is an accusation that creates anxiety. On second watch (after seeing the full explanation and the fix), the same line is heard as a problem with a known solution -- it shifts from threatening to motivating. This semantic reframe rewards re-watching because the viewer's emotional response to the hook changes with knowledge.
Loop Type: Semantic reframe + visual mirror
Visual and Production Notes
Camera Setup: Creator on camera for Beats 2 and 5 -- eye level, approximately 2-3 feet from the lens, slight lean forward during Beat 2. Phone close-up for Hook and Loop Frame -- mounted overhead or side-angle showing the screen clearly. B-roll (Beat 3) shot overhead.
Lighting: Soft, warm natural or ring light for creator shots. The phone scenes use the natural screen glow for atmosphere -- keep ambient light dim to make the screen glow visible and slightly eerie.
Subject Placement: Creator's face occupies the upper two-thirds of the frame during talking-head beats. Text overlays in center frame, no lower than 75% of frame height (to avoid caption UI overlap on TikTok).
Safe Zone: All text overlays confirmed within the middle 70% vertical band. No text within 15% of left or right edges (avoids TikTok engagement column on the right).
Cut Rhythm: Educational pacing -- cut at each narration pause point. Beats 1, 4, and 6 are the fastest-cut beats (visual interest without narration or with demonstration). Beats 2, 3, and 5 are single-shot or two-shot beats that hold longer for comprehension.
B-Roll Needed:
- Overhead hand reaching for phone, then stopping (Beat 3) -- 4 seconds
- Glass of water being picked up and drunk (Beat 4) -- 1.5 seconds
- Hand writing in a small notebook, single word visible (Beat 4) -- 1.5 seconds
- 5 jumping jacks in frame, fast and energetic (Beat 4) -- 2 seconds
Text Graphics Needed (5 overlays):
- "Your morning routine is backwards." -- Hook, large, center frame
- "Reactive mode = scattered all day" -- Beat 2, medium, center frame
- "First task = 2-3 hrs of cognitive tone" -- Beat 3, medium, center frame
- "1. Water 2. One priority 3. 5 min movement" -- Beat 4, list format, center frame
- "Focus before inbox." -- Beat 5, large and clean, center frame
Props / Setup Items: Smartphone with notification-heavy lock screen, glass of water, small notebook, pen
Estimated Takes: 3-4 takes for narration (to land the pacing and emphasis on "ruining"), 2-3 takes for demonstration beats (water/writing/movement must look natural, not staged), 1-2 takes for phone close-up (composition and light balance)