Writes a complete video script with hook, narrative structure, B-roll cues,
call-to-action placement, and speaker direction annotations. Use when the
user asks to write a video script, plan video content, create a YouTube
script, or draft narration for a video.
Do NOT use for storyboarding with shot lists (use video-storyboard),
voiceover-only scripts (use voiceover-script), or short-form video
under 60 seconds (use short-form-video-planning).
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Writes a complete video script with hook, narrative structure, B-roll cues,
call-to-action placement, and speaker direction annotations. Use when the
user asks to write a video script, plan video content, create a YouTube
script, or draft narration for a video.
Do NOT use for storyboarding with shot lists (use video-storyboard),
voiceover-only scripts (use voiceover-script), or short-form video
under 60 seconds (use short-form-video-planning).
Asks to write a video script for YouTube, a course platform, a brand channel, or an explainer video with both on-camera and supporting visual elements
Wants to structure a talking-head, narrated, or mixed-format video that runs between 60 seconds and 60 minutes
Needs a complete script with spoken dialogue, B-roll cues, speaker direction annotations, and CTA placement
Asks to plan a video that tells a story, teaches a concept, makes an argument, or promotes a product or service
Has a video topic but needs help determining narrative structure, pacing, hook strategy, or section organization
Is preparing a presenter for filming and needs director-level annotations alongside the dialogue
Wants to adapt existing written content (a blog post, a slide deck, a white paper) into a spoken video format
Do NOT use this skill when:
The user needs a shot-by-shot visual storyboard with frame composition and camera angles -- use video-storyboard
The output is voiceover-only with no on-camera presenter or talking-head element -- use voiceover-script
The video is under 60 seconds (Reels, Shorts, TikTok) -- use short-form-video-planning
The user needs a podcast script or audio-only format with no visual cues -- use a dialogue or interview script approach
The user needs post-production editing instructions, color grading notes, or timeline assembly -- this skill ends at pre-production
The user is looking for a video marketing strategy or channel planning -- address content strategy first before scripting
Process
Step 1 -- Collect the Full Video Brief Before Writing a Single Word
Never begin drafting until all of the following inputs are confirmed. Missing any one of them produces a script that will need to be entirely rewritten.
Topic and angle: A topic alone is not enough. "Email marketing" is a topic. "Why your open rates dropped after the iOS 15 update and how to fix them" is an angle. Push for the angle.
Target audience: Capture three dimensions -- who they are, what they already know about the topic (beginner/intermediate/expert), and what emotional state they are in when they press play (curious, frustrated, skeptical, motivated). These three dimensions control vocabulary, pacing, and tone.
Target duration: Ask for this in minutes. The word count formula is 130 words per minute for deliberate, educational delivery and 160 words per minute for energetic, conversational delivery. Use 150 wpm as the default midpoint. A 5-minute video targets 750 words of spoken dialogue (excluding annotation markers).
Video type: Classify into one of five types -- educational/explainer, tutorial/how-to, opinion/argument, promotional/brand, or storytelling/case study. The type determines the narrative structure chosen in Step 3.
Platform and publishing context: YouTube search-driven content needs a search-optimized title and a strong retention hook because the viewer chose the video. Course platform content can assume a warmer audience and skip the credibility-building step. Website explainers play to cold audiences and need faster value delivery.
Key message: The single sentence the viewer must leave with. If the user cannot state this in one sentence, facilitate it: "If your viewer remembers only one thing 30 days from now, what is it?" Do not proceed without this.
Desired outcome: What should the viewer DO after watching? Subscribe, purchase, book a call, share, change a behavior? The outcome determines CTA type and placement.
Existing assets: Does the user have a draft outline, a blog post, a deck, or research they want incorporated? Absorb these before structuring the script.
Step 2 -- Design the Hook Using the 15-Second Window Framework
The first 15 seconds of a video determine whether 60-70% of viewers continue watching, based on platform retention data. The hook is not an introduction -- it is a contract with the viewer that the next several minutes will be worth their time.
Hook architecture has three required components:
The Grab (0-5 seconds): The first thing the viewer hears and sees. Must create an open loop, trigger recognition of a pain point, or deliver a surprising claim. Exactly 15-25 words. No preamble, no "Hey guys, welcome back."
The Expand (5-10 seconds): One or two sentences that reinforce the grab and make the viewer feel seen. This is where you confirm the viewer is in the right place.
The Promise (10-15 seconds): An explicit commitment about what the video delivers. Be specific and falsifiable. "I'll show you three things" beats "We'll talk about some strategies."
The six proven hook formulas:
Counterintuitive claim hook: Challenge the dominant belief in the first sentence. "The reason your productivity system isn't working is that you're trying to manage time. That's the wrong problem entirely." Works best for opinion, educational, and argument-style videos.
Result-first hook: Open with the end state. "In this video, you're going to see how I went from 500 to 50,000 email subscribers in 14 months -- without paid ads." Works best for case studies and journey-format videos.
Problem-recognition hook: Describe the viewer's exact situation so precisely that they feel understood. Specificity is the key: "You've got 47 browser tabs open, a to-do list that grows faster than you cross things off, and you can't figure out why you're always busy but rarely done." Works for any format.
Question hook: Ask a question the viewer cannot immediately answer but wants to know. The question must create genuine curiosity, not a rhetorical formality. "Do you know what the most-deleted type of email is -- and how often you're sending it?"
Statistic shock hook: Lead with a number that reframes the viewer's understanding. The statistic must be surprising AND directly relevant. It cannot be a figure they already know. "The average employee loses 2.5 hours every day to interruptions they didn't initiate."
Story hook: Drop the viewer into a scene mid-action. "It was 11 PM on a Tuesday when I realized the campaign I'd just launched had a critical error -- and I had no way to stop it."
Hook mistakes to avoid:
Thanking the viewer, complimenting them, or mentioning their likes/comments before the hook is complete
Spending more than 3 seconds on a logo animation, intro bumper, or channel identification before the hook begins
Asking questions that are so easy the viewer does not feel compelled to find the answer
Promising something vague like "a lot of great tips" -- specificity holds attention
Step 3 -- Select and Map the Narrative Structure
Match the structure to the video type. Mismatching a structure to a type is the most common reason a script feels disorganized even when all the information is correct.
Problem-Agitate-Solve (PAS): The workhorse structure for educational and promotional videos. Name the problem (1-2 sentences). Agitate it by showing the consequences of NOT solving it -- this is where emotion enters (2-3 sentences). Then present the solution. The agitation step is what most amateur scripts omit, resulting in content that feels flat.
Journey (Before-Transformation-After): For storytelling and case study videos. Establish the "before" state in enough detail that the audience connects with it. Walk through the specific transformation process, including failures and setbacks (these build credibility). Arrive at the "after" state and make it tangible and measurable.
Concept-Evidence-Application: For deep educational content and thought leadership. State the core concept. Provide evidence that the concept is true (research, examples, case studies). Then shift to application -- how the viewer specifically acts on this concept. End sections with an explicit "So, what you do with this is..." transition.
Sequential how-to (Numbered steps): For tutorial and process videos. Number every step. Give each step a short, memorable name (not Step 1, Step 2 -- name it "The Inventory Step," "The Batch Step"). Include a brief rationale for why each step is in this order. Viewers who understand the why are more likely to follow through.
Comparison-and-Verdict: For review and recommendation videos. Establish evaluation criteria upfront. Apply the same criteria to each option. Give a clear verdict -- avoid "it depends on your situation" as a final conclusion, as it destroys retention and frustrates viewers.
Layered list (McKinsey structure): For complex multi-point content. Group points into categories (never more than three categories, never more than four points per category). Announce the structure at the start ("There are three reasons for this. Two are obvious. One will surprise you.").
After selecting the structure:
Create a section outline with a title for each section
Assign a word count target to each section based on its importance and the total target word count
Identify which sections need B-roll coverage and which can sustain on-camera talking head
Step 4 -- Write the Script Body with Full Annotations
Spoken dialogue rules:
Sentence length: 25 words maximum per sentence. Average target: 15-18 words. Short declarative sentences create punch. Long sentences require re-reading -- something viewers cannot do.
Paragraph length: 3-5 sentences maximum before a B-roll break, direction change, or visual element disrupts the pattern
Vocabulary: Match the audience's expertise level exactly. For beginners, define every piece of jargon the first time it appears. For experts, use the jargon without apology -- over-explaining insults them.
Contractions are mandatory: "You are" becomes "you're." "It is" becomes "it's." Scripts written in formal English sound unnatural when spoken and cause presenters to stumble.
Direct address: Use "you" and "your" throughout. Not "viewers," not "people," not "one." Always "you."
Transition phrases between sections: Make the transition a dual function -- it closes the previous section AND previews the next. "Now that you know why attention management beats time management, let's look at the three specific changes that make the biggest difference."
Annotation types and proper usage:
[B-ROLL: description] -- Place this marker where the camera should cut away from the on-camera presenter. The B-roll description should be specific enough that a producer can source or film it. "B-ROLL: stock footage" is useless. "B-ROLL: close-up of hands typing on a mechanical keyboard in a dimly lit office" is actionable. B-roll requirement: at minimum one cut every 60 seconds of dialogue, ideally every 30-45 seconds for YouTube content. B-roll sustains visual attention while the speaker continues speaking.
[TEXT ON SCREEN: exact text] -- Used for statistics, definitions, key terms, numbered points, and any information the viewer needs to read and retain. Write the exact text that will appear on screen, not a description of it. Keep on-screen text to 8 words or fewer per graphic.
[DIRECTION: instruction] -- Speaker coaching for the presenter. Include tone ("shift to a more urgent tone here"), physical cues ("lean forward, lower voice slightly"), gesture prompts ("hold up fingers to count the points"), and pace notes ("pause for 2 seconds after this sentence -- let it land"). Minimum 3 direction annotations per 5 minutes of script. Direction annotations save post-production re-shoots.
[PAUSE] -- Used after a key insight, a statistic, or a question posed to the viewer. A 1-2 second pause on camera is one of the most powerful rhetorical tools available and is chronically underused in scripts.
[SCREEN: show X] -- For tutorials where the speaker is demonstrating software or a physical process. Different from B-ROLL in that the screen recording or product shot is the primary visual, not a cutaway.
Engagement engineering -- embed these devices into the body:
Pattern interrupts: Every 2-3 minutes, shift something -- tone, visual style, delivery speed, or framing. Viewers' attention has a physiological refresh cycle of roughly 3 minutes; a pattern interrupt resets it.
Open loops: Plant a question or a "coming up" tease that is not immediately answered. "I'll come back to this in a moment because it connects to the third point in a way that surprised even me." This drives retention because unresolved curiosity keeps viewers watching.
Social proof integration: Where a claim is made, follow it with a concrete example, a case study reference, or a data point. Claims without evidence reduce trust; evidence without claims reduces clarity.
Viewer identification moments: Periodically confirm the viewer is in the right experience. "If you've been in this situation, this next part is specifically for you." This reduces drop-off at the 30-50% mark.
Step 5 -- Design and Place the Call-to-Action
CTAs are systematically misplaced in most video scripts. The end of the video is the worst time for a primary CTA because 40-60% of viewers have already left before a video finishes (this varies by platform and content type, but the pattern is consistent).
CTA placement by video duration:
Videos under 3 minutes: Single CTA at the 75-80% mark. No outro CTA -- the video is too short.
Videos 3-10 minutes: Primary CTA at the 70% mark. Secondary CTA in the outro. A mid-roll CTA mention at the 30-40% mark is optional for subscription drives.
Videos over 10 minutes: Three CTA touchpoints -- a soft mention at 25-30% ("by the way, if you want more of this, subscribe"), a primary CTA at 60-70%, and a secondary CTA in the outro.
CTA copy principles:
Always give the viewer a reason to act, not just an instruction to act. "Subscribe" is an instruction. "Subscribe so you get the follow-up video where I show you the exact email template that generated 400 replies -- dropping next week" is a reason.
Make the next action feel effortless and logical. The CTA should feel like the natural next step in the viewer's journey, not an interruption.
Prompt specific comments, not generic ones. "Leave a comment" gets nothing. "Comment and tell me: what's the single biggest attention killer in your workday?" gets hundreds of responses because the viewer has a specific, answerable question.
For course and brand videos, the CTA destination must match the viewer's stage of awareness. A cold audience seeing a brand video for the first time should be sent to a free resource, not a sales page.
Step 6 -- Write the Outro
The outro serves four functions: close the loop opened in the hook, deliver the key message one final time in different words, bridge to the next content, and sign off. It should run no longer than 30-45 seconds (75 words maximum).
Loop closure: Reference something from the hook. If you opened with a question, answer it explicitly in the outro. "So back to where we started -- why do most time management systems fail? Because they're solving the wrong problem."
Key message restatement: Say the key message in a new way. Do not repeat the exact words from earlier -- paraphrase it so it lands as a conclusion, not a repetition.
Content bridge: Tease the next video, related video, or next step in the series with a specific detail. Vague teasers ("I've got more great content") create zero urgency. Specific teasers ("Next week I'm going through the exact 90-minute work block setup that eliminated my Sunday anxiety") create anticipation.
Sign-off: Keep it brief and consistent. The sign-off is a brand signal, not additional content. Two to three words maximum: "See you next week," "Until then," "Take care."
Step 7 -- Calculate Timing and Complete Production Notes
Word count to time: Divide total spoken word count by the delivery rate (130 wpm for slow/educational, 150 wpm for standard, 160 wpm for energetic/conversational). This gives base spoken time.
Add pauses: Every [PAUSE] annotation adds 1-2 seconds. Every section transition adds 1-2 seconds. B-roll moments with no simultaneous speech add their own duration.
Adjustment factor: Add 12-15% to the base spoken time to account for natural delivery breathing, unscripted emphasis, B-roll duration that exceeds dialogue, and minor on-set variations.
Recording time estimate: Multiply the adjusted script duration by 2.5 for a first-time presenter, 1.75 for an experienced presenter. This accounts for retakes, resets, and water breaks.
B-roll inventory: Count every [B-ROLL:] annotation. Group them by type (footage to source vs. footage to film) to support production planning.
Graphics inventory: Count every [TEXT ON SCREEN:] annotation. Each represents one motion graphics element needed in post-production.
Output Format
Use this exact format for every script output. Customize section titles, section count, and timing based on the specific video.
## Video Script: [Full Video Title]
**Duration:** [X:XX estimated]
**Word Count:** [spoken words only, excluding annotations]
**Delivery Rate:** [130 | 150 | 160 wpm -- state which and why]
**Type:** [educational | tutorial | opinion | promotional | storytelling]
**Platform:** [YouTube | course platform | website | brand channel]
**Key Message:** [One sentence -- the single thing the viewer must remember]
**Primary CTA Goal:** [subscribe | purchase | book call | share | behavior change]
---
### HOOK (0:00 -- 0:15)
[DIRECTION: tone and physical setup instruction]
**THE GRAB (0:00 -- 0:05)**
"[15-25 words. The opening line. No preamble.]"
[B-ROLL: specific visual description for the grab moment]
**THE EXPAND (0:05 -- 0:10)**
"[Reinforce the grab. Make the viewer feel seen. 1-2 sentences.]"
**THE PROMISE (0:10 -- 0:15)**
"[Specific commitment about what this video delivers. Include a number if possible.]"
---
### SECTION 1: [Section Title] (0:15 -- X:XX)
**Estimated word count: [X words | X:XX at delivery rate]**
[DIRECTION: tone, posture, and energy instruction for this section]
"[Spoken dialogue. Sentences under 25 words. Contractions throughout. Direct address.]"
[PAUSE]
"[Continue dialogue.]"
[B-ROLL: specific description -- what, where, framing detail]
"[Dialogue continues during or after B-roll.]"
[TEXT ON SCREEN: exact text, 8 words or fewer]
"[Dialogue that references or extends the on-screen text.]"
**Section transition:** "[Closing sentence for this section + bridge to next section.]"
---
### SECTION 2: [Section Title] (X:XX -- X:XX)
**Estimated word count: [X words | X:XX at delivery rate]**
[DIRECTION: shift instruction -- note the change from previous section]
"[Spoken dialogue.]"
[B-ROLL: specific description]
"[Spoken dialogue.]"
[TEXT ON SCREEN: exact text]
[PAUSE]
"[Spoken dialogue.]"
**Section transition:** "[Bridge to next section.]"
---
### [Continue sections as needed following the same pattern]
---
### SOFT CTA MENTION (at ~30% mark for videos over 10 minutes)
[DIRECTION: casual, not salesy -- this is a brief aside]
"[1-2 sentence soft CTA. Give a specific reason.]"
---
### PRIMARY CTA (at 70% mark)
[DIRECTION: direct, warm, genuine -- not performative urgency]
"[Instruction + specific reason to act. 2-3 sentences maximum.]"
"[Comment prompt with a specific answerable question.]"
---
### SECTION [Final content section]: [Title] (X:XX -- X:XX)
**Estimated word count: [X words | X:XX at delivery rate]**
[DIRECTION: begin winding down energy -- reflective, conclusive tone]
"[Final content. This section ties back to the hook and resolves any open loops.]"
[B-ROLL: closing visual -- often a callback to the opening visual]
---
### OUTRO (X:XX -- end)
**Target length: 30-45 seconds max**
[DIRECTION: warm, direct, unhurried]
"[Loop closure -- reference the opening hook explicitly.]"
[PAUSE]
"[Key message restatement in new words.]"
"[Content bridge -- specific tease of next video or related content.]"
"[Secondary CTA -- one sentence.]"
"[Sign-off -- 2-3 words.]"
---
### Production Notes
| Element | Count | Notes |
|---|---|---|
| B-roll shots required | [X] | [Split: X to source, X to film] |
| On-screen text graphics | [X] | [Specify any that need animation] |
| Direction annotations | [X] | [Flag any requiring special setup] |
| Pause markers | [X] | [Total estimated pause time: X seconds] |
| Open loops planted | [X] | [List them and where they resolve] |
**Total spoken word count:** [X]
**Base spoken duration:** [X:XX at X wpm]
**Adjusted duration (+ 13%):** [X:XX]
**Estimated recording time -- experienced presenter:** [duration x 1.75]
**Estimated recording time -- new presenter:** [duration x 2.5]
**Primary CTA placement:** [X:XX -- confirm this falls at 70% of adjusted duration]
Rules
Never begin a video script with a greeting, channel intro, or acknowledgment of the audience. "Hey guys," "Welcome back," and "Thanks for clicking" are retention killers. The first words of the script must be the hook -- specifically the Grab element. If the client insists on a channel intro bumper, note it as a production element separate from the script, placed before the hook begins.
The hook must be completable in 15 seconds, not 5. The outdated "5-second hook" standard came from skippable ad formats. For standard video content, the hook has 15 seconds total (Grab, Expand, Promise). Within that 15 seconds, every sentence must earn its place. Delete any sentence that does not either create curiosity, build recognition, or make the promise more specific.
Never write a sentence longer than 25 words in spoken dialogue. Count the words. If a sentence exceeds 25 words, split it. Long sentences cause presenter stumbles, reduce viewer comprehension, and are harder to re-take on set. The 25-word rule is non-negotiable even when the content is complex -- complexity is handled through structure, not sentence length.
Every B-roll cue must be specific enough to brief a producer. Reject any B-roll annotation that uses generic terms alone: "stock footage," "relevant visuals," or "supporting images." Every B-roll cue must specify a subject, a context or setting, and ideally a framing or perspective detail. This prevents the production team from defaulting to irrelevant generic stock footage.
Place the primary CTA at no later than the 75% mark of total video duration. Calculate 75% of the adjusted duration and verify the CTA timestamp falls within 5 seconds of that mark. An outro-only CTA is a structural failure -- a significant portion of the audience will have left before they hear it.
The key message must appear explicitly in three places: the hook promise, the body conclusion, and the outro. Not paraphrased loosely -- the same core idea, restated in different language each time. Repetition with variation is the primary mechanism for viewer retention of the central point.
Transitions between sections must perform two jobs simultaneously. A transition phrase must close the preceding section (signal that it is complete) AND preview the next section (give a reason to keep watching). Single-function transitions like "Moving on..." are insufficient. Write the transition dialogue explicitly in the script -- do not leave it to the presenter to improvise.
Do not include apologies, qualifiers, or hedging language in the script. Phrases like "I might be wrong but," "This is just my opinion," and "You might feel differently" erode authority and trust. If a claim is uncertain, frame it as a specific condition: "If your audience is primarily mobile readers, this changes things significantly." Own the perspective -- hedging reads as incompetence.
Include at least one pattern interrupt per 3 minutes of script. A pattern interrupt is any deliberate shift in the viewer's experience: a change in delivery speed called out in a [DIRECTION:] annotation, a shift from monologue to a rhetorical question, a B-roll sequence that takes the viewer somewhere new, or a tone shift from analytical to story. Document the pattern interrupt type in the direction annotation.
The script stops at pre-production. Do not include editing cut points, color grade notes, music bed instructions, timeline markers, export settings, thumbnail recommendations, or SEO keyword placement notes. These belong in separate production and optimization workflows. The script covers what is spoken, what is shown visually as a cue, and how the presenter delivers it -- nothing beyond that scope.
Edge Cases
Multiple Hosts or Interview Format
When two or more people appear on camera, label every line of dialogue with the speaker's name in capital letters followed by a colon: HOST:, GUEST:, INTERVIEWER:, ALEX:. Include [CUT TO: speaker name] markers before each speaker's label whenever the camera angle switches. Script the interviewer's questions in full even if the interview will be conducted more loosely -- the scripted version is the briefing document that ensures the conversation covers the required territory. If the interview will use a run-and-gun style with no scripted dialogue, write a question framework document instead, not a full script, and note this distinction for the user.
Tutorial or Screen Recording Format with No On-Camera Presenter
Replace all [DIRECTION:] annotations with [PRESENTER ACTION:] or [SCREEN:] annotations. Maintain a strict separation between what the narrator says (spoken dialogue in quotation marks) and what happens on screen ([SCREEN: cursor moves to the Settings menu], [CLICK: Export button], [TYPE: yourfilename.csv]). Write the narration so that it leads the action slightly -- the spoken cue should come 1-2 seconds before the on-screen action, not simultaneously with it, because viewer comprehension requires hearing the instruction before seeing it executed.
Repurposing a Blog Post or Existing Written Content
Do not paste written content into a script and add annotations. Written content must be fully rewritten for spoken delivery. The conversion process: (1) identify the key claim of each paragraph, (2) discard the rest of the paragraph and rewrite the claim in spoken language, (3) identify which points need visual support and write those as B-roll cues. A 1,500-word blog post typically contains enough substance for a 6-8 minute video after the rewrite, but the spoken word count will be lower than the original because written language is denser than spoken language.
Highly Technical or Jargon-Heavy Topics
For expert-level audiences, jargon is not the problem -- mismatched jargon is. Define terms only when introducing concepts the audience may not have encountered, not when using established terminology in the field. For mixed-expertise audiences, use a layered approach: introduce the technical term, follow it immediately with a plain-language anchor ("That's the process of batch rendering -- essentially instructing the software to process multiple outputs simultaneously without manual intervention"), then use the technical term freely for the rest of the script. Insert a single [TEXT ON SCREEN:] for each new technical term defined this way.
Video Over 15 Minutes (Deep Dive, Documentary, or Course Module)
Add mid-video re-engagement markers at the 5-minute and 10-minute marks (and every 5 minutes thereafter for very long content). Each re-engagement marker is a brief 2-3 sentence device that: reminds the viewer what they have already gained, teases what is coming, and validates their investment in continuing to watch. Example: "You've now got the foundational piece -- and it's going to make the next section click much faster. What we're about to cover is the part most people skip, which is exactly why most people get stuck." Also add a chapter structure in the production notes with proposed timestamps for YouTube chapter markers.
Promotional or Sales-Driven Videos with a Skeptical Audience
Cold or skeptical audiences require a credibility architecture before any claim is made. In the body of the script, sequence every major claim in this order: claim, specific evidence, concrete example, implication for the viewer. Never lead with a promotional statement -- the promotional intent should emerge from the problem-solution structure, not be stated directly. The CTA in a sales-context video should offer a low-commitment next step (a free resource, a trial, a demonstration) rather than an immediate purchase, unless the video is explicitly a bottom-of-funnel asset for a warm audience.
No Clear Platform or Audience Specified
If the user cannot specify either the platform or the target audience, do not guess. Ask for the platform first because it constrains the format -- YouTube educational content has different structural requirements from a course platform module, which differs from a brand website explainer. Once the platform is established, ask for audience specifics. Building a script on assumed audience knowledge levels produces dialogue that either condescends to experienced viewers or loses beginners in the first two minutes.
Emotional or Sensitive Topics
For scripts covering mental health, grief, financial hardship, trauma, or similar sensitive subject areas, the narrative structure and the language within it require additional care. Avoid the Problem-Agitate-Solve structure's agitation phase -- deepening emotional pain as a rhetorical technique is inappropriate and potentially harmful for viewers in distress. Use the Concept-Evidence-Application structure instead, which leads with understanding rather than agitation. Replace urgency-based CTA language with supportive, low-pressure alternatives. Add a [DIRECTION: speak slowly, maintain steady eye contact, pause generously] annotation that runs across the full script rather than only at selected moments.
Example
Input: "Write a 7-minute YouTube video script about how to price freelance projects. My audience is freelancers who have been working for 1-2 years and keep undercharging. They know basic freelancing but feel anxious about pricing conversations."
Output:
Video Script: Stop Undercharging -- A Freelancer's Guide to Pricing Projects Correctly
Duration: 7:15 estimated
Word Count: 1,020 spoken words
Delivery Rate: 145 wpm (conversational but measured -- audience is anxious, pacing should feel steady and trustworthy, not rushed)
Type: educational / opinion
Platform: YouTube
Key Message: Freelancers undercharge because they price their time instead of their value -- and one mental shift fixes this permanently.
Primary CTA Goal: subscribe + comment engagement
HOOK (0:00 -- 0:15)
[DIRECTION: Look directly into the camera. Calm, knowing expression -- not aggressive. You've been in their shoes and you understand.]
THE GRAB (0:00 -- 0:05)
"You just sent a quote, and now you're wondering if you charged too much. That feeling means you're probably not charging enough."
[B-ROLL: Close-up of a hand hovering over a laptop trackpad, mouse cursor sitting on a "Send" button in an email window]
THE EXPAND (0:05 -- 0:10)
"That anxiety before you hit send? It's one of the most common experiences in freelancing. And almost every time it shows up, it's pointing in the wrong direction."
THE PROMISE (0:10 -- 0:15)
"In the next seven minutes, I'm going to show you the exact mental shift that changes how you price -- and walk you through a three-step framework you can use on your next project today."
[DIRECTION: Empathetic, unhurried. Nod slightly as you describe the patterns -- signal that you recognize these without judgment.]
"The undercharging problem almost never comes from a lack of skill. It comes from a specific type of math error. Most freelancers price their time."
[PAUSE]
"You figure out how many hours a project will take, multiply it by what feels like a reasonable hourly rate, and you land on a number. And then you knock 20% off it because you're worried the client won't say yes. Sound familiar?"
[B-ROLL: Split-screen -- left side shows a handwritten calculation on a notepad with an hourly rate multiplied by hours; right side shows the final number crossed out and a lower number written beneath it]
"Here's the problem with pricing your time. Clients don't actually care about your time. They care about what your work produces for them."
[TEXT ON SCREEN: Clients pay for outcomes, not hours]
"A logo that took you 3 hours is worth exactly the same as a logo that took you 30 hours -- to the business that needs it. Maybe more, if you've gotten faster because you're better."
"When you price your time, you're accidentally punishing yourself for getting good at what you do."
[PAUSE]
Section transition: "So if time-based pricing is the problem, let's talk about what to use instead -- and it starts with a question most freelancers never ask."
SECTION 2: The Value Question -- How to Find What a Project Is Actually Worth (1:45 -- 3:30)
Estimated word count: 235 words | 1:37 at 145 wpm
[DIRECTION: Shift slightly forward. More focused energy here -- this is the core insight. Slow down for the key question.]
"Before you quote a project, there's one question you need to answer. And most freelancers skip it entirely because it feels awkward to ask. The question is: what is this project worth to the client's business?"
[PAUSE]
"Not what's your budget. What is this worth."
"If you're writing email copy for an e-commerce brand and their email list generates $80,000 a month, a 10% improvement in their open rates is worth $8,000 a month to them. Every month. Permanently."
[B-ROLL: A whiteboard with a simple calculation -- "$80,000 x 10% = $8,000/month" written clearly, then "$8,000 x 12 months = $96,000/year"]
"Now tell me -- does $1,500 for that copywriting project feel like the right number? Or does it feel like you're leaving the client with a $94,500 gift?"
[TEXT ON SCREEN: Your fee as % of the value you create]
"Value-based pricing means your rate is anchored to the outcome you're helping the client achieve, not the hours you spend achieving it. You price yourself at a fraction of the value you create -- typically 10-20% for most service categories."
"Here's how to surface that value before you quote. You ask directly: 'Help me understand what success looks like for this project. If it goes really well, what does that mean for your business?'"
[DIRECTION: Slow down for the question. Consider reading it once, pausing, then repeating it.]
"That single question unlocks everything."
Section transition: "Once you know what the project is worth, you need a way to structure your quote so the number feels justified instead of arbitrary. That's where the three-part framework comes in."
SOFT CTA MENTION (at ~30% mark -- 2:10)
[DIRECTION: Casual, almost as an aside -- brief and genuine, not performative]
"Quick note -- if you're finding this useful, go ahead and subscribe. I put out new freelance business content every week, and next week's video is specifically about how to handle it when a client asks you to lower your rate. You're going to want that one."
SECTION 3: The Three-Part Pricing Framework (3:30 -- 5:45)
Estimated word count: 320 words | 2:12 at 145 wpm
[DIRECTION: Hold up fingers to count the three parts. Confident, structured energy. This is the practical core of the video -- deliver it with clarity and appropriate weight.]
"The framework has three components: anchor, scope, and protection. Let me walk you through each one."
"Anchor is your starting number. It's not your minimum -- it's where you start the conversation. Your anchor should be set based on the value question you asked. Take 15% of what the project is worth to the client's business. That's your anchor."
"If the email copy project is worth $96,000 a year, 15% is $14,400. That might be your anchor for a retainer arrangement. For a one-time project, you'd price differently -- but you've now got a value reference point instead of an hourly calculation."
[B-ROLL: Same whiteboard -- new calculation added: "$96,000 x 15% = $14,400"]
"Scope is every deliverable, revision round, and communication format included in your price. Document it completely. Not because clients are adversarial -- because vague scope is where undercharging happens invisibly. Projects that expand without boundary are why freelancers end up working three months for what they quoted as four weeks."
[TEXT ON SCREEN: Unscoped projects always expand]
[PAUSE]
"Write it down. Every line item. And include a clear out-of-scope clause: 'Additional rounds of revisions are billed at $X per round.'"
"Protection is your payment structure. For new clients, require 50% upfront before any work begins. This is not rude. It's standard practice in every professional service industry except freelancing, where it somehow became controversial."
[DIRECTION: Slightly more direct here -- this point gets pushback, deliver it with conviction]
"If a client refuses to pay a deposit, that's information about how they view the relationship. It's valuable information to have before you've done the work."
"Anchor. Scope. Protection. Those three components turn a quote into a professional proposal."
Section transition: "Now you've got the framework. But I want to spend a minute on the part that breaks down most often -- the actual conversation."
SECTION 4: Handling the Pricing Conversation (5:45 -- 6:40)
Estimated word count: 175 words | 1:12 at 145 wpm
[DIRECTION: Warm, practical. This section is about behavior change, not just information -- make it feel coaching-like.]
"When you send a quote or name your price out loud, the single biggest mistake is filling the silence that follows."
[PAUSE]
"You say your number. And then you wait. You don't say 'but I can be flexible.' You don't say 'let me know if that's too high.' You wait."
"Silence after a price is almost never rejection. It's processing. But freelancers trained on anxiety interpret it as rejection and start discounting before the client has had five seconds to think."
[B-ROLL: Two-person scene -- one person says something (no audio), then sits back with calm, open posture. The other person is thinking. 3 seconds of silence depicted visually.]
"If the client responds that the number is too high, ask: 'Can you help me understand what budget you're working with?' Not 'I can do it for less.' Get information first."
"Pricing conversations feel uncomfortable because you're not used to them. The more you have, the less uncomfortable they get."
PRIMARY CTA (6:40 -- 7:00)
[DIRECTION: Direct eye contact. Genuine. Not high-energy -- match the trust you've been building throughout.]
"If this video shifted how you're thinking about pricing, subscribe -- I cover the business side of freelancing every week, and next week I'm specifically getting into rate increase conversations with existing clients. That's a different kind of conversation, and there's a right way to handle it."
"Drop a comment and tell me: what's the project you most wish you'd charged more for? I read every comment, and a lot of next video ideas come directly from what you share."
OUTRO (7:00 -- 7:15)
[DIRECTION: Warm, unhurried. A breath before the final point. Not rushed.]
"Back to where we started -- that anxious feeling before you hit send on a quote. Now you know what it's telling you."
[PAUSE]
"Stop pricing your time. Price the outcome. Ask the value question, build your quote with anchor, scope, and protection -- and then hold the silence."
"Next week: how to raise your rates with clients you've had for years without losing them. It's a different conversation, and I'll walk you through exactly how to start it."
"Take care."
[B-ROLL: Callback to opening -- same hand at the laptop, same cursor on "Send" -- this time the hand clicks without hesitation]
Production Notes
Element
Count
Notes
B-roll shots required
6
4 to film (laptop hands, whiteboard, two-person scene, callback shot); 2 can be sourced from stock
On-screen text graphics
4
"Clients pay for outcomes not hours"; "Your fee as % of value"; "Unscoped projects always expand"; "1 Anchor 2 Scope 3 Protection" -- all need motion design
Direction annotations
9
Flag the "conviction" direction at 5:30 -- presenter should rehearse this section
Pause markers
6
Approximately 10-12 seconds of intentional silence total
Open loops planted
2
(1) "a question most freelancers never ask" planted at 1:40, resolved at 1:55; (2) next-week video tease planted at 2:10, reinforced at 7:05
Total spoken word count: 1,020
Base spoken duration: 7:03 at 145 wpm
Adjusted duration (+ 13%): 7:57 (rounds to 7:15 with tighter delivery -- test in read-through)
Estimated recording time -- experienced presenter: ~12-13 minutes
Estimated recording time -- new presenter: ~18-20 minutes
Primary CTA placement: 6:40 -- confirms at 91% of 7:15 duration -- ADJUST: move CTA start to 5:05 (70% of 7:15) if retention data shows significant drop-off after the 5-minute mark. The current placement assumes a high-retention audience typical of financially-motivated content.
Note on CTA timing: The heavy practical content in Section 3 (3:30-5:45) may sustain unusually high retention through the 70-80% mark for this audience type. Monitor first 30-day retention analytics and adjust CTA placement in the next script iteration if needed.