| name | talking-head-guide |
| description | Guide for editing speech-led videos where spoken delivery or conversation drives the cut — single-speaker talking-head / 口播, two- or multi-speaker interview / 访谈, video podcast, lecture, tutorial, course, and similar formats. Use for any non-trivial edit of those formats, including speech cleanup (剪口播 / 口播剪辑 / 去口癖 / clean up fillers / smooth speech), pause or repeated-take removal, motion graphics layered onto the footage (口播加 MG / 加动画), B-roll (加 B-roll / add B-roll), music, or captions.
|
| user-invocable | true |
Speech-Led Video Editing (Talking Head, Interview, Podcast) — Local Adaptation
Local Adaptations
This is a local adaptation of the talking-head guide for Local Video Editor. The following differences apply:
| Feature | Local Status |
|---|
| Tool prefix | All tools use local-video-editor_ prefix |
manage_design_style | ✅ Available: list_presets, apply_preset, get |
multicam_sync | ✅ Available locally |
search_fonts | ✅ Available |
Stock media (search_stock_media, browse_library) | ❌ Not available — use local files only |
Cloud screenshot (render_cloud_screenshot) | ❌ Replaced by local view_timeline_frames |
Workspace operations (pull_asset, push_asset) | ❌ Not available |
| MG visual rendering | ❌ Not available (code storage only — see create-motion-graphics) |
| ASR transcription | ❌ Not available (see transcription skill) |
| UI widget forms / Elicitation | ❌ Not available — use plain-text options |
Tool references throughout this skill omit the local-video-editor_ prefix for readability. Use the prefix when calling tools (e.g., local-video-editor_read_script, local-video-editor_edit_item).
What this skill covers
Required input: an existing speech-led source registered or imported into the project — for example a single-speaker talking-head / 口播, a two- or multi-speaker interview / 访谈, a video podcast, lecture, tutorial, or course.
When source media is provided as a local file path or attached in the conversation, use local-video-editor_import_media to import it into the project.
Independent treatments that can be applied to speech-led videos. Pick the ones that match what the user wants — not all are needed every time.
- A-roll editing (中文称 语音剪辑 / 含 去口癖、停顿、重复) — transcript-based speech editing. Common operations include cleanup, highlight extraction, restructure, opening hook, and others as needed for the aligned outcome.
- Motion graphics overlay (英文展示给用户时写全称 Motion Graphics,不要缩成 "MG";中文产品术语固定为 MG 动画——不要叫"动效""字幕条""动态字幕"等其它说法) — reinforce key information, structured content, and topic transitions with on-screen motion graphics
- B-roll (industry term — keep as "B-roll" in any language, do not translate) — cover jump cuts or visualize what's being said
- Background music (中文 背景音乐) — set mood and smooth micro-gaps
- Captions (中文 字幕) — on-screen text for accessibility
用户语言为中文时,在对话文案里严格使用上面括号里的产品术语——别自己再翻译一遍,会跟产品其它地方对不上。
What shapes the edit
Beyond picking treatments, a talking-head edit is shaped by several orthogonal variables. When the user's ask is vague, these are what's worth clarifying first:
- Target — platform (YouTube / TikTok / Shorts / ...), desired length, aspect ratio
- Which treatments to apply — the treatments above are optional; don't assume all of them apply
- Pacing / tone — tight / energetic / formal / casual; brand or voice preferences if stated. (For MG visual style, follow the Motion Graphics skill.)
When more than one of these variables is missing, ask concisely — don't run a fixed checklist.
Order of execution
When multiple treatments have been aligned with the user, they depend on each other and must be finalized in dependency order.
The speech timing (set by A-roll editing) anchors everything downstream — MG placement, B-roll cut-covers, music duration, and caption sync all reference the final speech timeline.
So: finalize A-roll editing before committing any visual, audio, or text layer. Don't write captions against pre-edit speech, don't cut music to pre-edit length, don't place MG against timing that will shift.
You must confirm the result with the user after each major step before starting the next, unless the user has explicitly asked to run end-to-end without stopping. Key checkpoints when multiple treatments apply: after A-roll editing finalizes the speech timing; before MG generation (confirm style and direction, and, when it isn't obvious, whether it sits over the video as an overlay or takes the whole frame); after MG generation; same pattern for B-roll, music, and captions. Don't bundle multiple checkpoints into one response — confirm each step separately.
A-roll editing
Scenario
In a talking-head workflow, the first step is usually A-roll editing: editing the original spoken footage.
A-roll edits are ultimately applied to the timeline and change what the viewer actually hears and sees. However, the editing decisions should usually start from the transcript, because the core question is: what spoken content should the viewer hear, and what should be removed, compressed, or reordered?
Common A-roll tasks
A-roll editing is not only cleanup. First decide what spoken-content task the user is asking for, then choose the editing strategy and tools.
Common tasks:
- Cleanup — remove mistakes, repeated attempts, verbal habits, filler words, and meaningless pauses so the speech becomes clearer and more natural.
- Highlight extraction — pull the most valuable, opinionated, emotional, or topic-relevant moments from longer footage.
- Restructure — reorder spoken content, such as moving the conclusion earlier, grouping by topic, or combining scattered parts into a clearer structure.
- Hook / short version — use a strong claim, result, conflict, or question from the source as the opening, or compress long content into a shorter version.
- Target-script / script alignment — match, keep, and reorder spoken content according to a user-provided target script, target paragraph, or desired content.
Shared A-roll principles
These principles apply to all A-roll tasks, not only cleanup.
- Decide the task before choosing the tool. Do not let tool availability change the editing strategy.
- Edit by complete semantic units. Whenever possible, move/delete/keep complete sentences, complete ideas, complete answers, or complete steps. Do not cut out a half-sentence just because a few words match.
- When the task names what to keep, trim to that boundary. The inverse of the rule above, for any task that specifies which content to keep — restoring a specific sentence, matching a target script, pulling a named highlight, building a version: keep exactly the requested span.
- Do not stitch unfinished fragments across retakes. Do not combine incomplete pieces from different attempts into one artificial sentence.
- Preserve connective tissue. List labels, contrast words, subjects, verbs, and adjacent source words are not filler when removing them makes a kept idea ungrammatical, abrupt, or misleading.
- Keep listening flow natural. The result should still have natural phrasing and breathing room.
- Be conservative when boundaries are uncertain. If unsure whether a cut harms meaning, logic, or listening flow, keep it or make a smaller cut.
- Confirm complex changes first. For complex restructuring, aggressive shortening, structural changes, or generated hooks, confirm with the user before editing.
- Explain content, never indices. You MUST NOT explain edits to the user with internal addresses such as
[sN], [cN], [gap], word indices, clip ids, or segment ids.
- Never name a screen position for a panel. When you invite the user to review or fine-tune the result, call it "the Transcript panel" (中文「文字稿面板」) — never a direction (left / right / side / 左侧 / 右侧).
Cleanup goals and decisions
What good cleanup means
Good cleanup does not mean making the video as short as possible, and it does not mean rewriting the speaker into a different script.
Good cleanup means:
- The logic stays coherent
- The expression becomes clearer
- The audio feels natural
- Obvious mistakes, repeated attempts, meaningless stalls, and filler are removed
- The speaker's intent, tone, and natural rhythm are preserved
Bad cleanup usually falls into two failure modes:
- Under-cleaning: obvious mistakes, repetition, long pauses, or filler remain.
- Over-cleaning: sentences are cut off, meaning is missing, rhythm becomes too hard, or the result sounds stitched together.
Default principle: remove defects without changing meaning; make speech smoother, not harder; prefer small local cuts over whole-sentence or whole-segment deletion; when unsure whether a cut harms meaning, keep it.
How to judge common cleanup cases
Meaningless filler words
Fillers fall into two categories.
The first category is clearly meaningless hesitation sounds. These are usually safe to remove:
When they do not carry special meaning, use local-video-editor_clean_script first for bulk cleanup.
The second category depends on context and must not be removed by word list alone:
so, like, 然后, 就是, 嗯, 啊, 那个, 那, 对, 所以, 但是
How to decide:
- If the word is only hesitation or padding, remove it.
- If it carries sequence, continuation, contrast, cause, reference, response, emphasis, or natural tone, keep it.
- If removing it makes the surrounding words sound hard-spliced, keep it or only compress the pause.
- If unsure, keep it.
Retakes and repeated attempts
A retake is when the speaker retries the same intended idea because they misspoke, got stuck, forgot words, or restarted. Retake cleanup is not "delete repeated text." The goal is to keep one complete, natural, logically coherent version of the intended idea.
Use this decision path:
- Decide whether it is really a retake. Treat it as a retake only when multiple attempts are trying to say the same intended idea.
- Define the complete version to keep. May include lead-in, connector, section marker, topic setup, contrast, qualifier, subject, object, or conclusion.
- Cut only the failed or covered part. Remove only words that are wrong, dangling, abandoned, or fully covered by the kept version.
- Choose the best complete attempt. If several attempts are complete, usually prefer the later one.
False starts and unfinished fragments
Only remove a fragment when it clearly does not form useful information.
Safe to remove:
- The speaker abandons the thought and a complete version appears later.
- The segment is only a dangling phrase.
- It is clearly the leftover beginning of a failed attempt.
Do not remove:
- A sentence that is imperfect but contains useful information.
- A lead-in that provides the subject, object, or context needed later.
- Content that provides setup, contrast, conclusion, emotion, or tone.
Pauses and breaths
Pause cleanup should default to compression, not zeroing out. Spoken video needs natural breathing room.
Default rules:
- Obvious long pauses over 0.8-1s: usually compress to about 0.3s.
- Between sentences: keep about 0.3-0.5s.
- Around topic shifts, contrast, or emphasis: keep slightly longer pauses.
- Short breaths inside one sentence: if they are normal breathing, do not remove them.
- Clear long pauses inside one sentence: compress them, but not so tightly that adjacent words sound glued together.
How to operate on pauses:
- For batch pause cleanup across the timeline or track, use
local-video-editor_clean_script.
- Use
local-video-editor_read_script({ showSilence: true }) only when you need to inspect or manually adjust a specific pause.
- After semantic edits, review the final clean
timeline.md. If the final pacing still has many long pauses, run clean_script only="silence"; if only one or two pauses feel wrong, use showSilence: true and adjust those manually.
Script gap primitive note:
- Do not create an accidental
[gap] on the primary video track as a pacing pause. A Script [gap] means no source is playing; on the only visible video track it renders as black.
Other A-roll task guidance
Building versions, highlights, and excerpts — stay on Script
Highlight, short version, excerpt, hook, restructure, and making several versions are all transcript-content tasks: drive them through Script (local-video-editor_read_script → edit timeline.md → local-video-editor_apply_script), never by looking up timestamps and placing source clips manually.
- Versions on the current timeline: trim or reorder
timeline.md and apply_script.
- A version on its own timeline:
manage_timelines action=duplicate, then read_script → trim → apply_script on it.
- To bring in source content the current cut no longer shows, read
library/<filename>.md, copy the needed [sN] line(s) into timeline.md, and apply_script.
- For multiple versions on one track: list every version's
[sN] segments in timeline.md in version order, then apply_script once.
- Never look up timestamps with
find_transcript and place spoken content with edit_item / split_item.
Check each version against its request. After assembling a version, confirm every requested sentence is present, in the requested order, with no extra source carried in.
A-roll / transcript-based editing workflow
Use this flow for any A-roll task driven by transcript meaning.
- Start with orientation. Call
local-video-editor_read_script, then read timeline.md once to understand the user's goal, the content structure, and whether fixed fillers or long pauses are present.
- For cleanup tasks, run the mechanical cleanup pass before semantic editing. Use
local-video-editor_clean_script for fixed hesitation sounds and batch pause compression.
- After
clean_script, always read the refreshed clean timeline.md before semantic editing. Then edit timeline.md with semantic judgment. For long transcripts, work one clear section at a time.
- Apply the edit with
local-video-editor_apply_script. If apply fails, fix the markdown error or stale state, re-read current timeline.md, and apply again.
- Review the edited result. Read the regenerated clean
timeline.md and check what the viewer will actually hear. If final pacing still needs batch pause adjustment, use clean_script only="silence".
What transcript editing actually changes
Editing timeline.md is not just changing displayed text. It describes which source media ranges should play on the timeline.
[sN] rows are ASR segments, not semantic units. A complete sentence, idea, retake, or transition may span several [sN] rows.
- Each spoken-text line maps to a playable source range.
- Inline
~~...~~ removes the corresponding audible audio range.
- Deleting a whole line removes that whole spoken segment.
- Moving/reordering lines changes playback order.
apply_script applies the result back to the timeline.
- Deleting words or pauses in the middle of a sentence splits the original clip into multiple new clips.
Tool boundaries
local-video-editor_clean_script: use for mechanical first-pass cleanup: bulk removal of fixed meaningless fillers and batch silence compression/adjustment.
local-video-editor_read_script + local-video-editor_apply_script: the main transcript-based editing surface for real semantic editing.
local-video-editor_manage_transcript action fix: only fixes ASR mistakes or speaker attribution. It does not cut audio.
local-video-editor_find_transcript: only locates when a phrase is spoken. It does not edit.
MG Overlay
Goal
Motion graphics layered into A-roll reinforce what the speaker is conveying — deepening the audience's impression of key points. Complete A-roll editing first; MG timing is based on the post-edit timeline.
⚠️ MG visual rendering is not available in Local Video Editor. MG code can be authored and stored (see create-motion-graphics skill) but cannot be rendered to video frames. The workflow below covers design, authoring, and placement — verification via frame inspection will not show MG visuals.
Visual identity
Design Style is the video's confirmed visual language. Resolve visual language before planning MG moments:
- Active Design Style — use it unless the user asks to change. If Project Context names an active style but no details, inspect with
local-video-editor_manage_design_style action="get".
- Specific user style / reference — follow it.
- Generic or vague direction — quality words such as "clean" or "premium" are goals, not a visual language. Use preset options.
- No visual direction — use
local-video-editor_manage_design_style with action: "list_presets", scenario: "talking-head", and the user's locale to show visual options.
- "Directly do" / no style input — choose a concrete temporary direction from the transcript and footage, then continue without confirmation.
Picker flow:
- Call
local-video-editor_manage_design_style with action: "list_presets", scenario: "talking-head" when clear, and the user's locale.
- Present returned presets as visual options.
- The picker is a turn boundary: after showing it, stop and wait for the user's selection.
- When the user picks, call
manage_design_style with action: "apply_preset" and the selected presetId, then inspect with action: "get" before authoring.
Where MG is useful
MG meaningfully helps comprehension or orientation when the content has:
- Identity / context labels — speaker name, role, product name, date, source
- Key information / quotes — a core concept, definition, statistic, conclusion
- Structured information — multiple points, steps, comparisons, rankings, lists
- Chapter / topic markers — opening titles, section titles, topic transitions
- Abstract concepts — cause-effect relationships, cycles, systems, frameworks
Per-MG decisions
For talking-head videos, do not start MG creation from transcript timing alone. Inspect the target frame first.
Before creating the MG, make four linked editor decisions:
| Decision | Question | Output |
|---|
| Content | What idea deserves a visual layer? | Message expressed by the MG |
| Timing | When should it land with the speech? | Timeline start, duration, read time |
| Form and placement | What kind of MG and where? | MG form / size, then placement |
| Background | Is this an overlay or its own moment? | Transparent overlay or opaque |
1. Content
Choose what the MG expresses, not just what text it repeats.
2. Timing
Use local-video-editor_find_transcript with includeWordTimestamps: true when the MG has internal rhythm.
3. Form and placement
Placement principles:
- Protect the subject and safe zones. Avoid the speaker's face, head, important objects, hand gestures, and captions.
- Keep the caption/subtitle area clear.
- Separate overlays from full-screen MGs.
- Keep the composition intentional.
4. Background
Default to Background: transparent for talking-head overlays. Use Background: opaque when the MG is its own visual surface.
Place and review
Since MG rendering is not available locally, placement and review are limited to:
- Create the MG asset with
local-video-editor_create_motion_graphic_from_code
- Place with
local-video-editor_edit_item (adds/updates)
- Verify placement metadata with
local-video-editor_read_project
- Visual verification of MG effects on rendered frames is not available
B-roll
Goal
Enrich visual layers and cover jump cuts left by A-roll editing.
Where B-roll is useful
- Cover jump cuts — when A-roll editing leaves visible jump cuts
- Visualize specific references — when the speaker mentions objects or scenes
B-roll depends on having suitable footage and adds production effort — treat it as an optional enhancement, not a default.
Sources
⚠️ Stock media libraries (search_stock_media, browse_library) are not available in Local Video Editor. B-roll footage must come from local files already in the project library or imported via local-video-editor_import_media.
How to place
Don't cut away in the first or last 3 seconds. For dense jump cuts (<3s apart), use one long cutaway covering multiple. Don't overlap with MG by default.
First decide the B-roll mode:
- Full-screen cutaway replaces the talking head for that moment.
- PiP / small-window overlay keeps the talking head visible.
PiP / small-window overlay placement
- Inspect the target timeline frame first. Exclude areas covering the A-roll's face/head, captions, existing overlays.
- Inspect the B-roll source frame(s). Identify the primary subject/action and protected information.
- Place the overlay at a useful size inside the chosen destination rectangle.
- Set the media item's native
borderRadius to 24-36 by default.
- Use
local-video-editor_view_timeline_frames to verify the affected frame before reporting success.
Full-screen cutaway placement
- Compare source aspect ratio with canvas aspect ratio before choosing fit.
- For substantially different ratios, inspect the source with
view_asset_frames first.
- Identify whether protected information would be cropped.
- Apply fit strategy per source asset, not as a batch default.
- Use
local-video-editor_view_timeline_frames to verify the final frame.
After editing, read back item ids with local-video-editor_read_project({itemId:"..."}) or use view:"track" for the affected track. If the result involved crop, fit, scale, or overlay placement, verify the affected frame with view_timeline_frames before reporting success.
Multicam (multiple camera angles of the same take)
When the user has two or more cameras recording the same moment, switching to another angle means the picture changes but the audio and lip-sync must stay matched.
✅ multicam_sync is available locally. Use the local-video-editor_multicam_sync tool — it runs the editor's audio-based alignment engine and repositions each angle clip so its picture matches the reference angle's audio.
Pass the angle clips' itemIds (the reference plus the follower angle(s)); optionally name the referenceItemId.
Key constraint: a single cutaway clip that spans a cut in the reference angle can't be aligned as one piece — split it at that cut with local-video-editor_split_item first, then pass both pieces to multicam_sync so each maps to the reference segment beneath it.
Track roles (turn on auto-ducking)
A track's role is the single declaration that drives the audio mix. Set it with local-video-editor_edit_track and the engine derives a seamless duck — followers dip under speech, then rise back in the gaps.
- The talking / interview / lecture track (and any voiceover / narration) is the anchor: set its
role to anchor. This is the track everything else ducks under.
- Background music, ambient beds, and b-roll audio beds → the follower: set
role: follower (auto-ducks under every anchor).
- Short sound effects (SFX), stingers, hits → leave their track role unset.
- Anything that should stay out of ducking → leave its
role unset (none).
Read the existing layout first. Before creating tracks or placing new clips, read the current track names and roles. If a track is already tagged for this content, put the new clip there and match its role; only make a new track when nothing fits. Organize before you assign — roles are per-track, so aim for one role per track.
After assigning, read the project back to confirm every track that should anchor/duck does, and that you left the music's base volume alone.
Background Music
Goal
Set the mood and smooth over micro-gaps in speech.
Principles
- Set the music track's
role to follower with local-video-editor_edit_track (and the talking track's role to anchor). That single pair turns on auto-ducking.
- Let
edit_track initialize audioRouting.duckDepthDb from the current timeline loudness.
- Keep the BGM clip's base
decibelAdjustment natural by default.
- Do not put short sound effects (SFX) or stingers on follower tracks by default.
- No prominent lyrics.
- Fade BGM in/out with
audioFadeIn / audioFadeOut in seconds, usually 1-2 seconds.
- Tone matches content.
Fit to duration
Fit BGM to the final video extent after A-roll timing is finalized.
- Unless the user specifies a different BGM start, start BGM at frame 0.
- If generated BGM is longer than the target, place one
audio item at the BGM start, set its duration to the target duration, and add a fade out.
- If generated BGM is shorter than the target, tile multiple
audio items until the target is covered.
- For tiled BGM, use alternating audio tracks so adjacent repeats can overlap by 1-2 seconds.
How the engine ducks music
Ducking is automatic once the music track's role is follower: the engine dips the track under audible anchor tracks, and lifts it back in pauses and the outro. This needs both halves — music track with role: follower and talking track with role: anchor (see Track roles above).
Captions
Goal
Improve accessibility and engagement with on-screen text.
Captions start from the source transcript. When the user asks for translation or bilingual captions, use local-video-editor_edit_captions action language_mode; its languageCode is the translation target.
Presets
Prefer built-in caption presets. Use only real built-in edit_captions preset names.
- For a general style request, first list the language-aware presets with
local-video-editor_edit_captions action template, then choose or offer relevant returned presets.
- Use custom
style / layout only when the user clearly requests a custom look or specific adjustment.
- For adjustments, start from the closest preset and change only the requested properties.