Use when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels), or asks to learn visual design patterns from a reference video/image. Trigger when the user mentions creating a knowledge video, narrated explainer, video podcast, or animated infographic-style video from a topic — even if they don't say "video podcast" explicitly. Also trigger when the user wants to regenerate, re-render, rebuild, update, or iterate on a narrated video this skill already produced — e.g. they edited the script/prompt, changed the visuals, or swapped the background music and want the final video remade (reuse the existing videos/{name}/ directory, never start a new project). Do NOT trigger for generic video editing, trimming, format conversion, color grading, or non-narrative video tasks. Produces 4K video via research → script → TTS → Remotion → MP4 + BGM.
Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.
Une commande directe contourne le prompt de vérification. Examinez la source avant de l'exécuter.
Use when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels), or asks to learn visual design patterns from a reference video/image. Trigger when the user mentions creating a knowledge video, narrated explainer, video podcast, or animated infographic-style video from a topic — even if they don't say "video podcast" explicitly. Also trigger when the user wants to regenerate, re-render, rebuild, update, or iterate on a narrated video this skill already produced — e.g. they edited the script/prompt, changed the visuals, or swapped the background music and want the final video remade (reuse the existing videos/{name}/ directory, never start a new project). Do NOT trigger for generic video editing, trimming, format conversion, color grading, or non-narrative video tasks. Produces 4K video via research → script → TTS → Remotion → MP4 + BGM.
If remotion-best-practices is not installed, minimum rules: chromium must be available, always wrap 4K content in <Scale4K>, use <TransitionSeries> with linearTiming, and treat audio as the master clock.
Video Podcast Maker
Automated pipeline for 4K Bilibili horizontal knowledge videos from a topic. Coding agent + TTS backend + Remotion + FFmpeg.
Resolve SKILL_DIR to the directory containing this SKILL.md:
Pi: the agent knows the skill path from the loaded skill list — set SKILL_DIR to that directory before running commands.
Claude Code: ${CLAUDE_SKILL_DIR} is auto-populated.
SKILL_DIR=""
python3
${SKILL_DIR:-${CLAUDE_SKILL_DIR}}
# Prerequisites (CLIs + backend env vars)
"${SKILL_DIR}/scripts/check_prereqs.py"
Updates flow through the plugin marketplace (/plugin update); direct git-clone installs use git pull per the README. This skill performs no update checks.
Prereqs failures — see README.md for setup. The check is backend-aware (resolves TTS_BACKEND env → user_prefs.jsonglobal.tts.backend → edge default), so only env vars required by the active backend are validated.
First video in a new project? Prefer reusing an existing Remotion project with node_modules/ already installed — creating a fresh project downloads ~2.2 GB of npm packages plus a 90 MB Chrome headless shell (one-time per project). If the user has a project from a previous video, use it. If a fresh project is necessary, run npm install in the background while you do Steps 1-4 (topic research and script writing).
All rendering goes into videos/{name}/ — every output.mp4, final_video.mp4, and thumbnail_*.png lands directly in the per-video directory. Never render to an out/ or dist/ directory; the --public-dir videos/{name}/ convention keeps everything self-contained.
TTS engine — all 11 backends (TTS_BACKEND=edge|azure|cosyvoice|doubao|tencent|baidu|minimax|xunfei|elevenlabs|openai|google) synthesize through the ttscn component skill, which is required: install it under ~/.claude/skills/ttscn or point TTSCN_HOME at its root (Agents365-ai/ttsCN). Each backend still needs only its own API keys (Edge needs none); check_prereqs.py validates both the install and the keys.
Pi users:ttscn is not bundled with Pi — install Agents365-ai/ttsCN as a Pi skill (its skills/ttscn/ layout is auto-detected) or set TTSCN_HOME; check_prereqs.py verifies the install before TTS.
Design Learning shortcut: If the user provides a reference video/image or asks to save/list/delete style profiles, see references/design-learning.md instead of running the workflow below.
Execution Modes
Detect Auto Mode (default) vs Interactive Mode at workflow start — the Auto-default decision table and per-request overrides are in references/workflow-script.md.
Regenerating an Existing Video
If videos/{name}/already exists and the user is iterating on a finished or in-progress video, reuse that directory. Do NOT start a new project or a new videos/{newname}/.
Pick the smallest re-run for what actually changed:
Changed
Re-run
Reuses (don't redo)
Narration script (podcast.txt)
Step 7 (TTS) → Step 8 preview → render+mix
topic research + section design
Visuals only (components, layout, colors)
Step 8 preview → render+mix
audio (podcast_audio.wav / timing.json)
Background music only
Re-mix BGM
output.mp4 (no re-render)
Subtitles only
Step 10.1 finalize
output.mp4 / video_with_bgm.mp4
Any re-run that changes what the viewer sees or hears re-enters the Step 8 gate: apply the change, let Studio hot-reload, and wait for a fresh explicit "render 4K" — the previous confirmation does not carry over. A script change shifts every downstream timestamp, so always regenerate timing.json through TTS — never hand-edit it. After any re-run, re-verify:
Iterating on a finished video? If videos/{name}/ already exists, see Regenerating an Existing Video above for the minimal re-run — do NOT start at Step 1.
At Step 1 start, create one task per step in your agent's tracker. Mark in_progress on start, completed on finish. Files in videos/{name}/ are the durable record — if interrupted, inspect the directory to determine where to resume.
Step 8 — Studio review. MUST launch npx remotion studio and wait for user feedback before rendering. NEVER render 4K until the user explicitly confirms ("render 4K" / "render final"). A reply containing adjustment requests is not confirmation — apply the changes, let Studio hot-reload, and ask again. Every round of adjustments needs its own fresh confirmation before Step 9.
Step 10 — verify_output.py. MUST pass before declaring the video done. Exit 0 = green; exit 2 = warnings still publishable. Auto-fixes common omissions (creates final_video.mp4 if missing). Validates publish info (title, description, tags, chapters) against the platform matrix — generate it in Steps 5.5 and 10.2. For machine-readable output add --format json.
Auto Mode: visual self-review. When running in Auto Mode (no user watching Studio), render 3-5 key frame stills before asking for render confirmation:
npx remotion still src/remotion/index.ts <CompositionId> videos/{name}/_review_001.png --public-dir videos/{name}/ --frame=<midpoint_frame>
Pick frames at: hero title (~10% in), a dense section midpoint, and the outro. Read the stills back as images and run the design-guide.md and visual-taste.md checklists against actual rendered output. Catch overflow, contrast, and layout regressions before the 4K render. Delete _review_*.png after review.
Validation Checkpoints
After Step
Check
7 (TTS)
podcast_audio.wav plays · timing.json covers all sections · SRT is UTF-8
9 (Render)
output.mp4 is 3840×2160 · audio-video sync · no black frames
10 (Verify)
verify_output.py exits 0 (or 2 with reviewed warnings)
Hard Rules
Rule
Requirement
Single Project
All videos under videos/{name}/ in user's Remotion project. NEVER create a new project per video.
4K Output
3840×2160 (or 2160×3840 vertical), use scale(2) wrapper over 1920×1080 design space
Audio Sync
Audio (podcast_audio.wav + podcast_audio.srt) is the master clock. timing.json MUST be generated from the real TTS output, never hand-estimated. Before rendering, final video duration must match audio within ±0.5s. See Audio-Master Clock & Sync.
Thumbnail
MUST generate both 16:9 (1920×1080) AND 4:3 (1200×900) — see design-guide.md
Studio Before Render
MUST launch remotion studio for review. NEVER render 4K until user explicitly confirms. Adjustment feedback ≠ confirmation — apply, hot-reload, ask again.
--public-dir
Every Remotion command uses --public-dir videos/{name}/. All output files (output.mp4, final_video.mp4, thumbnails) go directly into videos/{name}/ — never an out/ or dist/ dir.
Visual minimums (text sizes, content width, safe zones, animation safety) live in references/design-guide.md. MUST load before Step 8.
Audio-Master Clock & Sync
Golden rules
Audio is the master clock. Every slide start, subtitle, chapter, and animation beat is derived from podcast_audio.wav and podcast_audio.srt.
Generate timing from TTS, not from text estimates. Pipeline: podcast.txt → generate_tts.py → podcast_audio.wav + podcast_audio.srt + timing.json → composition → render.
Never hand-write timing.json before audio exists. If you already have curated slides, run align_timing_from_srt.py to anchor them to the real SRT.
Compensate TransitionSeries overlap.TransitionSeries renders sum(section.duration_frames) - (N-1) * transitionFrames frames. Scale every section proportionally to keep the rendered length equal to timing.total_frames. Do not stuff all overlap frames into the first section. The corrected pattern is in templates/Video.tsx.
Mandatory sync checkpoints
When
Check
After Step 7 (TTS)
timing.json.total_duration matches podcast_audio.wav within ±0.5s
Before render
Video.tsx scales all sections for transition overlap
After render
final_video.mp4 duration matches podcast_audio.wav within ±0.5s
Step 10 (verify)
verify_output.py exits 0 and reports green on audio/timing
All scripts are reachable through one dispatcher — start with python3 ${SKILL_DIR}/scripts/cli.py --help; full routes and envelope error codes: references/troubleshooting.md.
User Preferences
Mutable state (user_prefs.json, phonemes.json) lives in ~/.video-podcast-maker/ — safe from skill updates. Auto-migrated from the skill directory on first run. Run "show preferences" to view, or "set X Y" to change. Full commands: references/troubleshooting.md.