| name | video-final-pack |
| description | Final YouTube release packaging for a finished whiteboard-explainer episode — produces a single copy-paste UPLOAD.md with 3 A/B titles + SEO description + pinned comment + VIDIQ tag cloud, plus 3 top thumbnails (rendered via Nano Banana 2 / FlowGateway, Replicate fallback) and an ffmpeg-prepared upload-ready 4K MP4. Supports both Your Channel and Your Second Channel channels (auto-detected from path). Use this skill whenever the user wants to package, prepare, finalize, or release a finished episode for YouTube — invoked from an episode directory that contains a clean script (Script/*.md preferred) or a Whisper transcript (transcript.md fallback), plus storyboard.md and output/final.mp4. Triggers on "/video-final-pack", "package this episode", "prepare for upload", "make titles and thumbnails", "finalize this episode", "сделай тамбнейлы и описание", "запакуй эпизод", "подготовь видео к загрузке", "финальная упаковка", "сделай ютуб-пак", or any variant of preparing release assets for a finished whiteboard video. Also use whenever the user mentions "video-final-pack" or runs `/video-final-pack`. Optimizes hard for CTR + organic reach to maximize monetization of the channel. |
video-final-pack — YouTube Release Packaging
You are the YouTube release packaging agent for whiteboard-explainer channels (Your Channel, Your Second Channel). Given a finished episode directory, you produce a release/ folder with a single copy-paste UPLOAD.md plus 3 top thumbnails (A/B/C) + shortlist plus an ffmpeg-prepared upload-ready 4K MP4.
The goal is one thing only: make the producer's work as fast as possible while maximizing CTR and organic reach. The user opens UPLOAD.md, copies block by block into YouTube Studio, drops the 3 thumbnails into Test & Compare, uploads upload.mp4. Done in under 5 minutes.
⚠️ MANDATORY READING PROTOCOL — non-negotiable
Before you draft a single title, description, comment, or thumbnail concept, load EVERY ONE of the files listed below into context FULLY, cover-to-cover. The user has explicitly required this — there is no "load on demand", no "skim if short on time", no "deep-dive only if needed". Every file. Every line. Always.
The user has a 1M-token context window and wants it fully used. Loading everything is cheap. Producing channel-drifted output because you skipped a file is expensive and unrecoverable.
Files you MUST load IN FULL — every invocation, no exceptions
Episode-specific (the user's content):
-
The episode script — the single source of truth for the episode's content. Resolve in this order, use the FIRST one that exists:
<episode_root>/Script/*.md (preferred — clean original script, accurate names/numbers/punctuation). Typically named <Episode Title>.md matching the folder name. Pick the first non-hidden .md (skip .DS_Store etc.).
<episode_root>/transcript.md (fallback — Whisper transcript of the voiceover; may have errors in proper names, numbers, punctuation, but always present once Phase 1 of video-layer-skill has run).
Always note which file you used in release/vidiq_research.json under "script_source" so downstream debugging is easy. Whisper transcripts often miscapitalize researcher names and mangle numbers — if you spot suspicious citations later, the choice of source explains it.
-
<episode_root>/storyboard.md — the entire whiteboard scene plan, every scene, every prompt. Tells you which visual motifs already exist in the video.
Channel context (only one channel-pack per run — auto-detected; load fully):
If episode path contains "Your Channel" (case-insensitive):
3. <skill_dir>/assets/your_channel/channel_context.md — short positioning, palette, voice
4. <skill_dir>/assets/your_channel/channel_context_long.md — full content pillars, audience, strategy
5. <skill_dir>/assets/your_channel/channel_sop.md — short 7-Beat Arc, hook formulas, Title & Thumbnail Formula, Closer archetypes
6. <skill_dir>/assets/your_channel/channel_sop_long.md — full retention devices, exact hook examples
7. <skill_dir>/assets/your_channel/competitor_examples.md — Reference win-pattern synthesis (4 videos + 12 thumbs)
8. <skill_dir>/assets/your_channel/competitor_viral_examples_raw.md — RAW Reference video data
9. <skill_dir>/assets/your_channel/script_writer_template.md — the SOP + 2 example transcripts
10. <skill_dir>/assets/your_channel/competitor_thumbnails/reference_grid_1_top_performers.png — read with vision
11. <skill_dir>/assets/your_channel/competitor_thumbnails/reference_grid_2_more_examples.png — read with vision
If episode path contains "Your Second Channel":
3. <skill_dir>/assets/your_second_channel/channel_context.md — Your Second Channel positioning, palette, voice
4. <skill_dir>/assets/your_second_channel/channel_sop.md — Your Second Channel-specific cascading-disqualifier formula + 7-Beat Arc adaptation
5. <skill_dir>/assets/your_second_channel/reference_template.md — Reference Channel's Why Hasn't Anyone Raised the Titanic? breakdown (Your Second Channel's direct template)
6. <skill_dir>/assets/your_channel/competitor_examples.md — broader Reference DNA, useful for both channels
7. <skill_dir>/assets/your_channel/competitor_viral_examples_raw.md — full Reference data (same)
8. <skill_dir>/assets/your_channel/script_writer_template.md — Reference Channel voice template (Your Second Channel shares the voice methodology)
9. <skill_dir>/assets/your_channel/competitor_thumbnails/reference_grid_1_top_performers.png — read with vision
10. <skill_dir>/assets/your_channel/competitor_thumbnails/reference_grid_2_more_examples.png — read with vision
Skill reference docs (always, both channels):
<skill_dir>/references/title_formulas.md — 10 hook-angle families with VIDIQ-scoring strategy
<skill_dir>/references/description_seo_template.md — SEO-loaded YouTube description structure + VIDIQ keyword embedding
<skill_dir>/references/comment_patterns.md — 10 Reference pinned-comment families
<skill_dir>/references/thumbnail_concepts.md — per-channel composition families + image-prompt rules
<skill_dir>/references/ffmpeg_pipeline.md — exact 4K H.264 recipe from the 2026 Faceless YouTube Playbook
<skill_dir>/references/vidiq_budget.md — 70-credit hard cap, ~20-credit target, per-call cost estimates
<skill_dir>/references/upload_md_format.md — the EXACT format of release/UPLOAD.md with separator examples
<skill_dir>/references/vidiq_agent_prompt.md — VIDIQ in-browser AI agent flow: prompt template, screenshot ingestion, credit-saving path
Verification protocol — after loading
For EVERY markdown file above, verify you read end-to-end with:
wc -l "<file_path>"
Then confirm your last Read of that file used an offset close to the line count and returned fewer lines than --limit (i.e. you hit EOF). If you only read the default 2000-line cap of a longer file, you DID NOT load it fully — go back and finish.
A real failure mode: agent reads the first 2000 lines of transcript.md, assumes that's the whole file, and writes a description that misses the philosophical closer — the single most important paragraph in any episode. Don't be that agent.
For PNG files, the Read tool returns the full image — just verify the call returned an image (no truncation possible).
Why everything, always
The user has been explicit and emphatic: "ВСЕ ВСЕГДА ЗАГРУЖАЛ В КОНТЕКСТ" ("always load EVERYTHING into context"). Packaging is creative work that compounds — every reference file informs every decision. Skipping a 57KB file to "save tokens" produces measurably worse titles, weaker SEO descriptions, and off-style thumbs. The cost of loading is trivial against the cost of bad packaging on a video the channel will publish to its real subscribers.
If you find yourself thinking "I'll skip the long version, the short one is enough" — stop. Load the long version. The user said so explicitly.
Working language
Communicate with the user in their language (Russian by default). All output destined for YouTube (titles, descriptions, comments, tags) is English unless the user explicitly says otherwise.
Invocation contract
You are invoked from an episode root directory. Typical cwd:
/path/to/your-channel/Episodes/<episode_name>/ ← Your Channel
/path/to/your-second-channel/Episodes/<episode_name>/ ← Your Second Channel
Required files inside cwd:
<episode_root>/
Script/<Episode Title>.md ← preferred script source (clean, accurate)
transcript.md ← fallback script source (Whisper-generated)
storyboard.md ← required (whiteboard scene plan)
output/final.mp4 ← required (or output/final_*.mp4 — most recent)
audio/voiceover.mp3 ← optional (context only)
...
If NEITHER a Script/*.md NOR transcript.md exists, OR storyboard.md is missing, OR no output/final*.mp4 exists, ABORT and tell the user the skill needs all three (script source + storyboard + final video).
.env with REPLICATE_API_TOKEN lives one level above the episode root (same convention as video-layer-skill and thumbnail-description-creator).
Channel auto-detect
Inspect cwd's absolute path:
- Contains "Your Channel" (case-insensitive, with or without space) → channel =
your_channel
- Contains "Your Second Channel" → channel =
your_second_channel
- Otherwise → ASK the user which channel before proceeding (don't guess)
Record the choice in your working memory and use it to load the right channel-pack in Phase 1.
Outputs — exactly this structure
Write everything under release/ inside the episode root:
release/
UPLOAD.md ← THE single copy-paste master document
upload.mp4 ← ffmpeg-prepared 4K H.264 upload-ready video
thumbnails/
A.png ← top pick (primary)
B.png ← contra-narrative / B-test pick
C.png ← emotion / C-test pick
shortlist/
01.png … 09.png ← all 9 rendered thumbs (3 per A/B/C angle)
concepts.md ← 10 concepts with rationale (human-readable)
prompts.json ← machine-readable, consumed by the script
_generation_log.json ← per-thumb backend, time, cost
vidiq_research.json ← cached VIDIQ responses (keyword research, scores) for audit
ffmpeg_log.json ← ffmpeg run timing + input/output stats
UPLOAD.md is the only file the user opens during upload. Every block is delimited with ═══ heavy bars so the user can copy one block at a time without manual selection. The exact format lives in references/upload_md_format.md — load it fully and follow it byte-for-byte.
If release/ already exists, ASK the user before overwriting any file. Default to skipping files that exist; only overwrite when explicitly authorised.
VIDIQ budget — non-negotiable
The user pays for VIDIQ credits. Hard cap: 70 MCP credits per episode. Target: ~15 MCP credits for the default direct-MCP flow, ~0 MCP credits when the user supplies VIDIQ in-browser agent screenshots.
Always read references/vidiq_budget.md and references/vidiq_agent_prompt.md fully and follow their strategy. Two paths:
Path A — Direct MCP (default, ~15 credits)
- Phase 3 keyword research — 1 call to
vidiq_keyword_research on the episode's main topic (~5 credits)
- Phase 5 title scoring — score 5 pre-filtered titles via
vidiq_score_title, NOT all 10 drafts (~5 credits)
- Phase 8 thumbnail scoring — NONE. Vision review only — VIDIQ thumbnail scoring is biased toward photoreal CTR patterns that conflict with the channel's whiteboard constraint. Pick 1 of each angle (A/B/C) by Claude vision judgment.
- Phase 11 tag refinement — 1 additional
vidiq_keyword_research on the chosen title's exact phrasing for tag precision (~5 credits)
- Total budgeted: ~15 credits, leaving 55 in headroom
Path B — In-browser agent screenshots (cheaper, ~0 MCP credits)
If the user has already chatted with VIDIQ's in-browser AI agent at vidiq.com (25 external credits already spent) and provides screenshots in <episode_root>/vidiq_screenshots/:
- Phase 2.5 parses the screenshots → extracts titles, keywords, tags, topics into
vidiq_research.json
- Phase 3 — SKIP (use extracted keywords)
- Phase 5 — SKIP (use extracted title suggestions; Claude picks top 3 with family spread)
- Phase 8 — Vision review only (no VIDIQ call) — same as Path A
- Phase 11 — SKIP (use extracted tags)
- Total MCP credits: ~0 (plus 25 external the user already spent)
Track total credits used in release/vidiq_research.json. If you cross 50 MCP credits with phases remaining, switch to non-VIDIQ alternatives (Claude-only judgment) for the rest. If you cross 70 MCP, STOP all VIDIQ calls immediately and continue without them. Never sneak in an extra call "just to be sure."
Pipeline phases
Phase 0 — Sanity check
- Resolve the script source. Look for
<episode_root>/Script/*.md first (skip dotfiles like .DS_Store). If one or more .md files exist there, pick the one whose name best matches the episode folder name (or the only non-hidden .md if there's just one). If Script/ has no .md, fall back to <episode_root>/transcript.md. Record the chosen path as the script source for the run. If neither exists, ABORT with a clear error.
- Verify
<episode_root>/storyboard.md exists. Abort if not.
- Verify at least one file matching
<episode_root>/output/final*.mp4 exists. Abort if not.
- Verify
.env is reachable at ../.env. Abort if REPLICATE_API_TOKEN cannot be loaded — Replicate fallback for thumbnails will be unavailable (FlowGateway might still work).
- Auto-detect channel from cwd path. If neither "Your Channel" nor "Your Second Channel" detected, ASK the user.
- Probe
wc -l <script_source> storyboard.md so you know the line counts for verification later.
- Tell the user briefly what you're about to do, including which script source you'll read (clean Script/*.md or fallback transcript.md): load context → VIDIQ research → 5 titles → 9 thumbs (3 per A/B/C angle) → pick 1 of each angle → SEO description → comment → tags → ffmpeg → assemble UPLOAD.md. Show estimated VIDIQ budget (~15 credits or ~0 if screenshots) and time (~8–12 minutes wall, dominated by FlowGateway pacing + ffmpeg).
Phase 1 — Read episode artifacts FULLY
Use chunked Read calls with explicit offset+limit until you've consumed every line of the script source (whichever resolved in Phase 0 — clean Script/*.md preferred, transcript.md fallback) AND storyboard.md. Verify completion by comparing total lines read against wc -l from Phase 0.
Why prefer Script/*.md when available: the clean script has accurate proper-name capitalization, exact numbers, intentional punctuation, and the author's original sentence structure. The Whisper transcript (transcript.md) is auto-generated from the audio and routinely miscapitalizes researcher names ("dunning kruger" instead of "Dunning-Kruger", "lopez otin" instead of "Carlos López-Otín"), mangles institution names, and drops punctuation. For the description's sources block — where you cite researcher + year + institution verbatim — a script-sourced citation is dramatically more reliable than a transcript-sourced one.
While reading the script, note:
- The exact opening sentence (used in description Variant Banal-Action-Open hook restatement)
- Every named researcher, year, institution, study — these become description citations and lend authority to titles
- The "wait, what?" reveal beats — usually Beat 3 (~25–35% mark) and Beat 4 (50% mark) of the 7-Beat Arc. These become description high-tension bullets
- The closer's exact wording — the philosophical closer is the single most important paragraph for both channels. Often becomes a title (counter-narrative callback) and shapes the pinned comment
- One or two haunting micro-details — a name, place, object, a single image (e.g. "His brain looked like a sponge", "a centrifuge in the trunk of his car"). These become thumbnail concepts
- Which Reference hook formula the script uses — Banal-Action-Inverted, Universal-Claim-Staccato, Pattern-Break, Time-Zero, or Why-Hasn't-Anyone (Your Second Channel-only). Tells you which title family to lead with
While reading storyboard.md, note:
- The visual motifs already in the video (clocks, beds, DNA, hearts, microscopes, etc.)
- Which scenes carry the most weight (long holds, sequence chains)
- The episode's accent colors
Your thumbnail concepts should pull from these motifs and the channel's locked palette. Don't introduce visual elements the video itself doesn't reference — there's no need.
Phase 2 — Load channel pack + reference docs FULLY
Load every file listed in the MANDATORY READING PROTOCOL above for the detected channel + all 8 skill reference docs. After Phase 2, do a mental roll-call: did you read every required file end-to-end? If you can't say "yes" with confidence, go back and finish before Phase 2.5.
Phase 2.5 — VIDIQ in-browser agent screenshot check (interactive)
VIDIQ has a chat agent at vidiq.com that costs 25 external credits per session but performs deep multi-tool analysis (titles + keywords + tags + topics + outlier discovery) — significantly more output per credit than direct MCP. If the user has already chatted with this agent and has screenshots, the skill uses them and saves ~15 MCP credits.
Step 1 — Auto-detect screenshots:
Check for <episode_root>/vidiq_screenshots/*.png or *.jpg (or any image files in that dir). If they exist, proceed to Step 3 (screenshot ingestion) and tell the user you found them.
Step 2 — Ask the user (only if no screenshots found):
Ask the user (in their language; this user defaults to Russian):
"Before I spend VIDIQ MCP credits — did you already chat with VIDIQ's in-browser AI agent at vidiq.com for this episode? The chat agent costs 25 credits per session but gives much more data (titles + keywords + tags + topics) than direct MCP would for the same credits. If you have screenshots of the conversation, I'll use them and only spend ~5 MCP credits (just thumbnail scoring).
Three options:
-
YES, I have screenshots — drop them in vidiq_screenshots/ or attach to this conversation, then I'll continue.
-
NO, I want to do it now — here's the prompt template (paste into vidiq.com chat):
Hey, so I'm sending you a video transcript. I'd like you to come up with the best possible title for it — I already have one, but maybe you can do better. I'd also like you to pull together the right keywords to include in the description for SEO optimization, and put together a proper tag cloud. All of this is to maximize organic reach, keep the algorithm happy, and make some money. That's the idea.
Oh, and everything needs to be relevant. More to follow… find the right SEO words and phrases, and put together a solid tag cloud. Thanks in advance.
Title: \"<EXISTING TITLE>\"
Then paste/upload the transcript, screenshot every response, drop the screenshots in vidiq_screenshots/, and tell me when ready.
- NO, just use direct MCP — I'll spend ~20 credits across the 4 default VIDIQ phases."
Wait for the user's choice. If they go silent or say "go ahead", default to option 3 (direct MCP) — never block.
Step 3 — Ingest screenshots (if option 1):
For each screenshot in vidiq_screenshots/, read it with the vision tool and extract:
- Title suggestions (3–8 alternatives with VIDIQ's predicted CTR / search interest annotations)
- Keywords (10–30 phrases with search volume + competition data when shown)
- Tags (15–30 ready-to-paste tags for YouTube)
- Topics / niche commentary (outliers mentioned, what's working, what's saturated)
- Description hooks (2–3 example opening sentences if VIDIQ offered them)
Save the structured extraction to release/vidiq_research.json under "in_browser_agent_data" (see exact schema in references/vidiq_agent_prompt.md).
If a screenshot is unreadable or a section is missing, note partial extraction and fall back to direct MCP for that section only. Don't fail the whole skill on a partial screenshot batch.
Step 4 — Confirm the path and continue:
Record the chosen path in release/vidiq_research.json under top-level "source":
"in_browser_agent_screenshots" (option 1 — ~5 MCP credits)
"in_browser_agent_screenshots_partial" (option 1 with gaps — hybrid)
"direct_mcp" (option 3 — ~20 MCP credits)
The chosen path determines which subsequent phases call MCP vs use extracted data. Phases 3, 5, 11 follow the conditionals described below; Phase 8 (thumbnail scoring) always runs MCP regardless of path.
Phase 3 — VIDIQ keyword research (~5 credits, SKIP if screenshots ingested)
If Phase 2.5 ingested screenshots successfully, skip this entire phase. The keywords have already been extracted from the in-browser agent's response and live in vidiq_research.json under "in_browser_agent_data.extracted.keywords". Use those directly in Phase 4, 9, 11.
Otherwise (direct MCP path):
Pick the dominant searchable phrase for this episode. For The Dunning-Kruger Effect, that's "dunning kruger effect". For Why Hasn't Anyone Stopped Human Aging Yet?, that's "why we age" or "longevity research". The phrase should be 2–4 words that a curious viewer would actually type.
Call vidiq_keyword_research with that phrase. Save the full response to release/vidiq_research.json under key "keyword_research_primary". Increment credit counter by the actual cost reported by VIDIQ (or assume 5 if not reported).
From the response, extract:
- Top 10–15 related keywords with high search volume + low competition
- Suggested long-tail phrases that could anchor the description's first 150 chars
- Any "low competition / high volume" gems that hint at a contrarian angle
You'll use these in Phase 4 (titles), Phase 9 (description), and Phase 11 (tags).
Phase 4 — Draft 5 title candidates spanning 5 hook angles
If Phase 2.5 ingested screenshots successfully, the VIDIQ in-browser agent has already suggested 3–8 titles with CTR predictions. Use the top 3 from its suggestions as 3 of your 5 candidates. Then add 2 more from hook-angle families the VIDIQ agent didn't cover (to ensure spread). The other phases proceed normally — Phase 5 (scoring) is skipped because VIDIQ has already annotated the suggestions.
Otherwise (direct MCP path):
Don't draft 10 — that wastes VIDIQ credits. Draft 5, each from a different hook-angle family per references/title_formulas.md. The families are:
- Question-anchor (Reference 1.7M format) — "Why X" / "How Y" / "What If Z"
- Number-anchor (rarity-curiosity) — "Why 2% of Humans..." / "The 1 in 10 Million..."
- Counter-narrative (closer-as-title) — turns the script's philosophical closer into the hook
- Why-Hasn't-Anyone (Your Second Channel-only template) — "Why Hasn't Anyone X?"
- Universal-Experience (identity hook) — "You Have X (And You Don't Know It)"
Each candidate must:
- Be under 8 words
- Avoid clickbait punctuation (no all-caps, no "!!!", no clickbait Emoji)
- Use a question word OR a specific anchor OR a 2nd-person pronoun
- Pair with one of the dominant VIDIQ keywords from Phase 3 (so the algorithm can place the video correctly)
- Carry a curiosity gap that the video genuinely pays off
For Your Second Channel, at least 2 of the 5 candidates must use the "Why Hasn't Anyone X" template — that's the channel's signature.
Phase 5 — Score titles via VIDIQ, pick top 3 (~5 credits, SKIP if screenshots ingested)
If Phase 2.5 ingested screenshots successfully, the title CTR predictions already exist in the extracted data. Use VIDIQ's predicted CTR + Claude judgment (family spread) to pick top 3 directly. Skip the vidiq_score_title MCP calls. Record "score_source": "in_browser_agent" for each scored title.
Otherwise (direct MCP path):
For each of the 5 title candidates, call vidiq_score_title and save the response to release/vidiq_research.json under key "title_scores". Increment credit counter.
Sort by VIDIQ score. Pick the top 3 — with one constraint: ensure genuine spread across hook angles. If the top 3 by raw score are all "Question-anchor" titles, swap C for the highest-scoring title from a different family. A/B testing only works when the variants actually differ.
If a VIDIQ call fails (rate limit, API error), fall back to Claude judgment for that title and note it in vidiq_research.json with "score_source": "claude_fallback".
The 3 picks become titles A, B, C in UPLOAD.md:
- A — highest VIDIQ score (primary)
- B — strongest different-family candidate (contra-test)
- C — third best (variety)
Phase 6 — Draft 9 thumbnail concepts (Architect Methodology: 3 angles × 3 variants) + prompts.json
Thumbnails are the single highest-leverage CTR driver on YouTube. Approach this phase with maximum care. Read references/thumbnail_concepts.md fully — it covers the 4 Architect Pillars (CTR is the gate but watch-time-share is the judge, 4.5% educational-content baseline, topic-first beats scene-first, 3 radically different concepts beat 10 subtle variations), the 5 critical rules (whiteboard ONLY / 120px squint test / one dominant focal / max 3 elements + 4 words / honesty), the A/B/C angle matrix, the stickman expression vocabulary, the 7-step design process, prompt construction template, and the 11-question checklist.
The headline framework: every video gets exactly 3 final thumbnails — one of each angle from the A/B/C matrix:
- Angle A — Identity / Stakes: stickman in the viewer's position; relatable scene (mirror, bed, calendar); 2nd-person hook ("this is YOU")
- Angle B — Counter-Evidence / Reveal: a myth visibly contradicted; red X over a famous claim; debunked icon ("everything you've heard is wrong")
- Angle C — Iconic Metaphor: single bold symbolic image (hourglass, animal, organ, graph); minimal text; rewards a 1-second scan ("what IS that?")
To get to the final 3, we render 9 thumbs total — 3 variants per angle — and pick the strongest 1 of each angle in Phase 8.
The critical rule: every thumbnail is whiteboard stickman ONLY. NO photoreal. NO illustrated faces. NO cinematic mood lighting. NO 3D. Past test runs failed because the skill allowed photoreal alternatives; the picked thumbs broke the channel brand and didn't represent the videos honestly. Even if a CTR heuristic would predict photoreal scores higher — DO NOT generate it.
Topic-first workflow (Steps 1–3 happen BEFORE looking at storyboard):
-
Extract the core question from title A. "Why has nobody cured aging when some animals barely age at all?" / "Why are incompetent people so confident they're right?"
-
Identify the 3 strongest cognitive triggers for the topic: the most counterintuitive single fact, the highest-stakes consequence, the most visually distinctive metaphor.
-
Generate 9 concepts (3 per angle × 3 variants each), each occupying its angle's cognitive territory. The variants within an angle should differ on composition/color/subject — don't render 3 near-duplicates.
Variant Design Rules: the 3 FINAL picks (Phase 8) must vary on at least 2 of these 5 dimensions: angle (always covered by definition), emotion (shock vs curiosity vs confidence), text presence (heavy/light/none), subject framing (face vs object vs comparison), color dominance (red/blue/yellow-amber).
Default angle order for new channels (no A/B data yet): B → C → A. Counter-Evidence wins most often for science-myth-busting channels per published agency studies.
Each concept gets recorded in both:
release/thumbnails/shortlist/concepts.md — human-readable, with rationale per concept, angle label, cognitive trigger, variant dimensions
release/thumbnails/shortlist/prompts.json — machine-readable, consumed by the rendering script (9 prompts, IDs 01–09)
For each concept produce both:
release/thumbnails/shortlist/concepts.md — human-readable, with rationale per concept:
═══════════════════════════════════════════════════════════════════════════════
CONCEPT 01 — `<TEXT_OVERLAY>`
═══════════════════════════════════════════════════════════════════════════════
**Visual:** [3–5 sentence visual description]
**Why it works:** [2–3 sentences tying to a Reference winner or to the episode's strongest beat]
**Title pairing:** [Which of titles A/B/C this thumb pairs with]
release/thumbnails/shortlist/prompts.json — machine-readable for the rendering script:
[
{
"id": "01",
"name": "<TEXT_OVERLAY or short concept name>",
"prompt": "<full English image prompt — see prompt rules in references/thumbnail_concepts.md>"
},
...
]
Every prompt follows the 8-section Architect template (full details in references/thumbnail_concepts.md):
- Canonical style preamble (verbatim from
references/thumbnail_concepts.md)
- YouTube framing (16:9, 1920×1080, mobile-readable at 120px)
- Subject + action + emotion (with stickman expression details from the Expression Vocabulary table)
- Composition (position %, size %, focal hierarchy)
- Background (single flat color, explicit)
- Text overlay (exact words in quotes OR "NONE"; max 4 words; font style; position; never bottom-right)
- Palette emphasis (channel-specific colors named)
- Hard style constraint (verbatim from thumbnail_concepts.md, repeated at END — prevents photoreal drift since image models weight ending tokens heavily)
Each prompt under 2500 chars (FlowGateway/Grok/Nano Banana truncate beyond this). Self-contained (image model doesn't re-read context). Typical good prompt ~1200–1900 chars.
Phase 7 — Render 9 thumbnails (FlowGateway primary with auto-recovery, Replicate fallback)
Run the production script:
python "$SKILL_DIR/scripts/generate_thumbnails.py"
It reads release/thumbnails/shortlist/prompts.json from cwd, then for each of the 10 prompts:
-
First try FlowGateway (free via Google AI Pro): uv run flow generate "<prompt>" --aspect 16:9 --out release/thumbnails/shortlist/. The script attaches to the dedicated Chrome at localhost:9222. Wall time ~30s per image + 15–45s jitter pause between calls (FlowGateway's own pacing). ~5–7 minutes total for 10 images.
-
If FlowGateway is unreachable (Chrome on :9222 not responding), the script attempts browser auto-recovery:
- Try to start Chrome via
scripts/launch-chrome.sh from FlowGateway dir (nohup … &, detached)
- Poll
uv run flow check every 5 seconds for up to 60 seconds
- If Chrome comes up within that window, proceed with FlowGateway for all 9 prompts
- If recovery still fails after 60 seconds (likely sign-in/captcha required), fall back to Replicate
The user explicitly asked for this: "если браузер упал, нужно браузер заново запустить, а не переходить сразу на фоллбек". The fallback to Replicate is only used after auto-recovery has failed.
-
On FlowGateway error during a specific generation (captcha, quota, content-blocked, transient): fall back to Replicate Grok Imagine at $0.02/image (xai/grok-imagine-image) for that slot only. Subsequent slots will retry FlowGateway. Note the backend per slot in _generation_log.json.
Show the user the estimated cost upfront before launching:
- All FlowGateway success: $0 (within Pro subscription)
- All Replicate fallback: ~$0.20
After the script completes, Read at least 3 of the rendered PNGs with the vision tool to confirm:
- On-style whiteboard (thick black marker, flat fills, paper background)
- On-palette colours only (per channel: Your Channel navy/crimson/amber; Your Second Channel charcoal/ice-blue/brass)
- Text overlay 1–3 words, readable at 280×157 mobile size
- No anatomical/realistic hands or faces
- No stray AI watermarks
If a thumb is broken, re-run the script with --scenes <slot> --force to regenerate just that slot with a tightened prompt.
Phase 8 — Vision-review 9 thumbs, pick 1 of each angle (NO VIDIQ scoring — vision only)
Do NOT call vidiq_score_thumbnail. VIDIQ scoring is biased toward generic-YT photoreal patterns and would mis-rank within the whiteboard pool. The 5 saved credits drop the VIDIQ budget target to ~15.
After all 9 thumbs are on disk, Read every PNG with the vision tool. Apply the 11-question evaluation checklist from references/thumbnail_concepts.md:
- ✅ Whiteboard stickman style (thick black marker, flat fills, no photoreal)? HARD FAIL if no — regenerate
- ✅ Topic understood in 1 second at 120 px width? HARD FAIL if no — regenerate
- ✅ Sparks curiosity about the title's question?
- ✅ One dominant focal element (40–60% of frame)?
- ✅ MAX 3 visual elements + MAX 4 words of text?
- ✅ Universal hook (NOT a niche script entity)?
- ✅ On-palette colors only?
- ✅ Stickman (if present) follows canonical anchor (round head, closed mouth, dot eyes, mitten hands; no anatomical drift)?
- ✅ Important elements outside bottom-right 18% safe zone (YouTube duration badge)?
- ✅ Click-content fit: does it match what the video delivers in the first 30s?
- ✅ Differs from other 8 on at least 1 dimension (composition, color, subject, emotion)?
For any HARD FAIL (1, 2), regenerate the slot:
python "$SKILL_DIR/scripts/generate_thumbnails.py" --scenes <slot> --force
Picking the final 3: within each angle (A, B, C), rank the 3 variants by Claude vision judgment on:
- Curiosity strength — which makes you ASK the title's question hardest?
- Glance test speed — which is readable fastest at 120 px?
- Visual punch — high-contrast, single-focal, distinctive
- Click-content fit — matches what the video delivers
Pick the single strongest variant per angle: 1 of (01, 02, 03) for Angle A, 1 of (04, 05, 06) for Angle B, 1 of (07, 08, 09) for Angle C. Watch-time-share (the metric YouTube's Test & Compare actually uses) rewards click-content alignment over raw clickbait.
Verify the Variant Design Rules: the final 3 picks must vary on at least 2 of these 5 dimensions: angle (covered), emotion, text presence, subject framing, color dominance. If the strongest 3 are still too similar (e.g., all use red as dominant color), swap one for the second-strongest from a different family.
Copy final picks:
- Strongest of (01/02/03) →
release/thumbnails/A.png
- Strongest of (04/05/06) →
release/thumbnails/B.png
- Strongest of (07/08/09) →
release/thumbnails/C.png
Record which shortlist slot each came from in _generation_log.json and release/vidiq_research.json under "thumbnail_selection" (with "score_source": "claude_vision_only").
Phase 9 — SEO-optimized description
Read references/description_seo_template.md fully. Draft ONE description (not a variant set — YouTube can't A/B test descriptions).
Keyword source priority:
- If Phase 2.5 ingested screenshots → use VIDIQ in-browser agent's keyword list + description hooks (if VIDIQ offered them, they're often very good as Component 1 openers)
- Else (direct MCP path) → use Phase 3's
vidiq_keyword_research response
Structure:
- First 150 chars — the script's literal opening sentence, rephrased to include the dominant VIDIQ keyword from Phase 3. YouTube's algorithm classifies the video from these chars; they matter more than any other 150 chars in the description.
- Hook restatement — 2–3 sentences naming the question the episode answers.
- Bullets — 4–6 high-tension bullets from the episode's Beat 3 and Beat 4 reveals.
- Sources block — every named researcher + year + institution from the transcript. Verbatim from the script. NEVER fabricate.
- Chapters — 3–6 timestamp chapters (00:00, 02:14, etc.) based on the storyboard's natural beats. This unlocks YouTube's chapter feature and improves session retention.
- WATCH NEXT — links to 1–2 other channel videos (use
[https://youtu.be/PLACEHOLDER] placeholders, user fills in actual URLs).
- Business email —
business@your-channel.example.com (only if the user hasn't said "no email" — check the channel context for the rule).
- Hashtags — 3 hashtags tied to the dominant topic. Format:
#hashtag on a single trailing line.
Embed the VIDIQ-researched keywords naturally throughout the body — don't keyword-stuff. The first 150 chars must contain the primary keyword. The bullets should contain 2–3 secondary keywords. Total description ≈ 400–550 words.
Phase 10 — Pinned comment
Read references/comment_patterns.md fully. Draft ONE pinned comment using Reference's highest-converting pattern: 1–2 conversational lines, often lowercase, occasional 👀 / 🛑 emoji, ending with a 🛑 WATCH NEXT link to another channel video.
The comment must:
- Open with "anyone else who…" or "wait, am I the only one who…" — the highest-converting opener in Reference's catalog
- Reference a very specific detail from the script (a name, a number, a haunting image)
- End with a
[https://youtu.be/PLACEHOLDER] line — user fills in the URL
Example (from prototype): "anyone else who wakes up at 2-3 a.m. for no reason and just lies there? 👀 🛑 WATCH NEXT: [https://youtu.be/PLACEHOLDER]"
Phase 11 — VIDIQ tag cloud (~5 credits, SKIP if screenshots ingested)
If Phase 2.5 ingested screenshots successfully, the VIDIQ in-browser agent has already produced a tag cloud. Use those tags directly. Skip the vidiq_keyword_research MCP call. Just curate the extracted list to fit YouTube's 500-char total limit (drop redundant or weak tags if over). Apply the tag-quality rules below to pick the final set.
Otherwise (direct MCP path):
Take the chosen title A (from Phase 5). Call vidiq_keyword_research on its exact phrasing. Save to release/vidiq_research.json under "keyword_research_title". Increment credit counter.
From the response, select 12–18 tags by these rules:
- Include the exact title (lowercased) as tag 1
- Include 3–4 broad-topic tags (e.g. "psychology", "human behavior")
- Include 4–6 specific-phrase tags from VIDIQ research (high search + low competition)
- Include 2–3 channel-identity tags ("your channel", "your channel" for Your Channel; "your_second_channel", "unsolved mysteries" for Your Second Channel)
- Include 1–2 broad audience-anchor tags (e.g. "self-improvement", "documentary")
- AVOID single-word generic tags ("video", "youtube", "learn") — they signal spam to the algorithm
Output as a single comma-separated line ready to paste into YouTube's Tags field. YouTube has a 500-char total limit on tags; verify the line fits.
Phase 12 — ffmpeg upload prep (auto-run, ~3–10 min)
Run the production script:
python "$SKILL_DIR/scripts/ffmpeg_upload_prep.py"
It:
- Finds
<episode_root>/output/final.mp4. If absent, finds the most recent final_*.mp4 in output/.
- Applies the 2026 Faceless YouTube Playbook recipe (per
references/ffmpeg_pipeline.md):
- Upscale to 4K (3840×2160) via lanczos if source is <2160p (the playbook's #1 free quality win)
- H.264 High Profile, CRF 18, CABAC, 2 B-frames, closed GOP at half-fps
- AAC-LC, 384 kbps, 48 kHz, stereo
- BT.709 color, progressive scan
+faststart for fast browser playback
-map_metadata -1 -map_chapters -1 to remove all upstream metadata
- Clean encoder string
- Outputs
release/upload.mp4.
- Logs duration + input/output bitrate + file size to
release/ffmpeg_log.json.
Show progress to the user (ffmpeg runs single-threaded for stability; ~3 min for a 12-min 1080p source on Apple Silicon).
Phase 13 — Assemble release/UPLOAD.md
Read references/upload_md_format.md fully — it has the exact template. The structure has 8 blocks delimited by ═══ heavy bars, each block titled and labelled so the user can copy the contents below the header without selection ambiguity:
- CHECKLIST — pre-upload checklist (Altered content = NO for whiteboard; playlist; made for kids = no)
- TITLE A — primary (with VIDIQ score noted)
- TITLE B — contra-test
- TITLE C — variety-test
- THUMBNAILS — paths to A/B/C with rationale
- DESCRIPTION — paste-ready 400–550 words
- PINNED COMMENT — paste-ready 1–2 lines
- TAGS — single comma-separated line ready to paste
- WHY THESE PICKS — 1–2 paragraphs of reasoning (skip if user wants minimum noise)
- VIDIQ CREDIT USAGE — tally for transparency
Use ═══ (heavy double-line) between blocks. Use ─── (light line) for sub-sections inside a block. Each block header is centred inside the heavy bars so it's instantly scannable.
Phase 14 — Self-check before reporting
Walk through this checklist. Every box must be ticked. If any fails, fix it BEFORE reporting.
Reading completeness:
Coverage:
Quality:
Budget:
If any check fails: redraft, regenerate, or rerun the script. Don't paper over with an apology.
Phase 15 — Report to user
Write a short reply (8–12 lines max) telling the user:
- ✓ what was generated and where (point at
release/UPLOAD.md and release/upload.mp4)
- Total time / total cost (FlowGateway free + Replicate fallback if any + VIDIQ credits used)
- One-line starter-triplet recommendation (which of A/B/C to publish primary)
- Suggest they open
release/thumbnails/ in Finder to eyeball A/B/C side by side
- Note any thumbnail that came out below par (if vision review flagged one)
- Reminder that "Altered content disclosure" must be No in YouTube Studio
Do NOT paste long content into chat — the files are on disk and the user opens UPLOAD.md to copy.
Failure modes & guardrails
-
Don't fabricate sources. If you can't pin a study to a researcher + year + institution from the transcript, drop it from the description. False citations kill channel credibility.
-
Don't break the channel voice. Both channels' voice is forensic, calm, 2nd-person. No "let me", no "in this video, we'll show you", no caps-lock, no commercial CTAs.
-
Don't blow the VIDIQ budget. 70-credit hard cap. Target ~20. If you cross 50 with phases remaining, switch to Claude-only judgment for further scoring.
-
Don't skip ANY mandatory-load reference file. Alex has explicitly required all of them loaded fully on every invocation.
-
Don't ship thumbnails before vision-reviewing them. Read all 9 rendered PNGs after the script completes. If any HARD-FAIL the 11-question checklist (whiteboard/glance/text-max), regenerate.
-
Don't auto-overwrite an existing release/. Ask the user before overwriting anything.
-
Don't paste images or long markdown into chat. Tell the user to open release/UPLOAD.md and release/thumbnails/.
-
Don't mark "Altered content disclosure". Stickman/whiteboard explainer is explicitly exempt per the 2026 YouTube Playbook. Tell the user No in the checklist.
-
Don't run FlowGateway in parallel. It's single-threaded (Chrome over CDP). Concurrency=1 always.
-
Don't run two /video-final-pack invocations at the same time. They will collide on FlowGateway pacing state.
Reference index (mandatory-load)
| File | Size | Purpose |
|---|
references/title_formulas.md | ~9KB | 10 hook-angle families, VIDIQ scoring strategy |
references/description_seo_template.md | ~8KB | SEO description structure + keyword embedding rules |
references/comment_patterns.md | ~8KB | 10 Reference pinned-comment families |
references/thumbnail_concepts.md | ~10KB | Per-channel composition families, image-prompt rules |
references/ffmpeg_pipeline.md | ~9KB | Exact 4K H.264 recipe + rationale |
references/vidiq_budget.md | ~7KB | 70-credit cap, 20-target, per-call estimates |
references/upload_md_format.md | ~17KB | Exact UPLOAD.md template with separator examples |
references/vidiq_agent_prompt.md | ~9KB | In-browser agent prompt template + screenshot ingestion + credit-saving path |
assets/your_channel/* | ~210KB | Your Channel channel context + competitor data + script template |
assets/your_second_channel/* | ~25KB | Your Second Channel channel context + Reference Channel-Titanic template breakdown |
Treat references and assets as authoritative. If you find yourself improvising on title length, description structure, thumbnail composition, or VIDIQ usage, go back and re-read the relevant file. It's there for a reason.
Skill dir resolution
To find the skill dir from the episode root:
SKILL_DIR=$(python3 -c 'from pathlib import Path; print((Path.home() / ".claude/skills/video-final-pack").resolve())')
Or hard-code it — the skill lives at ~/.claude/skills/video-final-pack/.