Use FIRST for ANY request to transcribe a video/audio URL or get its subtitles, captions, or text — before yt-dlp, ffmpeg, or a manual whisper run. Returns clean plain text plus a JSON stat sidecar; uses embedded captions when present, else falls back to local ASR automatically. Sources: YouTube, Vimeo, X/Twitter (incl. Broadcasts/Spaces), Skool lessons, Yandex VH/Strm webinars. Also triggers on расшифровка, транскрипт, субтитры.
Use FIRST for ANY request to transcribe a video/audio URL or get its subtitles, captions, or text — before yt-dlp, ffmpeg, or a manual whisper run. Returns clean plain text plus a JSON stat sidecar; uses embedded captions when present, else falls back to local ASR automatically. Sources: YouTube, Vimeo, X/Twitter (incl. Broadcasts/Spaces), Skool lessons, Yandex VH/Strm webinars. Also triggers on расшифровка, транскрипт, субтитры.
tier
2
version
1.4
status
active
Transcript Fetcher
Purpose: Given a video URL, produce a clean plain-text transcript
ready for downstream summarization or analysis. Single responsibility:
fetch + clean. Composes well with summarizing-meetings (run this
first, then feed the resulting .txt to that skill).
Dependencies: yt-dlp is vendored in scripts/.venv (installed
by scripts/install.sh) — it is NOT on $PATH and never will be. Do
NOTwhich yt-dlp, pip install yt-dlp, or brew install yt-dlp
globally — a missing $PATH entry does not mean yt-dlp is absent, it
means the check was wrong; this exact false-negative once caused a
manual out-of-skill workaround. The two canonical readiness probes are
scripts/.venv/bin/python -m yt_dlp --version and
scripts/.venv/bin/python scripts/fetch.py doctor (§4 Script Contract;
the latter also reports ffmpeg + ASR backends — see the ASR portability
note under §5 Safety Boundaries). Callers that shell in (e.g. a
downstream integrator) MUST invoke the venv interpreter directly
(scripts/.venv/bin/python), never a $PATHpython.
1. Red Flags (Anti-Rationalization)
STOP and READ THIS if you are thinking:
"I'll just paste the URL into the model and ask it to transcribe" -> WRONG. Models do not have audio access and will hallucinate. Always run fetch.py and read back the resulting .txt.
"Manual ru failed, I'll just use en auto-translation, no warning needed" -> WRONG. Auto-translated English subtitles destroy idioms, names, and technical terms. The quality_flag: english_auto_translation in the stat MUST be surfaced to the user.
"I'll skip writing the JSON stat sidecar, the .txt is enough" -> WRONG. The sidecar records WHICH track was picked. Without it, downstream cannot tell whether the transcript is high-quality manual subs or low-quality auto-translation.
"The transcript has weird >> markers, I'll strip them" -> WRONG. Those are speaker-turn boundaries. Removing them collapses multi-speaker meetings into a single voice and ruins downstream attribution.
"Auto-generated Russian (ru-orig) is garbage, I'll prefer en instead" -> WRONG. ru-orig is the actual Russian audio transcribed; ru (without -orig) is often an English auto-translation back to Russian. ru-orig > ru > en.
"This talk is only on a proprietary player, so I'll ASR it" -> CHECK FOR A MIRROR FIRST. Conference and webinar recordings are routinely re-published by the same organiser on YouTube with auto-captions, even when the marketing page embeds its own player. Search the organiser's channel for the exact talk title before spending hours of ASR. Real case: a 9-talk Yandex AI Studio Series page embedded a caption-less Yandex VH/Strm player, but all 5 underlying streams were on the official Yandex Cloud YouTube channel with full-coverage ru-orig captions — ~7.6 h of ASR avoided. Prefer the mirror for cost, not quality: a spot-check on the same talk found the local MacWhisper output was actually cleaner than the YouTube captions (no rolling-caption duplication), it just costs hours. ASR is the fallback, not the opening move.
"I'll add ffmpeg to the pip deps just in case" -> WRONG. The caption path (WebVTT parsing) is pure Python and needs nothing extra; ffmpeg is a soft-optional external system tool, never a pip dependency. It IS genuinely required for the X ASR path on HLS sources (Broadcasts/Spaces) — yt-dlp uses it to extract a clean audio-only , and the skill fails fast (exit 7) when it is absent there — but it is detected at runtime (), not bundled. Do not pull heavy packages into .
2. Capabilities
Fetch YouTube and Vimeo captions via yt-dlp (no audio download).
Fetch X.com / Twitter — native status video AND Broadcasts/Spaces.
The X provider is captions-first: it reuses embedded
subtitles/automatic_captions when present, and only when none
exist does it download the smallest media and transcribe via ASR.
Fully automatic — no mode switch. ASR runs through a pluggable
backend chain: MacWhisper (mw) → Whisper CLI → whisper.cpp →
opt-in OpenAI/compatible cloud. ffmpeg is required for the X ASR
path on HLS sources (Broadcasts/Spaces): with it the smallest media
is extracted to a clean audio-only m4a; without it the skill fails
fast (exit 7) because yt-dlp's no-ffmpeg HLS output is not a valid
container the ASR engine can open. (For non-HLS progressive media, or
when embedded captions exist, ffmpeg is not needed.)
Fetch Yandex VH/Strm recordings (runtime.strm.yandex.ru/player/episode/<ID>,
frontend.vh.yandex.ru/player/<ID>) — the player behind Yandex Cloud
webinars and streams. ASR-only by design: that player carries no
caption track at all (verified five ways — no subtitle key in the config
JSON, zero #EXT-X-MEDIA tags in the HLS master, only video+lang="rus"
audio AdaptationSets in DASH, yt-dlp --list-subs reports none, and a
live player reports video.textTracks: []), so a caption ladder would be
dead code. yt-dlp has no extractor for this host, but its generic
extractor takes the signed manifest directly — the adapter is therefore
only a resolver (episode id -> config JSON -> signed DASH/HLS URL) and
reuses the shared download/ASR pipeline unchanged. No auth: the config
endpoint answers with zero headers. Two live traps the adapter handles —
the config's duration field lies (observed 4365 vs a real 3800 s,
and 43970 vs a real 8825 s), so the real duration is the #EXTINF sum
from the media playlist; and signed URLs are minted per request and
expire (~48 h), so they are never cached. ffmpeg is required here
(HLS/DASH-only source); the skill fails fast with exit 7 without it.
Before using this path, check for a YouTube mirror — see §1.
Fetch Skool lesson pages via a stdlib HTML scrape — public
communities work without auth; private/paid ones accept an optional
Netscape . Then embedded YouTube/Vimeo
videos to those adapters; capture author-supplied transcript field
when present.
3. Execution Mode
Mode: script-first
Why this mode: Fetching captions, parsing WebVTT, deduplicating rolling captions, and applying the fallback ladder are deterministic operations with > 5 lines of business logic. The CLI is the contract; SKILL.md is orchestration.
4. Script Contract
Install (one-time): creates the venv + yt-dlp, then reports which optional ASR components are present:
bash skills/transcript-fetcher/scripts/install.sh
# Optional ASR engines (for caption-less X media) — detect / install:
./scripts/.venv/bin/python scripts/install_components.py # status report
./scripts/.venv/bin/python scripts/install_components.py --install-whisper # pip openai-whisper into the venv
./scripts/.venv/bin/python scripts/install_components.py --system --run # brew/apt ffmpeg + whisper.cpp
Readiness check (doctor, no network, import-free): answers "is this skill ready?" without $PATH guessing — the canonical replacement for which yt-dlp:
./scripts/.venv/bin/python scripts/fetch.py doctor # human-readable report
./scripts/.venv/bin/python scripts/fetch.py doctor --json # {v, interpreter, in_venv, ready, components, remediation}
Reports the resolved interpreter, in-venv flag, yt-dlp version, ffmpeg, each ASR backend, and cloud opt-in state (key presence only, never the key). remediation names a flow-blocking gap for EVERY such gap — yt-dlp missing, ffmpeg missing (it back-stops X Broadcast/Space HLS ASR and whisper/whisper.cpp at runtime even though it is itself optional), or no ASR capability at all (no local backend AND no fully-configured cloud) — but never an individual missing ALTERNATIVE local ASR engine while ASR capability already resolves elsewhere (mw present, another local engine present, or cloud configured); those show only as informational → install-hint lines in the human report. A fully-configured cloud backend (--asr-allow-cloud + key) genuinely suppresses the no-local-ASR hint — it never appears in remediation, and the human report gets one informational note instead of a demand to install a local engine it does not need. Exit 0 when yt-dlp is present (regardless of remediation), 7 when it is not.
Single URL:
cd skills/transcript-fetcher
./scripts/.venv/bin/python scripts/fetch.py <URL> --out <path/to/output.txt>
Optional flags: --lang ru (default), --prefer manual|auto (default manual), , , , , (stage logging to stderr), (opt-in cloud ASR), , , (X: transcribe only the first N minutes — clips the download when ffmpeg is present, the only case where the download itself is clipped; also clips the media-timeout floor below in that case), (X ASR: do NOT strip long silences before transcription — silence removal is ON by default to cut Whisper hallucinated filler on silent lead-in/out), (per-host cookies), (X: load cookies from a local browser via yt-dlp), (X: parallel HLS fragment downloads for the media, default 8; = serial; values above 32 are capped at 32; CLI rejects with exit 2; env non-positive/malformed falls back to 8 — these are three DISTINCT layers, not one shared clamp), (X: per-attempt budget for the media download only, separate from the probe's ; default duration-derived s — capped at 6h, and derived from the -clipped duration when that flag is set AND ffmpeg is present (the only case where the download itself is clipped) — else s).
5. Safety Boundaries
Allowed scope: Reads from the network (yt-dlp HTTPS to YouTube/Vimeo; stdlib HTTPS to Skool). Writes ONLY to the user-specified --out / --out-dir path plus a .stat.json sidecar (and optionally a .description.md sidecar) next to it.
Default exclusions: Never downloads the video itself (only --skip-download + subtitle tracks / --write-info-json for metadata). Never writes to any path the user did not explicitly pass via --out / --out-dir.
Destructive actions: None. The script does not delete or modify any pre-existing files outside the chosen output paths. In batch mode, output path collisions are handled per --on-collision={error,skip,suffix} (default: error).
Optional artifacts: The JSON stat sidecar is mandatory in single-URL mode (it is the audit trail for which track was used). The .description.md sidecar is written only when --with-description is passed.
URL allowlist: Source dispatch is hostname-based against an explicit allowlist:
X / Twitter — x.com, www.x.com, mobile.x.com, twitter.com, www.twitter.com, mobile.twitter.com (status …/status/<id> and …/i/broadcasts/<id>).
Skool — skool.com, www.skool.com, app.skool.com; additionally URLs must match /<community>/classroom/<id>?md=<lesson-id>. Landing / /about / /calendar pages are rejected.
Yandex VH/Strm — runtime.strm.yandex.ru, frontend.vh.yandex.ru, strm.yandex.ru (episode …/player/episode/<id>). Other hosts (e.g. ) are NOT routed here.
URLs that merely contain a supported host as a substring elsewhere are rejected.
6. Validation Evidence
Local verification:
cd skills/transcript-fetcher
./scripts/.venv/bin/python -m unittest discover -s scripts/tests
All offline tests must pass without network. The end-to-end network test is gated behind TRANSCRIPT_FETCHER_E2E=1.
auto:ru-orig — YouTube auto-captions of the original Russian audio (good).
auto:ru — YouTube auto-translation TO Russian (often noisy if speech was in another language).
auto:en — English auto-captions as last resort (will set quality_flag = english_auto_translation).
For non-Russian content, pass --lang en (or another ISO code). For
non-Russian languages the lang-orig step is skipped — it is a YouTube
quirk that mainly matters for non-English speech.
Step 4: Run the CLI
Capture stdout (it carries the JSON stat). Read the stat to confirm
which track was used:
Open the generated <out>.txt.stat.json. If quality_flag is set
(currently only "english_auto_translation"), surface a warning to
the user before passing the transcript to a downstream summarizer:
⚠️ TRANSCRIPT QUALITY: only English auto-translation was available
for this URL. Idioms, proper names, and technical terms may be
distorted. Consider asking the user for a manual transcription.
Step 6: Hand off
The clean .txt is now ready for summarizing-meetings or any other
downstream consumer. Pass the path; do not paste the contents inline
(transcripts are often large).
8. Workflows
- [ ] Verify scripts/.venv/ exists (run install.sh otherwise)
- [ ] Decide single vs batch mode
- [ ] Run fetch.py with the chosen language and preference
- [ ] Read the JSON stat sidecar
- [ ] Surface quality_flag warning if set
- [ ] Hand .txt path to downstream skill
9. Best Practices & Anti-Patterns
DO THIS
DO NOT DO THIS
Always read the stat sidecar after fetching
Trust the .txt without checking which track was used
Prefer ru-orig over ru for Russian content
Pick the first track that returns text
Pass batch URLs through a file
Loop the CLI shell-side with arbitrary URLs
Surface quality_flag in any user-visible output
Silently downgrade to English auto-translation
Use --json-errors in CI/automation pipelines
Parse free-form stderr
Rationalization Table
Agent Excuse
Reality / Counter-Argument
"yt-dlp is on $PATH so I can just call it"
The skill invokes python -m yt_dlp from the per-skill venv. The system yt-dlp may be a different version with different output.
"The >> markers are clutter"
They are paragraph breaks for speaker turns. Downstream summarizers attribute statements by them.
"Rolling-caption dedup is overkill, just keep the longest cue"
The dedup IS keeping the longest cue. Without it, the same sentence would appear 3-4 times.
"I'll add a Vimeo adapter inline in fetch.py"
Add it as scripts/sources/vimeo.py. Each source is its own file.
10. Examples
See examples/:
example_input_url.txt — batch input format.
example_output_plain.txt — what a cleaned transcript looks like (excerpt).
example_output_stat.json — what the stat sidecar contains.
Composes well with summarizing-meetings: run transcript-fetcher first to get a clean .txt, then pass that file to summarizing-meetings for a structured Markdown summary. The two skills are intentionally separate — fetching is a deterministic file operation; summarizing is a prompt-first reasoning task.
m4a
install_components.py
requirements.txt
cookies.txt
delegate
Fall back through a configurable ladder (default for ru: manual ru -> auto ru-orig -> auto ru -> auto en).
Clean captions to plain text — WebVTT, and for X also SRT and TTML/DFXP (vtt/srt/ttml/best preference list, so a non-VTT track is no longer skipped to ASR): strip timestamps, inline timing tags, and rolling-caption overlap; decode HTML entities. TTML is parsed safely (DTD/entity declarations refused — XXE/billion-laughs guard). For X, captions-first is language-robust: if the requested --lang has no track but the post carries captions in another language, those are used (manual preferred, with a note) rather than dropping to ASR.
Preserve>> speaker-turn markers as paragraph breaks.
Emit a JSON stat sidecar (chosen track, char count, speaker-turn count, quality flag, plus optional title/uploader/duration metadata). For X media it also records transcript_origin (embedded-captions | macwhisper | whisper-cli | whisper-cpp | openai-api) so downstream skills know HOW the text was produced.
Optionally write<out>.description.md (YAML frontmatter +
Markdown body) when --with-description is passed — gives you a
ready-to-ingest description for RAG / Obsidian / human review.
Batch mode for processing multiple URLs from a text file.
Source-agnostic architecture: each platform is one file under scripts/sources/. The yt-dlp + ASR pipeline is shared (sources/_ytdlp_media.py + asr/), so a future TikTok/Twitch/Vimeo-ASR provider is one new file + one host entry — no pipeline changes. Zoom and podcast slots remain reserved.
Configurable + secrets-safe: a skill-local .env (see scripts/.env.example) externalises every endpoint, model, and tool path. The cloud ASR endpoint works with any OpenAI-compatible server (Groq, self-hosted whisper). A .env holding an API key is refused unless chmod 600 (and not a symlink); the key is sent only in an HTTP header, never on argv or in logs.
--with-description
--description-only
--cookies-file PATH
--json-errors
--debug
--asr-allow-cloud
--asr-model <id>
--asr-timeout-sec N
--max-duration-min N
--keep-silence
--auth-map PATH
--cookies-from-browser BROWSER
--concurrent-fragments N
1
<= 0
TRANSCRIPT_FETCHER_CONCURRENT_FRAGMENTS
--media-timeout-sec N
--timeout-sec
min(21600, max(600, duration*4))
--max-duration-min
1800
X.com / Twitter (captions-first, automatic ASR fallback; no mode switch). A Broadcast/Space usually has no captions → ASR via the first available local backend (MacWhisper, etc.):
cd skills/transcript-fetcher
./scripts/.venv/bin/python scripts/fetch.py \
"https://x.com/i/broadcasts/<id>" \
--out broadcast.txt --with-description --debug
# → broadcast.txt + .stat.json (source="x", transcript_origin="macwhisper",# chosen_track_kind="asr"). A status video WITH captions skips ASR# (transcript_origin="embedded-captions"). Use --cookies-file for# protected/age-gated media, or drop a Netscape cookies.txt at# ~/.transcript-fetcher/x.com-cookies.txt for zero-flag auth# (see the X cookie contract in §5 Safety Boundaries).
Skool lesson (cookies needed ONLY for private / paid communities — public ones work without):
Inputs: A YouTube / Vimeo / Skool-lesson URL (or a file with one URL per line). Empty lines and # comments in the batch file are ignored.
Outputs:
<out>.txt — clean plain text, UTF-8.
<out>.txt.stat.json — sidecar with the chosen track, quality flag, plus optional title / uploader / upload_date / duration_sec / embed_source / embed_url metadata.
<out>.description.md — only when --with-description is passed; YAML frontmatter + Markdown body.
One JSON stat record per URL on stdout.
Failure semantics: Non-zero exit. With --json-errors, stderr carries a single JSON line {v, error, code, type, details?}. Exit codes: 2 usage error (incl. malformed Skool URL, cookies file path missing on disk), 3 no transcript producible (no caption track in the ladder AND — for X — ASR produced nothing / every available backend failed, or the media download timed out — that case is transient/retryable and its remediation names --concurrent-fragments / --media-timeout-sec), 4 partial batch failure, 5 source-auth error (HTTP 401/403 — private Skool community needs cookies, X protected/suspended/age-gated media, or supplied cookies expired), 6 source rate-limit (HTTP 429), 7 missing dependency (yt-dlp absent, ffmpeg required-but-absent, or no ASR backend available for caption-less media — details.remediation carries the hint), 1 unexpected. When details carries a remediation key, it is ALSO printed as a second, plain stderr line (remediation: <text>) even WITHOUT --json-errors — the remedy is operator-visible either way.
Idempotency: Re-running overwrites the output file and sidecar. yt-dlp itself caches nothing the skill depends on; behaviour is reproducible given network availability.
Dry-run support: Not currently exposed as a flag. Inspect the fallback ladder via _build_ladder if needed.
*.yandex.ru
music.yandex.ru
Second-hop allowlist (Yandex only): the Yandex adapter resolves a stream URL out of a remote JSON document and hands it to yt-dlp, so that URL is untrusted input and is separately gated — https-only, against an exact-host set (strm.yandex.ru, runtime.strm.yandex.ru, frontend.vh.yandex.ru, vh.yandex.ru). Deliberately not a *.yandexcloud.net suffix rule: that is object storage where anyone can create a bucket, which would turn the allowlist into an open redirect. Every HTTP hop is fetched through the shared restricted opener, so a cross-host redirect is refused too (the pre-flight check on the original URL alone guarantees nothing about where the bytes came from), and each document read is capped at 8 MB.
Silence removal also runs on the Yandex ASR path (same default-on behaviour and --keep-silence opt-out documented for X below).
ASR backends (external, optional): For caption-less X media the skill shells out (argv arrays, never a shell string) to whichever local engine is present — MacWhisper mw, Whisper CLI, or whisper.cpp. None is a pip dependency; all are probed at runtime. ffmpeg (also external) is required to turn an X Broadcast's HLS stream into a valid audio file; the skill fails fast with exit 7 (clear remediation, before any large download) when ffmpeg is absent for an HLS source. bash scripts/install.sh reports which engines are available; scripts/install_components.py guides/installs them (incl. ffmpeg). If no ASR backend is available (and cloud is not opted in), the run also fails cleanly with exit 7, never a traceback. ASR portability: the fallback chain resolves in order mw → Whisper CLI → whisper.cpp → (opt-in) cloud (§2 Capabilities) — a caption-less Broadcast/Space genuinely REQUIRES ffmpeg and at least one of these; a box with neither (e.g. a bare Linux/CI runner) fails hard on exit 7. Remediate with scripts/install_components.py --install-whisper (in-venv Whisper CLI; ffmpeg itself needs the separate --system --run) or --asr-allow-cloud (+ an API key) to fall back to the cloud backend. Run scripts/fetch.py doctor (§4 Script Contract) before a long fetch to see which backends resolve, with zero downloads.
Cloud ASR egress (opt-in only): The OpenAI/compatible cloud backend is used only with --asr-allow-cloud (or TRANSCRIPT_FETCHER_ASR_ALLOW_CLOUD=1) AND an API key present. When used, the audio leaves the machine to the configured endpoint — disclosed here and in the stat notes. Local backends are always tried first; cloud is the last resort.
Silence removal before ASR (X): before transcribing, the X path runs ffmpeg silenceremove to trim leading silence and collapse long interior/trailing gaps — this cuts Whisper-family hallucinated filler (e.g. "Продолжение следует...") on silent lead-in/out. ON by default; --keep-silence (or TRANSCRIPT_FETCHER_SILENCE_REMOVAL=0) opts out; _THRESHOLD/_MIN_GAP_SEC/_KEEP_SEC tune it. Only true silence is removed (music/speech survive — a music-only intro can still trigger filler, see KNOWN_ISSUES TF-X-6). Never fatal: ffmpeg absent or a filter failure transparently falls back to the original audio. The stat notes record what was stripped (silence-removal: stripped ~Ns ...); the original media is kept for the ffprobe duration fill.
Secrets: The API key is read from OPENAI_API_KEY / TRANSCRIPT_FETCHER_OPENAI_API_KEY or a skill-local .env. A .env is loaded only at the CLI entry point and is refused if it is a symlink or not chmod 600 (group/world-readable). The key is sent only in an HTTP Authorization header — never on a command line, never logged. .env is git-ignored; only scripts/.env.example (placeholders) is committed.
Per-host cookies (~/.transcript-fetcher/): cookies for auth-walled media resolve (after an explicit --cookies-file) from a skill-local home folder — mirrors the html skill's ~/.html. An auth-map.json (--auth-map / TRANSCRIPT_FETCHER_AUTH_MAP / ~/.transcript-fetcher/auth-map.json) maps a host to its {cookies_file}, or the convention ~/.transcript-fetcher/<host>-cookies.txt is used — e.g. ~/.transcript-fetcher/x.com-cookies.txt for https://x.com/... URLs, ~/.transcript-fetcher/twitter.com-cookies.txt for https://twitter.com/... URLs. The convention lookup tries the EXACT URL hostname first, then — ONLY for the three well-known mirror-prefix labels www/mobile/m — the same file with that single label stripped (e.g. www.x.com and mobile.x.com both also resolve x.com-cookies.txt when the exact-host file is absent); distinct domains are never aliased to each other (x.com and twitter.com still need separate files — no generic parent-domain walk). A custom filename (e.g. x-cookies.txt) REQUIRES an auth-map.json entry; the convention path only matches the literal <host>-cookies.txt name (or its single-label-stripped mirror variant). Host match is label-boundary (a key x.com matches x.com/*.x.com, never evil-x.com); auth-map and convention files are hardened (symlink-reject + 0600). The resolved Netscape cookies.txt feeds yt-dlp's --cookies (and Skool's opener). --cookies-from-browser BROWSER loads cookies straight from a local browser via yt-dlp (opt-in — reads the browser's cookie store). On an X auth failure (SourceAuthError, exit 5) the message names the refresh path: the resolved --cookies-file when one was supplied, else the convention path to create — derived from the failing URL's own host (www./mobile. labels stripped), e.g. ~/.transcript-fetcher/x.com-cookies.txt for an x.com URL or ~/.transcript-fetcher/twitter.com-cookies.txt for a twitter.com URL; the convention lookup's mirror-prefix fallback above guarantees this hinted path is actually picked up on retry, for all 6 documented X hosts.
Temp-file hygiene: For X media, all intermediates (audio, VTT, info.json, .part, .m3u8, fragments) live under one tempdir removed in a finally block even on error — nothing is left behind.
Auth credentials: --cookies-file <path> accepts a Netscape cookies.txt and is ALWAYS OPTIONAL for every source. The file is read once at startup, never copied or re-emitted. For Skool, public communities (e.g. zero-one) serve lessons without auth; private / paid communities respond with HTTP 401/403 and the user then needs to supply cookies. YouTube/Vimeo optionally forward the file to yt-dlp's --cookies for age-gated or unlisted videos. The skill never blocks on missing cookies up-front — it tries the fetch and surfaces a SourceAuthError (exit 5) only if the source returns 401/403.