Skip to main content

faceless-channel

Use only when the user asks to produce a finished multi-scene narrator-led video and explicitly requests a faceless channel, YouTube automation, narrated explainer/story/education, documentary, storybook or myth retelling, or kids song video. Topic alone is never enough: a generic video of/about something, including a historical topic, uses ordinary video generation. Do not use for planning or ideas, single clips, silent animation, image-to-video, footage edits, ads, product demos, or UGC. This workflow requires a consistent non-photoreal style, reusable assets, narrator voiceover, and burned subtitles.

설치로 이동

소스 정보

저장소
openai/plugins
최근 소스 활동
2026년 8월 27일 20:29
감지된 SKILL.md 언어
영어
스타
7,086
포크
916

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
15 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
faceless-channel
description
Use only when the user asks to produce a finished multi-scene narrator-led video and explicitly requests a faceless channel, YouTube automation, narrated explainer/story/education, documentary, storybook or myth retelling, or kids song video. Topic alone is never enough: a generic video of/about something, including a historical topic, uses ordinary video generation. Do not use for planning or ideas, single clips, silent animation, image-to-video, footage edits, ads, product demos, or UGC. This workflow requires a consistent non-photoreal style, reusable assets, narrator voiceover, and burned subtitles.
# faceless-channel The channel factory for faceless, narrator-led video: five channel types on one motion pipeline (plus a stills pipeline and a song mode), any non-photoreal look, one voiceover, one finished file. Voice → the `narrator` skill; captions → the `subtitles` skill; everything else lives here. > HOW TO READ THIS FILE: execute the Phases 0→8 IN ORDER. Do not skip a phase, do not > reorder. Each phase has a **GATE** you must satisfy before the next. Long templates > live in the resolved helper skills — open them when the phase says so. Obey > every GOLDEN RULE. --- ## RUNTIME CONTRACT — Higgsfield sandbox only `sandbox_exec` is required for every download, probe, validator, audio measurement, assembly, transcription, caption burn, and upload PUT. Never run these commands in a client-local or built-in shell. - Before any paid generation, verify that `sandbox_exec`, `media_upload`, and `media_confirm` are callable. If the direct upload pair is unavailable, stop: `media_upload_and_confirm` cannot export a file that exists only in the sandbox. - Preflight once in Phase 0: ``` sandbox_exec({ restart:true, command:"set -e; for b in ffmpeg ffprobe python3 curl jq awk; do command -v \"$b\" >/dev/null; done; test -n \"$HF_WORKFLOWS\"; test -x \"$HF_WORKFLOWS/faceless-channel-video/scripts/finish_video.sh\"; mkdir -p work/{blocks,voices,frames,output}" }) ``` If subtitles are enabled, also require `python3 -c 'import faster_whisper'`. A failed preflight blocks the run. - Scripts are preinstalled at `${HF_WORKFLOWS}/faceless-channel-video/scripts/`, including validators, `narrator/`, `subtitles/`, both assemblers, and `finish_video.sh`. Pass these paths verbatim inside `sandbox_exec`; never paste, copy, or execute a helper skill's local `scripts/` copy. - The sandbox is ephemeral. Keep each command self-contained and idempotent: re-download missing inputs behind `[ -s file ] || curl …`, run the script, verify outputs, and export the deliverable before the sandbox expires. - Use `background:true` for assembly/captioning and poll the returned `log_path` immediately with the next `sandbox_exec` call. Never start a duplicate process while the first is alive. - Preserve deterministic manifest order (`block01`, `voice01`, `frame001`); never infer order from `ls` or job completion order. - `media_upload_and_confirm` is only for files attached by the ChatGPT user; it cannot read a sandbox path. To export a sandbox-created deliverable, call `media_upload` first, then append `curl -f -X PUT --upload-file …` to the SAME `sandbox_exec` command that creates the file. Call `media_confirm` only after that command reports HTTP 200. Deliver only the confirmed hosted URL. --- ## GOLDEN RULES (read first — violating any of these breaks the video) 1. **Models are LOCKED. Never substitute.** Assets/style key → `seedream_v5_pro` (image). Clips → `gemini_omni` (video). Voice → `seed_audio` (audio). Music bed (when one is due: Kids default, Fairy Tale & Myth default, or the user asked) → submit it through `generate_audio_batch` with `model:"sonilo_music"`, no voice id, and the exact video duration; instrumental only (mood by channel: Kids playful, Fairy Tale & Myth mysterious-calm — `${FACELESS_STYLES_DIR}/references/kids-styles.md §Kids music bed`, `${FACELESS_STYLES_DIR}/references/style-cinematic-storybook.md §Music`). No other model, ever. 2. **Every clip is ONE 10s shot-group of FIVE hard-cut shots (~2s each)** written into a single prompt (see Phase 4) — NO shot longer than 2.5s: a frame that hangs 3–5s reads as a slideshow. **Kids blocks use the FOUR-cut interplay pattern (2.5s)** (`${FACELESS_STYLES_DIR}/references/kids-styles.md`). Degradation on generation failure: a block that fails twice at its cut count drops ONE cut (5→4; Kids 4→3) — never the whole video. One `gemini_omni` call = one 10s block. Do NOT make separate clips per cut. 3. **Compose from the approved assets.** Every clip references the Phase-2 asset images (`medias`, role `image`), in the order **location → characters → props**. NEVER generate a clip/still from the style key alone. Frames are full staged scenes, never an object on a blank/white background. 4. **Pass `aspect_ratio` EXPLICITLY on every video call** (the chosen aspect; default `16:9`). It does NOT inherit from the style key. 5. **Preset-recommender handling:** a `gemini_omni` call (usually the first) may return a preset RECOMMENDATION instead of a job. Immediately resubmit the SAME call with `declined_preset_id` = that recommended preset's id (taken from the response). Never ask the user about it, never stop on a recommendation. 6. **NSFW is a ~50% probabilistic false-positive.** Use the RETRY LADDER (see below) — resubmit with a new seed, then reword. NEVER drop a block, NEVER deliver a gap. 7. **Characters never talk on screen** (no lip-sync). The voice is an external narrator added in post. Prompts say "characters only emote and gesture, they do NOT talk." Kids: characters DO visibly react to the narrator (wave, nod, look into the camera — the interplay in `${FACELESS_STYLES_DIR}/references/kids-styles.md`); reacting is gesture-only, never mouthed speech. 8. **Subtitle timing comes ONLY from Whisper on the final audio.** Never estimate from the script, never time-by-generating-per-phrase. 9. **Assembly fps = source fps** (probe `r_frame_rate`); never hardcode 30. 10. **Banned in prompts:** the tokens `child` / `kid` / `childlike` (use `naive` / `small` / `simple`); any real brand / studio / IP name (describe the look instead). 11. **Never expose mechanics to the user** — no model names, phase names, or studio names in chat, and no third-party brand/studio names inside `ask_user_input` texts either. The user sees only creative substance + the approval gates. 12. **Wait every job to a terminal state.** For OpenAI batch submissions use `jobs_wait` on the returned `{index, job_id}` pairs; `completed` = good, `failed`/`nsfw` = retry. Do not proceed on a non-`completed` job. 13. **Deliver ONE whole video file** (`final.mp4`). Concatenate ALL blocks + VO (+ subs) into a single file. NEVER split the output into `part1`/`part2` or hand back separate clips — the deliverable is exactly one video. 14. **Throughput: bulk media uses the headless batch tools.** Submit independent images, video blocks, and narration takes in ordered groups of at most SIX, then wait for that group together with `jobs_wait`. Never fan out parallel singleton `generate_*` calls in an OpenAI run. 15. **Duration is FIXED = N×10s (the target). NEVER shorten the video to fit short audio** (the "2:00 → 1:35" bug). Each block stays 10s. Write one dense voice line that naturally fills most of each block, then let the assembler center its detected speech. Minor underfill is acceptable after bounded narration retries. If detected speech exceeds its block, rewrite it shorter and regenerate. Never trim the video. 16. **Sync is by construction:** one line lives inside its own 10s block, so a line never bleeds into the next scene. **Never `atempo`/speed-change/pitch-shift** audio to fit — rewrite + regenerate instead. 17. **ONE voice everywhere.** Every audio chunk uses the SAME `voice_id` + `voice_type` (the one locked at intake). Never let chunks come out in different voices. 18. **NO time-stretching in post.** Never `atempo`/speed-up/slow-down/pitch-shift the audio to fit. If length is wrong, REWRITE + regenerate the beat. (`speech_rate` also untouched unless the user asks.) 19. **Style fidelity — clips MUST match the asset sheets 1:1.** Same character design, same palette, **same background treatment** (if assets are white/clean-bg webcomic, the video stays white/clean-bg webcomic). ONE consistent style across the whole video — no per-shot restyle, no object drift, no style scatter. 20. **Captions are tiny (if on):** ONE short line, ≤3 words / ≤15 chars, bottom ~12%, NEVER covering the subject or filling the frame. Clean CAPS + outline, no plate. 21. **No leading freeze.** Every block prompt demands motion from frame 1; the assembler adds no head padding and WARNS when a block's opening looks static — on that warning REGENERATE the block (never ship a still that "starts playing" a second later). 22. **No samey footage.** Vary shot SIZE and ANGLE on EVERY cut (WIDE / MEDIUM / CU / OTS / low / high) — do NOT reopen every block on the same establishing WIDE. **OTS is legal ONLY when a named on-screen character's shoulder/head is deliberately visible in the foreground.** For an object-only, diagram, empty-location, or otherwise characterless shot, OTS is FORBIDDEN — use overhead/top-down, low/high angle, macro, lateral, or another coverage angle instead. Never use OTS as a synonym for an angled view; it makes the video model invent a person. **Max ~20s (≈2 blocks) per location/distance**, then move (new location / coverage angle / variety insert). Rotate locations; never park the character back at the opening wide. 23. **Voice and subtitles are DELEGATED to installed skills; reference instructions are bundled into this skill. All executable work runs in `sandbox_exec`.** Phase 5 invokes `narrator`; Phase 7 invokes `subtitles`. Before Phase 0 resolve the directory containing this `SKILL.md` as `FACELESS_SKILL_DIR`, then set `FACELESS_FLOW_DIR`, `FACELESS_STYLES_DIR`, and `FACELESS_MODES_DIR` to that same directory. Their documents are all under `${FACELESS_SKILL_DIR}/references/`. Executable scripts always come from `${HF_WORKFLOWS}/faceless-channel-video/scripts/` inside the Higgsfield sandbox. For ordinary motion-video runs use `finish_video.sh` (internally `assemble_final.sh`); Picture Story uses its `--stills` route (internally `assemble_slides.sh`). These FFmpeg scripts are the canonical assembly paths, not fallbacks. Never call `explainer_video` merely to stitch completed clips and narration. Never hand-roll another FFmpeg path around the scripts; if an assembler errors, fix the inputs and rerun it. 24. **Generation widgets are stage summaries, never progress indicators.** Batch submission and waiting stay headless. Persist every `{index, job_id}` immediately and trust `jobs_wait`, not widget chrome. After the complete media stage is terminal, display its exact final ledger with `show_generation_by_ids`; never browse history with `show_generations`. 25. **Narration target is 9.4–9.8s of speech per full 10s block.** The narrator measures speech rather than padded file length, rewrites out-of-window lines, and retries each line up to its bounded attempt budget. After that budget, keep the closest completed non-overrunning take and report the miss; never time-stretch it. Stop with `AUDIO_GEN_FAILED` only when a take is missing, invalid, or still exceeds its block after retries. --- ## Types & format Output = **ONE video file**: a motion-clip montage for Explainer / History / Kids, or narrated STILLS for the Picture Story direction (`${FACELESS_MODES_DIR}/references/picture-flow.md`). Kids additionally has **SONG MODE** — a MUSIC VIDEO built on a real sung children's song (`${FACELESS_MODES_DIR}/references/kids-song.md`: song generated FIRST via `seed_audio` prompt-only, blocks choreographed to it, assembly via the assembler's `--song` mode; no narrator, no bed, no subtitles). Standalone image deliverables (slide decks, image sets) remain REMOVED — every direction ships a single video. | Channel type | Default style | Pacing | VO tone | |---|---|---|---| | **Kids** | **Baked-in Kids set** (`${FACELESS_STYLES_DIR}/references/kids-styles.md`, Studio 3D recommended; Fluffy Toy via preset widget) | FAST — 4 cuts per block (WIDE → CU character → ECU detail → MEDIUM), varied every block | warm teacher, direct address, catchphrases; narrator↔character↔viewer interplay is MANDATORY (see kids-styles.md) | | **History** | **Editorial Motion Graphics** (house style — `${FACELESS_STYLES_DIR}/references/style-editorial-collage.md`); named alternates: **Paper Diorama** (`${FACELESS_STYLES_DIR}/references/style-paper-diorama.md`), **Mannequin** (`${FACELESS_STYLES_DIR}/references/style-mannequin.md`); LONG-FORM direction: **Documentary 10+ min, Watercolor Chronicle** (`${FACELESS_MODES_DIR}/references/history-longform.md`) | slower, chronological (long-form: cold open → rewind → chapters) | witty, sarcastic, anachronistic storyteller (Mannequin: dry British; long-form: measured documentary narrator) | | **Explainer** | TWO main directions, offer both: **Editorial Motion Graphics** (first, recommended) / **Stickman Cartoon** (generic webcomic formula) | fast, rapid cuts | casual 2nd-person, deadpan, hook + promise | | **Picture Story** | Narrated STILLS (`${FACELESS_MODES_DIR}/references/picture-flow.md`): **Flat 2D Papercraft (recommended)** / Stickman / Hand-drawn Ink — one continuous narration drives a dense Whisper-timed microframe sequence | set by the audio timeline | free (kids-warm, history-witty, deadpan slice-of-life — tone follows the topic) | | **Fairy Tale & Myth** | **Cinematic Storybook** (`${FACELESS_STYLES_DIR}/references/style-cinematic-storybook.md`): lush hand-painted 2D-animation fairytale look, ANIMATED ON TWOS (`--stepped 12`); optional book-spread inserts | slower, atmospheric; 2–3 min default (12–18 blocks); 5 ~2s cuts per block | enchanting storyteller — hushed, warm, mysterious, unhurried, mythic (no jokes); MANDATORY mysterious-calm music bed | **Editorial Motion Graphics is the flagship house style** — the default for BOTH History and Explainer, pinned in `${FACELESS_STYLES_DIR}/references/style-editorial-collage.md` (STYLE FORMULA, palette lock, {MOTION} mapping, asset guidance) — no preset id, no picker needed. Alternates stay one chip away: **Paper Diorama** (History's named alternate — geopolitics/money/power or a "cinematic" ask; `${FACELESS_STYLES_DIR}/references/style-paper-diorama.md`), **3D Papercraft** (History) via the preset widget, and on Explainer the second MAIN direction **Stickman Cartoon** — described GENERICALLY in every prompt ("crude paint-program webcomic: thin wobbly black outlines, flat solid fills, egg-head dot-eye stick figures, plain flat-color backgrounds") and **never naming a real comic/brand/IP**. --- ## PIPELINE (Phases 0→8, in order) Resolve `FACELESS_SKILL_DIR` and the three reference aliases from rule 23 before reading bundled Markdown. Every executable command uses the sandbox-provided `$HF_WORKFLOWS`; never resolve a local scripts directory. > **PICTURE STORY runs use this same pipeline with the deltas in > `${FACELESS_MODES_DIR}/references/picture-flow.md`:** voice is one continuous narration generated > BEFORE the final frames; Whisper word timestamps create a dense ~0.7–1.2s > microframe timeline, and Phase 6 uses `${HF_WORKFLOWS}/faceless-channel-video/scripts/assemble_slides.sh` inside `sandbox_exec` (never > assemble_final.sh, never gemini_omni — nothing is animated). Rules 2/21/22 > (10s blocks, cuts, freeze probes) do not apply there; every other GOLDEN RULE does. > Its review order is necessarily AUDIO → IMAGES and it has no video checkpoint. ### OpenAI batch generation + stage review contract Use this contract for every ordinary OpenAI run. It replaces parallel singleton generation calls and all per-job progress widgets. 1. Submit independent work with `generate_image_batch`, `generate_video_batch`, or `generate_audio_batch`. Every request is `{index, params}`, every `index` is a stable non-negative script/asset number, and every `params.count` is `1`. Indices are unique across the ENTIRE media stage and never reset when a new submission group starts. A call carries 1–6 requests. 2. For more than six items, process sequential groups of at most six: submit one group, wait for it to finish, then submit the next. This avoids exceeding a workspace's concurrent-job limit and keeps every wait set within its limit. 3. Persist successful `{index, job_id}` pairs immediately. Never pass a `submission_failed` item without a `job_id` to `jobs_wait`. If submission fails with a concurrent-job/rate-limit error, finish the active group and retry only the rejected indices in a smaller later group. Honor a reported `concurrent_jobs_limit` as the next group size. 4. Call `jobs_wait` with the current group's pairs and `timeout_seconds:25`. If `all_terminal:false`, wait `poll_after_seconds`, then call it again only for active jobs and retryable `lookup_failed` jobs. Freeze completed indices. The 25 seconds are a per-call long-poll budget, not the total wait; repeat until terminal. If a group shows no status change for 20 minutes, stop and surface its pending indices/job ids instead of looping silently. A permanent lookup failure is terminal for that attempt and enters the normal retry/failure ladder. 5. Do not call `show_generation_by_ids`, `show_generations`, `job_display`, or `job_status` while the stage is running. The batch tools and `jobs_wait` are intentionally headless. Never use history-based `show_generations` for a batch stage. 6. After the ENTIRE media stage and its bounded retries are complete, build the final ledger with exactly one `completed` `{index, job_id}` per stage item. A successful retry replaces the failed job id at that same index. Sort the ledger by `index`, then interactive runs call `show_generation_by_ids` with those exact pairs. For 1–24 items call it once; the widget paginates locally in groups of 12. Only stages larger than 24 use consecutive display groups of at most 24. Never pass history arguments such as `type`, `size`, or `cursor`. 7. In interactive mode, immediately render one `ask_user_input` review question after that list and END THE TURN: - images: “The images are ready. Continue to video?” - videos: “The videos are ready. Continue to voiceover?” (SONG MODE: “The videos are ready. Assemble the final video?”) - audios: “The voiceover is ready. Assemble the final video?” Use two options: `Continue (recommended)` and `Stop here`. Localize the text to the user's language and use the same callable/raw GenUI fallback mechanism defined in Phase 0. Never start the next media stage in the turn that displays this question. 8. Explicit hands-off/auto runs and platform-dispatched jobs skip both `show_generation_by_ids` and the review question, then continue automatically after the stage is terminal. They use headless batches throughout.
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기