Skip to main content

reviewer-review-shot-inputs

Use when final grouped-task manifest inputs need independent integration, internal-cut and boundary review before video preparation or after scoped changes.

Zur Installation springen

Quellinformationen

Repository
wddxh/ShortVideoDirector
Letzte Quellaktivität
16. September 2026 um 15:06
Erkannte Sprache von SKILL.md
Englisch
Sterne
21
Forks
3

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
reviewer-review-shot-inputs
description
Use when final grouped-task manifest inputs need independent integration, internal-cut and boundary review before video preparation or after scoped changes.
user-invocable
false
agent
reviewer
allowed-tools
Read, Write, Edit, Glob, Grep, Bash, Task, Skill
model
opus
# Review Shot Inputs Read [review evidence rules](../_meta/rules/review-meta-rules.md), [input contract](../_meta/rules/shot-inputs.md), [visual context rules](../_meta/rules/visual-context.md) and [output language](../_meta/rules/output-language.md). Review actual prompt/media integration, necessary boundaries and changed details. Use current storyboard judgment for established narrative/camera decisions unless a concrete input conflict requires reopening them. This is not generated-video review or authority to prepare, submit or fix materials. Require a fresh independent Reviewer task, never a resumed producer/fixer or favorable summaries alone. Commission exact ep/task manifest targets, canonical config_path, script/storyboard, references, constraints and a designated `/tmp/opencode/...` preview directory. Derive scope from commissioned groups; partial selected-shot scope reports full membership/additional members, not silent expansion. Each task needs a full-group local MP4; missing necessary inputs remain unknown. The independent target owner stays purely textual for this entire task-video integration review: NEVER load images, frames or contact sheets, and never receive image attachments, including helper previews. Actual viewing belongs to fresh scoped Reviewer visual helpers using small, minimum necessary helper-thumbnail sets; NEVER Read original images. Further viewing/crops/sampling use new tasks, not resumed visual helpers. This owner/helper split applies to shot-input integration; a single-asset visual leaf may still view its bounded comparison set and write its own review. Only the target owner starts and finishes its canonical ep/kind/target round. Visual helpers return text observations, not target pass; they neither write that review nor start competing rounds or modify owner STATE. Independent ready target owners may run in parallel. Production cannot invent pass; unavailable required independent viewing leaves the affected target unknown. ## Evidence And Inspection Bind judgment to the current selected canonical manifest, its final prompt and resolved uploads. A candidate prompt, alternate media or proposed fix is feedback until the owner selects it in canonical materials; inspecting that candidate cannot justify pass for the old selected inputs. Reassess the selected version with current evidence after owner updates; ordinary drafts remain not ready under the existing contract. Review the final manifest with exactly `{shots,references,prompt}` and a nonblank string `prompt`. A two-key `{shots,references}` draft may resolve through the selected provider's material tool, but fails final review/readiness. The existing target fingerprint binds final prompt text. Inspect generic `storyboard-to-prompt.sh/.mjs --json STORYBOARD TASK_ID EP` as the exact model-facing manifest string, without text generation. Compare canonical sources, actual refs and the selected provider's material-tool output under its own pack/reference-syntax contract (Dreamina: [video.md](../creator-provider-dreamina/video.md#dreamina-authoring-materials)). Provider token counting, binding and rebasing are not universal slot rules. Judge source fidelity, completeness and integration: preserve narrative/action, exact dialogue and film wording, cuts, duration and audio in task-time prose. Check actual tokens, BOX/color/body/limb mapping and necessary final anatomy/action. Exclude only non-film source headings, internal IDs/paths and reviewer metadata; retain intentional screen titles and camera terms such as “wide shot”. Use semantic judgment, not a word scanner. Missing facts return to owners, not invented `use` facts. Submission does not append or rewrite accepted text. Each target owner uses review-round to write its commissioned `reviews/{ep}/task-inputs/taskNN.md`. Same ep/kind/target serialization protects output ownership; also order read/write and input dependencies, running distinct nonconflicting ready targets in parallel. Each round has scope=[manifest] and one completed result. Create the designated `/tmp/opencode/<task>` directory before start; allow helper STATE, reviewer payload and necessary previews/crops/sampled frames only there, with no workspace duplicate ledger. No production files, grants, receipts or task records may change. Before reading manifest/config, final prompt, timeline or source texts, or resolving dependencies, the text owner starts with explicit actual SVD_CONFIG. The inspection instructions above apply only after this capture. Keep kind=`shot-input`, target=`story/episodes/{ep}/task-inputs/taskNN.json`; STATE/PAYLOAD.json are distinct absolute files in the designated temp directory: Apply the shared multi-target capture rule to every consuming STATE before a shared reference is first read. Pass known valid project-relative sources as extras; `/tmp` diagnostics are feedback, not evidence substitutes. Delegate observations must identify precollected actual sources and viewing limits. Capture relied-upon semantic reviews, but avoid volatile status/review files added solely to prove readiness; schedule necessary dependencies to stabilize rather than dropping them. ```bash SVD_CONFIG="{config_path}" node "${CLAUDE_PLUGIN_ROOT}/scripts/review-round.mjs" start shot-input EP TARGET STATE [EXTRA_INPUT...] node "${CLAUDE_PLUGIN_ROOT}/scripts/review-round.mjs" add-input STATE PATH... bash "${CLAUDE_PLUGIN_ROOT}/scripts/storyboard-to-prompt.sh" --json "{storyboard}" "{task_id}" "{ep}" ``` start captures required dependencies and known extras; use add-input before reading every new semantic reference. STATE binds config and original hashes; finish rechecks them. Include manifest/config/script/storyboard, uploaded originals, sources, asset-visual dependencies and necessary continuity references. Inspect actual editable sources and needed project inputs; discovery does not prove semantic completeness. Preserve capture/discovery errors; finish records unknown on evidence issues, retaining first hashes. Compare final prompt + ordered references against reviewed members and user intent. Check unchanged membership/durations/dialogue/cuts, slot/use correspondence, declared roles, legibility and continuity. Final manifest has exactly shots/references/prompt, local PNG/MP4 entries and at least one complete MP4 covering every current consecutive task shot on the source-derived clock, including hybrid/concatenated outputs. PNGs or partial clips do not replace it. Holds, rigid translation and limited animation may intentionally abstract performance; judge their declared controls with actual use/final prompt, not hypothetical style leakage. Sources include actual SVG, scripts, fonts and other editable dependencies, not uploads; necessary controls belong in prompt/media. Fingerprint actual dependencies, not a whole-plan hash. Judge each medium by actual purpose/use. Identity stills primarily supply appearance, style, quality and materials; static shapes supply declared structure. Coarse video can supply a dynamic skeleton for perspective, occlusion, camera, layout, lighting, trajectories and cuts. Detailed references may retain adopted existing structure/materials or supply only space without copied action. Do not apply BOX placeholder exclusions to every video or grant undeclared authority because a model is detailed. Do not demand identical views or reopen harmless fine-detail differences. Explain concrete conflicts affecting plot-critical features/actions, explicit requirements, identity, quality or continuity with the observation, responsible control and impact. Here, "unchanged" means fidelity to the current reviewed canonical shots, including owner-coordinated revisions under the [original episode bounds](../director-orchestrate/SKILL.md#集总时长责任), not a superseded version or its total. In-range net growth needs no compensating cuts elsewhere. Grouping itself still cannot revise shots; current review evidence, scope/authorization, provider task limits and record/grant/inflight drift protections remain binding. Source revisions do not automatically renew evidence or protected records. As part of the existing integration review, consult the selected provider's video reference-syntax contract and check the converter's final prompt against its actual ordered resolved references and each input's purpose. Provider documentation supplies syntax knowledge, not execution authority. Required bindings must already be present for review; return concrete missing/misbound references to the owner rather than repairing, renumbering or assuming submission will append them. Token counts do not establish semantic or artistic quality; no separate review gate is added. Read and apply the full [global and inline binding rules](../_meta/rules/visual-prompt-craft-common.md#全局映射与实际使用处的引用). Across the entire final prompt, do actual tokens, stable entity names, proxies and purposes remain consistent with the ordered inputs? Report concrete drift or cross-use. Compare each `ref.use` and final prompt's global reference/purpose and source-video interval -> output task interval mapping against actual inputs/timeline. Read EVERY explicit video binding in final prose: does that use locally state a concrete source-video interval and adopted dimensions, with its output task interval clear? Even identical clocks/full-group mappings require these local details; an output passage label or “this/corresponding segment” does not supply the source interval. Stills need task applicability, not playback clocks. Check local action phases and entry/preparation/recovery/exit relations without extending them into unsupported loops; retain complete source clocks/actions and natural-performance responsibilities. Report the exact missing/misbound passage to Creator for semantic revision of existing prompt and affected uses before review. Judge meaning, not token counts or regex; do not require every sentence/pronoun to carry a reference, auto-edit prose or expect converter/submission additions. Use the existing review, with no parser or new gate. ### Whole-Frame Coarse Mapping Coverage Review the [existing-asset selection loop](../_meta/rules/visual-prompt-craft-common.md#已有资产选用闭环), not only correctness of selected refs. Compare authorized stable source objects and fresh helper observations with descriptions, relevant known cards/images, each using shot's header and actual provider slots. Identify source/visible objects omitted from prompt/use, and suitable existing identity/form/material assets left unselected where that omission affects the established appearance. “Not selected” does not prove “no image”; no applicable image may use established source description. Background classes may be grouped; do not require every object to have an upload/card or inspect the whole asset library. The owner stays pure-text. Helpers return observed objects, comparison facts and limits, never target pass. A newly needed card/image or source dependency returns unread for Director scope/stability coordination and owner start/add-input capture before a fresh helper reads it; this check grants no extra library access. Distinguish definite omissions/conflicts (needs_revision) from necessary unavailable evidence (unknown). Unsupported extra objects require media repair, not new assets to legitimize them. Route header findings via Director to Storyboarder and other source changes to their owners. After correction, check actual rerun material slots against the entire prompt and affected uses, including earlier bindings; an appended token is not a fix for stale prose. Assess changed source/bindings with scoped compatibility and current fingerprints under existing gates, not automatic reuse of old pass or blind hash refresh. Apply [whole-frame mapping](../_meta/rules/visual-prompt-craft-common.md#全画面粗代理映射) in this existing semantic review. Ask fresh visual helpers for actual distinguishable coarse objects/classes across the selected media, including environment components, floors, walls, background, furniture and items without asset images. Helpers describe observed colors/shapes/locations and viewing limits as text; the pure-text owner compares those facts with the opening global mapping, each actual `ref.use` and the complete final prompt. Grouped repeated classes are valid; omitted visual categories are not excused by absent asset images. No per-mesh quota, object registry or keyword parser. For each visible coarse object/class, is its actual identifier mapped to a final object, adopted spatial envelope (position, screen occupancy/scale, distance, orientation, route), and source-supported final shape/structure, softness/rigidity, material surface/light response and action-dependent changes? Are real uploaded identities explicitly bound, and imageless objects positively described from established source art rather than invented images/tokens? Does the prompt avoid unintended transfer of box silhouette, proportions, topology, rigidity or material while retaining camera/light/reveal/cut controls? Dedicated shape/topology and fine references keep their declared structure/material purposes; do not force fine modeling or new asset cards/images. Global stable mappings supply a basis, not a once-only rule. Local repetition or full introductions are valid for main objects, attention focus, possible coarse-shape/material transfer or concrete expression needs, not only new objects/changes. Check clarification of new objects, scene ambiguity, identity switches, shape/material states and purpose changes. Judge restatements by applicability, clarity, consistency and concrete impact; repetition alone never fails review or justifies deleting necessary emphasis. Avoid mechanical irrelevant whole-object lists without banning per-segment repetition or setting counts. Every explicit video binding still needs its concrete source interval, output task interval and adopted dimensions. Official advice supports material correspondence, story progression and necessary restatement; full-object coverage and locally concrete intervals are SVD refinements, not official mandates. All source action remains complete regardless of local limbs or fine performance. Report concrete unmapped categories, ambiguous targets, unsupported appearances, use/prompt conflicts or unintended shape/material inheritance with actual evidence. Missing media or necessary observation/design evidence remains unknown; established mapping omissions/conflicts require revision under existing statuses. Missing assets/unclear targets must be reported as found, not invented or silently exempted. Add no schema, gate or mandatory geometry detail. ### Completeness And Local Applicability Apply [per-task semantic authorship](../_meta/rules/visual-prompt-craft-common.md#每任务独立语义写作) to the entire selected `manifest.prompt` and every `ref.use`. Compare each passage with this group's actual source shots, refs, events, dialogue and task clock: does it retain all required facts while describing only applicable roles, actions, light, sound and controls? Identify imported events/characters/sounds absent from this group, irrelevant conditional branches, repetition that dilutes necessary instructions, and contradictions between general passages, timed prose and uses. Report the specific passage, source basis and impact; a new prompt with stale use text is still an integration problem. Each task expresses the applicable shared baseline; meaningful closing restatement of global style, identity consistency or common constraints is valid. Local rules stay local. Accurate common facts and necessary identical wording across tasks are valid; repetition alone is not failure. Materials still extract identical source style once, without restricting final prose to one mention. Judge task-specific meaning, not typing method: scripts may extract/present, number, protect/validate, batch-serialize independently authored complete drafts and read exact strings. They do not replace Creator's authorship with COMMON finished passages or script fill/template composition. Preserve important source actions and dialogue when removing irrelevant text. Use existing semantic review/statuses, without keyword blacklists, parsers, similarity thresholds or another gate. Apply [official guide scope](../_meta/rules/visual-prompt-craft-common.md#官方指南适用范围): recommendations 1–4/7 and only transitions from 5 take precedence within scope; 6 and extend operations are excluded. Review [task prompt organization](../_meta/rules/visual-prompt-craft-video.md#任务提示组织) for useful content: actual numbered materials/purposes, a one-sentence overview of this task's story, time/story progression and optional global closing restatement. Does the overview orient the actual subjects, setting, core event, genre/style and any designed special camera move? Does timed prose pair physical actions with source-supported key emotion, subtext or intent while preserving words, cuts and sound? Non-character passages need visual/narrative meaning, not invented psychology. Formula structure and character-detail guidance are writing aids, not mandatory headings, length, sentence counts or quotas. Character seven-dimension guidance and suggested 3–4 facial details adapt to actual identity, visible scope and non-human structure; bound identity need not be fully redescribed at each use. Judge completeness and applicability, not resemblance to a long example or mechanical coverage of every dimension. Request corrections for concrete lost meaning, ambiguity or conflict. ## Saved Spec And Actual Media Within the existing shot-input round, capture canonical config and the explicit selected complete clean MP4 with start/add-input before reading or probing them. Read `## 本地参考 epNN` and the exact qualified keys `epNN 本地参考宽度`, `epNN 本地参考高度`, `epNN 本地参考fps`. Check their final-video-settings basis or explicit user local override, and the single source-compatible episode CFR fps; an unset fps is not 30. Provider pixel mappings require evidence, not a universal 720p short-edge assumption. Run the measured check on one explicitly selected manifest-declared full-group clean video: ```bash SVD_CONFIG="{config_path}" node "${CLAUDE_PLUGIN_ROOT}/scripts/local-reference-media-check.mjs" EP TASK_ID --video PATH ``` Compare actual width/height, CFR fps and duration with saved spec and canonical group clock. Low-resolution drafts are allowed; formal selected clean must comply. Supplementary short videos need not all span the group. Check caption picture width/height, fps and duration against clean, with total height adding the band. Capture relied-upon caption media before inspection. Tool diagnostics establish measured properties, not semantic pass; visual viewing still belongs to fresh helpers. Reuse matching media; report concrete mismatches for scoped Creator export/conversion repair preserving camera/ratio/clock, not automatic remaking or metadata-only compliance. Missing saved spec goes through Director to the bounded config owner, separately from environment-report recovery; no new initialization probe or historical migration. Config changes affecting prior review fingerprints use scoped compatibility assessment, not blind refresh. Keep the existing kind, manifest schema and review round; add no automatic gate or extra review stage. Saved width/height must be positive even pixels and the media layer requires width >= 64. Positive integer/decimal/fraction fps is reduced exactly; canonical task duration × fps must be an integer matching the measured frame count. Probe failure leaves `actual:null`. The inherited CFR probe's 120-second execution timeout is not a video-duration cap. Config evidence binds the entire file, including episode specs; preserve authorized writer coordination and stable capture. ## Text Owner And Bounded Visual Handoffs After start/add-input, read complete final prompt, timeline, manifest and necessary source texts. Establish purpose before planning helper windows for space, paths, camera, light, reveals, cuts, reading/results, clock and external pairs. Every final prompt must contain complete source action regardless of limbs/wings or locally shown detail: necessary subject/part ownership, preparation/execution/recovery and contact changes, without invented actions or frame quotas. Helpers separately inspect rough limb/wing or fine mechanism signals and concrete conflicts; prefer removing existing signals before reuse, never reducing prose. Retained limbs/wings add imitation-risk checks. Independent action references need their selected interval's evidence; purpose never exempts final source semantics or requires creating fine rigs. Keep needed before/during/after relations together. Commission each fresh Reviewer with the specific observation question, each medium's actual purpose/adopted dimensions, concrete source-video and output task intervals or cut, captured source paths and relevant source intent, minimum necessary comparison references, designated temp preview directory, read/write scope and escalation conditions. The owner captures all actual inputs with start/add-input before the delegate reads; discovered new dependencies return as paths unread, pass Director dependency coordination and owner capture, then go to a fresh helper. No post-hoc snapshot substitutes for capture before viewing. Visual helpers return text only: actual observed facts and conflicts at times/frames, actual read source paths and media fingerprints, sampling method and preview-to-source/time/frame mapping (including helper JSON), coverage limits and necessary unread dependencies. Separate observed facts from source-inferred intent/geometry/timing; neither source code nor metadata is visual proof. Do not attach images/contact sheets or issue target-level pass. Helpers write only necessary previews and optional text feedback in their assigned temp scope. The text owner independently compares these facts with the whole final prompt/timeline and source intent, judges integration and coverage across windows, internal cuts/sound bridges and necessary external boundaries, and records the basis/limits. Do not sum local favorable findings into pass. Missing, contradictory or insufficient observations prompt another fresh bounded visual task when feasible; retain concrete conflicts as needs_revision and unresolved necessary evidence gaps as unknown. Sampling cannot prove unobserved interpolation or full-motion continuity; nonessential unseen detail alone is not a blocker. Within Director-coordinated stable dependencies and child scope, delegate directly when nesting works. After a confirmed depth failure or unavailable tool, remember the limit and send role/outcome/references/scope/constraints for Director sibling relay. Return real child handles, dependencies and outstanding read/write scope; Director retains subtree occupancy until actual completion. Relay actual text facts back by resuming this TEXT owner, never a visual helper. Pending/partial delegation is not a finished round or completed production commission; unavailable required isolation/viewing is a blocker, not permission for the owner or producer to view/self-pass. ## Temporal Judgment For selected transitions, compare A end state -> trigger -> process -> B start state across source, actual controls and final prompt, including necessary direction, scale, timing and sound. These are real views/subjects, not internal IDs. Hard cuts switch instantly and need no invented intermediate motion; sequential fade-black/appearance differs from simultaneous cross-dissolve. Preserve established choices unless a concrete integration conflict warrants owner revision. Do not require smoothing, a technique quota or extra transition seconds; all phases fit existing source windows mapped to task time. See [transition process](../_meta/rules/transition-craft.md#转场过程表达). Full-group completeness concerns the final MP4 timeline, not SVG coverage. Creator may combine existing asset PNGs, layered images, 2D animation, necessary 3D and existing clips per shot. SVG is an optional preferred planar material, not an every-shot requirement. Review actual material use, source facts, clock and cuts; do not require suitable reused media to be remade or judge tool purity. Apply [transition and film-text craft](../_meta/rules/transition-craft.md). Check source words/time facts, actual text, reading interval, contrast/cropping, competing attention and reveal order. Audience SUPERs and time/place/chapter/full-screen cards need no actor reading plane; diegetic UI still needs plausible stated use. Clean selected MP4 may contain designed film text/transition images, but no internal debug or rehearsal bottom-caption contamination. Standalone cards count in canonical shot/time limits; overlays do not double-count time and assembly adds no seconds. Text, black frames or simple cards are not automatic failures. Accept source/local controls and final prompt on their evidence, without promising generated text accuracy or final model quality. At independently generated TASK boundaries, apply the [strong default preference](../_meta/rules/shot-inputs.md#manifest) for motivated, visibly distinct camera/view/scale to reduce near-identical mismatch visibility. Judge relationships by source intent under Task Junction Coverage below. Assess actual needs for match/repeated compositions and uninterrupted essential contact/speech; justified choices need no new permission. This is not an every-shot change, angle quota, continuity guarantee or excuse for source-conflicting identity/state/action. Repetition alone is not failure; name concrete integration conflicts when reopening an established choice. Use shared constraint-authority and pose usability rules with the camera-language guidance below. Distinguish actual user requirements and confirmed source intent from self-chosen implementation details; inspect relevant authority when they conflict rather than strengthening an aesthetic constraint. Report necessary source conflicts even when geometry or timing checks pass. Read and apply the full [detailed prose and proxy authority rules](../_meta/rules/visual-prompt-craft-common.md#粗模控制与外观依据分离). Do both every video ref.use and Creator's actual final manifest.prompt distinguish declared controls from final performance, with source-sufficient action detail? Under [reference-specific exclusions](../_meta/rules/visual-prompt-craft-common.md#按实际参考限定非目标特征), compare actual visual-helper evidence, source intent and reference roles: are exclusions appropriate and specific ambiguities resolved in the relevant passages while retaining needed controls and positive targets? Judge meaning, not keyword counts. Assess misleading media and its concrete fidelity impact. For rough proxies check actor/prop separation, stable mappings, operating regions, space, camera, silhouettes and reveal timing. SVD limbless BODYBOX omits arms, palms, wings and fine mechanism animation; omitted hands are not floating errors. Every final prompt still follows official guidance for complete source action, with or without limbs or local performance. Inspect rough media/sources separately: simplify existing signals before reuse; retained limbs/wings add advice 7 imitation-risk checks, not a condition on completeness. Mechanism simplification is SVD's boundary. Detailed/static/action refs keep purpose-specific review without exempting final source semantics, regenerating identity PNGs or requiring fine rigs. Repair concrete conflicts, not all hypothetical behaviors. Exclude internal trajectories, coordinates, camera cones/debug labels; retain film text. TASK is a generation unit, not a mini-film, scene or act; it may contain multiple shots/scenes within constraints. Judge cuts by source intent under Task Junction Coverage. Within the same event, check necessary compatible direction, possession/contact, action progress, space and sound, not identical frames. Flag added stops/restarts only when actual controls or final prompt change the source event; an abstract held pose alone does not assert a stop. Hard cuts mid-action are valid. Keep essential uninterrupted contact/speech together when feasible; preserve model maximum, source durations, consecutive membership and grants. Source redesign belongs to Director/owners within the original budget. No mandatory dissolve. Read [source trigger-to-response and task-time integration](../_meta/rules/audiovisual-craft.md#从触发到人物反应). Compare source with final passages for complete action sequences, ownership/contact changes, trigger -> response, expression and ensuing speech/action on the task clock. Compare BOX staging only for its adopted spatial and temporal controls supporting those events. Judge framing, time and attention for the intended reaction after its trigger, including cuts, bridges and sound mix. A face is not mandatory for voice, posture or offscreen response. Rough media may omit facial/limb/contact performance; final prose always retains the complete source sequence regardless of those omissions. Retained limbs/wings add imitation-risk checks, and visible source contradictions need correction. Independently selected action references need adopted action evidence. Report concrete conflicts or necessary gaps under existing statuses, without a performance gate. When grouping affects dialogue overlap, sound bridges or audio timing, read the relevant [audiovisual craft](../_meta/rules/audiovisual-craft.md) guidance. Verify the source utterance ownership, actual voiced intervals and continuation across cuts; preserve one continuous utterance rather than repeating the full line for each view. Use [camera-language knowledge](../storyboarder-storyboard/camera-language.md) for staging/coverage, viewpoint, lens-distance-focus and continuity. Compare script beats, reviewed shots and final refs: has framing/occlusion hidden necessary evidence, focus/movement redirected attention, or a cut/overlap changed reveal/reaction timing? A close-up can remove the listener's only visible response even with all actions in prose. Reopen camera decisions for concrete integration conflicts, not preferred coverage. Local BOX previews establish camera/layout/whole-object trajectories and operating regions; final prose supplies fine body/contact performance. Media must still expose the source-clock facts required by their actual control responsibility. For books/pages, screens, photos and controls, distinguish audience readability from plausible actor use. Compare reader/operator or recipient, usable face, observer/camera side and actual actor view. A camera-facing surface conflicts when it prevents stated use; transparent displays and deliberate showing depend on intent, not a same-side test. Check POV eye position: tilting does not relocate a camera, and one's own face needs a supported mirror/feed. Coarse actor/eye position, operating region and overall object relationship suffice for spatial control; source holding, support or contact does not require hand/forearm proxies. Check those fine actions in final prose. Identify concrete geometry/story conflicts separately from possible fixes; side/high-angle, OTS, insert and POV are alternatives, not prescribed coverage or a dot-product gate. Fake audio and extra timed drafts remain optional diagnostics, not performance proof. Lightweight animatics/mixed references deliver a complete caption review MP4 from the complete clean task MP4 regardless of material mix, using the existing PLAN and `previs-preview.py --timecode` for human timing review. Check any relied-upon PLAN against all source dialogue/narration verbatim and their corresponding task-time windows, including source-intended cross-cut continuation, overlap and pauses. Full coverage does not mean filling silence with captions. Identify estimated windows as estimates, not heard speech. The complete caption version preserves clean duration, frame rate and cuts, and marks every actual internal cut with its task-global time using existing internal PLAN segments. A single-shot task with no internal cuts needs no invented markers; report that fact. Internal markers stay out of clean/final prompt and the caption version stays out of uploads; formal film text stays in clean. Capture relied-upon PLAN/media before reading/delegation. The owner remains pure-text; viewing uses fresh helpers/thumbnails. Report missing promised caption/cut-marker delivery to Creator/Director without adding a gate or review round. Captions/tones do not replace clean evidence or prove performance; optional fake audio's absence alone is not unknown, but necessary final timing/visibility gaps remain unknown. Apply BODYBOX and [pose usability](../_meta/rules/review-meta-rules.md#数值与姿态的可用性判断). Limbless holds/rigid translation may control composition, positions, paths, camera, light, reveals, cuts and clock. Every final prompt retains complete source body/contact, hand-change and mechanism sequences, whether references show them or have limbs. Missing local hands need no support evidence. Simplify existing rough signals, never prose; retained limbs/wings add imitation-risk checks even without adopted action. Repair wrong positions, paths, camera, reveals, UI timing and concrete contradictions; prose cannot reverse them. Independent action references bind adopted action facts. Unresolved conflicts or missing source-required descriptions are needs_revision; unknown names a necessary gap. Do not fail all potential behaviors or demand local fine performance. Creator chooses early picture checks, short motion drafts and export order to reduce revisions. Their presence or absence is not an extra review gate or required round; accept the actual selected complete task MP4 and final prompt under this existing independent process. A concrete 2D limitation may justify minimal Blender/hybrid repair, not forced 3D polish or redesign of source intent to fit 2D. Fresh pure-text owner, bounded visual helpers, thumbnails, five kinds and existing gates remain unchanged. Check the selected provider tool's documented material extraction, reference binding and timing behavior against canonical sources and actual uploads. Each explicit link must be declared by its own member header. Different single-line `视频风格` baselines return to owners. Judge the final prompt's complete task-time rendering of source facts, exact dialogue and local times, preserving valid overlap with one common style baseline and no internal IDs. Check coherent media/reference clock, internal cuts and sound bridging, alongside external boundaries. Reference use states controls/placeholders; boxes need readable framing, operating regions and overall relationships, with fine support/contact actions in final prose. For MP4 controls, fresh visual helpers inspect framing, scale, positions, whole-object trajectories, camera path and relevant timing. They record actual playback/sampling method, original media hash, sample times/frame mapping and temporal coverage/limits. Sample images go through review-image.py; only visual helpers read returned previews, with necessary detail crops in new contexts. Use the minimum necessary comparison set and return its evidence as text to the owner. Select samples for operating regions, surface/actor/camera relations, whole-object paths, reveals, light, UI-result timing and cuts. Inspect retained rough limbs/wings for concrete source contradictions and report their mapping to the final sequence; do not create every potential behavior to test it. Independently selected action references need intermediate evidence for their adopted action. Source pickup/handoff/mechanism detail alone does not require local animation. Collision/geometry checks are optional diagnostics, not visual proof or a required engine. Use fresh direct-MP4 helper previews without image sequences, fixed passes or phase artifacts; report necessary gaps below. Apply [task junction coverage](#task-junction-coverage) to actual selected final prompts and clean MP4 evidence. Capture each necessary neighbor manifest/storyboard/media/source/identity reference with start/add-input before reading or delegation. New dependencies return unread for coordination and capture before a fresh helper reads; no post-hoc snapshot or import registry. Explain actual pairs and limits in prose. Changes trigger scoped assessment, not automatic recursive re-rendering or a new kind. Endpoints or a few stills cannot certify interpolation/full-duration continuity. Cover adopted camera, paths, reveals and clock, rough-limb conflicts and independently selected actions. Necessary unassessable evidence is `unknown`, not pass from code/ffprobe/hashes; never claim unviewed video was observed. Missing media, unreadable inputs and evidence gates remain blocking. Omitted rough anatomy/contact/performance is not a local gap; every final prompt always carries complete source action regardless of limb presence, local performance or purpose. Retained limbs/wings add imitation-risk checks; prose cannot waive contradictory media. GIF remains unsupported. Definite conflicts or missing source-required descriptions are `needs_revision`; aesthetic improvements are optional. For changed bookkeeping or sources with unchanged rendered media, use a scoped independent compatibility assessment: inspect the actual diff, prior reviewed basis, current prompt/refs and media fingerprints. Explain why the changed dependency preserves the judgment before issuing a new round with current inputs. Do not blindly refresh hashes or require automatic full review. Any necessary new visual operation still uses a fresh task and thumbnails. For causal UI actions, compare the complete trigger, operation, feedback, result and reaction sequence in source/final prose. Rough-media helpers inspect operating regions, actor/camera visibility and displayed result states on the source-derived clock: a result cannot precede its trigger or keep changing after a source-defined stop. Fine touch/release and mechanism motion need not be locally performed. For composites, inspect the selected final MP4 around relevant seams for conflicts in adopted state, space, camera, light and timing; valid clips or successful assembly do not prove a usable composite. Use existing sampling/audiovisual guidance. Source code, collision checks, timestamps and encoding success are diagnostics, not visual proof. ## Task Junction Coverage For voice bridges, compare source utterance ownership and both prompts' incoming/outgoing voiced portions. The source line occurs once; continuation must not restart it. Disclose audio actually present and heard versus an unassessed stream, planned J/L continuity or fake timing. Independently generated matching voice descriptions do not prove continuous prosody, background or waveform. Judge whether the stated joining approach is supported by current inputs/capabilities; absent necessary evidence is unknown. Planned postproduction may be viable without future output existing, but is not an already-heard seamless bridge or permission to mix/cut. Reuse valid independent pair observations when their exact selected inputs, question and temporal coverage still apply and every consuming owner captured those inputs before the helper read them. Each owner independently relates the text observations to its own final prompt/target; neither another target's pass nor a favorable summary substitutes for this judgment. Do not require duplicate visual passes from both owners, reciprocal review-file dependencies or a pair ledger. Follow shared capture/compatibility rules for later consumers or changed inputs; new necessary viewing uses a fresh helper. Acceptance covers reference controls plus final prompts, not generated-video quality or automatic editing authority. The pure-text owner assesses every relevant junction by its intended type. Whole-episode scope covers all adjacent groups, including scene/act/time/place changes. Within the same continuous event, check source-required compatible action progress, possession/contact, space and sound continuity. At a scene/act/time/place jump, assess meaningful causal, emotional, information, thematic contrast or parallel relationships and viewer orientation appropriate to the intended reveal. Do not impose same-event criteria, same position, continuing action/sound or a bridge scene. Source-supported mystery, abruptness and hard cuts need no smoothing or immediate explanation. Underlying identity/world facts within the episode remain consistent except source-justified intentional changes. Partial scope covers necessary incoming/outgoing neighbor boundaries of selected groups, plus actual nonadjacent/cross-episode story dependencies. Do not require unrelated media or the whole plan. An unselected neighbor may be a read-only input without becoming a generation/review target; missing necessary neighbor evidence leaves the affected target unknown and names the junction, without expanding generation authority. After establishing each medium's actual purpose, plan a coherent tail/head comparison window around each needed junction. Fresh helpers inspect both selected clean MP4s through minimum necessary thumbnails and supported temporal viewing, with both final prompts' intent and source clock mappings. Same-event windows test adopted spatial/trajectory/reveal/timing controls and actual sound; the text owner checks complete action progress and possession/contact continuity in both prompts, requesting visual action evidence only where explicitly adopted. Scene/time/place jumps test narrative relation and viewer orientation without forcing physical/audio continuity. Compare source-supported identity/world facts within actual media purposes. Isolated endpoints cannot establish the cut; no fixed seconds, frame quota or obligatory dissolve.
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen