| name | generation |
| description | Generate MovScript creative outputs by calling generation tools: images, videos, voiceover, music, sound effects, subtitles, reference assets, storyboard images, keyframe images, and generated options. Before any generation tool executes, summarize the full generation context in user language and obtain explicit user confirmation. Also use when a user asks generically to generate text, images, video, or audio and the generation tool choice is unresolved, so Codex asks whether to use MovScript or another system before generating. Use internal content-production task, candidate, and selection project tools to preserve adoption gates, dependency checks, reference-shot inspection, generated artifact Resource preservation, and impact review. |
| toolGrants | ["mcp__movscript__movscript_runtime_status","mcp__movscript__system_model_list","mcp__movscript__generation_capability_list","mcp__movscript__generation_prepare","mcp__movscript__generation_submit","mcp__movscript__generation_job_get","mcp__movscript__generation_job_get_batch","mcp__movscript__generation_result_register","mcp__movscript__system_resource_library_query","mcp__movscript__system_resource_image_read","mcp__movscript__system_resource_image_transform_to_resource","mcp__movscript__system_resource_video_extract_frames","mcp__movscript__system_resource_video_probe","mcp__movscript__system_resource_video_extract_frame_to_resource","mcp__movscript__system_resource_video_extract_frames_to_resources","mcp__movscript__system_resource_video_trim_to_resource","mcp__movscript__system_resource_video_contact_sheet_to_resource","mcp__movscript__system_resource_video_extract_audio_to_resource","mcp__movscript__system_resource_image_annotate","mcp__movscript__system_resource_upload","mcp__movscript__system_resource_upload_batch","mcp__movscript__system_shot_library_query","mcp__movscript__system_external_resource_source_list","mcp__movscript__system_external_resource_search","mcp__movscript__domain_get_model","mcp__movscript__domain_overview","mcp__movscript__domain_query_entities","mcp__movscript__domain_query_assets","mcp__movscript__domain_query_remote_asset_groups","mcp__movscript__domain_query_remote_assets","mcp__movscript__domain_query_production_context","mcp__movscript__domain_read_project_context_snapshot","mcp__movscript__domain_upsert_project_standards","mcp__movscript__domain_read_script_source","mcp__movscript__domain_read_scene_moment_edit_plan","mcp__movscript__domain_derive_content_unit_artifact","mcp__movscript__domain_build_content_unit_backend_prompt","mcp__movscript__domain_read_content_unit_runtime_panel","mcp__movscript__domain_read_content_unit_generation_prompt","mcp__movscript__domain_read_content_unit_dependency_report","mcp__movscript__domain_read_content_unit_selection_validity","mcp__movscript__domain_upsert_content_unit","mcp__movscript__domain_update_content_unit_prompt","mcp__movscript__domain_register_raw_resource_as_content_unit_candidate","mcp__movscript__domain_create_content_candidate","mcp__movscript__domain_create_content_candidate_batch","mcp__movscript__domain_decide_content_unit_candidate","mcp__movscript__domain_select_content_unit_candidate","mcp__movscript__domain_select_content_unit_candidate_batch","mcp__movscript__domain_inspect","mcp__movscript__domain_review","mcp__movscript__domain_interpret","mcp__movscript__domain_production_status_summary","mcp__movscript__domain_regeneration_plan"] |
Generation
Use this skill when a user asks to generate or prepare creative outputs through MovScript, including images, videos, voiceover, music, sound effects, subtitles, reference images, 分镜图, 关键帧, or scene-beat videos. In user-facing language, "generation" means calling a generation tool to make the requested material. Think in production terms first: what output is being made, what references it depends on, whether existing generated options need adoption, and what review boundary protects downstream work. Before any generation execution tool runs, tell the user the full generation context and wait for explicit confirmation. Use internal content_unit/candidate/selection names only for tools and precise diagnostics; in ordinary replies, say 内容制作任务/制作项 or the exact item being made. If project locator, initialization, open/fetch state, or service availability is unclear, use the project skill's Project Management Gate before generation. If runtime ownership, provider gateway, or model capability availability is unclear, use the runtime skill first.
Production Contract
- Production step: calling generation tools, producing RawResources and generated options after planning and upstream selection gates are clear.
- Systems/config: Data Service/model gateway owns provider/model/job execution; Project Service owns prompt/candidate context; Resource service owns RawResource persistence; admin config owns routes/credentials/public access; runtime/daemon supplies endpoints.
- Blockers: unresolved generation tool choice, missing project locator, unresolved project initialization/open state, missing model capability/route/key, unavailable Data/Project service, no public/reference URL when provider requires it, stale or unselected upstream candidate, unsafe provider trust path, missing saved prompt backup, missing prompt compilation, missing full-context user confirmation, missing confirmed project style baseline for script-related visual generation, or a model-facing prompt that still depends on invisible/unresolved story context.
- Human review: generic text/image/video/audio generation requests require an explicit tool choice before generation. Every generation task, including image, video, voice, music, sound effect, subtitle, storyboard, keyframe, asset-reference, free-scope, and 内容制作任务 generation, requires a full context summary and explicit user confirmation before any
generation_* execution tool or external generation system is called. In user-facing replies, say 分镜图 for storyboard and 关键帧 for keyframe. Script-related image or video generation also requires a project style baseline that the user has already confirmed and that is saved in project standards; that baseline may be style reference images or a style prompt. Generated output is only an option; stable downstream use requires adoption/selection or an explicit user request for an unstable draft.
- Resource preservation: every materializable generated, imported, transformed, rendered, or external artifact should become a MovScript RawResource before creative quality judgment. Do not skip upload because the output looks bad, is unselected, exploratory, temporary, or only an intermediate draft.
- Output: report what will be made or was made, which saved materials were created, what still needs choosing, blockers/warnings, review URL, and model/job/resource IDs only when useful for debugging or user action.
Generation Posture
- Use concrete user-facing language in replies: image, video, shot, dialogue voice, narration voice, subtitle, sound effect, music, ambience, 分镜图, 关键帧, character reference, scene/place reference, prop reference, costume reference, saved material, generated result, chosen result, readiness, and impact review.
- For ordinary users, say "调用生成工具生成" instead of abstract generation-pipeline language. Make the task feel like: decide what to make, confirm references/style, call the tool, save every returned material, then choose a result.
- Before ordinary confirmation, result, blocker, or candidate-choice replies, open
../domain/references/user-facing-response.md and translate internal ids/tools into what will be made, saved, chosen, or blocked.
- Before uploading, importing, registering, or reporting generated artifacts, open
../domain/references/resource-discoverability.md and preserve user-readable title, purpose, placement, status, version, and provenance when available.
- Treat generic generation requests as a tool-choice gate before any generation. If the user asks to generate text, images, video, voice, music, sound effects, or other audio without explicitly choosing MovScript, first ask whether to use MovScript or another available system such as LibTV. Do not default to MovScript, do not call MovScript generation tools, and do not call non-MovScript generation tools until the user chooses.
- If the user chooses another system, stop using MovScript generation for that request and switch to the relevant tool/skill when available; if the other system is unavailable or uninstalled, explain the missing setup instead of silently falling back to MovScript.
- If the user chooses LibTV or another external generation system but the result must return to MovScript, open
references/external-generation-bridge.md. External outputs must be materialized into MovScript RawResources before they are referenced downstream.
- If the user explicitly says to use MovScript, or the request is already inside a MovScript project/output-task workflow, treat the tool choice as satisfied and continue with MovScript readiness, prompt, dependency, spend, and adoption gates.
- Treat prompt writing as the core generation task. Do not pass user prose, script excerpts, scene summaries, or project-internal shorthand straight to the model. First turn them into precise model-facing production direction.
- When project generation depends on unclear story, continuity, character, beat, dialogue, or narration context, read the script source before writing or compiling prompts. Use the script to recover project intent, then translate it into model-visible production direction.
- Before updating a generation prompt or asking the user to confirm generation, run a model-understandability audit: would a model that sees only this prompt plus resolved resources know exactly what to draw, move, frame, light, and hear? If not, rewrite the prompt before continuing.
- Keep model-facing prompts grounded in concrete visual/audio facts: subject identity, visible state, physical setting, action beats, camera/framing, timing, lighting, sound, style, continuity refs, and shot-specific negatives. Remove or translate abstract narrative such as backstory, relationship logic, hidden intention, "he realizes", "she feels betrayed", or "the family pressure peaks" unless it is expressed as observable gesture, dialogue, blocking, camera, or sound.
- Treat a successful generation as a new option, not as the chosen result. It becomes stable only after the user or workflow chooses
采纳/adopt or explicitly selects it.
- Separate Resource persistence from value judgment. Upload or verify every returned artifact as a RawResource first; express "good", "bad", "keep", "reject", or "maybe later" through candidate decisions, notes, or user-facing review, not by omitting the Resource.
- Stop before downstream generation when required upstream references have generated options but no adopted choice, unless the user explicitly asks for an unstable draft.
- Treat all generation execution as a confirmation gate. Unless the user has explicitly confirmed the specific job after seeing the full generation context, do not call
generation_prepare, generation_submit, generation_result_register, or any external generation trigger. Read-only discovery and context tools such as system_model_list, generation_capability_list, domain reads, resource reads, and domain_build_content_unit_backend_prompt may run before confirmation so the confirmation summary is accurate.
- Use internal names such as
content_unit, candidate, selection, scene_moment, and expression_unit only in tool calls, source edits, and concise technical diagnostics.
Generation Confirmation Gate
Use this gate before every image, video, audio, text, subtitle, asset, storyboard, keyframe, free-scope, or project-scoped 内容制作 task.
- Gather context with read-only tools only: project locator/context, script or story source, 内容制作任务 artifact bundle, dependency report, selected candidates, prompt refs, resources, model/capability options, style rules, aspect ratio, and known blockers.
- Write or refine source prompts when needed, then compile with
domain_build_content_unit_backend_prompt when an internal content_unit is involved. Do not call generation_prepare, generation_submit, or generation_result_register yet.
- Tell the user the full generation context in plain language: what material will be made, the prompt preview, which references/style will be used, candidate count or duration when relevant, what is still uncertain, where the result will be saved, and what the user will choose afterward. Include model/provider, cost, IDs, compiled prompt details, and route details only when they affect the user's approval or the user asked for them.
- Ask for explicit confirmation for that specific generation task/tool call. A prior request to "generate" is not enough if the context has not been summarized. Confirmation must happen after the context summary.
- Only after confirmation, call the needed generation execution tool. If the user changes context, model, prompt, refs, or scope, repeat this gate before execution.
Generated Artifact Preservation Gate
Use this gate after every confirmed generation, import, transform, render, external executor, or low-level/free-scope output.
- Collect every returned artifact, not only the best-looking one: media files, provider URLs, posters/previews, batch variants, extracted derivatives, generated text, subtitles, prompt snapshots, result JSON, job ids, and error/result metadata.
- Materialize anything that can be downloaded, uploaded, or represented as text/metadata. Save text/JSON outputs as UTF-8 artifacts when no media file exists.
- Ensure each materialized artifact becomes a MovScript RawResource. For internal output-task generation tool calls that auto-create candidates, verify the returned resource and candidate ids. For external, low-level, transformed, rendered, or manual outputs, use
system_resource_upload, system_resource_upload_batch, or the relevant export/import tool.
- Preserve discoverability and provenance on Resource/candidate metadata when available: user-readable title, purpose, project placement, status, version/batch index, model/provider, operation, prompt snapshot, source resource ids, external node/job id, original URL, parameters, and why the artifact was produced.
- If an uploaded artifact targets a story beat, shot/dialogue/narration/subtitle/sound item, reusable asset, 分镜图, 关键帧, audio cue, subtitle, style/reference batch, or another internal output task, register it as a candidate unless the 内容制作任务 generation monitor already did so. If it is exploratory or unscoped, still keep the RawResource discoverable with a title/status/provenance.
- Only after Resource preservation should the user or workflow decide
adopt, reject, or defer. reject and defer annotate the option; they must not delete or hide its RawResource.
- Do not leave a generated asset only in chat text, a local temp path, an external URL, a LibTV canvas node, a provider job page, or a renderer output folder when it belongs to MovScript production work.
Saved Prompt Gate
Use this gate before every project-scoped or script-related generation task, regardless of executor.
- Before any project-scoped generation executor runs, create or update the matching 内容制作任务 (
content_unit) and its edit_prompt as the durable prompt backup. This applies whether the actual executor is MovScript, LibTV, or another external tool.
- If the matching 内容制作任务 does not exist, create the smallest appropriate internal
content_unit first, such as scene_moment_ref, expression_unit_ref, asset_ref, storyboard_ref, keyframe_ref, audio_cue_ref, or a clearly named generic slot for project-level style/reference batches.
- Write the
edit_prompt in MovScript production language with semantic refs and model-visible direction. Do not leave the only prompt in chat text, provider-node state, a LibTV canvas, or an external task payload.
- Compile or read the saved prompt before generation when MovScript prompt compilation applies. For external tools, translate from the saved prompt into the provider prompt, then import the result back as a RawResource and generated option for that 内容制作任务.
- Bypass this gate only for an explicitly non-project throwaway generation that the user confirms should not return to MovScript. Do not use that bypass for project, script, story beat, 分镜图, 关键帧, asset, audio cue, subtitle, or style-reference work.
Project Style Gate
Use this gate before any script-related image or video generation, including 分镜图, 关键帧, asset image, story-beat image/video, and shot visual output. First decide whether the requested style is simple/unambiguous or special/ambiguous. Simple styles can be carried by a confirmed style prompt. Special, composite, uncommon, or ambiguous styles should be stabilized as global style reference images generated from a style prompt and chosen by the user. The only exception is generating/importing the style-reference image batch or writing the style prompt itself.
- Read
domain_read_project_context_snapshot and inspect prompt_preview, enabled style rules, core style fields, and style_reference_resource_ids.
- If neither a confirmed style prompt nor confirmed
style_reference_resource_ids exist in project standards, stop before visual generation. Do not generate script-related images or videos through MovScript, LibTV, or another external tool.
- If the style is simple and has no meaningful ambiguity, write the proposed reusable style prompt in the confirmation summary and wait for the user to approve it. Save the confirmed prompt in
project_standards.json with domain_upsert_project_standards, preferably in visual_style or project_style.custom_rules with key style_prompt, prompt_role style, enabled true.
- If the style is special, composite, uncommon, subjective, or likely to be interpreted inconsistently, first write a style prompt, then use that prompt to create/update a project-level style-reference 内容制作任务 and
edit_prompt. Summarize the full context and ask the user to confirm generating a batch of style reference image candidates. The confirmation summary must include script/project context, style intent, model/provider, prompt preview, candidate count, and adoption gate.
- After style-reference image generation/import, ask the user to choose the style image(s). Save the chosen RawResource IDs in
project_standards.json with domain_upsert_project_standards under project_style.custom_rules using key style_reference_images, enabled true, and a value that includes parseable IDs such as reference_resource_ids: [123,456] or resource#123, plus any useful project paths or labels.
- Run
domain_inspect and domain_interpret, then reread domain_read_project_context_snapshot. Only continue with script-related image/video generation after the snapshot exposes either the confirmed style prompt in prompt_preview/enabled style rules for simple styles, or confirmed style_reference_resource_ids for special/ambiguous styles.
- Once a style baseline exists, every later generation prompt must cite and use it. Include the confirmed style prompt in every saved
edit_prompt or the compiled model-facing prompt through the project context harness. When confirmed style_reference_resource_ids exist, pass them as global reference_resource_ids for all supported visual generation, in addition to semantic refs compiled from the saved prompt.
Agent Surface URLs
- When a generation, prompt, candidate, resource, or impact MCP result includes
surface.kind: "browser_url" and surface.url, include that URL in the user-facing response and tell the user to open it to continue.
- Map the URL to the concrete action: open the prompt page to edit/save the current prompt, open the generation progress page to monitor output and saved-material status, open the result page to choose
采纳/放弃/待定, or open the impact page before accepting stale dependencies.
- A returned URL is a handoff, not a completed user decision. Do not say generation output is accepted, selected, or stale impact is accepted unless the page action happened or the matching domain decision tool was called.
- If multiple surfaces are returned, lead with the primary
surface.url; mention secondary URLs only when useful for review. Use URLs exactly as returned.
Concepts
- MovScript MCP may be running as a cloud/external entrypoint, a local daemon-attached session, or a diagnostic/basic session. Generation depends on Data Service model/provider gateway, resource capabilities, Project Service prompt/candidate context, and MCP capability gating; do not require Desktop or cloud auth when
movscript_runtime_status shows the workflow is available locally.
- Generation tools do not infer project from session, cwd, route, or focus. Pass the intended
projectDir/project_dir or cwd for project-scoped generation, and include projectUid/project_uid when writing scoped backend candidate metadata.
- Current UI candidate and selection metadata lives in scoped project-data keyed by
projectUid plus runtime/app scope. Do not use top-level movscript candidate add/select, MOVSCRIPT_PROJECT_ID=..., or /api/v1/projects/:id/decisions as a fallback. If a scoped decisionStore is unavailable, diagnose projectUid, scope, auth, and runtime context, then stop instead of writing legacy project decisions.
- If the project is not clearly initialized, open/fetched, or able to return Project Service prompt context, switch to the
project skill's Project Management Gate before prompt compilation or generation. Use project init/create only when the user explicitly asks or confirms.
- User and organization identity are handled by MovScript app/frontend state and the MCP service. Do not pass
userId, user_id, orgId, or org_id to MCP tools.
- Generation outputs are MovScript resources first, before quality judgment. They become effective project state only when written as backend options and then adopted/selected when stability is required.
- Voiceover, music, sound effects, subtitle transcription, subtitle alignment, and subtitle translation are made by calling generation tools, but placing, trimming, mixing, burning-in, rendering, packaging, and exporting them are editing work.
input_resource_ids and reference_resource_ids accept MovScript RawResource IDs, not MCP resource URIs, local paths, or external provider URLs.
- Resource/media transform tools are intentionally business-neutral. Use
*_to_resource tools to create reusable RawResources for generation inputs, references, review artifacts, or later candidate writes; do not expect these tools to update 内容制作任务, candidates, or selections by themselves.
- Resource/media transform uploads persist generic derivative metadata (
operation, input_resource_ids, and params) on the created RawResource. Treat this as provenance, not domain acceptance.
- Video/audio editing, stitching, timeline render, HLS packaging, transcode, reframe, subtitle burn-in, audio mixdown, and export import are editing concerns. Use the
editing skill and editing_* tools for product editing; generation tools only create or prepare source resources.
- Output-task prompts may carry project-wide style prompts from
visual_style or enabled custom rules with prompt_role: style, and project-wide style reference images from project_standards.custom_rules[key=style_reference_images]. Treat confirmed style prompt text as the house-style text baseline for simple/unambiguous styles, and treat style_reference_resource_ids plus runtime_request.inputs[role=style_reference] as global house-style image references for special/ambiguous styles. Once either exists, every image/video/storyboard/keyframe/asset visual prompt must reference the confirmed project style baseline for consistency. Pass style image references as reference_resource_ids whenever they exist and the selected image/video model supports reference images.
- Content unit
edit_prompt supports MovScript prompt refs. Use {{asset::id}}, {{storyboard::id}}, {{keyframe::id}}, {{audio_cue::id}}, {{scene_moment::id}}, {{expression_unit::id}}, {{content_unit::id}}, {{candidate::id}}, or {{resource::123}} when the prompt should carry semantic dependency/selection semantics. The single-colon form is legacy-compatible, but the double-colon form is preferred in agent-authored prompts.
- Before generation, use semantic prompt refs for selected upstream assets/分镜图/关键帧/audio cues/output tasks. Do not write resolved resource numbers in place of semantic refs unless the user is explicitly providing loose RawResource inputs or a direct
{{resource::123}} ref. Stop before generation if a semantic ref has no adopted/selected upstream candidate. Unless the user explicitly asks to continue as an unstable draft, guide the user to adopt/select one of the existing asset, 分镜图, 关键帧, or audio cue candidates first.
- Prompt refs compile through backend decision metadata:
{{asset::id}} resolves the matching asset_ref 内容制作任务, {{storyboard::id}} resolves the matching storyboard_ref 内容制作任务, and the selected/adopted candidate resource becomes a resource mention such as @[resource:123] (legacy [[resource::123]] is also recognized). If the upstream candidate is not selected, stale, or has no resource, prompt build returns blockers and generation must stop unless the user explicitly asks for an unstable draft. Default response: explain which upstream candidate needs adoption/selection and ask the user to choose 采纳 before proceeding.
- Prompt refs are not mandatory for every raw resource. If the task only needs one or more RawResource IDs as loose model inputs or reference images, and does not need MovScript dependency tracking or selected-candidate semantics, pass them directly as
input_resource_ids / reference_resource_ids, or use direct {{resource::123}} refs.
- Before project-scoped generation, call
domain_read_project_context_snapshot. Use its prompt preview, enabled rules, negative rules, aspect ratio, and style_reference_resource_ids as the project context harness for generation decisions. For script-related image or video generation, missing confirmed style baseline is a hard blocker; establish and confirm a style prompt for simple/unambiguous styles, or generate a style-reference image batch from the style prompt for special/ambiguous styles.
- Do not change project standards during generation just because they are missing or weak. Only call standards write tools when the user explicitly asks to add, remove, or adjust project-wide rules, when the user confirms a reusable style prompt, or when the user has explicitly chosen style-reference images that must be persisted under
style_reference_images.
- If the conversation cannot finish all requested output/scope work, persist only stable, reusable, project-wide standards that the user has explicitly stated or clearly confirmed into
project_standards. Do not store transient task notes, job state, resource URLs, or unconfirmed guesses there.
- Content unit artifact bundles contain runtime panel, input version, dependency report, and selection validity. Derive or read these before changing generated content when they are relevant.
- For output-task image/video generation, first write or update the 内容制作任务
edit_prompt with semantic MovScript refs and the confirmed project style baseline, then call domain_build_content_unit_backend_prompt to compile and inspect semantic_ref_replacements, resource_ids, and blockers. A no-blocker prompt compile result is not enough permission to call a generation tool. For script-related image or video generation, also satisfy the Project Style Gate before summarizing the task. Summarize the full context and ask for explicit user confirmation before any generation_prepare or generation_submit call. Only after confirmation should you call generation_submit with scope: "content_unit", capability: "image_generation" | "video_generation", and a canonical operation such as text_to_image, reference_to_image, prompt_to_video, image_to_video, or reference_to_video. Use MovScript prompt refs for upstream assets/分镜图/关键帧/resources. The submit tool compiles the backend prompt again, submits generation with the resolved resource_ids, and automatically creates or refreshes content candidates when the monitored call succeeds.
- When a saved prompt is derived from script text, scene notes, or story-heavy user wording, open
references/content-unit-prompt-craft.md before writing the prompt. Do not copy the script into edit_prompt as the final prompt. Analyze characters/entities, setting, visible action, blocking, camera, lighting, performance/audio, continuity, and negatives, then write a prompt-ready description for the specific internal output type. For image outputs, also open references/image-prompt-craft.md and write a single-frame prompt with explicit purpose, subject, composition, setting, lighting, style, details, and restrictions.
- 内容制作任务 image/video generation is output-option generation, not naked provider generation. Treat returned
generation_mode: content_unit_candidate, candidate_policy: auto_create_on_success, will_auto_select: false, and requires_user_adoption: true as the contract.
- Use
generation_submit with scope: "free" for low-level prompt channels, debugging, or non-output-task workflows only after the same full-context confirmation gate. Do not use free scope as the primary path for 内容制作任务 outputs.
- Generated 内容制作任务 candidates and selections are backend decision metadata, not workspace source-file edits. 内容制作任务 generation tools create candidates automatically on successful job polling; inspect/review and interpret after candidate writes when downstream artifact tools need refreshed decision context.
domain_create_content_candidate is the preferred backend decision path for outputs anchored to internal output tasks, including story beats, shot/dialogue/narration/subtitle/sound items, reusable assets, 关键帧, 分镜图, and audio cues. Production-level assembled playback belongs in a production editing workspace, then returns through explicit editing export/import/candidate steps. Inline asset/keyframe candidates are compatibility paths for legacy source-entity candidate workflows.
- Use
domain_register_raw_resource_as_content_unit_candidate when the RawResource already exists from upload, transform, import, editing export, or low-level generation and should enter the 内容制作任务 candidate pool. Use domain_create_content_candidate when you need to record a richer manual candidate payload. Only use decision/selection tools when the candidate already exists and the task is to adopt/reject/defer/select it.
- Candidate decision is separate from candidate creation and Resource retention. Use
domain_decide_content_unit_candidate when the workflow or UI records adopt, reject, or defer; adopt makes the generated option the stable chosen result, while reject and defer preserve the option without making it a dependency.
- Terms are strict: a RawResource is the media/resource body; a candidate is a generated/imported option record pointing to one or more outputs; a selection is the currently chosen stable option/resource; adoption is the user/workflow action that writes selection. Do not use generated, selected, adopted, and candidate as interchangeable words.
- For completed generated 内容制作任务 outputs, omit
status when calling domain_create_content_candidate or domain_create_content_candidate_batch; the backend defaults it to succeeded. If you must pass status, use only queued, running, succeeded, failed, canceled, or imported; never use completed, ready, done, selected, or accepted.
- Do not select a candidate just because generation succeeded. Select only when the user or an explicit workflow asks to use, confirm, choose, lock, or set the output.
- When generation is happening inside the AI conversation UI, make the newly written 内容制作任务 candidate available for the decision card by preserving the returned
surfaceProjectKey or legacy route projectId, plus projectUid, contentUnitId, candidateId, and resourceId on the generated attachment/candidate metadata. The user-facing choices are 采纳/放弃/待定, corresponding to adopt/reject/defer.
- Before calling a generation tool, identify whether the output is a direct story-beat video or a concrete item inside that beat, such as a shot, dialogue voice, narration voice, subtitle, sound, 分镜图, or 关键帧; then classify it as
缺规划, 可补图, 缺选择, or 可生成.
- Treat
缺选择 as a hard stop for normal generation. If required asset/关键帧/分镜图 candidates exist but none is adopted/selected, do not generate downstream video or 关键帧. Summarize the generated options and recommend that the user click/choose 采纳; only proceed without adoption when the user explicitly requests an unstable draft despite the risk.
- Keep story-beat video generation short: internally a
scene_moment should normally be one atomic beat of about 10 seconds or less. If the requested output is longer or contains several independently reviewable actions, return to planning and split it into multiple story beats before generation.
- Default to one direct
scene_moment_ref video prompt only when the story beat is one coherent short output. For longer or multi-beat requests, generate multiple short story-beat video options and hand off composition to editing.
- Prefer
gpt-image-2 for non-person image generation, 分镜图, schematic composition guides, environments, props, and 关键帧-like visual anchors when the model is available in system_model_list. For reusable human/person identity images that will feed Seedance downstream video, prefer Seedream/Seedream 5.0 lite when available so the resulting RawResource can carry provider-generated trust/provenance acceptable to Seedance review.
- Before ordinary story-beat video production, prefer generating/adopting a
storyboard_ref candidate with gpt-image-2. Prompt 分镜图 as schematic or stylized animatic frames focused on blocking, camera, lighting, and composition; avoid photoreal real-person faces or real-person likenesses, and do not use these 分镜图 as final human identity assets.
- Generation prompts must be production-ready and model-understandable: specify subject identity, setting, action, camera/framing, motion, lighting, style, continuity references, timing, and negative constraints that matter. For narrative scenes, keep the story overview only as a short orienting layer, then translate abstract plot language into cast/blocking, visible actions, dialogue, lighting, camera, and audio layers. Do not rely on the model to guess important visual/audio elements or hidden project context. For script-derived internal output tasks, run
references/content-unit-prompt-craft.md first so the prompt carries analyzed scene details rather than pasted script text. For image prompts, open references/image-prompt-craft.md and match the prompt to the image output type. For video prompts, open references/video-model-prompt-routing.md after model selection, then open references/video-prompt-craft.md and run its prompt pass before writing or updating the saved prompt.
- Always write
scene_moment_ref prompts for video generation in video-prompt form, never as a bare scene summary, still-image description, or entity note. The prompt must direct motion over time and include action progression, camera/blocking, lighting, performance/audio when supported, and relevant negatives.
- Video prompts should read like scene direction over time, not image descriptions or keyword lists. Distinguish identity/continuity, action timeline, camera/blocking, lighting/color, performance/audio, and negative constraints. For image-to-video, treat the selected image/keyframe as the visual anchor and focus the prompt on what moves, changes, or performs.
- Keep provider-specific reference syntax out of MovScript saved prompts unless a provider adapter explicitly requires it. Translate outside examples such as
@image1 / @video1 / @audio1 into MovScript semantic refs, direct {{resource::123}} refs, or input_resource_ids / reference_resource_ids.
- If the selected model/provider is Seedance-like, or the user asks for 即梦 / Seedance 2.0 / 图生视频 / 运镜 / 音乐卡点 / 分镜板驱动 / 视频延长, open
references/seedance2-prompt-methods.md after references/video-model-prompt-routing.md and adapt its methods to MovScript refs and candidate semantics. When the video involves a reusable human/face identity, open references/provider-generated-artifact-trust.md and require one of the accepted Seedance face-reference paths before downstream video: certified virtual portrait, certified real-person portrait, or a Seedream 5.0 lite text-to-image identity image with valid RawResource provider_generated_artifact.
- For Seedance/即梦 private asset-library flows, use
domain_query_remote_asset_groups when the target remote group is not explicit, then use domain_certify_asset_provider only after the relevant asset image RawResource has a selected/adopted asset_ref candidate or the user explicitly names the RawResource to certify. Pass the concrete route provider, target model, and selected/explicit asset_group_id; certification is per provider account, group, and model, so Seedance 2.0 certification does not imply Seedance 2.0 fast/pro certification. The backend mirrors remote groups/assets/model certifications as first-class records; RawResource and asset-level certification metadata are compatibility mirrors only. Certification still needs a public image URL, either supplied as source_url or generated from the configured public backend URL. Do not confuse asset-library certification with provider_generated_artifact trust provenance. If a Seedance video needs a human face reference and no certified portrait or Seedream 5.0 lite identity image exists, stop and create/stabilize that upstream portrait asset first.
- If a concrete reusable screenplay/production entity or state, such as a character/person, prop, reusable location/scene space/set, instrument, costume, or voice identity, is involved in multiple generation tasks, or the user is dissatisfied with its look or sound, treat it as a continuity asset gate: open
references/continuity-asset-prompts.md, generate/refine an asset_ref 内容制作任务 first, and wait for adoption/selection before downstream generation. Do not create settings for abstract styles/rules/moods, and do not treat a reusable location called "场景" as a scene_moment.
- Asset generation should progress from simple to complex. Generate the base identity/shape/state first, then use the selected base candidate as a reference for white-background/clean-background multi-view sheets, state variants, and more specific assets.
- When generating a setting reference set, such as all views, expression rows, costume variants, place angles, environment states, prop details, or other grouped references for one setting, do not submit parallel
asset_ref jobs from text. Stabilize one source-of-truth base asset first, wait for adoption/selection, then generate derivative asset_ref candidates whose edit_prompt includes the selected base semantic ref such as {{asset::base_character}}, {{asset::base_room_layout}}, or {{asset::base_prop_shape}}. If no adopted/selected base exists, stop and ask the user to adopt/select or generate the base first.
- When the setting reference set is for a place, scene space, room, set, exterior location, stage, or environment, follow the scene reference pack order from
references/continuity-asset-prompts.md: generate/adopt base_scene_view first; derive topdown_layout_ref from that selected base; then generate optional clean plate, corner/cardinal views, depth/line/control maps, material details, or state variants. Do not generate a top-down/floor-plan/control image from text alone when a base scene is available, and do not use a derivative that contradicts the selected base layout.
- If composition, blocking, camera motion, subject placement, or rhythm matters but is underspecified, treat it as a visual anchor gate: prepare 分镜图 first, confirm before generation, wait for adoption/selection, then confirm again before generating selected start/end or other 关键帧 before video.
- Call generation tools only when the requested output's saved prompt/context and needed upstream visual/audio anchors are ready enough for the user's goal. For a low-consistency draft, a direct story-beat video task can be enough; for composed output, generate/select concrete shot/voice/subtitle/sound items before editing.
- If a downstream output depends on an upstream generated result, and the upstream result has no selection, stop before generation. Tell the user which asset/关键帧/分镜图/generated result should be adopted, and ask for a
采纳/selection decision. Continue without selection only when the user explicitly asks for an unstable draft path.
- If the user asks to mimic a specific shot or reference video, open
references/shot-imitation-workflow.md; analyze extracted frames and create 分镜图 before downstream video generation.
Workflow
- Resolve the intended source workspace from explicit user input, a passed
projectDir/project_dir/cwd, or a Project Service locator. Do not infer it from UI focus. If project initialization, open/fetch state, or Project Service prompt context is unclear, run the project skill's Project Management Gate before generation.
- Read workspace context when the request references project entities, scenes, script passages, keyframes, asset slots, or house style. If the request is unclear about story or continuity, read the script source before authoring the generation prompt. Call
domain_read_project_context_snapshot, then use domain_overview, domain_query_production_context, domain_query_assets, and 内容制作任务 artifact tools before reading many files. If the task is script-related image or video generation, verify a confirmed style prompt or confirmed style_reference_resource_ids, or route to the Project Style Gate before any executor runs.
- Identify the exact material to make: a scene video, shot image/video, dialogue voice, narration voice, subtitle, sound effect, music, reusable asset reference, 分镜图, or 关键帧. Map it internally to
scene_moment, expression_unit, asset_ref, storyboard_ref, or keyframe_ref only for tool calls.
- Decide whether the current goal needs strong consistency evidence or a fast draft path. Do not block a simple draft only because optional setting, asset, keyframe, or storyboard references are absent.
- If continuity assets are required, open
references/continuity-asset-prompts.md and prepare/refine the asset_ref prompt first. For a setting reference set, create/generate the source-of-truth base asset_ref first and wait for adoption/selection before derivative views or states. For a scene reference pack, generate/adopt base_scene_view before topdown_layout_ref, and make later angle/detail/state assets cite the base scene plus the selected top-down layout when available. Use selected simpler asset resources as references for more complex asset variants by writing refs such as {{asset::base_character}} or {{asset::base_scene_view}} in the downstream 内容制作任务 edit_prompt, then compile it with domain_build_content_unit_backend_prompt. Stop before asset_ref generation and summarize the full context for confirmation. After asset generation, stop before downstream generation until required asset candidates are adopted/selected and the compiled prompt has no blockers. If candidates already exist, recommend adoption/selection instead of regenerating or moving downstream.
- If visual anchoring is required and composition is underspecified, prepare 分镜图 candidates first and stop for full-context confirmation before generation. After 分镜图 generation, stop before 关键帧 or video generation until the 分镜图 candidate is adopted/selected. Reference selected storyboard outputs in downstream prompts with
{{storyboard::id}} rather than raw resource ids. If 分镜图 or 关键帧 candidates already exist, guide the user to adopt/select the best one before continuing.
- Call
system_model_list before generation unless the user or UI already provided a valid model_id. Prefer gpt-image-2 for storyboard/non-person image candidates and Seedream/Seedream 5.0 lite for human identity images intended for Seedance, but only use model IDs actually returned by the system/UI. For video generation, open references/video-model-prompt-routing.md after model discovery or selection to align prompt structure with model capabilities. If the model/provider or user request is Seedance-like, also open references/seedance2-prompt-methods.md.
- Use
system_resource_library_query when you need existing MovScript images/videos/text/audio. Only returned RawResource.ID values should be passed as input_resource_ids or reference_resource_ids.
- When you need to visually inspect an existing image RawResource, call
system_resource_image_read with its resource ID. When you need to inspect a video RawResource, call system_resource_video_extract_frames; use mode, timestamps_sec, range, or burst parameters for fine-grained frame selection, and do not request or read the original video blob for vision.
- When a frame, crop, resized image, contact sheet, trimmed clip, or extracted audio must be reused by generation or written as a candidate, create a RawResource with a neutral transform tool:
system_resource_video_extract_frame_to_resource, system_resource_video_extract_frames_to_resources, system_resource_image_transform_to_resource, system_resource_video_contact_sheet_to_resource, system_resource_video_trim_to_resource, or system_resource_video_extract_audio_to_resource.
- When a generation needs a simple visual instruction layer, call
system_resource_image_annotate with structured marks. The tool stores the annotated SVG at artifact_path and returns metadata only; call system_resource_upload with artifact_path so the guidance image becomes a RawResource.
- When selected shot/voice/subtitle/sound/music materials must be stitched into a story-beat or production edit, switch to
production-editing for production-bound workspaces, then use system_edit or remotion according to the opened workspace kind. For direct system editing, switch to the editing skill, create a MediaEditingProject with editing_project_create, render through editing_task_render_create, and only then explicitly import or write a candidate if the user/workflow asks for it.
- Use
system_shot_library_query for camera, composition, motion, narrative, emotion, and production-pattern references. Treat shot-library records as prompt guidance unless they also include a usable MovScript RawResource ID.
- Use
system_external_resource_source_list and system_external_resource_search only for external provider discovery. External results must be imported into MovScript before they can be used as generation resource IDs.
- For image output-task requests, first ensure the generation tool choice is satisfied when the user request was generic. Then satisfy the Saved Prompt Gate, open
references/image-prompt-craft.md; also open references/content-unit-prompt-craft.md when the image is derived from script/story material. Prefer supplementing 关键帧 when the shot is clear but visual anchors are missing. If the image is script-related, satisfy the Project Style Gate before generation. After the saved prompt, references, and model-understandability audit are clear enough, summarize the full generation context and ask for confirmation. Only after confirmation, use generation_submit with scope: "content_unit", capability: "image_generation", and operation: "text_to_image" | "reference_to_image" | "edit_image". Use generation_submit with scope: "free" only for free generation, debugging, or non-output-task workflows after the same confirmation gate.
- For asset references, open
references/continuity-asset-prompts.md and references/image-prompt-craft.md; prefer plain white or very clean backgrounds, multi-view/reference-sheet views when useful, and weakly plot-bound images unless the user asks for scene-specific imagery. Use selected base assets as references for more specific asset states or variants. When the request is for a setting reference set, generate/adopt the base first, then generate derivative views/states with semantic refs to that base; do not batch independent text-only prompts for the whole set. For scene/place reference packs, treat base_scene_view as the visual mother reference and topdown_layout_ref as the structural mother reference; both must be selected before relying on angle/control/detail/state derivatives for downstream video.
- For 分镜图 or shot-image requests, prefer supplementing 分镜图 when the shot is clear but visual anchor assets are missing. Open
references/image-prompt-craft.md, use gpt-image-2 when available, and prompt the result as a schematic/stylized storyboard frame without photoreal real-person likenesses.
- For reference-shot imitation, extract frames across the full reference clip, materialize useful reference frames or contact sheets as RawResources when they will condition generation, analyze the shot, create/update shot/分镜图/关键帧 structure, create a 分镜图 or
storyboard_ref internal output task, and require selection before dependent video generation.
- For video output-task requests, first ensure the generation tool choice is satisfied when the user request was generic. Then satisfy the Saved Prompt Gate, open
references/content-unit-prompt-craft.md when the source is script/story material, then open references/video-model-prompt-routing.md and references/video-prompt-craft.md, classify the prompt mode, and write or refine the saved edit_prompt using semantic refs, the confirmed project style baseline, plus the model/scenario routing prompt pass. Treat every scene_moment_ref video prompt as a video prompt with motion, timing, camera, blocking, lighting, and negatives, even when the source story beat is already well described. Run the model-understandability audit before prompt compilation: every sentence must either name a resolved reference, describe something visible/audible, direct camera/motion/timing, or constrain a concrete failure mode. Compile it with domain_build_content_unit_backend_prompt and inspect blockers before calling generation. If blockers show missing adoption/selection, stop and ask the user to adopt/select the upstream candidate. If this is script-related video and the project has neither a confirmed style prompt nor confirmed style references, stop and establish one of those style baselines first. If there are no blockers, summarize the full generation context and ask the user to explicitly confirm the specific video generation task/tool call before calling generation_submit. Only after that confirmation, use generation_submit with scope: "content_unit", capability: "video_generation", and a canonical video operation after the focused story beat or shot-video output is ready enough. Use generation_submit with scope: "free" for free generation, debugging, or non-output-task workflows only after the same confirmation gate.
- For voiceover requests, prepare the full context first, then ask for confirmation before
generation_submit with capability: "audio_generation" and operation: "text_to_speech". Pass script text in prompt, voice/language controls as explicit params, and selected reference audio only as typed reference_assets when the model supports it. Do not use transcription models such as gpt-4o-transcribe for voiceover.
- For speech-native conversation or omni audio reply requests, prepare the full context first, then ask for confirmation before
generation_submit with capability: "audio_generation" and operation: "speech_to_speech". Pass the user turn in prompt, optional source audio RawResource IDs in input_resource_ids, and voice/language/output-format controls in params. Treat the returned resource as generated audio plus optional transcript text.
- For music and sound-effect requests, prepare the full context first, then ask for confirmation before
generation_submit with capability: "audio_generation" and operation: "music_generation" or operation: "sound_effect_generation". Treat the result as source audio; do not place it on a timeline, mix it, or write it as final domain state until the user/workflow asks.
- For subtitle and speech-language requests, prepare the full context first, then ask for confirmation before
generation_submit with capability: "audio_generation" and operation: "speech_to_text" for transcription, operation: "speech_translate" for speech translation, or operation: "forced_alignment" for forced alignment. Pass source media or transcript RawResource IDs as typed reference_assets; do not burn subtitles into video in generation.
- Pass semantic upstream images/videos/audio/subtitles through saved prompt refs when dependency tracking and selected-candidate semantics matter. Compile the output task with
domain_build_content_unit_backend_prompt before generation so selected assets/分镜图/关键帧 become resolved resource inputs through backend decision metadata. Use explicit input_resource_ids / reference_resource_ids only for direct raw resources, uploaded guidance, low-level generation, or model-specific controls outside prompt compilation. Include the confirmed style prompt from domain_read_project_context_snapshot.prompt_preview/enabled style rules in the model-facing prompt, and pass project style/reference resources from style_reference_resource_ids as reference_resource_ids when they exist and the selected model supports them.
- Poll generation tool calls with
generation_job_get only for calls that were already confirmed and submitted, or when the user explicitly asks to inspect an existing call. Poll until terminal is true; successful terminal polls for output-task image/video calls automatically create or refresh content candidates. Verify returned resource/candidate ids, preserve discoverable names/status/provenance for returned materials, and preserve any additional materializable artifacts that were not auto-registered. Use verbosity: "summary" for repeated polling and verbosity: "debug" when inspecting prompt/provider details. When tracking multiple calls, use generation_job_get_batch instead of issuing one call per job.
- Do not manually call
domain_create_content_candidate after generation_submit output-task image/video calls; the 内容制作任务 generation monitor owns candidate creation. Use generation_result_register, domain_create_content_candidate, or batch only for transformed/imported/manual outputs or low-level generation tool calls that intentionally bypassed the content_unit scoped path. Preserve the RawResource first, then create/register the candidate when needed, with a user-readable title and status.
- For external generation systems such as LibTV, first satisfy the Saved Prompt Gate and, for script-related image/video work, the Project Style Gate. Never leave the external URL or canvas node as the only output. Download/materialize every generated text/image/video/audio result, including low-quality, unselected, alternate, poster/preview, metadata, and draft outputs, as agent-accessible artifacts; upload them to MovScript RawResource with
system_resource_upload or system_resource_upload_batch; and then manually register target outputs as 内容制作任务 candidates when they target a story beat, shot/dialogue/narration/subtitle/sound item, reusable asset, 分镜图, 关键帧, audio cue, style/reference batch, or subtitle.
- If the user is choosing among generated candidates, use
domain_decide_content_unit_candidate with decision: "adopt" | "reject" | "defer". Treat adopt as selection, reject as a rejected-but-discoverable candidate, and defer as unresolved candidate.
- Use
domain_select_content_unit_candidate or its batch variant only for legacy/explicit workflows that confirm the candidate should become the chosen output/reference without the richer decision status.
- Inline legacy candidate tools are not granted by the default generation skill. For compatibility/migration source-entity candidate flows only, use the CLI temporary fallback
movscript domain candidate legacy ... --json. For normal asset generation, create or use an asset_ref 内容制作任务 and write/select candidates through the standard candidate tools.
- Run
domain_inspect after content candidate/decision/selection writes, then run domain_interpret when downstream artifact tools need refreshed backend decision metadata. Use domain_regeneration_plan after interpret when the change may stale downstream generated media.
Notes
- MovScript MCP calls execute through the daemon MCP endpoint or a cloud/external runtime gateway; the provider-facing stdio adapter is only a protocol bridge. If tools report missing runtime, call
movscript_runtime_status and explain the missing Data/Project/Editing/Media/gateway capability instead of telling the user Desktop or cloud is mandatory.
- Pass
projectDir/cwd for generation and domain source reads/writes, and pass projectUid for scoped backend candidate/decision writes. MCP must not infer project.
- Candidate/selection writes for current UI projects must use scoped project-data through the domain candidate tools. Never recover from scoped-store errors by switching to top-level
movscript candidate ..., MOVSCRIPT_PROJECT_ID, or legacy project decisions.
- Prefer
model_id values returned by system_model_list; do not invent provider-specific model identifiers.
- Keep generation prompts grounded in project context, resource-library records, and shot-library references when available, but make the final model-facing prompt self-sufficient after refs resolve. Do not leave hidden source assumptions for the model to infer.
- Job prompt display may include both
compiled_prompt_text and provider_prompt_text. MovScript resource tokens should be preserved until Data Service resolves them into ordered placeholders such as 图片1 before the adapter call; use input_resource_ids, generation_intent.reference_assets, and semantic_ref_replacements for prompt debugging.
- In the final answer, briefly report the focused story beat or concrete output item's readiness and whether the next action is planning, supplementing 关键帧/分镜图, calling a generation tool, composing an edit plan, choosing a dependency, or choosing a generated result.
- Do not pass MCP resource URIs or external provider URLs to
input_resource_ids / reference_resource_ids; those fields accept MovScript RawResource IDs.
- Preserve UI review boundaries. Do not treat generated resources as final accepted domain state until the user/workflow records
adopt or a legacy selection, and domain_inspect plus domain_interpret refresh the relevant source or backend decision metadata.
- Use
domain_production_status_summary before broad production review/generation when you need a compact view of settings/assets, storyboards, keyframes, 内容制作任务 candidates, selections, stale hints, and blockers.
- Open
references/reference-index.md when choosing which generation reference files to load. Common routes include references/external-generation-bridge.md for LibTV/other external generation returning into MovScript, references/model-usage.md for direct-vs-composed generation decisions, references/prompt-mode-router.md for image/video mode routing, reference roles, quality scoring, and failure diagnosis, references/content-unit-prompt-craft.md for script-to-saved prompt conversion, references/image-prompt-craft.md for image prompts, references/video-model-prompt-routing.md plus references/video-prompt-craft.md for video prompts, references/seedance2-prompt-methods.md for Seedance-like requests, references/provider-generated-artifact-trust.md for Seedance/Seedream trusted-reference windows, references/continuity-asset-prompts.md for reusable assets, references/candidate-selection-flow.md for candidate writes/selections, references/resource-id-rules.md for mixed resource inputs, and references/shot-imitation-workflow.md for reference-shot imitation.