CutPilot is a local-first, Codex-first editing workflow with an embedded manual review timeline. Codex makes editorial decisions from the user's brief, narration, metadata, and contact sheets; CutPilot tools make those decisions reproducible and renderable, while the user can directly adjust supported timeline details.
-
plan_director_agent creates a cross-category, non-mutating beat plan from timeline captions or a user brief. Always show gaps and obtain explicit approval before apply_director_agent.
-
Long-running index, director-plan, and runtime-audit work can use the persistent task center. Read task state, cancel on request, retry only terminal tasks, and use recovery after an interrupted host session.
-
Seven director workflows share one evidence-based acceptance audit for online media, review gates, editable timelines, variants, export settings, and strict release readiness.
-
Multi-platform delivery packs derive independent editable 16:9, 9:16, 1:1, and 4:5 timelines and submit background exports that are pinned to each exact sequence.
-
Pure motion-graphics projects can turn user-authored title/data/list/CTA scenes into reviewed two-layer SVG scenes with beat alignment, editable keyframes, validated GLSL effects, transitions, music, and SFX cues without inventing copy or values.
-
Podcast projects can turn reviewed multicamera sync plus a speaker-labeled transcript into active-speaker cuts, overlap-wide shots, independent speaker cleanup, captions, chapters, reviewed guest lower-thirds, and separate In/Out-based short-clip timelines.
-
Wedding projects can build preparation/ceremony/portraits/speeches/reception structure from annotated footage, preserve transcribed vows as the anchor, add music ducking and reviewed titles, and create separate editable full-film and highlight timelines.
-
Product-promo projects can build a reviewed Hook/problem/benefit/proof/CTA structure, semantically match Hero and detail footage, align sections to available music beats, expose gaps, and render only human-approved advertising claims as editable brand MG.
-
Explainer information cards require explicit approve, edit, or reject decisions before safe SVG rendering; approved cards become editable Motion Graphic assets on V3, while unreviewed cards keep release locked.
-
The project model supports reusable assets, multiple timelines, multiple video/audio tracks, captions, markers, history, and atomic item edits.
-
Local whisper.cpp transcription supports segment and word timestamps; models live in the user's CutPilot cache.
-
Captions can be stored, exported as SRT/TXT, and burned into rendered video without uploading media.
-
Multitrack rendering composites stacked video tracks and mixes audio tracks with per-item dB, fades, and final limiting.
-
Mechanical filler removal and long-gap compression operate on timestamped transcript cues. Semantic speech editing remains an AI decision.
-
Codex-authored factual asset annotations feed deterministic narration-to-shot ranking with repetition penalties.
-
Word-timestamp Script edits remove phrases and rebuild ripple-closed speech clips and captions.
-
Track-level high-pass/FFT denoise, loudness normalization, anchor/follower routing, and sidechain ducking are supported.
-
Editable effect stacks support color, grayscale, blur, zoom, and .cube LUTs.
-
Fade/dip-to-black boundaries and per-item visual fades are rendered non-destructively.
-
Motion Graphic assets retain their source text, kind, and colors so Codex can regenerate them after changes.
-
Local macOS voice generation and provenance-aware registration cover generated voice, image, video, music, and SFX assets.
-
Immutable local snapshots and safety-preserving rollback make autonomous revisions reversible.
-
Render QA creates a contact sheet and reports sustained black/frozen segments.
-
FCPXML exports multitrack clip timing and audio gain; CMX 3600 EDL exports the primary video track.
-
Motion Graphic and visual items support editable natural-box placement, contain/cover fit, rotation, slide/fade entrances and exits, and subtle floating motion.
-
Persistent generation jobs separate submission, status polling, and asset materialization.
-
Offline procedural fixtures cover all media kinds; OpenAI image generation and configurable synchronous/asynchronous HTTP generator services are supported.
-
Direct-authored SVG templates expose editable text, number, color, and boolean properties with local safe rasterization.
-
Ordered x/y/scale/rotation/opacity keyframes render per frame with linear, ease-in, ease-out, ease-in-out, hold, and custom Bezier-Y interpolation.
-
Contiguous adjacent clips support a real overlapping cross-dissolve without changing project timing.
-
A token-protected localhost review UI runs inside the CutPilot MCP process and can be opened in Codex.
-
Timeline markers can be added, renamed, recolored, retimed, jumped to, and deleted by either Codex tools or the embedded editor.
-
Duplicate-shot review uses local representative-frame perceptual hashes; always present candidates and similarity before removing explicitly selected repeats.
-
The review UI reads the real multitrack project, plays a rendered preview, supports timeline dragging and precise item edits, edits captions, and changes SVG MG properties.
-
Every accepted UI write creates a project snapshot; invalid writes are rejected before persistence.
-
Source-continuous split/razor operations redistribute source in-points, fades, transitions, and relative keyframes.
-
Trim handles preserve source continuity; optional ripple out-trim closes time on the track.
-
Ripple delete removes the selected range across tracks, splitting media that spans the removed interval and remapping captions/markers.
-
Gap insertion splits crossing media and shifts the right side, captions, and markers.
-
Real FFmpeg showwavespic waveforms are cached locally and displayed on audio clips.
-
Eleven production transition types render locally: dissolve, dip, four wipes, radial reveal, four slides, and standard fade.
-
Wipe/radial transitions use frame-evaluated alpha masks; slides use time-evaluated overlay positions with source pre-roll.
-
Seedance, Kling, Mureka, and sound-effect bridge profiles have dedicated endpoint/token configuration, media-kind validation, provider identity, and async polling compatibility.
-
Persistent edit context stores up to 50 deduplicated asset, item, time, canvas-region, and transcript-range references in a local sidecar. MCP resolution returns exact objects plus concise AI prompt context, while the embedded review UI can select assets, clips, and timeline times without polluting Undo/Redo.
-
Nine versioned project starters cover blank editing, Vlog, talking head, vertical social short, video podcast, wedding, explainer, motion graphics, and product promo. Seven user-facing video types are selectable in onboarding; each materializes scenario-specific tracks, routing roles, bins, caption defaults, snapping, a bounded AI brief, and a recommended workflow.
-
The embedded source panel previews local video, audio, and images, marks source In/Out, selects a compatible target track, and performs append, ripple insert, or target-only overwrite. Source slicing preserves playback mapping and keyframe timing on retained clip remainders.
-
Persistent link groups keep camera video and original audio synchronized across append/insert/overwrite placement, group movement, and aligned trimming. The embedded source monitor can create linked AV, timeline selection highlights every linked member, linked drags/trims edit the group, and users can explicitly unlink before independent work.
-
Detached local export jobs cover video, audio, FCPXML, Premiere XML, EDL, and experimental Jianying drafts. Each job renders from an immutable project snapshot, persists queued/running/completed/failed/cancelled state, exposes coarse honest phases and verification results, survives UI closure, and appears in a polling embedded export queue.
-
Two-stage local pause removal uses FFmpeg silence detection to propose raw silence plus padded keep ranges without editing. Confirmation rebuilds source-continuous linked video/original-audio segments, remaps timed captions, creates a snapshot, and remains reversible. Threshold, minimum silence, breathing padding, and minimum retained segment are explicit.
-
Local FFmpeg scene detection identifies visual hard cuts, merges implausibly short scenes, extracts midpoint JPEGs, builds a dependency-light contact sheet, and persists searchable source subclips with names, tags, annotations, exact ranges, and thumbnails. Subclips load into the source monitor or place directly through MCP.
-
Local onset analysis estimates music tempo, distinguishes stronger accents, and persists editable beat markers. A two-stage cut-on-beat workflow builds a non-mutating montage proposal from assets or detected source subclips, then atomically applies reviewed cuts to a chosen video track; the embedded Beat panel exposes the same analyze, mark, plan, and confirm flow.
-
Subtitle interchange locally imports and strictly validates SRT, WebVTT, and ASS, retaining stable cue IDs and optional offsets. Exact-count or cue-ID translation alignment adds a second language without replacing the source text; original, translated, or bilingual SRT/VTT/ASS exports round-trip timing, and the renderer burns independently styled primary and secondary lines. The embedded Caption panel exposes paths, formats, translation rows, bilingual styling, and export variants.
-
Local H.264 proxy media supports 540p, 720p, and 1080p profiles with configurable quality, optional source audio, duration/dimension verification, and source size/mtime fingerprints. Ready proxies accelerate the embedded source monitor; missing or stale proxies safely fall back to originals, project-owned files can be detached securely, and timeline/final rendering deliberately retains original asset paths.
-
Local smart reframe samples grayscale video through FFmpeg, combines visual saliency with motion energy, smooths a normalized focus trajectory, and proposes editable focus keyframes without modifying the project. Confirmed trajectories drive per-frame FFmpeg crop positions, survive split/trim/range/overwrite operations through boundary interpolation, and remain manually editable in the embedded Reframe panel. This is explicitly motion/saliency tracking rather than unverified face recognition.
-
The Codex-embedded review surface includes a visual curve preview, editable keyframe rows, add-at-playhead, deletion, validation, snapshots, and project writeback.
-
Real GLSL fragment effects compile and render locally through Chrome WebGL1, with source/uniform persistence, four presets, MCP authoring, and an embedded shader editor.
-
JSX/React Motion Graphics retain editable component source and Props, receive frame/fps/progress, render transparently through Chrome, and composite as ordinary timeline assets.
-
FCPXML imports local assets, gaps, connected clips, nested media sequences, titles/captions, markers, lanes, source ranges, audio gain, and CutPilot metadata. Premiere xmeml imports reusable file references, video/audio tracks, source ranges, clip gain, and sequence markers. Both formats also export, and EDL export remains available.
-
Experimental JianyingPro v360000 plaintext draft export copies or references local media, maps video/audio timing and transforms, writes draft metadata, and validates every reference.
-
Rich captions support four templates, safe-area placement, color/outline/background controls, and word-timestamp karaoke highlighting with embedded UI editing.
-
Video/audio items support constant speed, eased editable speed curves, and reverse; video items also support deterministic freeze frames. Curves use integrated source mapping, shared audio/video timing, source-range validation, visual editing, and montage/hero presets.
-
Ordered visual effect stacks support chroma key, rectangular/elliptical masks with feather/invert, exposure/contrast/saturation/temperature/tint and shadow/highlight color balance, ten curve presets, vignette, LUT, blur, zoom, color, grayscale, and GLSL. The embedded Picture panel includes safe JSON editing and practical presets.
-
Ordered per-clip audio stacks support high/low-pass filters, up to twelve parametric EQ bands, compression, noise gating, de-essing, stereo balance/width/phase, duration-preserving pitch shifts, and limiting. The embedded Audio panel includes Voice Cleanup, Broadcast Voice, Music Wide, and Low Pitch presets.
-
Processed audio-only export supports 24-bit WAV, FLAC, 256 kbps MP3, and 192 kbps M4A/AAC with duration based only on audible audio tracks.
-
Projects expose first-class sequence creation, duplication, activation, rename, deletion, shared assets, isolated tracks/captions/markers, and per-sequence format. The embedded review UI can switch active sequences.
-
Per-sequence In/Out zones drive exact video or audio zone exports; speed curves, reverse source ranges, captions, words, markers, and audio remain synchronized after slicing.
-
Editable SVG and JSX/React Motion Graphics export independently as transparent ProRes 4444 MOV files with verified alpha pixel formats.
-
Text-based editing supports ordered source transcript segments: Codex or the embedded Transcript panel can move or delete passages, then atomically rebuild linked video/audio clips and word-level captions with exact source continuity.
-
The local asset library supports nested bins, assignment, searchable tags/names/annotations, type/generated/online filters, missing-media dependency reports, explicit relinking, and exact-filename recursive folder recovery. Relinking retains stable asset IDs and all editorial metadata.
-
Every semantic saveProject change enters a persistent, bounded 100-step Undo/Redo journal. It survives process restarts, supports branch invalidation after Undo, coexists with named snapshots, and is available through MCP plus embedded Review buttons.
-
Multicam sync locally cross-correlates scratch/reference audio envelopes, reports offsets and confidence before editing, then can build separate aligned angle/audio tracks or a source-continuous editable program cut from explicit camera switches.
-
Automatic multicam planning remains non-mutating until approved: stable/balanced/dynamic pacing scores quality and motion annotations, discourages immediate angle repetition, supports preferred cameras plus explicit directed holds, reports per-angle coverage, and feeds the editable program-cut tool.
-
Speaker-aware multicam planning consumes honestly speaker-labeled transcript cues, maps named speakers to synchronized cameras, selects an optional wide shot for overlapping speech, merges adjacent delivery, suppresses micro-shots, and reports missing mappings. It never claims diarization when labels are absent.
-
Category-first onboarding asks what kind of video the user is making before exposing the editor. Vlog is the first dedicated workflow, with a structured brief, purpose-built tracks/bins, five-part story planning, pace control, non-mutating review, and exact source-range application.
-
Before planning a Vlog, run analyze_vlog_coverage to score Hook, Setup, Journey, Payoff, and Outro coverage. Surface thin or missing sections and concrete pickup suggestions; the embedded Vlog workspace offers the same analyze, plan, review, and snapshot-backed apply flow.
-
For narration-led Vlogs, use plan_vlog_narration_broll after captions or a timestamped transcript are available. Review semantic matches, repeated assets, and explicit unmatched-line gaps before using apply_vlog_narration_broll on V2 · Cutaways.
-
After the story and cutaway tracks exist, run analyze_vlog_rhythm, then plan_vlog_rhythm. Apply only reviewed safe changes with apply_vlog_rhythm; long-take deletion and pickup decisions remain explicit manual-review items while visual prelaps, handoff transitions, audio fades, and anchor/follower ducking can be automated.
-
Use plan_vlog_sound for an assembled Vlog to build music zones, cutaway natural-sound moments, and explicitly identified SFX placements. Review missing-source warnings, then use apply_vlog_sound to write editable A2/A3/A4 clips while preserving voice anchor and music follower ducking.
-
Use plan_vlog_finishing before release to audit caption readability and build safe-area, title-card, chapter-marker, CTA, and export specifications. apply_vlog_finishing writes safe metadata and keeps title cards as explicit editable pending MG specs until their real render assets are created.
-
After applying finishing specs, use render_vlog_title_cards to create safe local SVG sources, transparent PNG assets, and editable V3 title clips. Review any cards skipped by overlap protection; text and colors remain editable through the SVG Motion Graphic workflow.
-
Before export, run analyze_vlog_release. Do not call a Vlog ready while blockers remain. apply_vlog_release_fixes may repair only mechanical settings such as audio roles, safe areas, caption enablement, and export presets; offline media, content gaps, and unrendered graphics require explicit resolution.
-
Editable sequences can switch among five social/broadcast aspect presets or custom 16-8192 pixel formats. Scale-layout mode remaps transforms, transform keyframes, mask geometry/feather, and caption sizing; output-only resolution and 24-60 fps overrides render from a clone and never mutate the project.
-
Video/audio tracks support explicit create, stable-ID rename, type-relative reorder, lock, mute/visibility, role/gain/processing, safe delete or non-overlapping item transfer. Frame-based snapping resolves deterministic guides for clip edges, markers, zones, playhead, and timeline start in MCP plus embedded drag/trim controls.
-
Use list_cutpilot_capabilities, read_cutpilot_gaps, and audit_runtime_readiness together. Remotion project rendering and concurrent local GPU Shader batches are implemented and tested. Paid commercial endpoints remain unverified until configured, and modern encrypted Jianying compatibility remains experimental and version-dependent.
-
For non-Vlog category projects, call get_category_workflow_status to read evidence-backed stages and the next recommended tool. Use analyze_category_release / apply_category_release_fixes for preflight and submit_category_release_export for the final gated MP4.
-
For a transcribed talking-head camera asset, use plan_talking_head_director first. Review every removed standalone filler cue and semantic B-roll match, then use apply_talking_head_director to rebuild linked picture/dialogue, B-roll, voice cleanup, karaoke captions, music ducking, and export settings.
-
For narration-led explainers, use plan_explainer_director to review primary visuals, secondary evidence, gaps, and fact-checkable information-card specs. Use apply_explainer_director only after review; pending information cards intentionally continue blocking release until rendered.