Skip to main content

studio-use

Record the screen and edit video in studio-use. Capture a display or window with its keystroke, click and cursor streams; import Screen Studio bundles or any footage; cut dead air, add spring-driven zooms, mask and caption, place Remotion code clips, author project compositions and effects, render frames and export, all through the attributed, undoable op surface the timeline UI shares. Use for ANY studio-use request — recording, opening or creating a project, editing a screen recording or any footage, tightening a demo, captioning, restyling, exporting.

Source facts

Repository
ShawnPana/studio-use
Last source activity
September 18, 2026 at 21:11
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
studio-use
description
Record the screen and edit video in studio-use. Capture a display or window with its keystroke, click and cursor streams; import Screen Studio bundles or any footage; cut dead air, add spring-driven zooms, mask and caption, place Remotion code clips, author project compositions and effects, render frames and export, all through the attributed, undoable op surface the timeline UI shares. Use for ANY studio-use request — recording, opening or creating a project, editing a screen recording or any footage, tightening a demo, captioning, restyling, exporting.
# Studio Use ## Boot — the first move, always One call answers everything: are we connected, what project is open, what it holds. Run it FIRST and report ready — when it succeeds, the setup checks further down are already satisfied; skip them. ```bash studio-harness <<'PY' s = api("GET", "/api/state") print("studio:", s.get("serverUrl")) print(outline()) print("compositions:", [c["name"] for c in s.get("compositions", [])]) print("watching:", s.get("uiViewers")) PY ``` No `which studio`, no `--help`, no directory hunting, no ffmpeg check before this — those diagnose a FAILED boot, they are not the boot. If this call fails, run `studio-harness doctor` — it names the broken piece — and read `install.md` (beside this file; the harness itself ships in `scripts/` next to it). First time with this user? `onboarding.md` is the walkthrough. Then boot again. The harness auto-opens the user's editor on your first edit, so there is nothing else to prepare: one call, one ready message, await direction. ## Principles 1. **Edit the document, not the pixels.** Every change is a named op applied to a JSON timeline. You never drive a UI and never shell out to ffmpeg to make an edit. Ops are the whole vocabulary — `studio ops` lists them, `studio ops <type>` prints the JSON Schema. 2. **Declare who you are.** History is the source of truth. The harness attributes you agent/external automatically; pass `--author agent --origin external` on every CLI `apply`. An edit logged as a human's is a lie the user can't detect later. 3. **One intent, one entry.** Each `apply` becomes one history entry the user can undo alone. Give it a `--reason` in their language, not yours: `"cut 4.2s of dead air at the start"`, not `"remove_range 0-4200"`. 4. **Cheap signals first, pixels last.** Read the timeline, the event streams and the transcript to narrow the search. Render frames to confirm. For privacy blurs and covers, inspect source frames to find the exact interval where the target is visible. 5. **Act, then report.** Don't ask permission for a reversible edit — make it and say what you did and how to undo it. Do ask before `export` (slow), before starting a screen recording, and before anything destructive to files on disk. 6. **The user owns taste.** Cut lengths, zoom levels, spring feel, where the logo goes — propose in plain language and defer. The rules below are correctness, not taste. ## Hard rules Deviating from these produces silent corruption, not an error message. 1. **Two clocks, never confused.** Clips are placed in **timeline** time (`startMs`). Zooms, masks and connected-track clips are stored in **source** time (`sourceStartMs`/`sourceEndMs`) — a position inside the recording. Passing a timeline time to `add_zoom` puts the zoom somewhere the user didn't ask for, and it will still validate. 2. **Apply ripple cuts LAST-FIRST.** `remove_range` with `ripple: true` shifts everything after it. Sort your cuts descending by start time before applying, or every cut after the first lands in the wrong place. (`studio cuts` already emits them in this order — preserve it.) 3. **Never cut inside a clip you haven't inspected.** Check `outline()` for what's actually on the timeline before removing a range. 4. **Verify with a frame, not an export.** A single frame costs ~2s; a full export costs minutes. Render one frame at the boundary you changed and look at it. 5. **Never write into the studio-use repo.** Compositions belong to the PROJECT — `compositions/<name>.tsx` in the bundle (see "Project compositions") — never to `packages/renderer/compositions/`, which is the product. Everything else — project JSON, media, recordings — lives in the user's project directory too. Media goes to `<project>/media/`, never a temp dir, because macOS purges `/var/folders` and the footage disappears. 6. **Say the fraction out loud.** If a cut list removes more than ~30% of the timeline, state the percentage before applying. The detector abstains above 95%, but between those it will happily do what you tell it. 7. **Don't fabricate targets.** `add_zoom` takes a normalized `{x, y}` in the capture frame. Get it from `cuts()`/click clusters or from looking at a frame — never guess a number that "looks central." 8. **Keep every privacy blur or cover as tight as possible.** Use the smallest area that fully obscures the target. Start on its first visible frame and end on the first frame where it is absent, including any fade frames. Split separate appearances instead of covering the gap. Verify the rendered frames immediately before and at each boundary; after an export, verify those boundaries in the MP4 too. Recheck after cuts, speed changes, reframing, or zooms. Never add guessed lead-in or trailing time. 9. **A recording is the user's screen.** Start a capture only when the user asked for one in this session, say when it starts and when it stops, and never leave one running when you hand back. ## Setup check On a cold start, verify rather than reinstall: `studio-harness doctor`. It prints the studio, the open project and who is watching, then `ready.`, or names the one thing that is broken. Fix only that, following `install.md`. Install nothing else. ## The harness The preferred way to drive a studio is `studio-harness` — thin like browser-harness: helpers are sugar, `api()` is the escape hatch to the ENTIRE server surface, and nothing is permission-gated. **The user watches everything.** Every edit you make must be visible in the user's open editor — this is a product principle, not a courtesy. The harness handles it automatically against a running studio. If you are driving a LOCAL project, launch its editor FIRST and keep it running: ```bash studio app projects/demo --port 5173 & # opens the user's browser; edits stream live ``` `5173` is the harness's default; any other port needs `STUDIO_URL=http://localhost:<port>`. Never edit a project nobody can see. ```bash studio-harness <<'PY' print(outline()) # the timeline, as text blocks = speech() # utterance boundaries, timeline ms apply({"type": "remove_range", "fromMs": 214000, "toMs": 290000, "ripple": True}, reason="cut the idle stretch") png = frame(21500) # render one frame, LOOK at it api("GET", "/api/state") # escape hatch: everything the UI can do PY ``` Helpers: `state doc outline history apply undo redo speech cuts frame export_video transcribe put_transcript phrases words find_words cut_span timeline_view wait_for_change show api`. The studio to drive is `$STUDIO_URL`, default `http://localhost:5173` — the editor `studio app` opened. A `studio_helpers.py` in the working directory (or `$STUDIO_HELPERS`) is auto-loaded: accumulate your own vocabulary there. The `studio` CLI below remains for one-shot commands and scripts. ## Recording The sidecar runs on this machine (macOS 14+ via ScreenCaptureKit, Linux via X11 + FFmpeg) and writes into the open project. The user records from the editor's top bar — **Record**, pick a display or window, mic and camera optional, **Stop** — and the clip lands on a new track with its keystroke, click and cursor streams on the same clock as the frames. When the user asks you to run the capture, the same three routes are yours: ```python devices = api("GET", "/api/devices") # displays, windows, cameras, microphones main = next((d for d in devices["displays"] if d.get("isMain")), devices["displays"][0]) api("POST", "/api/record/start", {"displayId": main["id"]}) # ... the user does the thing; api("GET", "/api/record/status") while you wait ... done = api("POST", "/api/record/stop", {}) print(done["recorded"]) # {durationMs, warnings}; the clip is on the timeline ``` `record/start` takes `displayId` OR `windowId`, plus optional `cameraDeviceId`, `microphoneDeviceId`, `systemAudio` and `captureKeystrokes` (default true — keep it; without keystrokes the idle detector reads only clicks and audio). One capture at a time; `stop` imports it as ONE history entry and returns the new state. The files land in `<project>/media/recording-<stamp>/` — never move them. Window capture on Linux follows the window id: keep the window mapped and its size fixed, or capture the display. **What you get from a recording, and what not.** A capture (or a `.screenstudio` bundle) carries event streams: `cuts()` reads idle spans from them, `auto-zoom` clusters clicks into zoom targets, and the cursor is re-rendered from the stream, so it survives zooms crisp. Outside footage has none of that — see "Bring in outside footage". **Side-by-side demos.** To show an automation next to the app it drives, give the app the left half and the tool the right half, no border or gutter, each centered within its half. Resize live windows to fill their halves where possible; otherwise preserve aspect and fit the whole content. Record both together, or align the two recordings' clocks and keep cuts synchronized. Inspect the middle boundary and both panes before exporting, and apply the tight privacy-mask rule to either side. ## Finding content in footage "Where does X happen?" has no op — footage is pixels and sound until you build a content layer for it. Do it in this order, cheapest first: 1. **Words — the studio's own capability, use it first:** ```python transcribe("as_xyz") # once per asset, cached forever print(phrases("as_xyz")) # the packed reading view, source time span = cut_span("as_xyz", "so basically let me start over") apply({"type": "remove_range", "fromMs": span["startMs"], "toMs": span["endMs"], "ripple": True}, reason="cut false start") ``` Quote what you READ in the packed view — the ASR's words, not the words you expect. A 404 from `phrases()` means NOT TRANSCRIBED, never "no speech" — absence is loud by design. A 501 from `transcribe()` means the studio has no whisper installed: run one yourself on the file in `<project>/media/` (`mlx_whisper file --word-timestamps True --output-format json`) and `put_transcript(assetId, words)` — the studio's cache is canonical whichever engine listened. Platform subtitles (`yt-dlp --write-auto-subs`) locate WHAT was said, never WHEN — they drift 0.3–0.7s. **Cut craft (numbers earned by renders — treat as physics):** never cut inside a word; `cut_span` pads 80ms and snaps to word boundaries for you; land cuts in silences ≥400ms; 150–400ms → look at `timeline_view(assetId, fromMs, toMs)` (filmstrip + waveform, one image) before committing; <150ms is unsafe. Extend past laughs and reactions, never through them. **Provenance first**: check the asset's `sourcePath` for the original URL before guessing — and when YOU download media, record its URL in `add_asset.sourcePath` so later sessions never have to guess. Only search a platform by title when the footage is KNOWN to be public; never send a private file's name to an external service on a hunch — for private footage, work local-only (whisper, sheets, silence). 2. **Events**: on a recording, `cuts()` reads idle spans from the keystroke, click and cursor streams — where the user did nothing is nearly free to find. 3. **Timing**: always re-measure against the audio itself — `ffmpeg -af silencedetect` on the local file (default floor 0.35s; drop to d=0.10 for rapid speaker handoffs). 4. **Visuals** (who is on screen, shot changes, where a secret appears): `timeline_view` for a window, or a contact sheet — one frame every N seconds via ffmpeg — then LOOK at it. Binary-search with finer sheets to land a boundary; it beats scrubbing. **Before calling an export done:** sweep the RENDERED file's cut boundaries (±1.5s each) with frames or a timeline view — visual jumps, audio pops, captions colliding, a mask that ends a frame early — at most three passes, then hand over. Verify the output, not your intentions. Other projects on the machine (histories, prior cuts of the same source) are HYPOTHESES, not answers: they solved a different request. Use them to narrow the search if you like, but re-verify every range against the actual footage — frames at the proposed boundaries, audio at the seams — before you cut. State in your history label how a range was determined. ## Workflow skills — crystallize what worked When a user asks for a STYLE of edit they'll plausibly want again ("tiktok style", "podcast clips", "demo tightening"), capture the workflow as its own skill once it works: - Write `~/.claude/skills/studio-use-<name>/SKILL.md` with a one-line `description:` in frontmatter saying when to use it. It becomes invocable as `/studio-use-<name>` immediately. - Record the RECIPE, not a transcript: the op sequence, the parameters that mattered, the measurement steps (never guessed values), the verification, and the judgment calls ("wides crop badly past 2×"). - Pair it with `studio_helpers.py` functions when code carries the weight; the skill says WHEN and WHY, the helper does HOW. - Improve the skill when a later run teaches something — these files are the user's accumulated taste, in their own machine, in their own git if they keep one. An existing example to follow: `studio-use-tiktok`. ## The surface ```bash studio new <name | project.json> [--name "..."] [--fps 60] [--size 1920x1080] studio import <bundle.screenstudio> [-o name | -o project.json] studio outline <project> # the timeline as you should read it studio ops [type] # the op surface, describing itself studio cuts <project> [--min-idle ms] [--noise dB] # analyse, change nothing studio speech <project> [--noise dB] [--min-pause ms] # utterance boundaries studio auto-zoom <project> [--level 2] # zooms from click clusters studio propose-cuts <project> # the same analysis as a reviewable take studio takes | accept <takeId> | reject <takeId> studio apply <project> '<op json>' --author agent --origin external --reason "..." studio history <project> [--undo | --redo] studio frame <project> <ms> -o frame.png # cheap: look at one frame studio export <project> -o out.mp4 [--from ms] [--to ms] # default ends at the last clip; --to may run PAST it — clip-less spans # render the background layer, black when the background is none studio app <project> [--port n] [--no-open] [--dev] # the editor; --dev hot-reloads the UI ``` `<project>` is a bundle directory (`projects/demo`) or a `project.json`. `studio cuts` and `studio auto-zoom` are *analysis* — they print what they'd do. To act, read their output and issue `apply`. ## The ops Run `studio ops <type>` for the exact schema. Grouped by what they're for: | doing | ops | |---|---| | structure | `add_track` `remove_track` `update_track` `move_track` `connect_track` `disconnect_track` | | media in | `add_asset` `add_clip` `add_code_clip` `update_code_clip` | | cutting | `split_clip` `trim_clip` `remove_clip` `remove_range` `move_clip` | | clip feel | `set_clip_speed` `set_clip_volume` `set_clip_fade` `set_clip_frame` | | camera | `add_zoom` `update_zoom` `remove_zoom` | | privacy | `add_mask` `update_mask` `remove_mask` | | wrap effects | `add_clip_effect` `update_clip_effect` `remove_clip_effect` | | look | `set_config` `set_canvas` | | project | `rename_project` | **Two track kinds, exactly.** `video` holds media clips; `code` holds Remotion compositions. There is no screen/camera/audio track kind — those differences live in FIELDS: a screen recording is a video asset carrying event streams; a camera bubble is a video clip with a `frameRect` (a normalized rect of the output frame — `set_clip_frame` moves or removes it); audio-only is a video clip whose asset has no picture (`add_asset` kinds are media types: `video`, `audio`, `image`). **The background is a layer, not config.** The bottom code track holds the built-in `background` composition on an OPEN-ENDED clip — `durationMs: null` means "runs to the content end", so it underlies whatever exists without retiming. A NEW project shows the capture edge to edge: that track starts muted and `paddingRatio`, `cornerRadius` and shadow are 0. For a backdrop, unmute the track and raise `paddingRatio` with `set_config`. Restyle it with `update_code_clip` on that clip (props: `type`, `color`, `gradientTo`). Never place other clips on the background track, and never put an outro there — it would render behind everything. Track order IS stacking order — first track composites at the bottom, last in front — so `move_track` changes what occludes what, not just tidiness. **Connected tracks — captions and overlays that ride the footage.** `connect_track {trackId, parentTrackId, side}` anchors a track to a media track's SOURCE material. From then on its clips' `startMs` (and a code clip's `durationMs`) are spans of the parent's footage, not timeline positions — cut, split or move the parent and every connected clip lands correctly on its own, with nothing to resync. Place captions by SOURCE time (where the words are in the recording, i.e. the times transcription gives you) and they survive every edit. `side: "front"` composites over the parent (captions), `"behind"` under it; the op keeps the child adjacent to its parent in the stack. A clip whose footage is cut projects nowhere — ORPHANED, hidden but intact; undoing the cut or re-trimming brings it back. `disconnect_track` freezes clips at their current on-screen positions. `remove_range` on a connected track is refused — cut the parent instead. `outline()` prints both clocks: `anchor 8.52s → 12.30-12.33`, or `orphaned (footage cut)`. `set_clip_fade` fades **sound and picture together** — one gesture, like an NLE's clip fade. A visual clip dissolves to whatever is behind it (for the footage, the background layer), so "fade out the video" at an outro is one op on the screen clip, not a composition. **Project compositions — your code, no deploy.** The project's `compositions/` directory is user space: a `.tsx` there IS a composition,
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub