Direct cinematic shots and coverage for AI-assisted content production. Use when an agent must translate story intent into framing, shot size, lens and perspective, camera position, movement, blocking, screen direction, lighting-aware coverage, continuity,…
calesthio/generative-media-skills
SkillsMP has collected 153 skills from calesthio/generative-media-skills. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 153
- GitHub stars
- 154
- GitHub forks
- 30
Skills in this repository
Showing 40 of 153 collected skills.
Provider-independent production guidance for converting observed video into precise, objective, temporally ordered language. Use for shot descriptions, searchable metadata, dataset captions, reference logs, generation-prompt handoffs, and analysis across…
Provider-independent quality assurance for AI-generated and AI-assisted media. Use when reviewing, accepting, revising, or reporting on images, video, audio, avatars, ads, product content, social clips, explainers, localization, mixed-source edits, captions,…
Provider-independent reference media analysis for generated-media production. Use when an agent must analyze reference images, videos, audio, style boards, product shots, brand assets, mood boards, storyboards, performances, prior cuts, or client examples and…
Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions. Use for pre-caption critique and post-caption revision, aspect-by-aspect factual review, critique precision/recall/constructiveness,…
Provider-independent workflow for producing terminal, IDE, documentation, browser, desktop, and application demonstration videos. Use when choosing authentic capture, browser automation, synthetic UI, or hybrid treatment; scripting actions and narration;…
Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Use for beat-, onset-, phrase-, energy-, silence-, or spectrum-reactive visualizers, edits, typography, and generated compositions;…
Provider-independent production guidance for deterministic Canvas 2D and p5.js animation. Use for particles, fields, trails, weather, procedural textures, generative geometry, and lightweight 2D simulations that need fixed media dimensions, seeded…
Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and…
Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and…
Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Use for TTS 1.0/2.0 selection, V3 bidirectional or unidirectional streaming, asynchronous long-text submit/query jobs, current voices and controls, timestamps and SSML caveats,…
Production guidance for reusable layered 2D character rigs in browser-rendered and video media. Use for SVG hierarchies, pivots and constraints, FK/IK planning, pose and facial libraries, acting, cycles, deterministic handoff, crop variants, and rig QA; not…
Production guidance for converting sourced data and approved claims into truthful, accessible, deterministic animated visualizations with D3. Use for data-driven charts, maps, networks, hierarchies, transitions, annotations, responsive video variants,…
Production guidance for authoring, integrating, and reviewing GSAP animation in browser-rendered media. Use for deterministic kinetic typography, SVG drawing and morphing, motion paths, FLIP transitions, responsive motion systems, or frame-seekable GSAP…
Production guidance for assessing, exporting, packaging, integrating, capturing, validating, and handing off Lottie vector animations. Use for Bodymovin/Lottie JSON, dotLottie archives, renderer and player compatibility, fonts and glyphs, image assets,…
Production guidance for planning, building, animating, capturing, and reviewing complete Three.js scenes for rendered media. Use for browser-native 3D product shots, title sequences, procedural worlds, glTF scene assembly, camera and lighting animation,…
Use this skill to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery. It covers requirements, topology repair, scale and axes, UVs, PBR baking, LODs, static…
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social content. Use when planning, mixing, repairing, mastering,…
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video recuts, training media, and localized content. Use when…
Provider-independent media quality control and delivery direction for generated images, video, audio, ads, product content, social clips, broadcast and streaming assets, localization packages, and mixed-source edits. Use when an agent must define final export…
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media generation; use when a task involves ComfyUI workflow JSON/API…
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables. Use when finalizing images, image sequences, video, audio, captions, overlays, social/export variants, checksums, manifests, delivery specs, and QA…
Provider-independent production workflow for AI agents assembling generated or source media into HyperFrames HTML/CSS/JS videos. Use for HyperFrames composition planning, scene architecture, media custody, animation/timing, captions/audio, deterministic…
Select, integrate, and operate multi-model image-generation gateways with model-specific schema discovery, version policy, asynchronous jobs, webhooks, spend approval, safe inputs and artifacts, data-governance review, and billing observability. Use for…
Plan, prompt, generate, reference, edit, motion-control, and quality-review videos with Kuaishou's direct Kling AI video platform and API, especially Kling VIDEO 3.0, 3.0 Omni, 3.0 Turbo, native audio, multi-shot, elements, start/end frames, and image/video…
Generate and edit synchronized audio-video with LTX-2.3 using the official hosted LTX API or official local/open-weight LTX-2 repository. Use for text-to-video, image-to-video, audio-to-video, retake, extend, HDR conversion, local inference, checkpoint…
Plan and direct Midjourney Video V1 image-to-video work through the official website or Discord, with human operator handoff, motion prompts, start/end frames, loops, extensions, resolution and batch budgeting, privacy, rights, safety, and delivery QA. Use…
Generate and operate Tencent Hunyuan video through the managed TokenHub HY-Video-1.5 API or official local HunyuanVideo repositories. Use for text-to-video, image-to-video, hosted job lifecycle, pricing and region planning, local checkpoint selection,…
Plan and integrate ShengShu Vidu Open Platform video generation with current model/mode selection, reference consistency, pricing, exact approval, async task handling, and safe artifact custody. Use for Vidu API text-to-video, image-to-video, start/end-frame,…
Use this skill for provider-independent reality capture that turns photographed or video-derived real places, objects, aerial sites, interiors, or people into photogrammetric meshes, NeRFs, or 3D Gaussian splats. It guides representation choice, capture…
Use this skill to plan, direct, finish, and QA provider-independent mono or stereo 180-degree, 360-degree, spherical, spatial, wide-FOV, and immersive video. It covers format choice, capture rigs, stitching, nadir and seam management, stereoscopic comfort,…
Use this skill to plan, direct, troubleshoot, and hand off provider-independent live or hybrid livestream event productions, including run of show, crew, capture, audio, graphics, switching, guests, encoding, ingest, transport choices, latency, redundancy,…
Provider-independent still-image post-production skill for RAW/rendered intake, nondestructive development, exposure, white balance, tone, color, ICC proofing, masking, repair, compositing, truthful retouching by genre, generative disclosure, sharpening,…
Plan, troubleshoot, and hand off provider-independent in-camera VFX and LED volume virtual-production work, with Unreal Engine/nDisplay/Live Link awareness. Use for ICVFX suitability, LED stage geometry, frustums, tracking, lens calibration, sync, render…
Use this skill when an agent needs production image or video understanding with Amazon Rekognition: labels, objects, scenes, OCR, moderation, image properties, Custom Labels or moderation adapters, stored-video analysis, conditional streaming-video workflows…
Use this skill when an agent needs Google Cloud Vision API for still-image understanding: labels, object localization, OCR, document text, SafeSearch, image properties, crop hints, web detection, batch annotation, Cloud Storage based pipelines, confidence…
Use DeepMotion Animate 3D for markerless video-to-3D human motion capture through the cloud portal, sales-gated API, or real-time SDK; covers capture planning, body/hand/face/multi-person limits, custom characters, retargeting, refinement, exports,…
Use this skill when planning, capturing, processing, validating, troubleshooting, or exporting markerless human motion capture with Move AI products, including Move One single-camera workflows and Genesis multi-camera studio systems. It covers provider…
Produce anime-style and stylized 2D animation with generative image and video tools as a craft discipline, not a filter. Use when a request asks for anime, manga-style, cel-shaded, sakuga, shonen/shojo/seinen/slice-of-life, mecha, or retro-90s-OVA animation;…
Produce full-length audiobooks and long-form narration with generative voice tools. Use when the task is to turn a manuscript or long text into hours of spoken audio: preparing the manuscript for narration (front/back matter, footnotes, tables, dialogue),…