Provider-independent character design and continuity direction for generated media. Use when an agent must create, preserve, revise, or QA a recurring character across AI-generated images, video clips, avatars, animation, comics, ads, product mascots,…
Skills in this repository
calesthio/generative-media-skills - Page 3
SkillsMP has collected 153 skills from calesthio/generative-media-skills. Open a skill to review its source and details.
calesthio/generative-media-skillsShowing 40 of 153 collected skills.
Provider-independent lighting direction for AI-generated images, video clips, ads, film scenes, avatars, product shots, interviews, beauty work, and social content. Use when an agent must specify, translate, maintain, troubleshoot, or QA motivated lighting,…
Provider-independent performance direction for generated video, avatar/spokesperson clips, animation, ads, film scenes, social content, and voice-led media. Use when an agent must cast or direct synthetic performers, avatars, animated characters, talking…
Provider-independent production design direction for AI-generated media. Use when translating narrative, brand, advertising, product, social, film, or worldbuilding intent into sets, locations, props, color and material palettes, era research, graphic and…
Provider-independent storyboard and previsualization direction for AI agents turning briefs, scripts, treatments, ads, films, social videos, and generated-media concepts into shot-by-shot boards, animatics, previs plans, camera/blocking/movement direction,…
Translate a creative brief, brand system, or reference set into an original, portable visual-direction system for image and video production. Use for art direction, reference analysis without imitation, design tokens, palette and contrast, typography,…
Provider-independent asset continuity and version management for generated-media production. Use when an agent must track generated or reference assets, prompts, seeds, model/tool versions, manifests, approvals, edit decisions, derivatives, localization…
Provider-independent provenance, rights, consent, disclosure, and release governance for AI-generated media. Use when producing or reviewing generated images, video, audio, ads, avatars, product content, social clips, documentaries, localization, music,…
Provider-independent color grading and finishing direction for generated video, ads, trailers, product films, social clips, explainers, and mixed-source edits. Use when planning, directing, reviewing, or QAing color correction, shot matching, exposure,…
Provider-independent editing and montage direction for AI agents producing generated videos, ads, trailers, social clips, explainers, product videos, documentary-style pieces, and music- or beat-driven cuts. Use when planning, revising, QAing, or handing off…
Provider-independent motion graphics direction for AI agents producing title sequences, lower thirds, explainer graphics, product UI overlays, social ads, kinetic typography, diagrams, data callouts, brand films, and video post. Use when planning, prompting,…
Provider-independent title design and kinetic typography direction for video, animation, ads, explainers, social hooks, brand films, lyric/text videos, title cards, main titles, lower thirds, captions-as-design, and other motion-led text. Use when an agent…
Provider-independent VFX compositing direction for generated video, ads, trailers, product films, social clips, explainers, mixed-source edits, and image/video composites. Use when an agent must plan, brief, generate, integrate, repair, or quality-check…
Provider-independent production workflow for creating Manim-based explainer animations. Use when an agent must plan, code, render, QA, or hand off precise math, science, data, diagram, algorithm, or concept animations with Manim, including storyboard, scene…
Provider-independent production workflow for assembling generated or sourced media into React/Remotion videos. Use when an agent must plan, build, render, review, variant-render, or hand off Remotion compositions with media custody, captions, audio,…
Use NVIDIA Maxine / NVIDIA AI for Media audio effects for speech cleanup and enhancement in live or offline media pipelines, including Background Noise Removal, denoise+dereverb, room echo removal, acoustic echo cancellation, audio super-resolution, Studio…
Use D-ID to plan, generate, stream, localize, and QA avatar/talking-head videos, including V2 Photo Avatar Talks, V3 Pro/Instant Avatar Clips, V4 Expressive Avatar Scenes, D-ID Agents, voices, consent, lifecycle, safety, and artifact custody.
Produce Hedra character, talking-avatar, lip-sync, motion-avatar, and live avatar work. Use when an AI agent needs to choose Hedra models, prepare image/audio/script inputs, direct character performance, call or plan around the Hedra API or Studio, evaluate…
Produce HeyGen avatar videos with Direct Video, Video Agent, Digital Twin, Avatar Realtime, and related avatar/voice/asset APIs. Use for provider routing, avatar and voice selection, script-to-video workflows, lip-sync from audio, backgrounds/scenes,…
Produce presenter-led AI avatar videos with Synthesia Studio and Synthesia API. Use for Synthesia-specific avatar video planning, avatar/voice/language selection, scripted scenes, template personalization, API job lifecycle, localization/dubbing,…
Use for producing Tavus AI-human videos and real-time avatar conversations with Tavus Faces/Replicas, PALs/Personas, async Video Generation, CVI conversations, consent-safe likeness workflows, webhooks, backgrounds, localization, and avatar QA.
Use Adobe Firefly Services to generate, edit, expand, fill, match, composite, and upscale still images through the current REST APIs. Apply when selecting Firefly Image 5 versus Image 3/4 or custom models, implementing authenticated asynchronous image…
Plan, prompt, call, edit, iterate, and productionize Alibaba Cloud Model Studio image generation with current Wan 2.7 Image, Qwen-Image 2.0, and Z-Image Turbo models. Use for Alibaba/DashScope text-to-image, multi-reference generation, instruction editing,…
Produce and review production image-generation and image-editing workflows with Amazon Nova Canvas on Amazon Bedrock, including native InvokeModel payloads, safe authentication, validation, retries, cost controls, provenance, and lifecycle migration checks.…
Plan, prompt, execute, troubleshoot, and quality-control Black Forest Labs FLUX image generation and editing across the BFL direct API and licensed local/open-weight deployments. Use for FLUX.2 model selection, text-to-image, single- or multi-reference…
Build and operate rights-aware Bria FIBO and FIBO Lite image-generation workflows with structured prompts, reference images, asynchronous status handling, webhooks, cost gates, and safe artifact downloads. Use when a user asks for Bria/FIBO generation,…
Build and operate production image generation and natural-language image editing with ByteDance Seedream through first-party Volcengine Ark (China) or BytePlus ModelArk (global), including model and region selection, multi-reference and grouped outputs,…
Plan, generate, edit, and quality-check still images with Google's Gemini native image models (Nano Banana), including Gemini 3.1 Flash Image, Gemini 3.1 Flash-Lite Image, Gemini 3 Pro Image, and legacy Gemini 2.5 Flash Image. Use for model and API-route…
Generate, remix, edit, inpaint, reframe, background-process, describe, layerize, and upscale images with the Ideogram Developer API. Use when Codex must integrate Ideogram 4.0 or 3.0, render reliable typography or structured layouts, use style or character…
Plan, generate, edit, reference, and quality-control still images with Kling AI's hosted IMAGE surfaces or Kuaishou's open Kolors checkpoints. Use for Kling IMAGE 3.0/3.0 Omni/O1/2.1 web, official Kling CLI/MCP or Open Platform integration, and local Kolors…
Create, edit, guide, upscale, and quality-control still images with Leonardo.Ai's official Production API, including native Lucid and Phoenix models, uploaded or generated references, image-to-image, ControlNet guidance, realtime-canvas inpainting, Pro and…
Use Luma AI's first-party Photon and Photon Flash image API for text-to-image, image/style/character references, image modification, asynchronous generation, secure artifact handling, production prompting, cost control, and policy-aware delivery. Trigger for…
Plan, prompt, iterate, edit, migrate, and quality-check still-image work made with Midjourney's documented website and Discord interfaces. Use for Midjourney model and parameter selection, image/style/Omni references, Draft or Conversational workflows,…
Generate, edit, composite, stream, and production-review still images with OpenAI GPT Image models. Use when an agent must choose between GPT Image 2 and legacy GPT Image/DALL-E integrations, select the Image API or Responses API, build prompts and…
Design and produce raster images, native SVG artwork, brand-controlled visuals, and documented image edits with Recraft's API. Use for Recraft model and route selection, exact request construction, image-to-image, inpainting, outpainting, background work,…
Generate, edit, and iterate still images with Runway's official API, especially native Gen-4 Image and Gen-4 Image Turbo reference workflows. Use when implementing Runway text-to-image, reference-driven image generation or natural-language image edits, task…
Operate Stability AI image generation, image-to-image, edit, control, background, and upscale APIs safely and reproducibly. Use when selecting or calling Stable Image Ultra, Core, Stable Diffusion 3.5, v2beta image edit/control services, or Stability…
Generate and edit production images with xAI's first-party Grok Imagine API, including model selection, multiple references, synchronous and Batch API workflows, durable output handling, prompting, iteration, QA, cost and rate controls, privacy, safety, and…
Use ACE-Step and ACE-Step 1.5 for local or hosted AI music generation, including text-to-music, lyrics-to-song, instrumental beds, covers, repainting, stem/track extraction, track completion, LoRA personalization, REST/Python/Gradio workflows, rights review,…
Generate and iterate music with ElevenLabs Eleven Music for production deliverables. Use when planning, prompting, API-calling, editing, inpainting, reviewing, or licensing AI-generated songs, instrumental beds, ad music, soundtrack cues, video-to-music…