Skip to main content
在 Manus 中运行任何 Skill
一键导入
GitHub 仓库

generative-media-skills

generative-media-skills 收录了来自 calesthio 的 153 个 skills,并提供仓库级职业覆盖和站内 skill 详情页。

已收集 skills
153
Stars
92
更新
2026-07-14
Forks
12
职业覆盖
6 个职业分类 · 已分类 100%
仓库浏览

这个仓库中的 skills

cinematic-shot-direction
特效艺术家和动画师

Direct cinematic shots and coverage for AI-assisted content production. Use when an agent must translate story intent into framing, shot size, lens and perspective, camera position, movement, blocking, screen direction, lighting-aware coverage, continuity, shot specifications, or diagnose why generated or filmed shots do not cut together.

2026-07-14
precise-video-description
音视频技术员

Provider-independent production guidance for converting observed video into precise, objective, temporally ordered language. Use for shot descriptions, searchable metadata, dataset captions, reference logs, generation-prompt handoffs, and analysis across subject, scene, motion, spatial, and camera aspects; not for deciding what to shoot, accessibility captions, creative interpretation, or provider-specific video analysis APIs.

2026-07-14
generated-media-qa
影片与视频编辑

Provider-independent quality assurance for AI-generated and AI-assisted media. Use when reviewing, accepting, revising, or reporting on images, video, audio, avatars, ads, product content, social clips, explainers, localization, mixed-source edits, captions, accessibility, provenance, model metadata, safety/policy, and delivery readiness.

2026-07-14
reference-media-analysis
影片与视频编辑

Provider-independent reference media analysis for generated-media production. Use when an agent must analyze reference images, videos, audio, style boards, product shots, brand assets, mood boards, storyboards, performances, prior cuts, or client examples and translate them into safe, non-copying direction, prompts, QA criteria, provenance records, and handoff notes for image, video, audio, avatar, or post-production agents.

2026-07-14
video-description-oversight
影片与视频编辑

Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions. Use for pre-caption critique and post-caption revision, aspect-by-aspect factual review, critique precision/recall/constructiveness, second-stage peer review, calibration, appeals, versioned triplets, provenance, and acceptance reporting; not for pixel/audio QA, accessibility captioning, model training, or creative shot direction.

2026-07-14
screen-demo-production
影片与视频编辑

Provider-independent workflow for producing terminal, IDE, documentation, browser, desktop, and application demonstration videos. Use when choosing authentic capture, browser automation, synthetic UI, or hybrid treatment; scripting actions and narration; resetting demo state; protecting secrets and personal data; directing cursor/callouts; creating crop variants; and reviewing workflow truth, readability, accessibility, and provenance.

2026-07-13
audio-reactive-video-composition
影片与视频编辑

Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Use for beat-, onset-, phrase-, energy-, silence-, or spectrum-reactive visualizers, edits, typography, and generated compositions; not for music-video concept direction, audio generation, or a runtime-specific recipe.

2026-07-13
procedural-canvas-animation
网页开发工程师

Provider-independent production guidance for deterministic Canvas 2D and p5.js animation. Use for particles, fields, trails, weather, procedural textures, generative geometry, and lightweight 2D simulations that need fixed media dimensions, seeded repeatability, transparent compositing, aspect variants, performance QA, or frame-addressable rendering.

2026-07-13
kling-advanced-lip-sync
影片与视频编辑

Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy review, sync QA, and repair. Do not use for still-image avatar generation or ordinary Kling video generation.

2026-07-13
byteplus-seed-speech-tts
音视频技术员

Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and concurrency, authorized replicated voices, privacy, error handling, and output QA. Do not use for mainland-China Volcengine Doubao Speech endpoints.

2026-07-13
volcengine-doubao-speech-tts
软件开发工程师

Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Use for TTS 1.0/2.0 selection, V3 bidirectional or unidirectional streaming, asynchronous long-text submit/query jobs, current voices and controls, timestamps and SSML caveats, expiring results, quota/error handling, authorized cloned voices, privacy, and audio QA. Do not use for international BytePlus endpoints.

2026-07-13
2d-character-rig-animation
特效艺术家和动画师

Production guidance for reusable layered 2D character rigs in browser-rendered and video media. Use for SVG hierarchies, pivots and constraints, FK/IK planning, pose and facial libraries, acting, cycles, deterministic handoff, crop variants, and rig QA; not for character concept continuity, 3D armatures, or motion-capture solving.

2026-07-13
d3-animated-data-visualization
特效艺术家和动画师

Production guidance for converting sourced data and approved claims into truthful, accessible, deterministic animated visualizations with D3. Use for data-driven charts, maps, networks, hierarchies, transitions, annotations, responsive video variants, frame-by-frame browser rendering, and visualization QA; not for generic dashboards, unsourced decorative charts, or ordinary motion graphics without a data contract.

2026-07-13
gsap-animation-composition
特效艺术家和动画师

Production guidance for authoring, integrating, and reviewing GSAP animation in browser-rendered media. Use for deterministic kinetic typography, SVG drawing and morphing, motion paths, FLIP transitions, responsive motion systems, or frame-seekable GSAP timelines inside HTML video composition and React-based renderers; not for ordinary CSS transitions or generic frontend development.

2026-07-13
lottie-animation-delivery
特效艺术家和动画师

Production guidance for assessing, exporting, packaging, integrating, capturing, validating, and handing off Lottie vector animations. Use for Bodymovin/Lottie JSON, dotLottie archives, renderer and player compatibility, fonts and glyphs, image assets, markers and segments, deterministic video capture, responsive sizing, optimization, accessibility, and cross-platform QA; not for general After Effects animation craft or live-action/video delivery.

2026-07-13
threejs-scene-composition
特效艺术家和动画师

Production guidance for planning, building, animating, capturing, and reviewing complete Three.js scenes for rendered media. Use for browser-native 3D product shots, title sequences, procedural worlds, glTF scene assembly, camera and lighting animation, shaders, post-processing, deterministic frame export, and Three.js render QA; not for modeling standalone assets, WebXR interaction, or generic website decoration.

2026-07-13
3d-asset-production
特效艺术家和动画师

Use this skill to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery. It covers requirements, topology repair, scale and axes, UVs, PBR baking, LODs, static and skeletal readiness, rig handoff, optimization, glTF/USD/FBX delivery, validation, QA, and rights/provenance. Do not use it for provider-specific generation prompts, full scene/world building, mocap operation, or exact CAD/manufacturing work.

2026-07-12
audio-mixing-mastering
音视频技术员

Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social content. Use when planning, mixing, repairing, mastering, QCing, or delivering dialogue, music, ambience, and sound effects, including loudness/true-peak targets, intelligibility, accessibility, stems, stereo/immersive decisions, platform/client specs, and final audio QA.

2026-07-12
captions-media-accessibility
音视频技术员

Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video recuts, training media, and localized content. Use when planning, authoring, reviewing, localizing, burning in, exporting, or QAing captions, subtitles, SDH, transcripts, audio description, flashing/motion safety, caption readability, or accessible media handoff files.

2026-07-12
media-qc-delivery
影片与视频编辑

Provider-independent media quality control and delivery direction for generated images, video, audio, ads, product content, social clips, broadcast and streaming assets, localization packages, and mixed-source edits. Use when an agent must define final export specs, inspect technical and perceptual quality, check captions/audio/color/accessibility/provenance/rights, create delivery sheets, version filenames, checksums, manifests, package deliverables, set acceptance or rejection criteria, and perform final delivery QA before handoff to a platform, broadcaster, streamer, client, or ad trafficker.

2026-07-12
comfyui-media-workflows
特效艺术家和动画师

Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media generation; use when a task involves ComfyUI workflow JSON/API graphs, model and custom-node inventories, reproducibility, provenance, safety/rights review, local or cloud execution hygiene, or production QA.

2026-07-12
ffmpeg-media-finishing
音视频技术员

Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables. Use when finalizing images, image sequences, video, audio, captions, overlays, social/export variants, checksums, manifests, delivery specs, and QA with ffmpeg/ffprobe, including transcode versus stream-copy decisions, scaling, frame-rate handling, color metadata caution, loudness normalization, subtitle burn-in or sidecars, concat/trim, GIFs, batch variants, and reproducible command reporting.

2026-07-12
hyperframes-video-composition
特效艺术家和动画师

Provider-independent production workflow for AI agents assembling generated or source media into HyperFrames HTML/CSS/JS videos. Use for HyperFrames composition planning, scene architecture, media custody, animation/timing, captions/audio, deterministic preview/render QA, accessibility/flashing checks, provenance ledgers, and delivery handoff.

2026-07-12
image-generation-gateways
软件开发工程师

Select, integrate, and operate multi-model image-generation gateways with model-specific schema discovery, version policy, asynchronous jobs, webhooks, spend approval, safe inputs and artifacts, data-governance review, and billing observability. Use for comparing or building against fal.ai, Replicate, or Together AI image APIs, including controlled failover; do not use for direct model-provider APIs, local inference, training, dedicated endpoint provisioning, video generation, or general image editing.

2026-07-12
kling-video
软件开发工程师

Plan, prompt, generate, reference, edit, motion-control, and quality-review videos with Kuaishou's direct Kling AI video platform and API, especially Kling VIDEO 3.0, 3.0 Omni, 3.0 Turbo, native audio, multi-shot, elements, start/end frames, and image/video references. Use for first-party Kling video production or API integration, not Kolors image generation, avatars, standalone audio, consumer-only membership advice, or third-party Kling gateways.

2026-07-12
ltx-2-video
软件开发工程师

Generate and edit synchronized audio-video with LTX-2.3 using the official hosted LTX API or official local/open-weight LTX-2 repository. Use for text-to-video, image-to-video, audio-to-video, retake, extend, HDR conversion, local inference, checkpoint selection, quantization, and LTX-specific production planning.

2026-07-12
midjourney-video
软件开发工程师

Plan and direct Midjourney Video V1 image-to-video work through the official website or Discord, with human operator handoff, motion prompts, start/end frames, loops, extensions, resolution and batch budgeting, privacy, rights, safety, and delivery QA. Use when creating Midjourney videos or when determining whether a requested API, automation, plan, or production workflow is supported.

2026-07-12
tencent-hunyuanvideo
软件开发工程师

Generate and operate Tencent Hunyuan video through the managed TokenHub HY-Video-1.5 API or official local HunyuanVideo repositories. Use for text-to-video, image-to-video, hosted job lifecycle, pricing and region planning, local checkpoint selection, hardware and acceleration, prompt design, licensing, safety, provenance, and delivery QA.

2026-07-12
vidu-video
软件开发工程师

Plan and integrate ShengShu Vidu Open Platform video generation with current model/mode selection, reference consistency, pricing, exact approval, async task handling, and safe artifact custody. Use for Vidu API text-to-video, image-to-video, start/end-frame, multi-subject reference, media-reference, or Q2 multi-frame work; do not conflate the API with vidu.com consumer plans or Vidu-S1 streaming digital humans.

2026-07-12
neural-reality-capture
特效艺术家和动画师

Use this skill for provider-independent reality capture that turns photographed or video-derived real places, objects, aerial sites, interiors, or people into photogrammetric meshes, NeRFs, or 3D Gaussian splats. It guides representation choice, capture planning, calibration and scale, reconstruction or training, cleanup, compression and LOD, coordinate and interchange handoff, visual/geometric QA, archival records, and property, location, people, and cultural rights checks. Do not use it for text-to-3D generation, synthetic asset invention, guaranteed metrology, or broad downstream finishing after handoff.

2026-07-11
immersive-spatial-video-production
音视频技术员

Use this skill to plan, direct, finish, and QA provider-independent mono or stereo 180-degree, 360-degree, spherical, spatial, wide-FOV, and immersive video. It covers format choice, capture rigs, stitching, nadir and seam management, stereoscopic comfort, attention direction, stabilization, ambisonic and spatial audio, titles, captions, accessibility, edit, color, VFX, projection metadata, Vision Pro, YouTube, headset delivery, device QA, and privacy, location, and likeness rights. Do not use it for interactive XR apps, volumetric/world-model generation, game-engine experiences, or provider-specific video generation prompts.

2026-07-11
livestream-event-production
影片与视频编辑

Use this skill to plan, direct, troubleshoot, and hand off provider-independent live or hybrid livestream event productions, including run of show, crew, capture, audio, graphics, switching, guests, encoding, ingest, transport choices, latency, redundancy, accessibility, moderation, monitoring, incidents, recording, and delivery.

2026-07-11
still-image-retouching-finishing
平面设计师

Provider-independent still-image post-production skill for RAW/rendered intake, nondestructive development, exposure, white balance, tone, color, ICC proofing, masking, repair, compositing, truthful retouching by genre, generative disclosure, sharpening, noise, resampling, variants, metadata, export, QA, rights, and ethical finishing. Use when an agent must plan, perform, direct, or review still photography retouching and finishing for print, web, social, archive, publication, or client delivery; do not use for image-generation APIs, video grading, or campaign-specific production strategy.

2026-07-11
virtual-production-icvfx
特效艺术家和动画师

Plan, troubleshoot, and hand off provider-independent in-camera VFX and LED volume virtual-production work, with Unreal Engine/nDisplay/Live Link awareness. Use for ICVFX suitability, LED stage geometry, frustums, tracking, lens calibration, sync, render nodes, color, lighting/reflections, artifact mitigation, rehearsal, live-composite fallback, monitoring, records, QA, and handoff. Do not use for ordinary offline post compositing, generic storyboards, or full game production.

2026-07-11
amazon-rekognition
软件开发工程师

Use this skill when an agent needs production image or video understanding with Amazon Rekognition: labels, objects, scenes, OCR, moderation, image properties, Custom Labels or moderation adapters, stored-video analysis, conditional streaming-video workflows for existing eligible accounts, searchable media libraries, confidence evaluation, S3/IAM/event architecture, privacy, biometric consent, cost control, lifecycle management, and QA.

2026-07-11
google-cloud-vision
软件开发工程师

Use this skill when an agent needs Google Cloud Vision API for still-image understanding: labels, object localization, OCR, document text, SafeSearch, image properties, crop hints, web detection, batch annotation, Cloud Storage based pipelines, confidence evaluation, privacy, quotas, cost, and QA. Do not use it for Gemini multimodal reasoning, video analysis, image generation, custom model training, product catalog search design, or human reference/authenticity review.

2026-07-11
deepmotion-animate-3d
特效艺术家和动画师

Use DeepMotion Animate 3D for markerless video-to-3D human motion capture through the cloud portal, sales-gated API, or real-time SDK; covers capture planning, body/hand/face/multi-person limits, custom characters, retargeting, refinement, exports, DCC/game-engine handoff, QA, pricing, privacy, consent, and rights checks.

2026-07-11
move-ai-motion-capture
特效艺术家和动画师

Use this skill when planning, capturing, processing, validating, troubleshooting, or exporting markerless human motion capture with Move AI products, including Move One single-camera workflows and Genesis multi-camera studio systems. It covers provider selection, setup, calibration, wardrobe and lighting constraints, documented multi-person operation, processing, cleanup, skeleton and retargeting choices, DCC/game-engine export, data management, consent, and privacy. Do not use it for general keyframe animation, non-Move mocap systems, 3D asset generation, or generic animation cleanup unrelated to Move AI outputs.

2026-07-11
anime-animation-production
特效艺术家和动画师

Produce anime-style and stylized 2D animation with generative image and video tools as a craft discipline, not a filter. Use when a request asks for anime, manga-style, cel-shaded, sakuga, shonen/shojo/seinen/slice-of-life, mecha, or retro-90s-OVA animation; when planning shots, motion, and cutting rhythm to anime convention; when holding a character on-model and preventing style drift across shots or episodes; when choosing prompting vocabulary that actually controls anime look; when pairing voice, music, and impact SFX to stylized action; and when navigating the cultural, IP, and platform-policy sensitivities of AI anime (studio/artist style imitation, fan-art and IP boundaries, disclosure and monetization). This is provider-neutral; models are named only as illustrative options.

2026-07-11
audiobook-production
音视频技术员

Produce full-length audiobooks and long-form narration with generative voice tools. Use when the task is to turn a manuscript or long text into hours of spoken audio: preparing the manuscript for narration (front/back matter, footnotes, tables, dialogue), casting a single narrator or full cast, controlling pronunciation and voice consistency across a whole book, running a proofing/QC listen, meeting a retailer's technical delivery specs (RMS, peak, noise floor, room tone, chapterized files, metadata), and choosing a distribution route under each platform's current AI-narration policy. Not for short TTS clips, single-line voiceover, podcast production, or captioning existing video.

2026-07-11
当前展示该仓库 Top 40 / 153 个已收集 skills。