بنقرة واحدة
video-automator-skills
يحتوي video-automator-skills على 43 من skills المجمعة من bachdyon، مع تغطية مهنية على مستوى المستودع وصفحات skill داخل الموقع.
Skills في هذا المستودع
Download public Facebook Reel videos with yt-dlp into the repository raw asset pool or a specific video job. Use when a user provides a facebook.com/reel, facebook.com/watch, Facebook video, or fb.watch URL and asks to download, save, or ingest the video as a local asset.
Search LyricFind for song lyrics and lyric context, especially to correct Vietnamese music STT before writing subtitle captions; use the bundled Python client instead of generic web search when a song title or artist is known.
Default low-cost lipsync workflow using WaveSpeedAI InfiniteTalk. Use when the user asks for lipsync, lip sync, talking photo/avatar, digital human from one image plus audio, cheap avatar video generation, or asks to animate a person speaking from an image and audio unless they explicitly choose another provider.
Create AI videos with KIE.AI Google Veo 3.1. Use when the user wants text-to-video, image-to-video with first/last frames, reference/material-to-video, extending Veo 3.1 tasks, polling status, or upgrading/downloading 1080p/4K results through the KIE Veo3 API.
Convert a finished video job into a reusable project template and Codex skill. Use when the user asks to turn a Remotion job into a template, standardize a copied job, create a reusable video template skill, parameterize a job, or make future video requests token-light while preserving visual quality and render reliability.
Prepare effect/overlay videos so they can be composited over another video, especially stock overlays with black backgrounds, dust, light particles, petals, smoke, flares, or non-alpha MP4/MOV assets. Use when Codex needs to check whether overlay videos have real alpha, convert ordinary overlay footage into Remotion-ready assets, normalize fps/size, preserve color for screen blending, avoid black halos from bad alpha keying, or generate reusable overlay files and compositing snippets.
Search and download Klipy GIFs, stickers, clips, and memes for video reaction visuals. Use when Codex needs meme/reaction assets for Threads videos, editorial recap videos, short-form commentary, scene beats, or any request to find GIF/sticker/clip/meme media from Klipy.
Create standalone animated HTML from local SVG artwork. Use when Codex needs to animate an SVG with CSS/inline browser-safe code, preserve the original visual style, group paths semantically, choose motion that matches the drawing content, and verify the result in a browser without external assets.
Sinh TOML semantic index cho ảnh/video raw để các skill phía sau map vào kịch bản. Ưu tiên đọc từ asset-index SQLite vector DB (mỗi file gọi Gemini đúng 1 lần trong toàn bộ project), chỉ fallback về probe + vision pass tươi khi không có index.
Làm việc với AusyncLab voice và Text-to-Speech API khi có `AUSYNCLAB_API_KEY`: liệt kê voice, gợi ý voice cho creative plan, lưu setting voice ưa thích, sinh audio narration mp3 hoặc wav, và hậu xử lý tốc độ WAV sau TTS bằng pydub. Nếu không có `AUSYNCLAB_API_KEY`, dùng `$free-tts`.
Dùng CapCut common task client (`K07VN/capcut-tts-api`, `capcut_common_task_client.py`) để tạo/query Text-to-Speech task, upload audio/video, hoặc chạy STT/subtitle recognition qua flow CapCut đã phân tích; dùng khi user yêu cầu rõ CapCut TTS, giọng CapCut, CapCut subtitle/STT, hoặc muốn thử workflow CapCut thay cho AusyncLab/free TTS.
Sinh ảnh AI bằng fal.ai (model nano-banana / flux) từ prompt, hỗ trợ multi-reference image để cố định nhân vật. Lưu ảnh vào jobs/<id>/input/raw_assets/images/ai_generated/ để watcher asset-index tự pickup và pipeline tiếp tục như asset thật.
Tạo giọng đọc miễn phí/local bằng VieNeu-TTS khi người dùng yêu cầu tự động tạo voice/narration/giọng đọc và repo-root `.env` không có `AUSYNCLAB_API_KEY`, hoặc khi user yêu cầu rõ dùng free TTS/VieNeu-TTS thay AusyncLab.
Viết hoặc cập nhật video_content.md cho video chia sẻ kiến thức/giáo dục đời sống, với voiceover kể chuyện, câu mở đầu kéo người xem vào vai chính, luận điểm có câu hỏi phụ, và gợi ý visual AI không chữ.
Phân tích khung hình bằng vision LLM không phải Gemini để đề xuất vị trí text/image/video overlay tránh che chủ thể và tránh vùng unsafe, dùng khi cần quyết định placement cho creative plan hoặc render plan.
Convert local PNG images to SVG files through the Convertio API. Use when Codex needs an SVG asset from a PNG source, especially for logos, icons, overlays, templates, or design assets where an SVG file is required and local conversion is not sufficient.
Khớp transcript hoặc scene intents với asset ảnh/video đã được index ngữ nghĩa, sinh TOML timeline mapping với start, end, file_path, đoạn source được chọn và lý do.
Giải quyết coverage shortage và lặp asset trong baseline semantic mapping bằng cách ra quyết định kiểu editor sáng tạo (cutaway B-roll, hold + Ken Burns, slowdown). Thiết kế cho AI agent (không gọi LLM nội bộ) — agent đọc context, áp khung quyết định, và viết decisions JSON để 1 helper script nhỏ apply lên TOML.
Làm việc với SocialKit API bằng Python để lấy transcript, summary, stats, comments, channel stats, search, video listing và YouTube download cho YouTube, TikTok, Instagram, Facebook hoặc video file URL; dùng script Python đi kèm, không dùng curl.
Short-form subtitle normal + PUNCH đồng bộ word-level; chunk cố định bằng script; punch do coding agent gán bằng suy luận ngữ nghĩa (cấm heuristic chọn punch trong code). Punch 2 lớp z-index, ~7–8 từ/chunk. Dùng khi build Remotion TikTok/Reels hoặc chuẩn hóa punch_segments.
Send Telegram bot messages and media through the official Bot API using channel-specific `.env.<channel>` files. Use when the user asks to send Telegram messages, videos, photos, documents, check bot identity, or inspect bot updates/chat IDs.
Choose distinctive, high-quality typography for visual work. Use when creating or revising websites, apps, Remotion videos, captions, thumbnails, social posts, decks, landing pages, brand visuals, or any artifact where font choice, font pairing, scale, weight, or typographic style affects perceived quality.
Compress local MP4/MOV videos to fit under 25MB for APIs and messaging platforms such as Zernio upload-direct, WhatsApp/Messenger inbox attachments, or other small-file upload limits. Use when Codex needs to reduce video file size while preserving duration, resolution, frame rate, and broad H.264/AAC compatibility when possible.
Biến yêu cầu video mới (kèm Video Design Specification tùy chọn) thành creative plan, kịch bản, scene intents, overlay text và asset requirements ở mức sẵn sàng đưa vào pipeline video short-form.
Xây dựng Video Design Specification (VDS) tái sử dụng được từ video gốc, giữ nguyên DNA phong cách và xóa thông tin nhận diện cá nhân. Dùng khi user yêu cầu tạo, chuyển đổi, chuẩn hóa, hoặc tái sử dụng phong cách video short-form cho TikTok/Reels/Shorts.
Download TikTok videos or slide images through the TikWM form API and save them into raw_assets/ or jobs/<id>/input/raw_assets/ for the video pipeline.
Tạo và quản lý các video job sản xuất biệt lập, gồm metadata yêu cầu, asset đầu vào, artifact theo stage, status, theo dõi stale, và path chuẩn cho pipeline video.
Điều phối toàn bộ pipeline video short-form dài hạn từ video mẫu, yêu cầu sáng tạo mới, raw asset, sinh voice, transcribe, semantic mapping, render plan, đến render cuối cùng.
Audit code-first cho Remotion source và render plan (ưu tiên overlay readability + safe-area), gom toàn bộ lỗi mỗi pass rồi áp dụng batch fix một thể, lặp tối đa 3 pass và xuất báo cáo TOML + HTML tiếng Việt.
Chuyển VDS, creative plan, transcript, và semantic asset mapping thành TOML edit decision list chi tiết với crop, timing, text overlay, subtitle, motion, audio, transition, và chỉ thị render.
Render video short-form cuối cùng từ TOML render plan, voice audio, source asset, subtitle, overlay, và quy tắc style VDS bằng renderer của project (Remotion hoặc FFmpeg).
Direct distinctive visual design for videos rendered from HTML, CSS, JavaScript, or Remotion. Use when creating or revising video visuals, scene layouts, motion systems, backgrounds, captions, design briefs, render plans, or HTML/CSS/JS video compositions where the output should avoid generic AI/template aesthetics and feel intentionally art-directed.
Trích xuất transcript narration có timestamp cấp câu và cấp từ; script mặc định dùng local faster-whisper, có fallback OpenAI Whisper khi cần.
Download public YouTube videos or Shorts quickly with yt-dlp, always save an MP4 into raw_assets/videos/downloaded and a compatible MP3 audio file into raw_assets/audio/downloaded. Use when the user gives a YouTube URL and asks to download the video, get the audio, export MP3, or optimize the workflow for speed and low token usage.
Add a finished VAS job to the public showcases repository by copying only its output/ folder. Use when the user asks to publish, mirror, archive, or document a job under showcases and update showcase docs.
Work with FilePost files through the FilePost API: upload local files to permanent public CDN URLs, list uploaded files, get file details, and delete files. Use when Codex needs durable hosted media/file URLs beyond short-lived temporary upload services, or when the user asks for FilePost list/get/upload/delete operations.
Create HeyGen image-to-video/photo-avatar talking-head videos from a person's image plus custom audio. Use when the user wants a personal brand video, video nhân hiệu, talking avatar, lip-synced photo video, or a HeyGen video generated from one image and one audio file.
Create AI videos with KIE.AI Bytedance Seedance 2.0 Fast. Use when the user wants text-to-video, image-to-video with first/last frames, multimodal reference-to-video, optional generated audio, polling task status, or downloading Seedance results through the KIE Market API.
Upload images, videos, audio, or PDFs to HeyGen and return reusable asset IDs. Use when the user wants to upload local files for HeyGen, prepare assets for Video Agent, avatars, image-to-video, video translation, or any HeyGen endpoint that accepts asset_id file inputs.
Create or edit images with KIE.AI GPT Image 2. Use when the user wants text-to-image generation, image-to-image editing, or prompt-based image creation with optional reference images through the KIE Market API, including task polling and downloading generated images.