Skip to main content

Skills في هذا المستودع

calesthio/generative-media-skills - الصفحة ٤

جمع SkillsMP عدد ١٥٣ من skills من calesthio/generative-media-skills. افتح أي skill لمراجعة مصدره وتفاصيله.

calesthio/generative-media-skills

عرض ٣٣ من أصل ١٥٣ skills مجمعة.

المهنة
مطوّرو البرمجيات
الوصف

Use Google Lyria music generation models through Google Cloud / Vertex AI / Gemini Enterprise Agent Platform for production music beds, songs, vocal tracks, image-conditioned music, short-form social audio, ad music, and video scoring. Covers Lyria 2, Lyria 3…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Produce music with MiniMax Music 2.6 and MiniMax cover/lyrics APIs for songs, instrumentals, AI-generated lyrics, reference-audio covers, video/social/ad soundtracks, artifact custody, rights checks, and production QA.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Produce, direct, generate, QA, and integrate non-speech audio with ElevenLabs Text to Sound Effects / Sound Effects API. Use when an agent needs custom SFX, foley, ambience, loops, stingers, impacts, UI sounds, musical one-shots, trailer braams, or…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use Microsoft Azure Speech in Foundry Tools for media-production speech workflows: speech-to-text, fast and batch transcription, diarization, captions/subtitles, real-time transcription, text-to-speech neural and HD voices, SSML, custom/personal voices,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use for Deepgram speech and voice production workflows: speech-to-text transcription, live captions, diarization, audio intelligence, Aura text-to-speech, Flux and Voice Agent live audio, model selection, cost/limits/privacy checks, artifact custody, and…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use for ElevenLabs provider-specific dubbing, localization, voice changer, speech-to-speech, and voice isolation/enhancement workflows. Applies when an agent must prepare source media, choose ElevenLabs dubbing versus voice conversion, manage…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use this skill when a media-production agent needs Google Cloud speech and voice services for transcription, captions/subtitles, long-form audio analysis, live caption planning, Text-to-Speech narration, Chirp 3 / Gemini-TTS voice selection, consented custom…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use NVIDIA Speech NIM microservices for speech production and voice workflows, including self-hosted ASR/STT, TTS, text translation, speech-to-speech pipeline design, Riva Python client integration, deployment planning, GPU/runtime sizing, licensing/privacy…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Produce and understand audio with OpenAI request-based audio APIs and audio-capable chat models, including text-to-speech, transcription, translation, multimodal audio input/output, model routing, prompt and performance direction, artifact custody, approval…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use Amazon Transcribe for AWS-based speech-to-text production: batch S3 transcription, real-time streaming, captions/subtitles, diarization, channel identification, custom vocabularies, vocabulary filters, language identification, PII/PHI handling, toxicity…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use this skill when an agent needs AssemblyAI for speech-to-text or speech-understanding work in media production, including pre-recorded, synchronous short-file, and real-time streaming transcription; speaker diarization or speaker identification; captions…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
فنيو الصوت والفيديو
الوصف

Use ElevenLabs Scribe for speech-to-text production workflows: transcribing audio or video files, diarization and speaker/channel handling, word/character timestamps, captions/subtitles, keyterm prompting, entity detection/redaction, webhooks, realtime STT…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Use Amazon Polly for production text-to-speech work: selecting Standard, Neural, Long-form, or Generative engines and compatible voices; authoring SSML; creating speech marks for captions, word highlighting, or lip-sync; managing pronunciation lexicons;…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Use Cartesia Sonic and related Cartesia voice APIs for production speech: text-to-speech, realtime WebSocket TTS, voice selection, instant and professional voice cloning, pronunciation/language/emotion controls, voice localization, voice changer,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Produce, direct, integrate, and quality-control ElevenLabs text-to-speech for narration, character performance, multilingual media, long-form voiceover, and streamed or batch AI-content audio. Use when selecting ElevenLabs models or voices; preparing spoken…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Use Hume Octave for emotionally expressive speech and voice production: text-to-speech, voice design, voice cloning, voice conversion, streaming/realtime TTS, multilingual narration, dialogue continuity, timestamps/lip-sync, safety/rights review, and…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Use this skill when producing speech, narration, dubbing, localization, advertising voice, voice-clone previews, or interactive voice with MiniMax speech/audio APIs. It covers MiniMax T2A HTTP, WebSocket streaming, async long-form TTS, Speech 2.8/2.6/02 model…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Produce text-to-speech with Alibaba/Qwen Qwen3-TTS through DashScope/Model Studio or open-weight Qwen3-TTS checkpoints. Use when an agent must choose Qwen3-TTS models, voices, realtime versus non-realtime synthesis, voice cloning, voice design,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Use Resemble AI Chatterbox for text-to-speech and voice-cloning workflows, including local open-weight Chatterbox, Chatterbox Multilingual, Chatterbox Turbo, and Resemble-hosted Chatterbox API routes. Apply when an agent must choose a Chatterbox variant,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Plan, implement, and review Alibaba Wan video generation using Alibaba Cloud Model Studio/DashScope hosted APIs or official Wan 2.1/2.2 open-weight checkpoints. Use for Wan text-to-video, image-to-video, first/last-frame or continuation, reference-to-video,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Produce and operate Amazon Nova Reel video-generation jobs through Amazon Bedrock. Use for Nova Reel text-to-video, image-conditioned animation, automated or manual multi-shot storyboards, async S3 delivery, cost and approval gates, production continuity,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Generate, reference, and conversationally edit short videos with Google's Gemini Omni Flash through the Gemini Developer API or Gemini Enterprise Agent Platform. Use when a task specifically needs Gemini Omni video, multimodal reference roles, time-directed…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Direct production with Google DeepMind's Veo video-generation family across the Gemini Developer API and Google Cloud Gemini Enterprise Agent Platform (formerly Vertex AI). Use for Veo route and model selection, text/image/reference/first-last-frame/extension…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Direct and operate Luma Ray video generation through the current Luma Agents API and distinguish it from the consumer Luma App and legacy Dream Machine API. Use for Ray 3.2 text/image/multi-keyframe video, edits, extends, reframes, HDR/EXR, pricing and…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Use MiniMax's first-party Hailuo video API safely and reproducibly across global and mainland-China platforms. Covers current text-to-video, image-to-video, first/last-frame, and subject-reference models; prompt and camera craft; region/account routing; exact…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Select, run, and govern NVIDIA Cosmos world-video generation across Cosmos 3 Generator, Predict2.5, Transfer2.5, downloadable checkpoints, self-hosted NIMs, and hosted preview surfaces. Use for text/image/video-to-world, controlled world transfer, multiview…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Build and operate production-safe Runway API video generation, video editing, and character-performance workflows with native Runway models, explicit paid-call approval, duplicate-create protection, asynchronous task handling, secure media transfer, and…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Direct ByteDance Dreamina Seedance 2.0 Standard, Fast, and Mini video production across BytePlus ModelArk and verified gateways. Use for text/image/reference-to-video, native synchronized audio and dialogue, multimodal image-video-audio reference,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Produce, edit, extend, and govern short videos with xAI's direct Grok Imagine Video API. Use for text-to-video, image-to-video, multi-reference video, natural-language video edits, video continuation, exact media-cost approval, asynchronous request recovery,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Build and evaluate low-latency spoken, multimodal, and translation experiences with Google Gemini Live API and Gemini Enterprise Agent Platform Live API. Use when a media-production or voice-agent workflow needs real-time audio input/output, barge-in, voice…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Build production realtime voice agents with Hume's Empathic Voice Interface (EVI): speech-to-speech sessions, empathic/prosodic response design, EVI 3 versus EVI 4-mini selection, WebSocket and SDK integration, tools/function calling, interruption/barge-in,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Build production OpenAI Realtime voice agents and low-latency spoken interactions with live audio sessions, WebRTC or WebSocket transport, voice activity detection, tool/function calling, prompt design, logging, consent, privacy, safety, latency, and cost…

لغة النص الأصلي: الإنجليزية

آخر تحديث
عرض ٣٣ من أصل ١٥٣ skills مجمعة.