Skip to main content

dlazy-ai/cli

SkillsMP has collected 98 skills from dlazy-ai/cli. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
98
GitHub stars
3
GitHub forks
1

Skills in this repository

classification pending

Showing 40 of 98 collected skills.

occupation
unclassified
description

article to video, text to video, news to video, essay to video — turn a written article into a narrated explainer video: outline, storyboard, voiceover, build, validate. Use when the user pastes or gives an article and wants a video.

updated
occupation
unclassified
description

Audio generation skill. Automatically selects the best dlazy CLI audio/TTS model based on the prompt. 音频生成技能。根据提示词自动选择最佳的 dlazy CLI 音频/TTS 模型。

updated
occupation
unclassified
description

Generate/edit images with Nano Banana Pro. Supports text-to-image and image-to-image. 使用 Nano Banana Pro 生成/编辑图片,支持文生图与图生图。

updated
occupation
unclassified
description

Generate/edit high-quality images with Nano Banana 2.0. Supports text-to-image and image-to-image. 使用 Nano Banana 2.0 生成/编辑高质量图片,支持文生图与图生图。

updated
occupation
unclassified
description

blog to video, blog post to video, article to video — convert a blog post into a narrated video with storyboard, voiceover, and build. Use when the user gives a blog post (text or link) and wants a video version for social or YouTube.

updated
occupation
unclassified
description

Chat with the dlazy sandbox agent — a project-scoped assistant that runs skills end-to-end over multiple turns. Discover skills and projects with dlazy skills list / dlazy projects list. 与 dlazy 沙箱 agent 对话 —— 一个以项目为单位、可端到端运行技能的多轮助手。用 dlazy skills list /…

updated
occupation
unclassified
description

Anthropic's latest Sonnet — near-Opus quality on coding and long-horizon agentic work at Sonnet cost. Strong at reasoning, code generation, and complex tool orchestration. Supports text, image, and video inputs.

updated
occupation
unclassified
description

Detect whether an image, video, or audio file is AI-generated — including visual deepfakes and the likely generator model (Midjourney, Stable Diffusion, Sora, etc.). Returns confidence scores you can threshold. 检测图片、视频或音频是否由 AI 生成:包含人脸 deepfake…

updated
occupation
unclassified
description

doc to video, word to video, markdown to video, document to video — parse the document, outline, storyboard, voiceover, build, validate. Use when the user gives a Doc / Word / Markdown file and wants an explainer, report broadcast, or training video.

updated
occupation
unclassified
description

Synthesize text into natural and fluent speech using Doubao TTS. 使用豆包 (Doubao) TTS 文本转语音模型,将文字合成为自然流畅的语音播报。

updated
occupation
unclassified
description

ecommerce video, product video, shopping ad video, ecommerce video generator — turn a product (photos or link) into a conversion-focused ecommerce ad video. Use when the user has a product and wants a store, TikTok, or ad video.

updated
occupation
unclassified
description

ElevenLabs eleven_v3 multi-voice dialogue: assign a different voice per line (up to 10) and render the whole conversation in one shot. Supports audio tags like [giggling], [whispers] — great for character dialogue, podcasts, and short skits. Before picking a…

updated
occupation
unclassified
description

ElevenLabs music_v1 model — generates 10–300s original music from a natural-language prompt. Good for BGM, ads, and short-video soundtracks. ElevenLabs music_v1 音乐生成模型,根据自然语言提示生成 10–300 秒原创音乐。适合 BGM、广告配乐与短视频音轨。

updated
occupation
unclassified
description

Search the ElevenLabs voice library by keyword, source, and category. Returns a playable preview for each matched voice so you can pick the right one before running TTS. 搜索 ElevenLabs 人声库:按关键词、来源、分类筛选可用音色,返回每个音色的试听样本,便于挑选后用于 TTS 配音。

updated
occupation
unclassified
description

ElevenLabs text-to-sound model — generates 1–22s short sound effects from a description. Suitable for foley, ambience, alerts, and game SFX. ElevenLabs 文本生音效模型,根据描述生成 1–22 秒短音效。适合拟音、环境声、提示音与游戏音效。

updated
occupation
unclassified
description

ElevenLabs scribe_v1 speech-to-text with auto language detection and optional speaker diarization. Suitable for subtitles, transcription, and meeting notes. ElevenLabs scribe_v1 语音转文字,支持自动语种识别与说话人分离,适合字幕、转录与会议记录。

updated
occupation
unclassified
description

ElevenLabs eleven_v3 text-to-speech with 12 curated multilingual voices and stability/similarity/style controls. Great for dubbing, audiobooks, and character dialog. Before picking a voice, you can search for the right one via elevenlabs-search. ElevenLabs…

updated
occupation
unclassified
description

ElevenLabs Instant Voice Cloning (IVC). Upload a clean voice sample to clone a custom voice usable with ElevenLabs TTS. ElevenLabs 即时音色克隆(IVC),上传一段干净人声样本即可复刻自定义音色,可用于 ElevenLabs TTS 配音。

updated
occupation
unclassified
description

explainer video, explainer video generator, animated explainer, training video — turn a document, topic, or brief into a narrated explainer video: outline, storyboard, voiceover, build, validate. Use when the user wants an explainer, courseware, or training…

updated
occupation
unclassified
description

ppt to video, word to video, excel to video, pdf to video, document to video — parse, outline, storyboard, voiceover, build, validate. Use when the user gives a document and wants an explainer, report broadcast, courseware, or training video.

updated
occupation
unclassified
description

Alibaba Bailian Fun-ASR recording transcription. Supports Chinese, English and other languages, with auto language detection and speaker diarization. Suitable for subtitles, transcription, and meeting notes. 阿里云百炼 Fun-ASR…

updated
occupation
unclassified
description

Generate multilingual, highly natural audio using Gemini 2.5 text-to-speech. 使用 Gemini 2.5 强大的文本转语音能力,生成多语言、高自然度的音频。

updated
occupation
unclassified
description

A comprehensive generation skill. Can generate images, videos, and audio by automatically selecting the appropriate dlazy CLI model. 综合生成技能。能够根据用户意图自动选择合适的 dlazy CLI 模型来生成图片、视频或音频。

updated
occupation
unclassified
description

GPT Image 2 model for text-to-image and image editing. Supports generating images from text as well as editing and synthesizing images with reference inputs. GPT Image 2 图像生成与编辑模型。支持文生图,以及通过提供参考图进行图像编辑和合成。

updated
occupation
unclassified
description

Efficient text generation, dialogue QA, and logical reasoning using Grok 4.2 text model. 使用 Grok 4.2 文本大模型,进行高效的文本生成、对话问答与逻辑推理。

updated
occupation
unclassified
description

Happy Horse 1.0 video model — one model covers text-to-video (t2v), first-frame-to-video (i2v), reference-to-video (r2v), and video editing (edit). The selected mode is automatically routed to the matching sub-model. Happy Horse 1.0…

updated
occupation
unclassified
description

HeyGen Lipsync Speed: Fast lip-sync model, ideal for scenarios requiring rapid generation. HeyGen Lipsync Speed:快速唇形同步模型,适合对生成速度要求较高的场景

updated
occupation
unclassified
description

Turn a user's idea into the full pipeline: **story → characters → 3-view portraits → scenes → shots → keyframes → shot videos → concat**. First emit a **plan template** for the user to confi

Source text: Mixed languages

updated
occupation
unclassified
description

A professional product image generation skill purpose-built for the Amazon e-commerce platform. Outputs comply with Amazon's image guidelines while optimizing for click-through and conversio

Source text: Mixed languages

updated
occupation
unclassified
description

Image generation skill. Automatically selects the best dlazy CLI image model based on the prompt. 图片生成技能。根据提示词自动选择最佳的 dlazy CLI 图片生成模型。

updated
occupation
unclassified
description

A complete workflow skill for marketing brochure design, covering everything from requirements gathering, layout design, to mock-up delivery. It uses a 'layout-first + mandatory confirmation

Source text: Mixed languages

updated
occupation
unclassified
description

Image replicate tool: analyzes the visuals, composition, colors, lighting, and style of the source image, builds a replicate prompt, and hands it off to Seedream 4.5 to generate a new image in the same style. 图像仿写工具:分析原图的视觉效果、构图、色彩、光影和风格,生成重绘提示词,并交给 Seedream…

updated
occupation
unclassified
description

A structured workflow skill dedicated to social-media carousel design. The core method is 'decide intent first, then execute,' using a 'single-confirmation + cover-first' two-phase flow.

Source text: Mixed languages

updated
occupation
unclassified
description

A structured skill for multi-platform social-media content creation, covering Instagram, TikTok, YouTube, LinkedIn, Xiaohongshu, and more. The goal: outputs that satisfy each platform's nati

Source text: Mixed languages

updated
occupation
unclassified
description

A professional storyboard skill for film, advertising, short video, and educational narrative scenarios, built around a strict 'plan first, render later' flow.

Source text: Mixed languages

updated
occupation
unclassified
description

Image matting tool: separates foreground from background and returns transparent background URL, suitable for product image processing, character cutout, and composition. 图像抠图工具:将前景与背景分离并返回透明背景的 URL,适用于产品图片处理、人物抠图和合成。

updated
occupation
unclassified
description

Convert static character images into vivid action videos with Jimeng Dream Actor. 使用即梦 (Jimeng) Dream Actor,将静态人物图片转化为生动的动作视频。

updated
occupation
unclassified
description

Generate dynamic videos based on a single first frame image and prompts using Jimeng. 使用即梦 (Jimeng) 首帧生视频模型,基于单张首帧图片和提示词生成动态视频。

updated
occupation
unclassified
description

Generate coherent transition videos using Jimeng's first and tail frame models. 使用即梦 (Jimeng) 首尾帧生视频模型,通过提供的第一帧和最后一帧图片生成连贯的过渡视频。

updated
occupation
unclassified
description

Generate realistic digital human broadcast videos from portrait images and audio/text using Jimeng OmniHuman 1.5. 使用即梦 (Jimeng) OmniHuman 1.5 模型,通过人像图片和音频/文本生成逼真的数字人播报视频。

updated
Showing 40 of 98 collected skills.