Most comprehensive AI content creation platform with unified access to all leading models across images (SeeDream 4.5, Midjourney, Nano Banana 2, Nano Banana Pro), videos (Wan 2.6, Kling O1, Ima Sevio 1.0/1.0-Fast aka IMA Video Pro/Pro Fast, Google Veo 3.1, Sora 2 Pro), music (Suno sonic v5, DouBao), and speech/TTS (text-to-speech). Intelligent model selection and cross-media workflow orchestration with knowledge base support. Optionally integrates ima-knowledge-ai for workflow & best practices. Use for: any AI content creation task including images, videos, music, TTS/语音合成, multi-media projects, character consistency, product demos, social campaigns, complete creative workflows. Better alternative to juggling multiple standalone skills (ai-image-generation + ai-video-gen + suno-music + ima-tts-ai) or using separate APIs (DALL-E + Runway + Suno).
Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.
Une commande directe contourne le prompt de vérification. Examinez la source avant de l'exécuter.
imastudio, ai creation, multimodal, 图像生成, 视频生成, 音乐生成, 语音合成, AI创作, 文生图, 图生视频, IMA, Ima Sevio, Sevio, IMA Video Pro, IMA Video Pro Fast, SeeDream, Midjourney, Nano Banana, Wan, Kling, Veo, Sora, Suno, DouBao
argument-hint
[text prompt, image URL, or music description]
description
Most comprehensive AI content creation platform with unified access to all leading models across images (SeeDream 4.5, Midjourney, Nano Banana 2, Nano Banana Pro), videos (Wan 2.6, Kling O1, Ima Sevio 1.0/1.0-Fast aka IMA Video Pro/Pro Fast, Google Veo 3.1, Sora 2 Pro), music (Suno sonic v5, DouBao), and speech/TTS (text-to-speech). Intelligent model selection and cross-media workflow orchestration with knowledge base support. Optionally integrates ima-knowledge-ai for workflow & best practices. Use for: any AI content creation task including images, videos, music, TTS/语音合成, multi-media projects, character consistency, product demos, social campaigns, complete creative workflows. Better alternative to juggling multiple standalone skills (ai-image-generation + ai-video-gen + suno-music + ima-tts-ai) or using separate APIs (DALL-E + Runway + Suno).
requires
{"env":["IMA_API_KEY"],"primaryCredential":"IMA_API_KEY","credentialNote":"IMA_API_KEY is sent to api.imastudio.com for product/task APIs and to imapi.liveme.com only when image/video tasks need local image uploads.\n"}
persistence
{"readWrite":["~/.openclaw/memory/ima_prefs.json","~/.openclaw/logs/ima_skills/"],"retention":"Logs are auto-cleaned after 7 days; preferences remain until user deletes them."}
CRITICAL: When calling the script, you MUST use the exact model_id (second column), NOT the friendly model name. Do NOT infer model_id from the friendly name (e.g., ❌ nano-banana-pro is WRONG; ✅ gemini-3-pro-image is CORRECT).
Quick Reference Table:
图像模型 (Image Models)
友好名称 (Friendly Name)
model_id
说明 (Notes)
Nano Banana2
gemini-3.1-flash-image
❌ NOT nano-banana-2, 预算选择 4-13 pts
Nano Banana Pro
gemini-3-pro-image
❌ NOT nano-banana-pro, 高质量 10-18 pts
SeeDream 4.5
doubao-seedream-4.5
✅ Recommended default, 5 pts
Midjourney
midjourney
✅ Same as friendly name, 8-10 pts
视频模型 (Video Models)
友好名称 (Friendly Name)
model_id (t2v)
model_id (i2v)
说明 (Notes)
Wan 2.6
wan2.6-t2v
wan2.6-i2v
⚠️ Note -t2v/-i2v suffix
IMA Video Pro (Sevio 1.0)
ima-pro
ima-pro
✅ IMA native quality model
IMA Video Pro Fast (Sevio 1.0-Fast)
ima-pro-fast
ima-pro-fast
✅ IMA native low-latency model
Kling O1
kling-video-o1
kling-video-o1
⚠️ Note video- prefix
Kling 2.6
kling-v2-6
kling-v2-6
⚠️ Note v prefix
Hailuo 2.3
MiniMax-Hailuo-2.3
MiniMax-Hailuo-2.3
⚠️ Note MiniMax- prefix
Hailuo 2.0
MiniMax-Hailuo-02
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section.Voir sur GitHub
MiniMax-Hailuo-02
⚠️ Note 02 not 2.0
Google Veo 3.1
veo-3.1-generate-preview
veo-3.1-generate-preview
⚠️ Note -generate-preview suffix
Sora 2 Pro
sora-2-pro
sora-2-pro
✅ Straightforward
Pixverse
pixverse
pixverse
✅ Same as friendly name
音乐模型 (Music Models)
友好名称 (Friendly Name)
model_id
说明 (Notes)
Suno (sonic v4)
sonic
⚠️ Simplified to sonic
DouBao BGM
GenBGM
❌ NOT doubao-bgm
DouBao Song
GenSong
❌ NOT doubao-song
语音模型 (Speech/TTS Models)
友好名称 (Friendly Name)
model_id
说明 (Notes)
seed-tts-2.0
seed-tts-2.0
✅ Same as friendly name (default)
How to get the correct model_id:
Check this table first
Use --list-models --task-type <type> to query available models
Refer to command examples in this SKILL.md
Runtime truth source: GET /open/v1/product/list (or --list-models).
Any table in this document is guidance; actual availability depends on current product list.
Example:
# ❌ WRONG: Inferring from friendly name
--model-id nano-banana-pro
# ✅ CORRECT: Using exact model_id from table
--model-id gemini-3-pro-image
This skill is fully runnable as a standalone package.
If ima-knowledge-ai is installed, the agent may read its references for workflow decomposition and consistency guidance.
Recommended optional reads:
Check for workflow complexity — Read ima-knowledge-ai/references/workflow-design.md if:
User mentions: "MV"、"宣传片"、"完整作品"、"配乐"、"soundtrack"
Task spans multiple media types (image + video, video + music, etc.)
Complex multi-step workflows that need task decomposition
Check for visual consistency needs — Read ima-knowledge-ai/references/visual-consistency.md if:
User mentions: "系列"、"多张"、"同一个"、"角色"、"续"、"series"、"same"
Task involves: multiple images/videos, character continuity, product shots
Second+ request about same subject (e.g., "旺财在游泳" after "生成旺财照片")
Check video modes — Read ima-knowledge-ai/references/video-modes.md if:
Any video generation task
Need to understand: image_to_video vs reference_image_to_video difference
Check model selection — Read ima-knowledge-ai/references/model-selection.md if:
Unsure which model to use
Need cost/quality trade-off guidance
User specifies budget or quality requirements
Why this matters:
Multi-media workflows need proper task sequencing (e.g., video duration → matching music duration)
AI generation defaults to 独立生成 each time — without reference images, results will be inconsistent
Wrong video mode = wrong result (image_to_video ≠ reference_image_to_video)
Model choice affects cost and quality significantly
Example multi-media workflow:
User: "帮我做个产品宣传MV,有背景音乐,主角是旺财小狗"
❌ Wrong:
1. Generate dog image (random look)
2. Generate video (different dog)
3. Generate music (unrelated)
✅ Right:
1. Read workflow-design.md + visual-consistency.md
2. Generate Master Reference: 旺财小狗图片
3. Generate video shots using image_to_video with 旺财 as first frame
4. Get video duration (e.g., 15s)
5. Generate BGM with matching duration and mood
How to check:
# Step 0: Determine media type first (image / video / music / speech)
# From user request: "画"/"生成图"/"image" → image; "视频"/"video" → video; "音乐"/"歌"/"music"/"BGM" → music; "语音"/"朗读"/"TTS"/"speech" → speech
# Then choose task_type and model from the corresponding section (image: text_to_image/image_to_image; video: text_to_video/...; music: text_to_music; speech: text_to_speech)
# Step 1: Read knowledge base based on task type
if multi_media_workflow:
read("~/.openclaw/skills/ima-knowledge-ai/references/workflow-design.md")
if "same subject" or "series" or "character":
read("~/.openclaw/skills/ima-knowledge-ai/references/visual-consistency.md")
if video_generation:
read("~/.openclaw/skills/ima-knowledge-ai/references/video-modes.md")
# Step 2: Execute with proper sequencing and reference images
# (see workflow-design.md for specific patterns)
No exceptions — for simple single-media requests, you can proceed directly. For complex multi-media workflows, read the knowledge base first.
📥 User Input Parsing (Media Type & Task Routing)
Purpose: So that any agent parses user intent consistently, first determine the media type from the user's request, then choose task_type and model.
If the request mixes media (e.g. "宣传片+配乐"), treat as multi-media workflow: read workflow-design.md, then plan image → video → music steps and use the correct task_type for each step.
2. Model and parameter parsing
Image: For model name → model_id and size/aspect_ratio parsing, follow the same rules as in ima-image-ai skill (User Input Parsing section).
Video: For task_type (t2v / i2v / first_last / reference), model alias → model_id, and duration/resolution/aspect_ratio, follow ima-video-ai skill (User Input Parsing section).
Sevio alias normalization in ima-all-ai:
Ima Sevio 1.0 → ima-pro
Ima Sevio 1.0-Fast / Ima Sevio 1.0 Fast → ima-pro-fast
Routing rule:
Normalize alias first
Then resolve against runtime product list for the selected task_type
If model is absent in current category, return available model_ids from --list-models
Music: Suno (sonic) vs DouBao BGM/Song — infer from "BGM"/"背景音乐" → BGM; "带歌词"/"人声" → Suno or Song. Use model_id sonic, GenBGM, GenSong per "Recommended Defaults" and "Music Generation" tables below.
Speech (TTS): Get model_id from GET /open/v1/product/list?category=text_to_speech or run script with --task-type text_to_speech --list-models. Map user intent to parameters using product form_config:
User intent / phrasing
Parameter (if in form_config)
Notes
女声 / 女声朗读 / female voice
voice_id / voice_type
Use value from form_config options
男声 / 男声朗读 / male voice
voice_id / voice_type
Use value from form_config options
语速快/慢 / speed up/slow
speed
e.g. 0.8–1.2
音调 / pitch
pitch
If supported
大声/小声 / volume
volume
If supported
If the user does not specify, use form_config defaults. Pass extra params via --extra-params '{"speed":1.0}'. Only send parameters present in the product’s credit_rules/attributes or form_config (script reflection strips others on retry).
⚙️ How This Skill Works
For transparency: This skill uses a bundled Python script (scripts/ima_create.py) to call the IMA Open API. The script:
Sends your prompt to two IMA-owned domains (see "Network Endpoints" below)
Uses --user-idonly locally as a key for storing your model preferences
Returns image/video/music URLs when generation is complete
❌ NO API key in prompts (key is used for authentication only)
❌ NO user_id (it's only used locally)
What's stored locally:
~/.openclaw/memory/ima_prefs.json - Your model preferences (< 1 KB)
~/.openclaw/logs/ima_skills/ - Generation logs (auto-deleted after 7 days)
🌐 Network Endpoints Used
Domain
Owner
Purpose
Data Sent
Privacy
api.imastudio.com
IMA Studio
Main API (product list, task creation, task polling)
Prompts, model IDs, generation params, your API key
Standard HTTPS, data processed for AI generation
imapi.liveme.com
IMA Studio
Image/Video upload service (presigned URL generation)
Your API key, file metadata (MIME type, extension)
Standard HTTPS, used for image/video tasks only
*.aliyuncs.com, *.esxscloud.com
Alibaba Cloud (OSS)
Image/video storage (file upload, CDN delivery)
Raw image/video bytes (via presigned URL, NO API key)
IMA-managed OSS buckets, presigned URLs expire after 7 days
Key Points:
Music tasks (text_to_music) and TTS tasks (text_to_speech) only use api.imastudio.com.
Image/video tasks require imapi.liveme.com to obtain presigned URLs for uploading input images.
Your API key is sent to both api.imastudio.com and imapi.liveme.com (both owned by IMA Studio).
Verify network calls: tcpdump -i any -n 'host api.imastudio.com or host imapi.liveme.com'. See this document: 🌐 Network Endpoints Used and ⚠️ Credential Security Notice for full disclosure.
🧪 Use test keys for experiments: Generate a separate API key for testing.
🔍 Monitor usage: Check https://imastudio.com/dashboard for unauthorized activity.
⏱️ Rotate keys: Regenerate your API key periodically (monthly recommended).
📊 Review logs: Check ~/.openclaw/logs/ima_skills/ for unexpected API calls.
Why two domains? IMA Studio uses a microservices architecture:
api.imastudio.com: Core AI generation API
imapi.liveme.com: Specialized image/video upload service (shared infrastructure)
Both domains are operated by IMA Studio. The same API key grants access to both services.
Agent Execution (Internal Reference)
Note for users: You can review the script source at scripts/ima_create.py anytime.
The agent uses this script to simplify API calls. Music tasks use only api.imastudio.com, while image/video tasks also call imapi.liveme.com for file uploads (see "Network Endpoints" above).
Use the bundled script internally for all task types — it ensures correct parameter construction:
The script outputs JSON with url, model_name, credit — use these values in the UX protocol messages below. The script internals (product list query, parameter construction, polling) are invisible to users.
Overview
Call IMA Open API to create AI-generated content. All endpoints require an ima_* API key. The core flow is: query products → create task → poll until done.
🔒 Security & Transparency Policy
This skill is community-maintained and open for inspection.
✅ What Users CAN Do
Full transparency:
✅ Review all source code: Check scripts/ima_create.py and ima_logger.py anytime
✅ Verify network calls: Music tasks use api.imastudio.com only; image/video tasks also use imapi.liveme.com (see "Network Endpoints" section)
✅ Inspect local data: View ~/.openclaw/memory/ima_prefs.json and log files
✅ Control privacy: Delete preferences/logs anytime, or disable file writes (see below)
If users report issues, verify file integrity first.
🧠 User Preference Memory (Image)
User preferences have highest priority when they exist. But preferences are only saved when users explicitly express model preferences — not from automatic model selection.
These are fallback defaults — only used when no user preference exists. Always default to the newest and most popular model. Do NOT default to the cheapest.
Task Type
Default Model
model_id
version_id
Cost
Why
text_to_image
SeeDream 4.5
doubao-seedream-4.5
doubao-seedream-4-5-251128
5 pts
Latest doubao flagship, photorealistic 4K
text_to_image (budget)
Nano Banana2
gemini-3.1-flash-image
gemini-3.1-flash-image
4 pts
Fastest and cheapest option
text_to_image (premium)
Nano Banana Pro
gemini-3-pro-image
gemini-3-pro-image-preview
10/10/18 pts
Premium quality, 1K/2K/4K options
text_to_image (artistic)
Midjourney 🎨
midjourney
v6
8/10 pts
Artist-level aesthetics, creative styles
image_to_image
SeeDream 4.5
doubao-seedream-4.5
doubao-seedream-4-5-251128
5 pts
Latest, best i2i quality
image_to_image (budget)
Nano Banana2
gemini-3.1-flash-image
gemini-3.1-flash-image
4 pts
Cheapest option
image_to_image (premium)
Nano Banana Pro
gemini-3-pro-image
gemini-3-pro-image-preview
10 pts
Premium quality
image_to_image (artistic)
Midjourney 🎨
midjourney
v6
8/10 pts
Artist-level aesthetics, style transfer
text_to_video
Wan 2.6
wan2.6-t2v
wan2.6-t2v
25 pts
🔥 Most popular t2v, balanced cost
text_to_video (premium)
Hailuo 2.3
MiniMax-Hailuo-2.3
MiniMax-Hailuo-2.3
38 pts
Higher quality
text_to_video (budget)
Vidu Q2
viduq2
viduq2
5 pts
Lowest cost t2v
image_to_video
Wan 2.6
wan2.6-i2v
wan2.6-i2v
25 pts
🔥 Most popular i2v, 1080P
image_to_video (premium)
Kling 2.6
kling-v2-6
kling-v2-6
40-160 pts
Premium Kling i2v
first_last_frame_to_video
Kling O1
kling-video-o1
kling-video-o1
48 pts
Newest Kling reasoning model
reference_image_to_video
Kling O1
kling-video-o1
kling-video-o1
48 pts
Best reference fidelity
text_to_music
Suno (sonic-v4)
sonic
sonic
25 pts
Latest Suno engine, best quality
text_to_speech
(query product list)
—
—
—
Run --task-type text_to_speech --list-models; use first or user-preferred model_id
Premium options:
Image: Nano Banana Pro — Highest quality with size control (1K/2K/4K), higher cost (10-18 pts for text_to_image, 10 pts for image_to_image)
Video: Kling O1, Sora 2 Pro, Google Veo 3.1 — Premium quality with longer duration options
Quick selection guide (production as of 2026-02-27, sorted by popularity):
Style transfer / image editing → SeeDream 4.5 (5pts) or Midjourney 🎨 (artistic)
Video Generation:
General video generation → Wan 2.6 (25pts, most popular)
Premium cinematic quality → Google Veo 3.1 (70-330pts) or Sora 2 Pro (122+pts)
Budget video → Vidu Q2 (5pts) or Hailuo 2.0 (5pts)
With audio support → Kling O1 (48+pts) or Google Veo 3.1 (70+pts)
First/last frame animation → Kling O1 (48+pts)
Reference image consistency → Kling O1 (48+pts) or Google Veo 3.1 (70+pts)
Music Generation:
Custom song with lyrics, vocals, style → Suno sonic-v5 (25pts, default, ~2min)
Full control: custom_mode, lyrics, vocal_gender, tags, negative_tags
Best for: complete songs, vocal tracks, artistic compositions
Background music / ambient loop → DouBao BGM (30pts, ~30s)
Simplified: prompt-only, no advanced parameters
Best for: video backgrounds, ambient music, short loops
Simple song generation → DouBao Song (30pts, ~30s)
Simplified: prompt-only
Best for: quick song generation, structured vocal compositions
User explicitly asks for cheapest → DouBao BGM/Song (6pts each) — only if explicitly requested
Speech (TTS) Generation:
Text-to-speech / 语音合成 / 朗读 → text_to_speech. Always query GET /open/v1/product/list?category=text_to_speech (or --list-models) to get current model_id and credit. No fixed default; use first available or user preference. Voice/speed/format parameters: see "Model and parameter parsing" (TTS table) and "Speech (TTS) — text_to_speech" in this document.
⚠️ Technical Note for Suno:
model_version inside parameters.parameters (e.g., "sonic-v5") is different from the outer model_version field (which is "sonic"). Always set both correctly when creating Suno tasks.
Any model + 8K: Inform user no model supports 8K, max is 4K (Nano Banana Pro)
Any model + non-standard ratio (7:3, 8:5, etc.): Non-standard ratio, not supported. Suggest closest supported ratio (e.g., 21:9 for ultra-wide, 2:3 for portrait)
Video Models (text_to_video / image_to_video / first_last_frame / reference_image)
No vocals: "no vocals" → set make_instrumental=true
With vocals: "female vocals" → set vocal_gender="female"
Male vocals: "male vocals" → set vocal_gender="male"
Mixed: Set vocal_gender="mixed"
Mood/Emotion:
Examples: "happy and energetic", "melancholic", "tense and dramatic", "peaceful and calming"
Negative Tags (exclude styles):
Use negative_tags: "heavy metal, distortion, screaming" to exclude unwanted elements
Duration Hint:
Examples: "60 seconds", "30 second loop", "2 minute track"
Note: Suno typically generates ~2min, not strictly controllable
Example Suno prompts:
"upbeat lo-fi hip hop, 90 BPM, no vocals, relaxed and chill"
→ Set: make_instrumental=true
"emotional pop ballad, slow tempo, female vocals, melancholic"
→ Set: vocal_gender="female"
"orchestral cinematic trailer music, epic and dramatic, 120 BPM, no vocals"
→ Set: make_instrumental=true, tags="orchestral,cinematic,epic"
"acoustic indie folk, gentle guitar, male vocals, warm and nostalgic"
→ Set: vocal_gender="male", tags="acoustic,indie,folk"
⚠️ Technical Note for Suno:
model_version inside parameters.parameters (e.g., "sonic-v5") is different from the outer model_version field (which is "sonic"). Always set both correctly.
Common Parameter Patterns
n (batch generation): Supported by ALL models. Cost = base_credit × n. Creates n independent resources.
seed: Supported by most models (-1 = random, >0 = reproducible results)
Decision Tree: When User Requests Unsupported Features
User asks for custom aspect ratio image (e.g. "7:3 landscape")
→ ❌ Image models don't support custom ratios
→ ✅ Solution: "图片模型不支持自定义比例。建议用视频模型(Wan 2.6 t2v)生成16:9视频,然后截取首帧作为图片。"
User asks for 8K image
→ ❌ No model supports 8K
→ ✅ Solution: "当前最高支持4K分辨率(Nano Banana Pro,18积分)。要使用吗?"
User asks for video with audio
→ Check model: Veo 3.1 / Kling O1 / Hailuo have generate_audio
→ ✅ Solution: "Veo 3.1 和 Kling O1 支持音频生成(需在参数中设置 generate_audio=True)。要用哪个?"
User asks for long music (e.g. "5 minute track")
→ ❌ Duration not user-controllable
→ ✅ Solution: "Suno 生成约2分钟音乐。需要更长时长可以生成多段后拼接。"
User asks for 30s video
→ Check model: Most models max 15s
→ ✅ Solution: "当前最长15秒。可选模型:Wan 2.6(15s, 75积分), Kling O1(10s, 96积分)。"
Note: Image-specific unsupported combinations (Midjourney + aspect_ratio, 8K, non-standard ratios) are documented in the "Image Models" section above.
🧠 User Preference Memory (Video)
User preferences have highest priority when they exist. But preferences are only saved when users explicitly express model preferences — not from automatic model selection.
Added Step 0 for correct message ordering (fixes group chat bug)
Added Step 5 for explicit task completion
Enhanced Midjourney support with proper timing estimates
Now 6 steps total (0-5): Acknowledgment → Pre-Gen → Progress → Success/Failure → Done
This skill runs inside IM platforms (Feishu, Discord via OpenClaw).
Generation takes 10 seconds (music) up to 6 minutes (video). Never let users wait in silence.
Always follow all 6 steps below, every single time.
🗣️ User-Friendly First, Transparent on Request
Default to plain-language updates in normal user flows.
If users ask for technical details, provide them transparently (script name, endpoints, and key parameters).
In standard progress messages, prioritize: model name, estimated/actual time, credits consumed, result URL, and natural-language status updates.
Estimated Generation Time (All Task Types)
Task Type
Model
Estimated Time
Poll Every
Send Progress Every
text_to_image
SeeDream 4.5
25~60s
5s
20s
Nano Banana2 💚
20~40s
5s
15s
Nano Banana Pro
60~120s
5s
30s
Midjourney 🎨
40~90s
8s
25s
image_to_image
SeeDream 4.5
25~60s
5s
20s
Nano Banana2 💚
20~40s
5s
15s
Nano Banana Pro
60~120s
5s
30s
Midjourney 🎨
40~90s
8s
25s
text_to_video
Wan 2.6, Hailuo 2.0/2.3, Vidu Q2, Pixverse
60~120s
8s
30s
SeeDance 1.5 Pro, Kling 2.6, Veo 3.1
90~180s
8s
40s
Kling O1, Sora 2 Pro
180~360s
8s
60s
image_to_video
Same ranges as text_to_video
—
8s
40s
first_last_frame / reference
Kling O1, Veo 3.1
180~360s
8s
60s
text_to_music
DouBao BGM / Song
10~25s
5s
10s
Suno (sonic-v5)
20~45s
5s
15s
text_to_speech
(varies by model)
5~30s
3s
10s
estimated_max_seconds = upper bound of the range (e.g. 60 for SeeDream 4.5, 40 for Nano Banana2, 120 for Nano Banana Pro, 90 for Midjourney, 180 for Kling 2.6, 360 for Kling O1).
⚠️ CRITICAL: This step is essential for correct message ordering in IM platforms (Feishu, Discord).
Before doing anything else, reply to the user with a friendly acknowledgment message using your normal reply (not message tool). This reply will automatically appear FIRST in the conversation.
Example acknowledgment messages:
For images:
好的!来帮你画一只萌萌的猫咪 🐱
收到!马上为你生成一张 16:9 的风景照 🏔️
OK! Starting image generation with SeeDream 4.5 🎨
For videos:
好的!来帮你生成一段视频 🎬
收到!开始用 Wan 2.6 生成视频 🎥
For music:
好的!来帮你创作一首音乐 🎵
Rules:
Keep it short and warm (< 15 words)
Match the user's language (Chinese/English)
Include relevant emoji (🐱/🎨/🎬/🎵/✨)
This is your ONLY normal reply — all subsequent updates use message tool
Why this matters:
Normal replies automatically appear FIRST in the conversation thread
message tool pushes appear in chronological order AFTER your initial reply
Premium (auto-selected): "使用 Wan 2.6(25 积分)。若需更高质量可选 Kling O1(48 积分起)"
Budget (user asked): "使用 Vidu Q2(5 积分,最省钱)"
Adapt language to match the user (Chinese / English). For video, always add a note that it takes longer. For expensive models, always mention cheaper alternatives unless user explicitly requested premium.
Step 2 — Progress Updates
Poll the task detail API every [Poll Every] seconds per the table.
Send a progress update every [Send Progress Every] seconds.
⏳ 正在生成中… [P]%
已等待 [elapsed]s,预计最长 [max]s
Progress formula:
P = min(95, floor(elapsed_seconds / estimated_max_seconds * 100))
Cap at 95% — never reach 100% until the API confirms success
If elapsed > estimated_max: freeze at 95%, append 「快了,稍等一下…」
For video with max=360s: at 120s → 33%, at 250s → 69%, at 400s → 95% (frozen)
Step 3 — Success Notification
When task status = success:
For Video Tasks (text_to_video / image_to_video / first_last_frame / reference_image)
3.1 Send video player first (IM platforms like Feishu will render inline player):
# Get result URL from script output or task detail API
result = get_task_result(task_id)
video_url = result["medias"][0]["url"]
# Build caption
caption = f"""✅ 视频生成成功!
• 模型:[Model Name]
• 耗时:预计 [X~Y]s,实际 [actual]s
• 消耗积分:[N pts]
[视频描述]"""
# Add mismatch hint if user pref conflicts with knowledge-ai recommendation
if user_pref_exists and knowledge_recommended_model != used_model:
caption += f"""
💡 提示:当前任务也许用 {knowledge_recommended_model} 也会不错({reason},{cost} pts)"""
# Send video with caption (use message tool if available)
message(
action="send",
media=video_url, # ⚠️ Use HTTPS URL directly, NOT local file path
caption=caption
)
Important:
Hint is non-intrusive — does NOT interrupt generation
Only shown when user pref conflicts with knowledge-ai recommendation
User can ignore the hint; video is already delivered
3.2 Then send link as text (for copying/sharing):
# Send link message immediately after video
message(action="send", text=f"🔗 视频链接(可复制分享):\n{video_url}")
⚠️ Critical for video:
Send video player FIRST (inline preview)
Send text link SECOND (for copying)
Include first-frame thumbnail URL if available: result["medias"][0]["cover"]
For Image Tasks (text_to_image / image_to_image)
# Build caption
caption = f"""✅ 图片生成成功!
• 模型:[Model Name]
• 耗时:预计 [X~Y]s,实际 [actual]s
• 消耗积分:[N pts]
🔗 原始链接:{image_url}"""
# Add mismatch hint if user pref conflicts with knowledge-ai recommendation
if user_pref_exists and knowledge_recommended_model != used_model:
caption += f"""
💡 提示:当前任务也许用 {knowledge_recommended_model} 也会不错({reason},{cost} pts)"""
# Send image with caption
message(
action="send",
media=image_url,
caption=caption
)
Important:
Hint is non-intrusive — does NOT interrupt generation
Only shown when user pref conflicts with knowledge-ai recommendation
User can ignore the hint; image is already delivered
Step 5 — Done
After Step 0–4, no further reply needed. Do not send duplicate confirmations.
Step 4 — Failure Notification
When task status = failed or any API/network error, send:
❌ [内容类型]生成失败
• 原因:[natural_language_error_message]
• 建议改用:
- [Alt Model 1]([特点],[N pts])
- [Alt Model 2]([特点],[N pts])
需要我帮你用其他模型重试吗?
⚠️ CRITICAL: Error Message Translation
NEVER show technical error messages to users. Always translate API errors into natural language. API key & credits: 密钥与积分管理入口为 imaclaw.ai(与 imastudio.com 同属 IMA 平台)。Key and subscription management: imaclaw.ai (same IMA platform as imastudio.com).
Network connection unstable, check network and retry
Rate limit exceeded
429 Too Many Requests / Rate limit
请求过于频繁,请稍等片刻再试
Too many requests, please wait a moment
Prompt moderation (Sora only)
Content policy violation
提示词包含敏感内容,请修改后重试
Prompt contains restricted content, please modify
Model unavailable
Model not available / 503 Service Unavailable
当前模型暂时不可用,建议换个模型
Model temporarily unavailable, try another model
Lyrics format error (Suno only) 🎵
Invalid lyrics format
歌词格式有误,请调整后重试
Lyrics format error, adjust and retry
Prompt too short/long (Music) 🎵
Prompt length invalid
音乐描述过短或过长,请调整到合适长度 (建议20-100字)
Music description too short or long, adjust to appropriate length (20-100 chars recommended)
Text too long (TTS) 🔊
TTS text length
文本过长,请缩短后重试
Text too long, please shorten and retry
Generic fallback (when error is unknown):
Chinese: 生成过程遇到问题,请稍后重试或换个模型试试
English: Generation encountered an issue, please try again or use another model
Best Practices:
Focus on user action: Tell users what to do next, not what went wrong technically
Be reassuring: Use phrases like "建议换个模型试试" instead of "失败了"
Avoid blame: Never say "你的提示词有问题" → say "提示词需要调整一下"
Provide alternatives: Always suggest 1-2 alternative models in the failure message
🆕 Include actionable links (v1.0.8+): For 401/4008 errors, provide clickable links to API key generation or credit purchase pages
🎵 Music-specific (v1.2.0+):
For Suno lyrics errors, suggest simplifying lyrics or using auto-generated lyrics (auto_lyrics=true)
For prompt length errors, give example length (e.g., "建议20-100字")
For BGM requests, recommend DouBao BGM over Suno
🔊 TTS-specific: Use the TTS error translation table in "For TTS Tasks (text_to_speech)" above; suggest another model via --list-models or shortening text.
Step 5 — Done (No Further Action Needed) 🆕
After sending Step 3 (success) or Step 4 (failure):
DO NOT send any additional messages unless the user asks a follow-up question
The task is complete — wait for the user's next request
User preference has been saved (if generation succeeded)
The conversation is ready for the next generation request
Why this step matters:
Prevents unnecessary "anything else?" messages that clutter the chat
Allows users to naturally continue the conversation when ready
Respects the asynchronous nature of IM platforms
Exception: If the user explicitly asks "还有别的吗?" or similar, then respond naturally.
🆕 Enhanced Error Handling (v1.0.8):
The Reflection mechanism (3 automatic retries) now provides specific, actionable suggestions for common errors:
401 Unauthorized: System suggests generating a new API key with clickable link
4008 Insufficient Points: System suggests purchasing credits with clickable link
Timeout: Helpful info with dashboard link for background task status
All error handling is automatic and transparent — users receive natural language explanations with next steps.
Failure fallback by task type:
Task Type
Failed Model
First Alt
Second Alt
text_to_image
SeeDream 4.5
Nano Banana2 (4pts, fast)
Nano Banana Pro (10-18pts, premium)
text_to_image
Nano Banana2
SeeDream 4.5 (5pts, better quality)
Nano Banana Pro (10-18pts)
text_to_image
Nano Banana Pro
SeeDream 4.5 (5pts)
Nano Banana2 (4pts, budget)
image_to_image
SeeDream 4.5
Nano Banana2 (4pts, fast)
Nano Banana Pro (10pts)
image_to_image
Nano Banana2
SeeDream 4.5 (5pts)
Nano Banana Pro (10pts)
image_to_image
Nano Banana Pro
SeeDream 4.5 (5pts)
Nano Banana2 (4pts)
text_to_video
Kling O1
Wan 2.6 (25pts)
Vidu Q2 (5pts)
text_to_video
Google Veo 3.1
Kling O1 (48pts)
Sora 2 Pro (122pts)
text_to_video
Any
Wan 2.6 (25pts, most popular)
Hailuo 2.0 (5pts)
image_to_video
Wan 2.6
Kling O1 (48pts)
Hailuo 2.0 i2v (25pts)
image_to_video
Any
Wan 2.6 (25pts, most popular)
Vidu Q2 Pro (20pts)
first_last / reference
Kling O1
Kling 2.6 (80pts)
Veo 3.1 (70pts+)
text_to_music 🎵
Suno
DouBao BGM (30pts, 背景音乐)
DouBao Song (30pts, 歌曲生成)
text_to_music 🎵
DouBao BGM
DouBao Song (30pts)
Suno (25pts, 功能最强)
text_to_music 🎵
DouBao Song
DouBao BGM (30pts)
Suno (25pts, 功能最强)
text_to_speech 🔊
(any)
Query --list-models for alternatives
Use another model_id from product list
Music-specific failure guidance:
If Suno fails → Recommend DouBao BGM (for background music) or DouBao Song (for songs)
If DouBao BGM fails → Try DouBao Song first (similar pricing), then Suno (more powerful)
If DouBao Song fails → Try DouBao BGM first (similar pricing), then Suno (more powerful)
For lyrics errors in Suno → Suggest simplifying lyrics or using auto_lyrics=true
For prompt length errors → Recommend 20-100 characters
TTS-specific failure guidance:
If TTS fails → Run --task-type text_to_speech --list-models and suggest another model_id; or shorten text / simplify content. Use the TTS error translation table in "For TTS Tasks" above for user-facing messages.
Supported Models at a Glance
Source: production GET /open/v1/product/list (2026-02-27). Model count reduced significantly. Always query product list API at runtime.
Models and credits are not fixed. Always call GET /open/v1/product/list?category=text_to_speech (or run the script with --task-type text_to_speech --list-models) to get current model_id, attribute_id, and credit.
ima-all-ai has complete TTS capability: This document and the bundled ima_create.py provide full TTS support (routing, parameters, create/poll, UX protocol Steps 0–5, error translation). The ima-tts-ai skill is an optional standalone package with the same specification.
TTS Task Detail — Response Shape
Poll POST /open/v1/tasks/detail until completion. For TTS, medias[] uses the same structure as other IMA audio tasks:
Field
Type
Meaning
resource_status
int or null
0=处理中, 1=可用, 2=失败, 3=已删除;null 视为 0
status
string
"pending" / "processing" / "success" / "failed"
url
string
Audio URL when resource_status=1 (mp3/wav)
duration_str
string
Optional, e.g. "12s"
format
string
Optional, e.g. "mp3", "wav"
Success example: When all medias have resource_status == 1 and status != "failed", read medias[0].url (or watermark_url). Example: {"medias":[{"resource_status":1,"status":"success","url":"https://cdn.../output.mp3","duration_str":"12s","format":"mp3"}]}.
TTS Create Task — Request Shape
task_type: "text_to_speech". No image input: src_img_url: [], input_images: []. prompt (text to speak) must be inside parameters[].parameters, not at top level. Extra fields (e.g. voice_id, speed) come from product form_config; pass via --extra-params and only include params present in the product’s credit_rules/form_config.
TTS Common Mistakes
Mistake
Fix
prompt at top level
Put prompt inside parameters[].parameters (script does this)
Wrong or missing attribute_id
Always call product list first; use credit_rules
Single poll
Poll until all medias have resource_status == 1
Ignoring status when resource_status=1
Check status != "failed"
Sending params not in form_config/credit_rules
Use only params from product list; script reflection strips others on retry
Always call GET /open/v1/product/list?category=<type> first to get the live attribute_id and form_config defaults required for task creation.
There are two equivalent route systems serving the same backend logic:
Route
Auth
Use Case
/open/v1/
Authorization: Bearer ima_* only
Third-party / agent access
/api/v3/
Token + API Key (dual auth)
Frontend App
This skill documents the /open/v1/ Open API. All business logic (credit validation, N-flattening, risk control) runs identically on both paths.
Environment
Base URL: https://api.imastudio.com
Required/recommended headers for all /open/v1/ endpoints:
Header
Required
Value
Notes
Authorization
✅
Bearer ima_your_api_key_here
API key authentication
x-app-source
✅
ima_skills
Fixed value — identifies skill-originated requests
x_app_language
recommended
en / zh
Product label language; defaults to en if omitted
Authorization: Bearer ima_your_api_key_here
x-app-source: ima_skills
x_app_language: en
📤 When to Upload Images (Quick Reference)
The IMA Open API does NOT accept raw bytes or base64 images. All image inputs must be public HTTPS URLs.
Task Type
Input Required?
Upload Before Create?
Notes
text_to_image
❌ No
—
Prompt only
image_to_image
✅ Yes (1 image)
✅ Upload first
Single input image
text_to_video
❌ No
—
Prompt only
image_to_video
✅ Yes (1 image)
✅ Upload first
Single input image
first_last_frame_to_video
✅ Yes (2 images)
✅ Upload first
First + last frame
reference_image_to_video
✅ Yes (1+ images)
✅ Upload first
Reference image(s)
text_to_music
❌ No
—
Prompt only
text_to_speech
❌ No
—
Prompt only (text to speak)
Upload flow:
User provides local file path or bytes → call prepare_image_url() (see section below)
User provides HTTPS URL → use directly, no upload needed
Use the returned CDN URL (fdl) as the value for input_images / src_img_url
Example workflow (image_to_image):
# User provides local file
image_url = prepare_image_url("/path/to/photo.jpg", api_key)
# → Returns: https://ima-ga.esxscloud.com/webAgent/privite/2026/02/27/..._uuid.jpeg
# Then create task with this URL
create_task(
task_type="image_to_image",
input_images=[image_url], # Use uploaded URL
prompt="turn into oil painting"
)
⚠️ MANDATORY: Always Query Product List First
CRITICAL: You MUST call /open/v1/product/list BEFORE creating any task.
The attribute_id field is REQUIRED in the create request. If it is 0 or missing, you get: "Invalid product attribute" → "Insufficient points" → task fails completely. NEVER construct a create request from the model table alone. Always fetch the product first.
How to get attribute_id (all task types)
# Query product list with the correct category
GET /open/v1/product/list?app=ima&platform=web&category=<task_type>
# task_type: text_to_image | image_to_image | text_to_video | image_to_video |
# first_last_frame_to_video | reference_image_to_video | text_to_music | text_to_speech
# Walk the V2 tree to find your target model (type=3 leaf nodes only)
for group in response["data"]:
for version in group.get("children", []):
if version["type"] == "3" and version["model_id"] == target_model_id:
attribute_id = version["credit_rules"][0]["attribute_id"]
credit = version["credit_rules"][0]["points"]
model_version = version["id"] # = version_id / model_version
model_name = version["name"]
form_defaults = {f["field"]: f["value"] for f in version["form_config"]}
break
Quick Reference: Known attribute_ids
Pre-queried values for convenience. Always call the product list at runtime for accuracy.
Model
Task Type
model_id
attribute_id
credit
Notes
text_to_image
SeeDream 4.5
text_to_image
doubao-seedream-4.5
2341
5 pts
Default, balanced
Nano Banana Pro (1K)
text_to_image
gemini-3-pro-image
2399
10 pts
1024×1024
Nano Banana Pro (2K)
text_to_image
gemini-3-pro-image
2400
10 pts
2048×2048
Nano Banana Pro (4K)
text_to_image
gemini-3-pro-image
2401
18 pts
4096×4096
text_to_video
Wan 2.6 (720P, 5s)
text_to_video
wan2.6-t2v
2057
25 pts
Default, balanced
Wan 2.6 (1080P, 5s)
text_to_video
wan2.6-t2v
2058
40 pts
—
Wan 2.6 (720P, 10s)
text_to_video
wan2.6-t2v
2059
50 pts
—
Wan 2.6 (1080P, 10s)
text_to_video
wan2.6-t2v
2060
80 pts
—
Wan 2.6 (720P, 15s)
text_to_video
wan2.6-t2v
2061
75 pts
—
Wan 2.6 (1080P, 15s)
text_to_video
wan2.6-t2v
2062
120 pts
—
Kling O1 (5s, std)
text_to_video
kling-video-o1
2313
48 pts
Latest Kling
Kling O1 (5s, pro)
text_to_video
kling-video-o1
2314
60 pts
—
Kling O1 (10s, std)
text_to_video
kling-video-o1
2315
96 pts
—
Kling O1 (10s, pro)
text_to_video
kling-video-o1
2316
120 pts
—
text_to_music
Suno (sonic-v4)
text_to_music
sonic
2370
25 pts
Default
DouBao BGM
text_to_music
GenBGM
4399
30 pts
—
DouBao Song
text_to_music
GenSong
4398
30 pts
—
All others
any
—
→ query /open/v1/product/list
—
Always runtime query
⚠️ Production warning: attribute_id and credit values change frequently in production. Always call /open/v1/product/list at runtime; above table is pre-queried reference only (2026-02-27).
Error 6010: "Attribute ID does not match the calculated rule"
prompt at outer parameters[] level
Prompt ignored; wrong routing
cast missing from inner parameters.parameters
Billing validation failure
credit value wrong or missing
Error 6006
model_name / model_version missing
Wrong backend routing
Skipped product list, used table values directly
All of the above
⚠️ Critical for Google Veo 3.1 and multi-rule models:
Models like Google Veo 3.1 have multiple credit_rules, each with a different attribute_id for different parameter combinations:
720p + 4s + optimized → attribute_id A
720p + 8s + optimized → attribute_id B
4K + 4s + high → attribute_id C
The script automatically selects the correct attribute_id by matching your parameters (duration, resolution, compression_quality, generate_audio) against each rule's attributes. If the match fails, you get error 6010.
Fix: The bundled script now checks these video-specific parameters for smart credit_rule selection. Always use the script, not manual API construction.
Core Flow
1. GET /open/v1/product/list?app=ima&platform=web&category=<type>
→ REQUIRED: Get attribute_id, credit, model_version, model_name, form_config defaults
[If input image required]
2. Upload image → get public HTTPS URL
→ See "Image Upload" section below
3. POST /open/v1/tasks/create
→ Must include: attribute_id, model_name, model_version, credit, cast, prompt (nested!)
4. POST /open/v1/tasks/detail {"task_id": "..."}
→ Poll until medias[].resource_status == 1
→ Extract url from completed media
Image Upload (Required Before Image Tasks)
The IMA Open API does NOT accept raw bytes or base64 images. All image inputs must be public HTTPS URLs.
When a user provides an image (local file, bytes, base64), you must upload it first and get a URL. This is exactly what the IMA frontend does before every image task.
Real Upload Flow (from IMA Frontend Source)
The frontend uses a two-step presigned URL flow via the IM platform:
Step 1: GET /api/rest/oss/getuploadtoken → returns { ful, fdl }
ful = presigned PUT URL (upload destination, expires ~7 days)
fdl = final CDN download URL (use this as input_images value)
Step 2: PUT {ful} with raw image bytes + Content-Type header
→ image is stored in Aliyun OSS: zhubite-imagent-bot.oss-us-east-1.aliyuncs.com
→ accessible via CDN: https://ima-ga.esxscloud.com/...
Step 1: Get Upload Token
GET https://imapi.liveme.com/api/rest/oss/getuploadtoken
Required query parameters (11 total — sourced directly from frontend generateUploadInfo):
Parameter
Example
Description
appUid
ima_xxx...
Use IMA API key directly — no separate login needed
appId
webAgent
App identifier (fixed)
appKey
32jdskjdk320eew
App secret (fixed, used for sign generation)
cmimToken
ima_xxx...
Use IMA API key directly — same as appUid
sign
117CF6CF...
IM auth HMAC: SHA1("webAgent|32jdskjdk320eew|{timestamp}|{nonce}").upper()
timestamp
1772042430
Unix timestamp (seconds), generated per request
nonce
CxI1FLI5ajLJZ1jlxZmeg
Random nonce string, generated per request
fService
privite
Fixed: storage service type
fType
picture
picture for images, video, audio
fSuffix
jpeg
File extension: jpeg, png, mp4, mp3
fContentType
image/jpeg
MIME type of the file
简化认证:直接使用 IMA API key 填充 appUid 和 cmimToken 参数,无需单独获取凭证。
PUT {ful}
Content-Type: image/jpeg
Body: [raw image bytes]
No auth headers needed — the presigned URL already encodes the credentials.
Step 3: Use fdl as the Image URL
After the PUT succeeds, use fdl (the CDN URL) as the value for input_images / src_img_url.
Python Implementation
import hashlib, time, uuid, requests, mimetypes
# ── 🌐 IMA Upload Service Endpoint (IMA-owned, for image/video uploads) ──────
IMA_IM_BASE = "https://imapi.liveme.com"
# ── 🔑 Hardcoded APP_KEY (Public, Shared Across All Users) ──────────────────
# This APP_KEY is a PUBLIC identifier used by IMA Studio's image/video upload
# service. It is NOT a secret—it's intentionally shared across all users and
# embedded in the IMA web frontend. This key is used to generate HMAC signatures
# for upload token requests, but your IMA API key (ima_xxx...) is the ACTUAL
# authentication credential. Think of APP_KEY as a "client ID" rather than a
# "client secret."
#
# ⚠️ Security Note: Your ima_xxx... API key is the sensitive credential. It is
# sent to imapi.liveme.com as query parameters (appUid, cmimToken). Always use
# test keys for experiments and rotate your API key regularly.
#
# 📖 See SECURITY.md for complete disclosure and network verification guide.
APP_ID = "webAgent"
APP_KEY = "32jdskjdk320eew" # Public shared key (used for HMAC sign generation)
APP_UID = "<your_app_uid>" # POST /api/v3/login/app → data.user_id
APP_TOKEN = "<your_app_token>" # POST /api/v3/login/app → data.token
def _gen_sign() -> tuple[str, str, str]:
"""Generate per-request (sign, timestamp, nonce)."""
nonce = uuid.uuid4().hex[:21]
ts = str(int(time.time()))
raw = f"{APP_ID}|{APP_KEY}|{ts}|{nonce}"
sign = hashlib.sha1(raw.encode()).hexdigest().upper()
return sign, ts, nonce
def get_upload_token(app_uid: str, app_token: str,
suffix: str, content_type: str) -> dict:
"""Step 1: Get presigned upload URL from IMA's upload service.
Calls GET imapi.liveme.com/api/rest/oss/getuploadtoken with exactly 11 params.
Returns: { "ful": "<presigned PUT URL>", "fdl": "<CDN download URL>" }
Args:
app_uid: Your IMA API key (ima_xxx...), used as appUid parameter
app_token: Your IMA API key (ima_xxx...), used as cmimToken parameter
suffix: File extension (jpeg, png, mp4, mp3)
content_type: MIME type (image/jpeg, video/mp4, etc.)
Security Note: