一键导入
image-generation-super
图片生成与编辑(超级版),调用 GPT-Image-2 模型生成和编辑图片。需要 AI 画图、生成图片、编辑图片、多图融合、背景替换、风格转换、电商商品图合成、海报设计、插画创作时优先使用该工具。纯像素操作(文字叠加、加水印、裁剪、缩放)请改用 Pillow,不要触发本工具。
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
图片生成与编辑(超级版),调用 GPT-Image-2 模型生成和编辑图片。需要 AI 画图、生成图片、编辑图片、多图融合、背景替换、风格转换、电商商品图合成、海报设计、插画创作时优先使用该工具。纯像素操作(文字叠加、加水印、裁剪、缩放)请改用 Pillow,不要触发本工具。
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Generate short videos from a static image using the Kling AI image-to-video API. Use this skill whenever the user wants to animate an image, create a video from a photo, turn a picture into a video clip, or generate AI video from an uploaded image.
基于文本描述生成短视频(5s/10s),适用于电商营销、创意宣传、教育讲解等场景,异步轮询获取结果。
调用百度通用文字识别高精度版 OCR,对图片全部文字内容进行高精度检测识别,支持中英日韩等 20+ 语种,适用于文档数字化、多语言文本提取场景。
根据用户输入的主题自动生成完整 PPT 文件,返回封面图和下载链接;适用于办公汇报、教学课件、产品展示等需要快速生成 PPT 的场景。
识别飞机行程单图片,结构化提取24个字段(乘客、航班、票价、税费等);适用于差旅报销、财务管理、行程记录场景。
百度千帆AI搜索,实时检索全网网页/视频内容,返回结构化引用列表(标题/URL/摘要/日期)。用户需要实时信息、新闻、网页检索时使用。
| name | image-generation-super |
| description | 图片生成与编辑(超级版),调用 GPT-Image-2 模型生成和编辑图片。需要 AI 画图、生成图片、编辑图片、多图融合、背景替换、风格转换、电商商品图合成、海报设计、插画创作时优先使用该工具。纯像素操作(文字叠加、加水印、裁剪、缩放)请改用 Pillow,不要触发本工具。 |
| license | MIT |
调用 GPT-Image-2 模型进行 AI 图像生成与编辑,支持通过自然语言描述生成高质量图片,以及上传多张图片进行 AI 编辑融合。
| 属性 | 值 |
|---|---|
| Plugin ID | e480d4b6-835c-45f8-a494-d38da962b394 |
| 认证模式 | platform_managed(密钥由平台注入) |
| 密钥来源 | process.env["INTEGRATIONS_API_KEY"] |
| Auth Header | X-Gateway-Authorization: Bearer <key> |
| 支持平台 | Web、MiniProgram |
| 响应格式 | SSE 流(text/event-stream),心跳保活 + 最终结果事件中 Base64 编码内嵌于 data[].b64_json |
接口列表:
| 接口 | 方法 | Endpoint | 说明 |
|---|---|---|---|
| 创建图片 | POST | http://app-bo4w33bsdqm9-api-wLNdpny6ZpVa-gateway.appmiaoda.com/v1/images/generations | 根据文本描述生成图片 |
| 编辑图片 | POST | http://app-bo4w33bsdqm9-api-baBw3XMNVmv9-gateway.appmiaoda.com/v1/images/edits | 上传 1–3 张图片进行 AI 编辑融合 |
核心能力:
prompt 描述生成全新图片,支持多种尺寸和数量配置revised_prompt,展示模型自动优化后的提示词平台差异概览:
| 平台 | Edge Function 返回 | 前端获取图片方式 |
|---|---|---|
| Web(创建图片) | SSE 流式响应(心跳 + 结果) | 用 fetch + getReader() 消费 SSE 流,收到 type: "result" 事件后解析 |
| Web(编辑图片) | SSE 流式响应(心跳 + 结果) | 用 fetch + getReader() 消费 SSE 流,收到 type: "result" 事件后解析 |
| MiniProgram | JSON(含 Base64) | 解析 JSON,写临时文件后用 <image> 组件展示 |
详细参数说明、代码示例及两平台完整实现见:
references/image-generations-api.md — 创建图片接口references/image-edits-api.md — 编辑图片接口调用本工具前,先判断场景是否真的需要 AI 生成:
| 场景 | 推荐方案 |
|---|---|
| 根据文字描述生成全新图片 | ✅ 本工具(文生图) |
| 上传图片 + 提示词做风格转换或内容编辑 | ✅ 本工具(图生图) |
| 多张图片融合 / 背景替换 / 海报合成 | ✅ 本工具(多图编辑) |
| 在图片上叠加文字 / 水印 | ❌ 改用 Pillow(速度快、可离线) |
| 裁剪、缩放、格式转换、像素级操作 | ❌ 改用 Pillow |
| 图片内容审核 / 质量评分 | ❌ 改用视觉模型直接分析,无需生成 |
底层模型(GPT-Image-2)对英文提示词的理解和图像质量通常优于中文,请优先将用户需求改写为英文后再提交 API。
写作原则:
"a ginger cat sitting in a sunlit garden" 好于 "可爱的猫""no background",改写 "isolated on pure white background"high quality, detailed, 8k, photorealistic文生图模板:
[Subject], [Action/Pose/State], [Scene/Environment], [Lighting], [Style], [Quality]
示例:
A golden retriever puppy, sitting and looking up curiously, in a cozy living room with warm afternoon lighting, watercolor illustration style, high quality, detailed
图生图 / 多图编辑额外建议:
"convert to anime style" 或 "oil painting style""use image 1 as background, place the product from image 2 in the center"在调用 API 之前,先将用户需求翻译/改写为英文提示词,GPT-Image-2 模型对英文输入的图像质量明显优于中文。
两个接口均为同步调用,直接返回 Base64 编码图片数据,不含 URL。获得响应后必须立即将 Base64 解码保存为图片文件。
const apiKey = process.env["INTEGRATIONS_API_KEY"]!;
interface CreateImageResult {
created: number;
data: Array<{
b64_json: string;
revised_prompt: string;
}>;
background: string;
output_format: string;
quality: string;
size: string;
model: string;
}
/** 创建图片(文生图) */
async function createImage(
prompt: string,
size?: string,
n?: number
): Promise<CreateImageResult> {
const response = await fetch(
"http://app-bo4w33bsdqm9-api-wLNdpny6ZpVa-gateway.appmiaoda.com/v1/images/generations",
{
method: "POST",
headers: {
"Content-Type": "application/json",
"X-Gateway-Authorization": `Bearer ${apiKey}`,
},
body: JSON.stringify({
model: "gpt-image-2",
prompt,
size,
n,
}),
}
);
if (!response.ok) throw new Error(`HTTP error: ${response.status}`);
const json = await response.json();
if (json.error) throw new Error(`API error: ${JSON.stringify(json.error)}`);
return json;
}
interface EditImageResult {
created: number;
data: Array<{
b64_json: string;
revised_prompt: string;
}>;
background: string;
output_format: string;
quality: string;
size: string;
model: string;
usage?: {
input_tokens: number;
input_tokens_details: { image_tokens: number; text_tokens: number };
output_tokens: number;
output_tokens_details: { image_tokens: number; text_tokens: number };
total_tokens: number;
};
}
/** 编辑图片(多图融合/编辑) */
async function editImage(
prompt: string,
images: File[],
size?: string,
n?: number
): Promise<EditImageResult> {
const formData = new FormData();
formData.append("model", "gpt-image-2");
formData.append("prompt", prompt);
if (size) formData.append("size", size);
if (n) formData.append("n", String(n));
images.forEach((file, index) => {
formData.append(`image[${index}]`, file);
});
const response = await fetch(
"http://app-bo4w33bsdqm9-api-wLNdpny6ZpVa-gateway.appmiaoda.com/v1/images/edits",
{
method: "POST",
headers: {
"X-Gateway-Authorization": `Bearer ${apiKey}`,
},
body: formData,
}
);
if (!response.ok) throw new Error(`HTTP error: ${response.status}`);
const json = await response.json();
if (json.error) throw new Error(`API error: ${JSON.stringify(json.error)}`);
return json;
}
生成期文件保存(必须执行):
两个接口均返回 Base64 编码图片,数据仅存在于当次响应中。获得 Base64 后,必须立即使用 Bash 工具将其解码并保存到本地,以便用户查看结果。
echo "<base64_data>" | base64 -d > <本地路径>.png
完整生成期工作流(含保存步骤):
createImage 或 editImage)json.data[0].b64_json 提取 Base64 数据echo "<b64_json>" | base64 -d > <本地路径>.pngrevised_prompt注意:Base64 数据仅存在于当次响应中,必须及时保存,否则数据丢失。
空间位置描述(生成期 Prompt 增强):
在提示词中加入空间位置词可显著提高构图准确性:
| 位置关键词 | 说明 | 示例 |
|---|---|---|
centered / in the center | 主体居中 | "a red rose, centered, white background" |
in the top-left / bottom-right corner | 角落定位 | "logo in the top-left corner" |
in the foreground / background | 前景/背景层次 | "flowers in the foreground, mountains in the background" |
on the left side / right side | 左右分布 | "person on the left, product on the right" |
filling the entire frame | 占满画面 | "texture filling the entire frame" |
应用内通过 Edge Function 安全调用上游 API,密钥不暴露给前端。
安全合约:
Deno.env.get("INTEGRATIONS_API_KEY") 读取密钥X-Gateway-Authorization: Bearer ${apiKey}429(配额超限)和 402(余额不足)错误体原样透传给前端Edge Function 实现:
image-generations:代理创建图片接口,处理 JSON 请求,解析 body 后强制注入 body.model = "gpt-image-2",采用 SSE 流式响应,每 15 秒发送心跳事件防止中间路由超时断开连接,最终以 type: "result" 事件返回结果image-edits:代理编辑图片接口,必须用 req.formData() 解析请求体(不能用 req.arrayBuffer() 直接透传),检查并注入 model 字段(formData.set("model", "gpt-image-2")),然后采用 SSE 流式响应,每 15 秒发送心跳事件防止中间路由超时断开连接,最终以 type: "result" 事件返回结果完整 Edge Function 代码和前端调用代码详见:
references/image-generations-api.md(创建图片的 Edge Function + 前端代码)references/image-edits-api.md(编辑图片的 Edge Function + 前端代码)| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
model | string | 是 | 固定值:gpt-image-2 |
prompt | string | 是 | 图片生成描述词 |
size | string | 否 | 输出尺寸:1024x1024、1536x1024、1024x1536、2848x1152 |
n | integer | 否 | 生成数量,默认 1 |
| 参数 | 类型 | 必填 | 说明 |
|---|---|---|---|
model | string | 是 | 固定值:gpt-image-2 |
prompt | string | 是 | 图片编辑描述词 |
size | string | 否 | 输出尺寸 |
n | integer | 否 | 输出数量,默认 1 |
image[0] | file | 是 | 主图片文件 |
image[1] | file | 否 | 附加图片文件 |
image[2] | file | 否 | 附加图片文件 |
| 字段路径 | 类型 | 说明 |
|---|---|---|
created | number | 创建时间戳 |
data | array | 生成结果列表 |
data[].b64_json | string | Base64 编码图片内容 |
data[].revised_prompt | string | 模型自动优化后的提示词 |
usage | object | Token 消耗统计(仅编辑接口返回) |
INTEGRATIONS_API_KEY 仅可在 Edge Function 服务端读取,严禁暴露到前端代码或客户端环境变量中。base64 -d 保存为文件。image[0] 必填,image[1]、image[2] 可选),需确保图片格式和大小符合上游要求。model 字段(值为 gpt-image-2)是上游网关路由的必要字段,缺失会导致 403。对于 image-generations,Edge Function 在解析 JSON body 后注入 body.model = "gpt-image-2";对于 image-edits,Edge Function 在解析 FormData 后检查并补充 model 字段。即使前端忘记传递 model,Edge Function 也必须兜底注入,确保上游请求中始终包含该字段。429 — 配额已用尽402 — 余额不足400 — 请求参数错误401 — 认证失败enable_billing: false),但仍需确保 API Key 有效且配额充足。当生成质量不满足要求时,判断是否可以用 Pillow 完成:
可降级到 Pillow 的场景:
Pillow 基础示例(生成期 Agent 可直接运行):
from PIL import Image, ImageDraw, ImageFont
# 在图片上叠加文字水印
def add_text_watermark(image_path: str, text: str, output_path: str) -> None:
img = Image.open(image_path).convert("RGBA")
draw = ImageDraw.Draw(img)
# 使用系统字体,或指定 .ttf 路径
try:
font = ImageFont.truetype("/System/Library/Fonts/PingFang.ttc", size=36)
except OSError:
font = ImageFont.load_default()
# 右下角绘制文字
w, h = img.size
bbox = draw.textbbox((0, 0), text, font=font)
tw, th = bbox[2] - bbox[0], bbox[3] - bbox[1]
draw.text((w - tw - 20, h - th - 20), text, fill=(255, 255, 255, 180), font=font)
img.convert("RGB").save(output_path)
add_text_watermark("input.png", "© 2024 My Brand", "output.png")
决策流程:
error 或 HTTP 429(配额超限)/402(余额不足)→ 向用户说明错误