用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/pluginagentmarketplace/custom-plugin-prompt-engineering --skill multi-modal命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
基于 SOC 职业分类
正在显示 SKILL.md
| name | multi-modal |
| description | Multi-modal prompting with vision, audio, and document understanding |
| sasmp_version | 1.3.0 |
| bonded_agent | 07-advanced-techniques-agent |
| bond_type | PRIMARY_BOND |
Bonded to: advanced-techniques-agent
Skill("custom-plugin-prompt-engineering:multi-modal")
parameters:
modality:
type: enum
values: [vision, document, audio, video]
required: true
task_type:
type: enum
values: [analysis, extraction, generation, qa]
default: analysis
detail_level:
type: enum
values: [low, medium, high]
default: medium
Analyze this image and provide:
1. Main subjects and objects
2. Actions or activities
3. Setting and context
4. Notable details
5. Overall interpretation
Be specific and descriptive.
Look at the image carefully.
Question: {question}
Provide a detailed answer based only on what you can see in the image.
Analyze this chart/graph:
1. Type of visualization
2. Axes and labels
3. Key data points
4. Trends or patterns
5. Main insights
6. Limitations or caveats
Extract the following from this document:
- Title and headers
- Key information: {fields}
- Tables (if any)
- Important dates/numbers
Output as structured JSON.
extraction_schema:
document_type: "invoice|form|contract"
fields:
- name: vendor
type: string
- name: date
type: date
- name: total
type: currency
- name: line_items
type: array
Transcribe and enhance:
1. Accurate transcription
2. Speaker identification
3. Timestamps for key points
4. Summary of main topics
5. Action items (if applicable)
best_practices:
image_prompts:
- Be specific about what to look for
- Request structured output
- Ask for confidence levels
document_prompts:
- Define extraction schema
- Handle multi-page documents
- Validate extracted data
audio_prompts:
- Specify language if known
- Request speaker diarization
- Ask for timestamps
| Issue | Cause | Solution |
|---|---|---|
| Hallucinated details | Over-interpretation | Ask for visible-only info |
| Missed text in images | Low resolution | Request higher detail |
| Wrong document parsing | Complex layout | Break into sections |
| Inaccurate transcription | Audio quality | Acknowledge limitations |
See: GPT-4V documentation, Claude Vision Guide