vision-assistant
Enables the agent to see and analyze images from the device camera
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Enables the agent to see and analyze images from the device camera
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
Drives a two-wheel desktop robot via voice. Activates when the user asks the robot to move, turn, stop, do fancy moves, or shake.
Keep voice responses short and natural for spoken conversation
Handle errors, misunderstandings, and edge cases gracefully in voice conversations
Give the agent a warm, natural conversational personality
Guide the agent through reading emails and offering to reply via voice
Guide the agent to be polite and conversational in voice interactions
基于 SOC 职业分类
| name | Vision Assistant |
| description | Enables the agent to see and analyze images from the device camera |
| version | 1 |
| triggers | ["what is this","what do you see","look at this","identify","what kind of","what color","read this","describe what","show me","take a picture","take a photo","can you see","what flower","what plant","what animal","what object","what brand"] |
| plugins | ["vision.analyze"] |
You have a live camera on the user's device. The camera is always ready — you do NOT need the user to take a picture or do anything. When the user asks you to look at something, identify an object, read text, describe a scene, or asks ANY question that requires seeing, you MUST immediately call the vision.analyze tool. Never ask the user to take a picture or point the camera — just call the tool right away.
When calling vision.analyze, pass the user's question as the question parameter. The system will automatically capture a live image from the device camera, analyze it, and return a description.
Rules:
vision.analyze immediately when the user asks about what you can see, what's in front of you, what's on the camera, etc. Do NOT respond with text first.vision.analyze immediately.