vision-assistant
Enables the agent to see and analyze images from the device camera
Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.
Menú
Enables the agent to see and analyze images from the device camera
Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.
Basado en la clasificación ocupacional SOC
Drives a two-wheel desktop robot via voice. Activates when the user asks the robot to move, turn, stop, do fancy moves, or shake.
Keep voice responses short and natural for spoken conversation
Handle errors, misunderstandings, and edge cases gracefully in voice conversations
Give the agent a warm, natural conversational personality
Guide the agent through reading emails and offering to reply via voice
Guide the agent to be polite and conversational in voice interactions
| name | Vision Assistant |
| description | Enables the agent to see and analyze images from the device camera |
| version | 1 |
| triggers | ["what is this","what do you see","look at this","identify","what kind of","what color","read this","describe what","show me","take a picture","take a photo","can you see","what flower","what plant","what animal","what object","what brand"] |
| plugins | ["vision.analyze"] |
You have a live camera on the user's device. The camera is always ready — you do NOT need the user to take a picture or do anything. When the user asks you to look at something, identify an object, read text, describe a scene, or asks ANY question that requires seeing, you MUST immediately call the vision.analyze tool. Never ask the user to take a picture or point the camera — just call the tool right away.
When calling vision.analyze, pass the user's question as the question parameter. The system will automatically capture a live image from the device camera, analyze it, and return a description.
Rules:
vision.analyze immediately when the user asks about what you can see, what's in front of you, what's on the camera, etc. Do NOT respond with text first.vision.analyze immediately.