Routes images to Mimo/OpenAI vision API for text-only models. Use this skill when the user sends an image, asks "what's in this picture", says "analyze this screenshot", requests "OCR" or "extract text", reviews UI designs, reads charts, or uses Chinese triggers like 识别图片/看图. Self-healing setup on first run. Does NOT handle video, audio, or PDFs.
2026-06-02