Skip to main content

vision

You **MUST** use the vision skill when when the your model is text-only (e.g. glm-5.2, deepseek-v4-pro) AND: (1) the user's message contains images; (2) OR the user's message contains URLs or paths to images; (3) OR the user asks to visually verify/check something ("visually verify", "screenshot shows", "centered/visible/hidden", "looks right", "matches the design"); (5) OR a tool result contains an image attachment the current model cannot see (attachments[].mime = "image/png", url = "data:image/png;base64,..."); (4) OR you think it is necessary to READ any visual contents; Triggers on screenshots from chrome-devtools_take_screenshot, playwright_browser_take_screenshot, cua-driver_get_window_state/zoom/take_screenshot and cannot see images itself. Extracts the visual intent from context, designs a prompt-local JSON response template for that specific task, delegates, and parses the returned JSON.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
WeZZard/opencode-vision
آخر نشاط في المصدر
٢٣ يوليو ٢٠٢٦ في ٢٢:٢٢
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٢٦
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.