Skip to main content
Run any Skill in Manus
with one click

multimodal-structured-extraction

Stars0
Forks0
UpdatedMay 23, 2026 at 12:34

Use this skill when extracting structured data from images (receipts, menus, statements, IDs, business cards, forms, screenshots) using a vision-capable LLM. Covers schema-first design with zod, validation gates and retry, locale handling (currency / dates / romanization), confidence signalling, and the anti-patterns that cause silent garbage extraction. Triggers on "extract from image", "OCR with structure", "vision model JSON", "Gemini multimodal", "GPT-4 vision extraction", "parse receipt", "parse business card", "image to schema", "structured output from image".

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly