Skip to main content

multimodal-page-grounding

Use this skill when the user wants tasks where the page layout, screenshot, specific visual position, layout adjacency in tables, or physical drag-and-drop motions actually matter. Trigger it for requests like “make it rely on what’s visible on the page,” “extract the data from this complex table layout,” “buttons and layout should dictate the choice,” or “force the agent to navigate based on spatial coordinates.” It is the right skill for browser tasks where perception must combine HTML-like structure with visual grounding to execute continuous mouse events or parse complex semi-structured layouts (like grids or key-value structures).

Aller à l'installation

Informations de source

Dépôt
Dingxingdi/paper_fast_search_backup
Dernière activité de la source
10 avril 2026 à 01:27
Langue détectée de SKILL.md
anglais
Étoiles
0
Forks
0

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.