Skip to main content

extract-html

Extract structured JSON from HTML using Schematron3B, deterministic table extraction, and Vision-based media OCR.

Informations de source

Dépôt
grahama1970/agent-stack-public
Dernière activité de la source
24 septembre 2026 à 15:51
Langue détectée de SKILL.md
anglais
Étoiles
0
Forks
0

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
19 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
extract-html
description
Extract structured JSON from HTML using Schematron3B, deterministic table extraction, and Vision-based media OCR.
triggers
["extract html","html to json","scrape html","table extraction"]
provides
["extract-html"]
composes
["task-monitor","agentic-evals"]
disciplines
["extraction"]
# Extract-HTML Skill A robust skill for converting HTML documents into strictly valid JSON based on a user-provided JSON Schema. ## Capabilities 1. **Schema Compliance**: Guarantees output conforms to the provided JSON Schema (using Schematron-3B + validation loop). 2. **Deterministic Tables**: Extracts HTML tables using `pandas.read_html` and injects them as context, preventing hallucination of data. 3. **Media Text Extraction**: Identifies images, filters by pixel size, and optionally uses a Vision API (OpenAI-compatible) to extract text/OCR. 4. **Self-Correction**: Validates model output and retries with error feedback if schema validation fails. ## Usage ### Basic Conversion (Local Only) ```bash ./run.sh convert \ --html input.html \ --schema target.schema.json \ --out result.json ``` ### Advanced (With Vision & Remote Fetch) ```bash ./run.sh convert \ --html input.html \ --schema target.schema.json \ --out result.json \ --fetch-remote-media \ --vision-api-base "https://glhf.chat/api/openai/v1" \ --vision-api-key "sk-..." \ --vision-model "gpt-4o-mini" ``` ## Options - `--max-attempts <int>`: Number of self-correction retries. - `--extract-tables / --no-extract-tables`: Toggle deterministic table extraction. - `--extract-media-text`: Enable image processing. - `--min-image-px`, `--max-image-px`: Filter images by size. - `--include-sections`: detailed H1-H6 hierarchy in context. ## Dependencies - Ollama running `schematron-3b` (or compatible model). - Python 3.11+ - See `pyproject.toml` for python deps.
Voir sur GitHub