Skip to main content

extract-html

Extract structured JSON from HTML using Schematron3B, deterministic table extraction, and Vision-based media OCR.

来源信息

仓库
grahama1970/agent-stack-public
最近来源活动
2026年9月24日 15:51
检测到的 SKILL.md 语言
英语
星标
0
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
19 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
extract-html
description
Extract structured JSON from HTML using Schematron3B, deterministic table extraction, and Vision-based media OCR.
triggers
["extract html","html to json","scrape html","table extraction"]
provides
["extract-html"]
composes
["task-monitor","agentic-evals"]
disciplines
["extraction"]
# Extract-HTML Skill A robust skill for converting HTML documents into strictly valid JSON based on a user-provided JSON Schema. ## Capabilities 1. **Schema Compliance**: Guarantees output conforms to the provided JSON Schema (using Schematron-3B + validation loop). 2. **Deterministic Tables**: Extracts HTML tables using `pandas.read_html` and injects them as context, preventing hallucination of data. 3. **Media Text Extraction**: Identifies images, filters by pixel size, and optionally uses a Vision API (OpenAI-compatible) to extract text/OCR. 4. **Self-Correction**: Validates model output and retries with error feedback if schema validation fails. ## Usage ### Basic Conversion (Local Only) ```bash ./run.sh convert \ --html input.html \ --schema target.schema.json \ --out result.json ``` ### Advanced (With Vision & Remote Fetch) ```bash ./run.sh convert \ --html input.html \ --schema target.schema.json \ --out result.json \ --fetch-remote-media \ --vision-api-base "https://glhf.chat/api/openai/v1" \ --vision-api-key "sk-..." \ --vision-model "gpt-4o-mini" ``` ## Options - `--max-attempts <int>`: Number of self-correction retries. - `--extract-tables / --no-extract-tables`: Toggle deterministic table extraction. - `--extract-media-text`: Enable image processing. - `--min-image-px`, `--max-image-px`: Filter images by size. - `--include-sections`: detailed H1-H6 hierarchy in context. ## Dependencies - Ollama running `schematron-3b` (or compatible model). - Python 3.11+ - See `pyproject.toml` for python deps.
在 GitHub 查看