docling
Convert PDFs, Office files, HTML, Markdown, images, and other documents into structured outputs with Docling; includes install, conversion, batch, chunking, advanced PDF guidance, and DOCX/OOXML interoperability fallbacks.
ソース情報
- リポジトリ
- Rycen7822/docling-skill
- ソースの最終更新活動
- 2026年6月4日 00:56
- 検出された SKILL.md の言語
- 英語
- スター
- 0
- フォーク
- 0
インストール方法
デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。
ソースファイルを確認
インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。
SKILL.md を表示中
SKILL.md
ソースの指示 · 読み取り専用プレビュー- name
- docling
- description
- Convert PDFs, Office files, HTML, Markdown, images, and other documents into structured outputs with Docling; includes install, conversion, batch, chunking, advanced PDF guidance, and DOCX/OOXML interoperability fallbacks.
- version
- 1.1.1
- author
- Imported local docling pack, adapted by Hermes Agent
- license
- Unknown
- metadata
- {"hermes":{"tags":["docling","documents","pdf","conversion","ocr","chunking"],"category":"user-imported","related_skills":["obsidian","clip"]}}
## Hermes compatibility note
- This local pack was installed into Hermes as `docling`.
- Task documents were mirrored into `references/` so they are directly accessible via `skill_view` linked files.
- Reusable helpers remain under `scripts/`.
- The original imported folder layout (`tasks/`, `examples/`, `scripts/`) was preserved.
# Docling Skill (Install • Convert • Export • Chunk • Batch • Advanced PDF)
Use this skill when you need to convert PDFs, images, Office files, HTML, Markdown, CSV, or XML into structured outputs with **Docling**.
This pack is based on the official Docling documentation, especially the installation, usage, CLI reference, and example pages.
## Core operating rules
1. **Prefer `DocumentConverter` first.** The official docs recommend the high-level `DocumentConverter` API for most work. Drop to specific backends only when you have a clear reason.
2. **Single-document workflow:** convert one source -> export the formats you need -> inspect the output -> adjust pipeline options only if quality requires it.
3. **Batch workflow:** use `convert_all(..., raises_on_error=False)` so you can inspect partial failures instead of stopping on the first bad file.
4. **Chunking workflow:** convert to a Docling document first, then run `HybridChunker`; for embeddings, prefer the text returned by `chunker.contextualize(...)`.
5. **Remote services are opt-in.** If a pipeline talks to an external service, you must explicitly enable it.
6. **Offline use is first-class.** You can pre-download models and point Docling to a local artifacts directory.
## Quick start
```bash
# Install the base package
pip install docling
# Optional: for hybrid chunking
pip install transformers
# Convert one file with the packaged helper
python scripts/docling_convert.py input.pdf --output-dir out --to md json text
# DOC/DOCX fallback helpers (LibreOffice / pandoc / OOXML ZIP workflows)
python scripts/docling_docx_tools.py convert old.doc docx --output-dir converted
python scripts/docling_docx_tools.py extract revised.docx revised.md --track-changes all
python scripts/docling_docx_tools.py unpack report.docx unpacked_report
# Batch convert a folder
python scripts/docling_batch.py docs/*.pdf --output-dir out_batch --to md json
# Chunk a document into JSONL
python scripts/docling_chunk.py input.pdf --output chunks.jsonl --contextualize
```
## What this pack covers
- Installation and environment setup
- Single-document conversion
- CLI usage and key flags
- Advanced PDF pipeline configuration
- Batch conversion and export
- Hybrid chunking for RAG pipelines
- Common troubleshooting notes
- DOCX / OOXML interoperability fallbacks for Word-native edge cases
## DOCX / OOXML interoperability and boundaries
Compared with Anthropic's `docx` skill, the biggest gap was not generic format conversion but Word-native workflows around legacy `.doc`, tracked changes, raw OOXML inspection, and page-preview rendering.
This `docling` skill now includes a lightweight supplement for those cases:
- convert legacy `.doc` files to `.docx` or PDF with LibreOffice
- on WSL, if Linux `soffice` is unavailable but Windows Microsoft Word is installed, the helper can use Word automation as a conversion fallback
- extract `.docx` content with tracked changes visible via `pandoc --track-changes=...`
- unpack / repack `.docx/.pptx/.xlsx` archives for raw XML inspection
- render Word documents to page-preview JPEGs via PDF + `pdftoppm`
Still intentionally out of scope:
- full-fidelity Word document authoring with `docx-js`
- deep OOXML schema validation / redline surgery / comment editing
- replacing a dedicated Word production workflow when the user needs polished `.docx` deliverables rather than extraction
Rule of thumb:
- if the goal is extraction / structured export / chunking, start with Docling
- if the goal is Word-native inspection or preprocessing before Docling, use the DOCX interoperability helper next
## Task index
Start with the smallest task that answers your need. In Hermes, use the mirrored files under `references/` for direct access via `skill_view`:
- `references/install.md` - installation, optional extras, offline models
- `references/convert.md` - one-document conversion and export
- `references/cli.md` - official CLI patterns and important flags
- `references/advanced_pdf.md` - OCR, tables, remote services, model artifacts, streams
- `references/batch.md` - `convert_all`, partial failures, batch export
- `references/chunk.md` - `HybridChunker`, tokenizer choice, contextualized text
- `references/troubleshooting.md` - common failure modes and operational notes
- `references/docx-interoperability.md` - legacy `.doc`, tracked changes, OOXML unpack/repack, and preview-image fallbacks
The original imported `tasks/` directory is still preserved on disk.
## Package map
- `SKILL.md` - overview and routing
- `manifest.txt` - machine-readable file list
- `requirements.txt` - suggested Python dependencies
- `references/*.md` - Hermes-accessible task docs mirrored from the imported `tasks/` directory
- `references/docx-interoperability.md` - docx-specific supplement inspired by comparison with Anthropic's `docx` skill
Original imported task docs preserved on disk:
- `tasks/install.md`
- `tasks/convert.md`
- `tasks/cli.md`
- `tasks/advanced_pdf.md`
- `tasks/batch.md`
- `tasks/chunk.md`
- `tasks/troubleshooting.md`
Scripts:
- `scripts/docling_common.py` - shared helpers for import checks, converter setup, and export
- `scripts/docling_convert.py` - convert one source and export multiple formats
- `scripts/docling_batch.py` - batch conversion with summary JSON
- `scripts/docling_chunk.py` - convert + chunk to JSONL
- `scripts/docling_docx_tools.py` - LibreOffice / pandoc / OOXML ZIP helpers for Word-native edge cases
Examples:
- `examples/basic_convert.py`
- `examples/custom_pdf_pipeline.py`
- `examples/multi_format_convert.py`
- `examples/hybrid_chunking.py`
- `examples/offline_models.py`
- `examples/stream_convert.py`
## Practical defaults
- Start with `DocumentConverter()` unless you need explicit PDF tuning.
- For PDFs with tables, enable table structure and keep `TableFormerMode.ACCURATE` unless speed matters more than quality.
- For OCR-heavy scans, explicitly turn OCR on and select an OCR backend only after you have installed its dependencies.
- For large corpora, export at least Markdown and JSON.
- For RAG, keep both raw chunk text and contextualized chunk text.
## Deliverable expectations
- Clear installation path for CPU-only and standard environments
- Explicit handling for local/offline artifacts
- Multiple export formats (`md`, `json`, `html`, `text`, `doctags`, optional `yaml`)
- Batch summaries that distinguish success, partial success, and failure
- No hidden reliance on remote services
GitHubで見る