Unified PDF skill — generate, reformat, fill, and read PDFs. Covers: text-to-PDF (reports, resumes, proposals, 可视化报告), LaTeX thesis, Markdown→PDF conversion, PDF form filling, and PDF reading/extraction/OCR. Trigger on any task with PDF as primary input or output. Not for DOCX or PPT.
소스 정보
- 저장소
- MiniMax-AI/minimax-code
- 최근 소스 활동
- 2026년 9월 18일 11:25
- 감지된 SKILL.md 언어
- 영어
- 스타
- 589
- 포크
- 67
설치 방법
기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.
소스 파일 검토
설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.
파일 탐색기
84 개 파일SKILL.md 표시 중
SKILL.md
소스 지침 · 읽기 전용 미리보기- name
- description
- Unified PDF skill — generate, reformat, fill, and read PDFs. Covers: text-to-PDF (reports, resumes, proposals, 可视化报告), LaTeX thesis, Markdown→PDF conversion, PDF form filling, and PDF reading/extraction/OCR. Trigger on any task with PDF as primary input or output. Not for DOCX or PPT.
- descriptions
- {"zh-Hans":"生成、重排、填写和读取 PDF,支持报告、简历、提案、Markdown/LaTeX 转 PDF、表单和 OCR。"}
- metadata
- {"version":"3.0","category":"document-pdf"}
# pdf
Unified PDF skill. The model chooses the route based on user intent; this SKILL.md is an index. Each
route has its own guide in `docs/`. Read the guide before authoring or running anything.
## Operational rules — read before doing anything
> **1. Match user query against [`docs/pitfalls-index.md`](docs/pitfalls-index.md) FIRST.** It
> contains 10 production-ready **canonical query templates** (P1–P10), each with a
> `Match signatures` block (sample queries) and a complete executable prompt that already encodes
> every known pitfall, verification gate, and fall-back path. Workflow:
>
> 1. Scan the Quick lookup table — match user's query keywords to a row.
> 2. **Copy the matching canonical query verbatim**, substitute the `Slots` (e.g. `{PDF_PATH}`,
> `{OUTPUT_PATH}`) with the user's actual values, and execute step-by-step.
> 3. Multiple partial matches → fuse: take the strictest verification from each, never relax a
> constraint.
> 4. No match → fall back to the Routes / route guides below.
>
> Do NOT skip verification steps in the canonical queries — they exist because past evaluation runs
> shipped wrong outputs without them.
> **2. Locate before bulk-extracting any non-trivial PDF.** Three independent thresholds, all
> enforced together:
>
> | Threshold | Rule |
> | ---------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
> | **>20 pages** AND user wants a specific datum (not the whole document) | locate-first is **mandatory** — do not run pdfplumber over every page; build a heading index first (pypdf outline → printed TOC → keyword grep) |
> | **2 blind grep passes** without landing on the target | stop and build the heading index, regardless of page count — a 3rd / 4th / 5th keyword search is the most common time sink |
> | **>200 pages** | always build a heading index up-front, even before the first grep — at this size 6-8 blind greps balloon into 30+ shell calls |
>
> See [`docs/read-guide.md`](docs/read-guide.md) §3 for the actual outline / TOC / grep recipes.
> **3. Chart pages and complex financial tables — vision, ONE PAGE PER CALL, mandatory.** Any page
> with a chart / diagram / info-graphic whose values matter, **or** any page with a complex
> financial / regulatory table (balance sheet, income statement, cash-flow, debt schedule —
> multi-level headers, merged cells, footnoted sub-totals), MUST be read visually, **one page per
> call, never a range** — pdfplumber returns scrambled fragments on these layouts even when the PDF
> is text-native, and packing neighbour pages into one image makes the model mis-attribute values to
> the wrong page. **Pick the visual path by model capability:**
>
> - **Vision-capable model (e.g. M3 — can natively accept image input):** rasterise that single page
> to PNG (`pdftoppm` / `pypdfium2`, or `scripts/render/page_rasterize.py`) and read the PNG
> directly with the **Read tool**. Faster, offline, no upstream LLM call — this is the default
> now.
> - **Text-only model (e.g. M2.7 — cannot see images):** fall back to `read_pdf_vision.py` invoked
> with `--pages N` (a single page), which ships the page to native Matrix vision.
> **4. Verify HTML→PDF page size and chart presence after every render.** Always pass `--format A4`
> or `--format Letter` explicitly to `make.sh render` — Chromium overrides CSS `@page { size }` when
> the CLI flag is missing. After render:
>
> ```bash
> pdfinfo out.pdf | grep "Page size" # must match user intent
> pdfimages -list out.pdf | tail -n +3 | wc -l # ≥ 1 per chart / logo
> ```
>
> `pdftotext` cannot see images and will silently pass a chart-less deck. Both checks are mandatory.
> **5. Don't suppress stderr.** `2>/dev/null` is **never** the right choice in this skill
> (`make.sh render`/`reformat`/`fill`, `read_pdf_vision.py`, `pdfinfo`, `pdftotext`, `pdfimages`,
> `qpdf`). On failure you lose the only signal that explains why and have to rerun blind. If output
> is too noisy, redirect to a log file and grep on demand:
>
> ```bash
> python3 -m scripts.read_pdf_vision --input report.pdf --pages 5 \
> 2>/tmp/vision.log
> # If the result looks wrong, only then:
> # grep -in "error\|trace\|fail\|502\|413" /tmp/vision.log | head -20
> ```
> **6. Always serialise JSON with `ensure_ascii=False`.** When this skill writes a JSON config /
> manifest that a downstream step parses (chart data, content manifests, form values), use
> `json.dumps`, never hand-concatenate strings. CJK / smart quotes / em-dashes in data are the most
> common reason a "looks fine" JSON file fails to `json.load()`:
>
> ```python
> Path("content.json").write_text(
> json.dumps(payload, ensure_ascii=False, indent=2),
> encoding="utf-8",
> )
> ```
> **7. AcroForm fill — copy the one canonical pypdf snippet.** In pypdf ≥ 4 the only working pattern
> is `PdfWriter(clone_from=src)` + `update_page_form_field_values(...)`
>
> - `set_need_appearances_writer(True)`. `clone_reader_document_root`, direct `/Annots` patching,
> and `append_pages_from_reader` all _silently_ produce a PDF with no values written — there is no
> error to debug. See [`docs/forms-guide.md`](docs/forms-guide.md) §B.
> **8. Header/footer discipline for generated PDFs.** For any formal or multi-page PDF (contracts,
> reports, proposals, forms, manuals, translated documents), decide the header/footer strategy
> before rendering: preserve source headers/footers when present; otherwise add a conservative
> running header/footer or explicitly justify why none is appropriate (e.g. cover-only one-pager).
> Reserve print-space so running elements do not collide with body content, tables, signatures, or
> charts. Verification must include a visual check of at least one body page and the
> final/signature/table-heavy page, not only `pdftotext`. Implementation details live in
> [`docs/html-pdf-spec.md`](docs/html-pdf-spec.md) §3.3.
> **9. DOCX→PDF is a DOCX-native render/export task, not an HTML task.** When the user asks to
> convert a Word/DOCX file to PDF while preserving the Word document, route to `docx` / the DOCX
> renderer first (e.g. `scripts/docx_to_pdf.py` or LibreOffice/soffice export). DOCX already has
> native page geometry, styles, sections, headers/footers, fields, numbering, and table layout;
> converting DOCX → Markdown/HTML → PDF just to make a PDF is a fidelity bug. Verify the native PDF
> with `pdfinfo`, `pdftotext`, and visual spot checks. Use HTML→PDF only for explicit
> redesign/recomposition, when the native render is visibly unacceptable, or when the requested
> deliverable is a newly authored web/print design. If HTML is used, say it is a recomposition
> route, not the default DOCX→PDF conversion path.
> **10. Every PDF output needs clickable TOC/index navigation, regardless of route.** This is a
> global delivery contract for CREATE, REFORMAT, LATEX_THESIS, FILL/overlay, MUTATE/merge/split,
> DOCX-native export handoff, and any read→write chain. Any multi-page PDF produced, transformed,
> merged, or substantially reformatted by this skill must include a visible TOC / index that maps
> major sections to their destination pages and is clickable in the final PDF. For HTML→PDF,
> implement TOC rows as internal anchors (`<a href="#section-id">`) and give every target section a
> stable, unique `id`. For LaTeX, `hyperref` is mandatory and `\tableofcontents` plus any manual
> `\addcontentsline` targets must resolve to live links. For markdown/text reformatting, generate or
> preserve a TOC before rendering; do not ship a flat prose PDF without navigable section links. For
> filled forms, official one-page forms may omit TOC, but multi-page filled packets must preserve
> existing bookmarks/links or add an index/outline without altering the form semantics. For
> merged/split/watermarked/pypdf/reportlab-built outputs, preserve existing links where possible and
> add/update PDF outline/bookmarks plus `/Link` annotations when the visible TOC cannot be generated
> by the renderer alone. Exceptions are only single-page forms/posters/certificates or
> source-faithful official forms where adding pages/visual TOC would invalidate the document; the
> delivery note must explicitly say why and what navigation was preserved instead. Verification must
> include checking link or outline annotations and spot-clicking several TOC/index entries to
> confirm they land on the intended pages.
## Routes — pick by user intent, then read the guide
### WRITE a PDF
| Intent | Guide | Entry |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ |
| **CREATE** — author a polished PDF from scratch (cover, charts, KPI grids, branded report) | [`docs/create-guide.md`](docs/create-guide.md) + [`docs/design-guide.md`](docs/design-guide.md) + [`templates/INDEX.md`](templates/INDEX.md) | `bash scripts/make.sh render --in page.html --out out.pdf` |
| **REFORMAT** — restyle markdown / text / pdf as a clean PDF (no charts, no design recomposition) | [`docs/reformat-guide.md`](docs/reformat-guide.md) | `bash scripts/make.sh reformat --input src.md --out out.pdf` |
| **FILL** — write values into a PDF form (AcroForm or visual overlay) | [`docs/forms-guide.md`](docs/forms-guide.md) | `bash scripts/make.sh fill probe form.pdf` |
| **LATEX_THESIS** — typeset an academic thesis/dissertation with LaTeX (Chinese university template, GB/T 7714 bibliography, cover merge) | [`docs/latex-academic-thesis-guide.md`](docs/latex-academic-thesis-guide.md) | tectonic / xelatex + qpdf merge |
| **LATEX_TECHNICAL_BOOK** — typeset a Chinese technical book / engineering monograph / source-code reading book with LaTeX (B5, O'Reilly-like cover, code, diagrams) | [`docs/latex-technical-book-guide.md`](docs/latex-technical-book-guide.md) + [`templates/latex-technical-book/README.md`](templates/latex-technical-book/README.md) | latexmk -xelatex |
| **MUTATE** — merge / split / rotate / crop / watermark / encrypt / annotate / sign / replace text | [`docs/advanced-reference.md`](docs/advanced-reference.md) | qpdf / pypdf / reportlab cookbook (no in-skill route) |
> Mechanical contract for any HTML→PDF authoring (page geometry, page-break rules, Chart.js settle,
> CJK cascade, color fidelity) lives in [`docs/html-pdf-spec.md`](docs/html-pdf-spec.md). Read it
> before writing any new HTML.
### READ a PDF
| Intent | Guide | Entry |
GitHub에서 보기이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기