- name
- org-rendering
- description
- Render books and long-form documents from PDF, EPUB, or HTML into semantically faithful Org. Use when converting a document to Org; repairing flattened paragraphs, footnotes, bibliographies, or indexes; or designing source-agnostic Org output conventions. Use after OCR/extraction, not to choose an OCR engine.
# Source-independent Org rendering
## Purpose
Treat PDF, EPUB, and HTML as different **evidence surfaces**, not as different
final formats. The enduring task is to recognize document structure and render
it into Org without turning structural material into prose.
The pipeline is:
```text
source evidence (HTML DOM / PDF text+geometry / OCR layout)
-> structural intermediate representation
-> common Org renderer
-> validation against the source
```
Do not make a source extractor's linear text stream the final Org document.
In particular, raw PDF text is reliable enough for much prose, but it erases
column order, hanging indents, and page-footnote placement.
Use this skill after text extraction is available. For a scanned source, read
and follow `scanbook` first to obtain OCR/layout evidence; then apply this
skill's rendering and validation rules.
## Decide the output contract first
Before parsing, name the Org forms the work must preserve:
| Source structure | Default Org representation |
| --- | --- |
| part, chapter, section | Org headings at the corresponding hierarchy |
| ordinary paragraphs | one logical paragraph, reflowed only after classification |
| quotation, verse, code-like matter | suitable Org block; preserve line breaks where meaningful |
| ordered/unordered list | Org list, including nesting and item order |
| figure and caption | local asset link plus caption/name when available |
| note reference and note definition | Org footnote reference and definition with a stable identifier |
| bibliography entry | one entry per logical record; retain its source wording |
| index | a lossless, searchable index section with its reading order and subentry indentation |
| uncertain material | lossless raw/example block plus a diagnostic, never silent deletion |
Do not invent metadata, BibTeX fields, footnote links, or heading depth merely
because a presentation pattern resembles one. Faithful structure outranks a
prettier-looking Org file.
The repository's existing EPUB conventions are useful precedent. Read
`epub2org/PATTERNS.org` when its relevant pattern applies, but do not run its
source-specific regular expressions blindly over PDF output.
## Workflow
1. **Inventory the source.** Record page/section extent, native text versus
OCR, assets, links/anchors, page numbering, and likely back matter. Preserve
the source and make the conversion reproducible.
2. **Build a structural intermediate representation.** It may be a checked
artifact or in-memory data, but it must distinguish content roles from their
extracted strings. Keep source coordinates, DOM anchors, page numbers, and
column order when those establish meaning.
3. **Classify before normalizing.** Classify headings, prose, lists, figures,
notes, bibliography, index, and unknown material. Only ordinary prose is a
candidate for line-unfilling/reflowing.
4. **Render each role to the output contract.** Render source-independent Org,
with local asset paths and stable identifiers. Retain a source-page or
anchor trace in diagnostics where identity would otherwise be ambiguous.
5. **Validate the rendered document and its boundaries.** Test the generated
Org mechanically, then compare representative pages/sections with the
source—especially transitions into and out of back matter.
For patterns and examples for the three high-risk structures, read
`references/structure-patterns.md`.
## Source-specific evidence
### EPUB and HTML
Prefer semantic evidence: heading elements, list elements, figure/caption
pairing, IDs, `href` links, and explicit note anchors. A note link is safe to
convert when its source reference and definition resolve to each other. Keep
the original anchor identity in the intermediate representation.
### Born-digital PDF
Start with the embedded text layer, but use layout/bounding-box data whenever
the text stream loses structure. Geometry is evidence, not a cosmetic detail:
- x-position and repeated indentation reveal hanging bibliography entries and
index subentries;
- compare x-position to the **page/column baseline**, not a global coordinate:
bound books can alternate their inner and outer margins on facing pages;
- y-position near a page footer, a note marker, and a separated smaller block
reveal page-footnote definitions;
- columns must be read column-by-column, then top-to-bottom within each column;
- headers, footers, and printed page numbers must be identified separately from
body text.
Do not use a raw-text paragraph joiner across these boundaries.
### Scanned PDF / OCR
OCR is evidence with uncertainty. Follow the `scanbook` skill for acquisition
and correction. Once regions/text are available, use the same intermediate
roles and rendering rules as native PDF. Preserve questionable regions for
review rather than laundering an OCR guess into a confident structural edit.
## Non-negotiable rendering rules
- **Prose-only reflow:** Never unfill text until structural roles have been
detected. Headings, lists, captions, quotations, tables, code, bibliography,
index, and note definitions retain their own line/indent semantics.
- **References before cosmetics:** Resolve footnote reference → definition from
source anchors or corroborating layout evidence. If a match is uncertain,
retain the visible marker and definition, report the candidate, and do not
manufacture an Org `[^id]` link.
- **Reading order is meaning:** For multi-column material, serialize one column
completely before the next, using each page's own column baselines. Never
accept a text extraction order merely because it is linear.
- **Losslessness wins on uncertainty:** Prefer an `#+begin_example` or other
plainly labeled preservation form over a clever but unverifiable conversion.
- **Back matter is first-class:** Bibliography and index need their own
classifiers. They are not failed prose paragraphs.
## Required validation
Check the following before presenting a conversion as ready:
- Outline count, order, and heading depth agree with the source's contents and
visible section boundaries.
- No prose reflow has merged list items, captions, bibliography entries, index
entries, notes, or neighboring headings.
- Every bibliography entry begins as a distinct logical record; hanging-indent
continuation lines remain attached to that record.
- Index columns were not interleaved; alphabetic sequence is plausible within
each source column; subentry indentation survives.
- Every converted footnote reference resolves to exactly one definition; report
unmatched references and definitions instead of silently dropping either.
- Asset links exist and captions remain associated with their figures.
- Emacs can parse the Org document (`org-element-parse-buffer` at minimum), and
representative difficult pages match a visual/source comparison.
Record exceptions and unresolved candidates beside the generated work so the
next run can reproduce the judgment rather than rediscover it.
在 GitHub 查看