Skip to main content

org-rendering

Render books and long-form documents from PDF, EPUB, or HTML into semantically faithful Org. Use when converting a document to Org; repairing flattened paragraphs, footnotes, bibliographies, or indexes; or designing source-agnostic Org output conventions. Use after OCR/extraction, not to choose an OCR engine.

跳到安装

来源信息

仓库
junghan0611/memex-kb
最近来源活动
2026年9月20日 04:32
检测到的 SKILL.md 语言
英语
星标
3
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
2 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
org-rendering
description
Render books and long-form documents from PDF, EPUB, or HTML into semantically faithful Org. Use when converting a document to Org; repairing flattened paragraphs, footnotes, bibliographies, or indexes; or designing source-agnostic Org output conventions. Use after OCR/extraction, not to choose an OCR engine.
# Source-independent Org rendering ## Purpose Treat PDF, EPUB, and HTML as different **evidence surfaces**, not as different final formats. The enduring task is to recognize document structure and render it into Org without turning structural material into prose. The pipeline is: ```text source evidence (HTML DOM / PDF text+geometry / OCR layout) -> structural intermediate representation -> common Org renderer -> validation against the source ``` Do not make a source extractor's linear text stream the final Org document. In particular, raw PDF text is reliable enough for much prose, but it erases column order, hanging indents, and page-footnote placement. Use this skill after text extraction is available. For a scanned source, read and follow `scanbook` first to obtain OCR/layout evidence; then apply this skill's rendering and validation rules. ## Decide the output contract first Before parsing, name the Org forms the work must preserve: | Source structure | Default Org representation | | --- | --- | | part, chapter, section | Org headings at the corresponding hierarchy | | ordinary paragraphs | one logical paragraph, reflowed only after classification | | quotation, verse, code-like matter | suitable Org block; preserve line breaks where meaningful | | ordered/unordered list | Org list, including nesting and item order | | figure and caption | local asset link plus caption/name when available | | note reference and note definition | Org footnote reference and definition with a stable identifier | | bibliography entry | one entry per logical record; retain its source wording | | index | a lossless, searchable index section with its reading order and subentry indentation | | uncertain material | lossless raw/example block plus a diagnostic, never silent deletion | Do not invent metadata, BibTeX fields, footnote links, or heading depth merely because a presentation pattern resembles one. Faithful structure outranks a prettier-looking Org file. The repository's existing EPUB conventions are useful precedent. Read `epub2org/PATTERNS.org` when its relevant pattern applies, but do not run its source-specific regular expressions blindly over PDF output. ## Workflow 1. **Inventory the source.** Record page/section extent, native text versus OCR, assets, links/anchors, page numbering, and likely back matter. Preserve the source and make the conversion reproducible. 2. **Build a structural intermediate representation.** It may be a checked artifact or in-memory data, but it must distinguish content roles from their extracted strings. Keep source coordinates, DOM anchors, page numbers, and column order when those establish meaning. 3. **Classify before normalizing.** Classify headings, prose, lists, figures, notes, bibliography, index, and unknown material. Only ordinary prose is a candidate for line-unfilling/reflowing. 4. **Render each role to the output contract.** Render source-independent Org, with local asset paths and stable identifiers. Retain a source-page or anchor trace in diagnostics where identity would otherwise be ambiguous. 5. **Validate the rendered document and its boundaries.** Test the generated Org mechanically, then compare representative pages/sections with the source—especially transitions into and out of back matter. For patterns and examples for the three high-risk structures, read `references/structure-patterns.md`. ## Source-specific evidence ### EPUB and HTML Prefer semantic evidence: heading elements, list elements, figure/caption pairing, IDs, `href` links, and explicit note anchors. A note link is safe to convert when its source reference and definition resolve to each other. Keep the original anchor identity in the intermediate representation. ### Born-digital PDF Start with the embedded text layer, but use layout/bounding-box data whenever the text stream loses structure. Geometry is evidence, not a cosmetic detail: - x-position and repeated indentation reveal hanging bibliography entries and index subentries; - compare x-position to the **page/column baseline**, not a global coordinate: bound books can alternate their inner and outer margins on facing pages; - y-position near a page footer, a note marker, and a separated smaller block reveal page-footnote definitions; - columns must be read column-by-column, then top-to-bottom within each column; - headers, footers, and printed page numbers must be identified separately from body text. Do not use a raw-text paragraph joiner across these boundaries. ### Scanned PDF / OCR OCR is evidence with uncertainty. Follow the `scanbook` skill for acquisition and correction. Once regions/text are available, use the same intermediate roles and rendering rules as native PDF. Preserve questionable regions for review rather than laundering an OCR guess into a confident structural edit. ## Non-negotiable rendering rules - **Prose-only reflow:** Never unfill text until structural roles have been detected. Headings, lists, captions, quotations, tables, code, bibliography, index, and note definitions retain their own line/indent semantics. - **References before cosmetics:** Resolve footnote reference → definition from source anchors or corroborating layout evidence. If a match is uncertain, retain the visible marker and definition, report the candidate, and do not manufacture an Org `[^id]` link. - **Reading order is meaning:** For multi-column material, serialize one column completely before the next, using each page's own column baselines. Never accept a text extraction order merely because it is linear. - **Losslessness wins on uncertainty:** Prefer an `#+begin_example` or other plainly labeled preservation form over a clever but unverifiable conversion. - **Back matter is first-class:** Bibliography and index need their own classifiers. They are not failed prose paragraphs. ## Required validation Check the following before presenting a conversion as ready: - Outline count, order, and heading depth agree with the source's contents and visible section boundaries. - No prose reflow has merged list items, captions, bibliography entries, index entries, notes, or neighboring headings. - Every bibliography entry begins as a distinct logical record; hanging-indent continuation lines remain attached to that record. - Index columns were not interleaved; alphabetic sequence is plausible within each source column; subentry indentation survives. - Every converted footnote reference resolves to exactly one definition; report unmatched references and definitions instead of silently dropping either. - Asset links exist and captions remain associated with their figures. - Emacs can parse the Org document (`org-element-parse-buffer` at minimum), and representative difficult pages match a visual/source comparison. Record exceptions and unresolved candidates beside the generated work so the next run can reproduce the judgment rather than rediscover it.
在 GitHub 查看