Skip to main content

org-rendering

Render books and long-form documents from PDF, EPUB, or HTML into semantically faithful Org. Use when converting a document to Org; repairing flattened paragraphs, footnotes, bibliographies, or indexes; or designing source-agnostic Org output conventions. Use after OCR/extraction, not to choose an OCR engine.

معلومات المصدر

المستودع
junghan0611/memex-kb
آخر نشاط في المصدر
٢٠ سبتمبر ٢٠٢٦ في ٠٤:٣٢
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
2 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
org-rendering
description
Render books and long-form documents from PDF, EPUB, or HTML into semantically faithful Org. Use when converting a document to Org; repairing flattened paragraphs, footnotes, bibliographies, or indexes; or designing source-agnostic Org output conventions. Use after OCR/extraction, not to choose an OCR engine.
# Source-independent Org rendering ## Purpose Treat PDF, EPUB, and HTML as different **evidence surfaces**, not as different final formats. The enduring task is to recognize document structure and render it into Org without turning structural material into prose. The pipeline is: ```text source evidence (HTML DOM / PDF text+geometry / OCR layout) -> structural intermediate representation -> common Org renderer -> validation against the source ``` Do not make a source extractor's linear text stream the final Org document. In particular, raw PDF text is reliable enough for much prose, but it erases column order, hanging indents, and page-footnote placement. Use this skill after text extraction is available. For a scanned source, read and follow `scanbook` first to obtain OCR/layout evidence; then apply this skill's rendering and validation rules. ## Decide the output contract first Before parsing, name the Org forms the work must preserve: | Source structure | Default Org representation | | --- | --- | | part, chapter, section | Org headings at the corresponding hierarchy | | ordinary paragraphs | one logical paragraph, reflowed only after classification | | quotation, verse, code-like matter | suitable Org block; preserve line breaks where meaningful | | ordered/unordered list | Org list, including nesting and item order | | figure and caption | local asset link plus caption/name when available | | note reference and note definition | Org footnote reference and definition with a stable identifier | | bibliography entry | one entry per logical record; retain its source wording | | index | a lossless, searchable index section with its reading order and subentry indentation | | uncertain material | lossless raw/example block plus a diagnostic, never silent deletion | Do not invent metadata, BibTeX fields, footnote links, or heading depth merely because a presentation pattern resembles one. Faithful structure outranks a prettier-looking Org file. The repository's existing EPUB conventions are useful precedent. Read `epub2org/PATTERNS.org` when its relevant pattern applies, but do not run its source-specific regular expressions blindly over PDF output. ## Workflow 1. **Inventory the source.** Record page/section extent, native text versus OCR, assets, links/anchors, page numbering, and likely back matter. Preserve the source and make the conversion reproducible. 2. **Build a structural intermediate representation.** It may be a checked artifact or in-memory data, but it must distinguish content roles from their extracted strings. Keep source coordinates, DOM anchors, page numbers, and column order when those establish meaning. 3. **Classify before normalizing.** Classify headings, prose, lists, figures, notes, bibliography, index, and unknown material. Only ordinary prose is a candidate for line-unfilling/reflowing. 4. **Render each role to the output contract.** Render source-independent Org, with local asset paths and stable identifiers. Retain a source-page or anchor trace in diagnostics where identity would otherwise be ambiguous. 5. **Validate the rendered document and its boundaries.** Test the generated Org mechanically, then compare representative pages/sections with the source—especially transitions into and out of back matter. For patterns and examples for the three high-risk structures, read `references/structure-patterns.md`. ## Source-specific evidence ### EPUB and HTML Prefer semantic evidence: heading elements, list elements, figure/caption pairing, IDs, `href` links, and explicit note anchors. A note link is safe to convert when its source reference and definition resolve to each other. Keep the original anchor identity in the intermediate representation. ### Born-digital PDF Start with the embedded text layer, but use layout/bounding-box data whenever the text stream loses structure. Geometry is evidence, not a cosmetic detail: - x-position and repeated indentation reveal hanging bibliography entries and index subentries; - compare x-position to the **page/column baseline**, not a global coordinate: bound books can alternate their inner and outer margins on facing pages; - y-position near a page footer, a note marker, and a separated smaller block reveal page-footnote definitions; - columns must be read column-by-column, then top-to-bottom within each column; - headers, footers, and printed page numbers must be identified separately from body text. Do not use a raw-text paragraph joiner across these boundaries. ### Scanned PDF / OCR OCR is evidence with uncertainty. Follow the `scanbook` skill for acquisition and correction. Once regions/text are available, use the same intermediate roles and rendering rules as native PDF. Preserve questionable regions for review rather than laundering an OCR guess into a confident structural edit. ## Non-negotiable rendering rules - **Prose-only reflow:** Never unfill text until structural roles have been detected. Headings, lists, captions, quotations, tables, code, bibliography, index, and note definitions retain their own line/indent semantics. - **References before cosmetics:** Resolve footnote reference → definition from source anchors or corroborating layout evidence. If a match is uncertain, retain the visible marker and definition, report the candidate, and do not manufacture an Org `[^id]` link. - **Reading order is meaning:** For multi-column material, serialize one column completely before the next, using each page's own column baselines. Never accept a text extraction order merely because it is linear. - **Losslessness wins on uncertainty:** Prefer an `#+begin_example` or other plainly labeled preservation form over a clever but unverifiable conversion. - **Back matter is first-class:** Bibliography and index need their own classifiers. They are not failed prose paragraphs. ## Required validation Check the following before presenting a conversion as ready: - Outline count, order, and heading depth agree with the source's contents and visible section boundaries. - No prose reflow has merged list items, captions, bibliography entries, index entries, notes, or neighboring headings. - Every bibliography entry begins as a distinct logical record; hanging-indent continuation lines remain attached to that record. - Index columns were not interleaved; alphabetic sequence is plausible within each source column; subentry indentation survives. - Every converted footnote reference resolves to exactly one definition; report unmatched references and definitions instead of silently dropping either. - Asset links exist and captions remain associated with their figures. - Emacs can parse the Org document (`org-element-parse-buffer` at minimum), and representative difficult pages match a visual/source comparison. Record exceptions and unresolved candidates beside the generated work so the next run can reproduce the judgment rather than rediscover it.
عرض على GitHub