| name | kordoc-korean-document-parser |
| description | Parse HWP, HWPX, and PDF Korean documents to Markdown using kordoc — supports CLI, programmatic API, and MCP server integration. |
| triggers | ["parse hwp file to markdown","convert korean document to text","extract text from hwpx","parse pdf to markdown with kordoc","compare two hwp documents","extract form fields from korean document","set up kordoc mcp server","convert hwp hwpx pdf to structured data"] |
kordoc Korean Document Parser
Skill by ara.so — Daily 2026 Skills collection.
kordoc is a TypeScript library and CLI for parsing Korean government documents (HWP 5.x, HWPX, PDF) into Markdown and structured IRBlock[] data. It handles proprietary HWP binary formats, table extraction, form field recognition, document diffing, and reverse Markdown→HWPX generation.
Installation
npm install kordoc
npm install pdfjs-dist
npx kordoc document.hwpx
Core API
Auto-detect and Parse Any Document
import { parse } from "kordoc"
import { readFileSync } from "fs"
const buffer = readFileSync("document.hwpx")
const result = await parse(buffer.buffer)
if (result.success) {
console.log(result.markdown)
console.log(result.blocks)
console.log(result.metadata)
console.log(result.outline)
console.log(result.warnings)
} else {
console.error(result.error)
console.error(result.code)
}
Format-Specific Parsers
import { parseHwpx, parseHwp, parsePdf, detectFormat } from "kordoc"
const fmt = detectFormat(buffer.buffer)
const hwpxResult = await parseHwpx(buffer.buffer)
const hwpResult = await parseHwp(buffer.buffer)
const pdfResult = await parsePdf(buffer.buffer)
Parse Options
import { parse, ParseOptions } from "kordoc"
const result = await parse(buffer.buffer, {
pages: "1-3",
ocr: async (pageImage, pageNumber, mimeType) => {
return await myOcrService.recognize(pageImage)
}
})
Working with IRBlocks
import type { IRBlock, IRBlockType, IRTable, IRCell } from "kordoc"
for (const block of result.blocks) {
if (block.type === "heading") {
console.log(`H${block.level}: ${block.text}`)
console.log(block.bbox)
}
if (block.type === "table") {
const table = block as IRTable
for (const row of table.rows) {
for (const cell of row) {
console.log(cell.text, cell.colspan, cell.rowspan)
}
}
}
if (block.type === "paragraph") {
console.log(block.text)
console.(block.)
.(block.)
}
}
Convert Blocks Back to Markdown
import { blocksToMarkdown } from "kordoc"
const markdown = blocksToMarkdown(result.blocks)
Document Comparison
import { compare } from "kordoc"
const bufA = readFileSync("v1.hwp").buffer
const bufB = readFileSync("v2.hwpx").buffer
const diff = await compare(bufA, bufB)
console.log(diff.stats)
for (const d of diff.diffs) {
console.log(d.type, d.blockA?.text ?? d.blockB?.text)
}
Form Field Extraction
import { parse, extractFormFields } from "kordoc"
const result = await parse(buffer.buffer)
if (result.success) {
const form = extractFormFields(result.blocks)
console.log(form.confidence)
for (const field of form.fields) {
console.log(`${field.label}: ${field.value}`)
}
}
Markdown → HWPX Generation
import { markdownToHwpx } from "kordoc"
import { writeFileSync } from "fs"
const markdown = `
# 제목
본문 내용입니다.
| 구분 | 내용 |
| --- | --- |
| 항목1 | 값1 |
| 항목2 | 값2 |
`
const hwpxBuffer = await markdownToHwpx(markdown)
writeFileSync("output.hwpx", Buffer.from(hwpxBuffer))
CLI Usage
npx kordoc document.hwpx
npx kordoc document.hwp -o output.md
npx kordoc *.pdf -d ./converted/
npx kordoc report.hwpx --format json
npx kordoc report.hwpx --pages 1-3
npx kordoc watch ./incoming -d ./output
npx kordoc watch ./docs --webhook https://api.example.com/hook
MCP Server Setup
Add to your MCP config (Claude Desktop, Cursor, Windsurf):
{
"mcpServers": {
"kordoc": {
"command": "npx",
"args": ["-y", "kordoc-mcp"]
}
}
}
Available MCP Tools
| Tool | Description |
|---|
parse_document | Parse HWP/HWPX/PDF → Markdown + metadata + outline + warnings |
detect_format | Detect file format via magic bytes |
parse_metadata | Extract only metadata (fast, no full parse) |
parse_pages | Parse a specific page range |
parse_table | Extract the Nth table from a document |
compare_documents | Diff two documents (cross-format supported) |
parse_form | Extract form fields as structured JSON |
TypeScript Types Reference
import type {
ParseResult, ParseSuccess, ParseFailure,
ErrorCode,
IRBlock, IRBlockType, IRTable, IRCell, CellContext,
DocumentMetadata, OutlineItem,
ParseWarning, WarningCode,
BoundingBox,
InlineStyle,
ParseOptions, FileType,
OcrProvider,
WatchOptions,
DiffResult, BlockDiff, CellDiff, DiffChangeType,
FormField, FormResult,
} from "kordoc"
Common Patterns
Batch Process Files with Error Handling
import { parse, detectFormat } from "kordoc"
import { readFileSync } from "fs"
import { glob } from "glob"
const files = await glob("./docs/**/*.{hwp,hwpx,pdf}")
for (const file of files) {
const buffer = readFileSync(file)
const fmt = detectFormat(buffer.buffer)
if (fmt === "unknown") {
console.warn(`Skipping unknown format: ${file}`)
continue
}
const result = await parse(buffer.buffer)
if (!result.success) {
if (result.code === "ENCRYPTED") {
console.warn(`Encrypted, skipping: ${file}`)
} else if (result.code === "IMAGE_BASED_PDF") {
console.warn(`Image-based PDF needs OCR: ${file}`)
} else {
console.()
}
}
.()
}
Extract All Tables from a Document
import { parse } from "kordoc"
import type { IRTable } from "kordoc"
const result = await parse(buffer.buffer)
if (result.success) {
const tables = result.blocks.filter(b => b.type === "table") as IRTable[]
tables.forEach((table, i) => {
console.log(`\n--- Table ${i + 1} ---`)
for (const row of table.rows) {
const cells = row.map(cell => cell.text.trim()).join(" | ")
console.log(`| ${cells} |`)
}
})
}
OCR with Tesseract.js
import { parse } from "kordoc"
import Tesseract from "tesseract.js"
const result = await parse(buffer.buffer, {
ocr: async (pageImage, pageNumber, mimeType) => {
const blob = new Blob([pageImage], { type: mimeType })
const url = URL.createObjectURL(blob)
const { data } = await Tesseract.recognize(url, "kor+eng")
URL.revokeObjectURL(url)
return data.text
}
})
Watch Mode Programmatic API
import { watch } from "kordoc"
const watcher = watch("./incoming", {
output: "./converted",
webhook: process.env.WEBHOOK_URL,
onFile: async (file, result) => {
if (result.success) {
console.log(`Converted: ${file}`)
}
}
})
watcher.stop()
Troubleshooting
buffer.buffer vs Buffer — kordoc requires ArrayBuffer, not Node.js Buffer. Always pass readFileSync("file").buffer or use .buffer on a Uint8Array.
PDF tables not detected — Line-based detection requires pdfjs-dist installed. Install it: npm install pdfjs-dist. For borderless tables, kordoc uses cluster-based heuristics automatically.
"IMAGE_BASED_PDF" error — The PDF contains scanned images with no text layer. Provide an ocr function in parse options.
"ENCRYPTED" error — HWP DRM/password-protected files cannot be parsed without the decryption key. No workaround.
Korean characters garbled in output — Ensure your terminal/file uses UTF-8 encoding. kordoc outputs UTF-8 Markdown by default.
Large files are slow — Use pages option to parse only needed pages: parse(buf, { pages: "1-5" }). Metadata-only extraction is faster: parse_metadata MCP tool or check result.metadata directly.
HWP table columns wrong — Update to v1.6.1+. Earlier versions had a 2-byte offset misalignment in LIST_HEADER parsing causing column explosion.