| name | docx |
| description | Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files). Triggers include: any mention of "Word doc", "word document", ".docx", or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. |
DOCX creation, editing, and analysis
Overview
A .docx file is a ZIP archive containing XML files.
Quick Reference
| Task | Approach |
|---|
| Read/analyze content | pandoc or unpack for raw XML |
| Create new document | Use docx-js - see Creating New Documents below |
| Edit existing document | Unpack -> edit XML -> repack |
Reading Content
pandoc --track-changes=all document.docx -o output.md
Converting .doc to .docx
python scripts/office/soffice.py --headless --convert-to docx document.doc
Creating New Documents
Generate .docx files with JavaScript. Install: npm install -g docx
Setup
const { Document, Packer, Paragraph, TextRun, Table, TableRow, TableCell, ImageRun,
Header, Footer, AlignmentType, PageOrientation, LevelFormat, ExternalHyperlink,
TableOfContents, HeadingLevel, BorderStyle, WidthType, ShadingType,
VerticalAlign, PageNumber, PageBreak } = require('docx');
const doc = new Document({ sections: [{ children: [] }] });
Packer.toBuffer(doc).then(buffer => fs.writeFileSync("doc.docx", buffer));
Critical Rules for docx-js
- Set page size explicitly — defaults to A4; use US Letter (12240 x 15840 DXA) for US docs
- Landscape: pass portrait dimensions — docx-js swaps internally
- Never use
\n — use separate Paragraph elements
- Never use unicode bullets — use
LevelFormat.BULLET with numbering config
- PageBreak must be in Paragraph — standalone creates invalid XML
- ImageRun requires
type — always specify png/jpg/etc
- Always set table
width with DXA — never use WidthType.PERCENTAGE (breaks in Google Docs)
- Tables need dual widths —
columnWidths array AND cell width, both must match
- Use
ShadingType.CLEAR — never SOLID for table shading
- TOC requires HeadingLevel only — no custom styles on heading paragraphs
- Override built-in styles — use exact IDs: "Heading1", "Heading2", etc.
- Include
outlineLevel — required for TOC (0 for H1, 1 for H2, etc.)
Editing Existing Documents
Follow all 3 steps in order.
Step 1: Unpack
python scripts/office/unpack.py document.docx unpacked/
Step 2: Edit XML
Edit files in unpacked/word/. Use "Claude" as the author for tracked changes and comments.
Use the Edit tool directly for string replacement. Do not write Python scripts.
Step 3: Pack
python scripts/office/pack.py unpacked/ output.docx --original document.docx
XML Reference
Tracked Changes
Insertion:
<w:ins w:id="1" w:author="Claude" w:date="2025-01-01T00:00:00Z">
<w:r><w:t>inserted text</w:t></w:r>
</w:ins>
Deletion:
<w:del w:id="2" w:author="Claude" w:date="2025-01-01T00:00:00Z">
<w:r><w:delText>deleted text</w:delText></w:r>
</w:del>
Dependencies
- pandoc: Text extraction
- docx:
npm install -g docx (new documents)
- LibreOffice: PDF conversion
- Poppler:
pdftoppm for images