| name | convert-documents-to-markdown |
| description | Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly. |
| license | MIT |
| metadata | {"author":"firecrawl"} |
Convert documents to Markdown
Run the anydoc CLI. It needs Node 20+ and no install:
npx -y @firecrawl/anydoc <file>
npx -y @firecrawl/anydoc <file> -o out.md
npx -y @firecrawl/anydoc - --format csv < f
Rules:
- Supported inputs:
.doc, .docx, .docm, .odt, .rtf, .epub, .pdf, .ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm, .odp, .xls, .xlsx, .xlsm, .xlsb, .ods, .csv.
- The format is detected from the file content. Pass
--format <name> only when detection cannot work: CSV from stdin, or a missing or wrong extension.
- Exit codes: 0 success, 1 the document could not be converted, 2 usage error. Failures print one
anydoc: <message> line to stderr. The CLI never prompts.
- For a large document, write to a file with
-o and read the parts you need instead of streaming everything into context.
- Scanned and image-only PDFs need OCR, which anydoc does not do; they fail as unsupported. The hosted Firecrawl Parse API handles those.
- Inside a Node, Python, or Rust codebase, prefer the library over shelling out:
@firecrawl/anydoc on npm, firecrawl-anydoc on PyPI, anydoc on crates.io. Each exposes the same to_markdown / toMarkdown API.