用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/malue-ai/dazee-small --skill mineru-pdf命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Analyze and process Excel/CSV files using pandas and openpyxl. Supports data summary, filtering, pivot tables, and chart generation.
Playwright browser automation — navigate, read, and interact with web pages using text snapshots and ref-based targeting. Supports keyboard, dialogs, file upload, JS evaluation, console/network debugging, and PDF export. Login state persists across sessions. Use when user wants to open a URL, fill a web form, scrape page content, or operate any website that requires clicking/typing.
Local web search (Tavily/Exa, requires API Key). For quick searches. If no Key configured or deep research needed, use cloud_agent instead.
基于 SOC 职业分类
正在显示 SKILL.md
| name | mineru-pdf |
| description | Parse PDF documents locally into structured Markdown/JSON using MinerU. CPU-only, privacy-first. |
| metadata | {"xiaodazi":{"dependency_level":"lightweight","os":["common"],"backend_type":"local","user_facing":true,"python_packages":["magic-pdf"]}} |
本地解析 PDF 文档为结构化 Markdown 或 JSON,保留标题层级、表格、列表等结构。CPU 运行,数据不出本机。
| 工具 | 擅长 | 局限 |
|---|---|---|
| nano-pdf | 简单文本提取、PDF 元数据 | 不保留结构 |
| pdf-toolkit | 合并/拆分/加密/水印 | 不做内容解析 |
| mineru-pdf | 结构化解析(标题/表格/列表) | 安装包较大 |
优先使用 mineru-pdf 做内容提取,pdf-toolkit 做文件操作。
pip install magic-pdf
magic-pdf -p /path/to/document.pdf -o /path/to/output/ -m auto
参数说明:
-p:输入 PDF 路径-o:输出目录-m:模式选择
auto:自动判断(推荐)txt:纯文本 PDFocr:扫描件 PDFfrom magic_pdf.data.data_reader_writer import FileBasedDataWriter, FileBasedDataReader
from magic_pdf.pipe.UNIPipe import UNIPipe
reader = FileBasedDataReader("")
writer = FileBasedDataWriter(output_dir)
pdf_bytes = reader.read(pdf_path)
pipe = UNIPipe(pdf_bytes, model_list=[], image_writer=writer)
pipe.pipe_classify()
pipe.pipe_analyze()
pipe.pipe_parse()
md_content = pipe.pipe_mk_markdown(image_dir, drop_mode="none")
解析后在输出目录生成:
*.md:Markdown 格式的结构化内容images/:提取的图片*.json:结构化元数据 引用