用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/lamm-mit/scienceclaw --skill pdf命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
正在显示 SKILL.md
基于 SOC 职业分类
| name | |
| description | Extract text, tables, and metadata from scientific PDF papers and reports |
| metadata | null |
PDF processing toolkit for extracting text, tables, and metadata from scientific papers, supplementary data files, and technical reports. Uses pdfplumber for high-fidelity text and table extraction with layout preservation, falling back to pypdf when pdfplumber is unavailable.
Supports page range selection for large documents and targeted extraction modes (text-only, tables-only, metadata-only) for efficient processing.
# Extract everything from a PDF
python3 skills/pdf/scripts/pdf_extract.py --file /path/to/paper.pdf
# Extract only text from pages 1-5
python3 skills/pdf/scripts/pdf_extract.py --file /path/to/paper.pdf --pages "1-5" --extract text
# Extract tables only
python3 skills/pdf/scripts/pdf_extract.py --file /path/to/supplementary.pdf --extract tables
# Extract metadata only
python3 skills/pdf/scripts/pdf_extract.py --file /path/to/paper.pdf --extract metadata
{
"file": "/path/to/paper.pdf",
"text": "Abstract\n\nWe present a novel approach to protein structure...",
"tables": [
[["Gene", "Expression", "p-value"], ["BRCA1", "2.4x", "0.001"]],
[["Compound", "IC50 (nM)"], ["Compound A", "12.3"]]
],
"metadata": {
"title": "Novel Approach to Protein Structure Prediction",
"author": "Smith et al.",
"creation_date": "2024-01-15",
"pages": 12
},
"page_count": 12
}
Install with pip:
pip install pdfplumber pypdf