用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/cxcscmu/SkillLearnBench --skill document-text-extraction命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
正在显示 SKILL.md
基于 SOC 职业分类
| name | document-text-extraction |
| description | Extracts text from PDF, DOCX, and PPTX files using Python. |
Use libraries like PyMuPDF (fitz), python-docx, and python-pptx to extract text from documents.
pip install PyMuPDF python-docx python-pptx
import fitz
def extract_pdf(file_path):
doc = fitz.open(file_path)
text = ""
# Just read the first few pages to save time/memory for classification
for page in doc[:3]:
text += page.get_text()
return text
from docx import Document
def extract_docx(file_path):
doc = Document(file_path)
return "\n".join([p.text for p in doc.paragraphs[:50]])
from pptx import Presentation
def extract_pptx(file_path):
prs = Presentation(file_path)
text = ""
for slide in prs.slides[:5]:
for shape in slide.shapes:
if hasattr(shape, "text"):
text += shape.text + "\n"
return text