用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/cxcscmu/SkillLearnBench --skill docx-pptx-extraction命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
正在显示 SKILL.md
基于 SOC 职业分类
| name | docx-pptx-extraction |
| description | Extract text from DOCX and PPTX files for content analysis |
Extract text content from Microsoft Office files (.docx, .pptx) for subject classification.
pip install python-docx python-pptx
from docx import Document
def extract_docx_text(docx_path, max_chars=5000):
"""Extract text from DOCX files"""
try:
doc = Document(docx_path)
text = ""
for para in doc.paragraphs:
text += para.text + "\n"
if len(text) > max_chars:
break
return text[:max_chars]
except Exception as e:
return f"Error reading DOCX: {str(e)}"
from pptx import Presentation
def extract_pptx_text(pptx_path, max_chars=5000):
"""Extract text from PPTX files"""
try:
prs = Presentation(pptx_path)
text = ""
for slide_num, slide in enumerate(prs.slides):
if slide_num >= 5: # First 5 slides
break
for shape in slide.shapes:
if hasattr(shape, "text"):
text += shape.text + "\n"
if len(text) > max_chars:
return text[:max_chars]
return text[:max_chars]
except Exception as e:
return f"Error reading PPTX: {str(e)}"
def extract_text_by_type(file_path):
"""Route to appropriate extraction method based on file extension"""
ext = file_path.lower().split('.')[-1]
if ext == 'pdf':
return extract_pdf_text(file_path)
elif ext == 'docx':
return extract_docx_text(file_path)
elif ext == 'pptx':
return extract_pptx_text(file_path)
else:
return ""