用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/cxcscmu/SkillLearnBench --skill pdf-calendar-extractor命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
正在显示 SKILL.md
| name | pdf-calendar-extractor |
| description | Extract text and identifying colored regions (e.g., rectangles) from a PDF using pdfplumber. |
This skill provides a way to extract text elements and graphical shapes (like rectangles) from a PDF, including their positions and color properties. This is useful for processing visual calendars where appointments are represented by colored blocks.
pdfplumberimport pdfplumber
def extract_calendar_data(pdf_path):
with pdfplumber.open(pdf_path) as pdf:
page = pdf.pages[0]
# Extract text with positions
text_elements = page.extract_text_full() # or extract_words()
words = page.extract_words()
# Extract rectangles (appointments)
rects = page.rects
# For each rectangle, you can get:
# rect['top'], rect['bottom'], rect['left'], rect['right']
# rect['non_stroking_color'] (fill color)
return words, rects
# Identifying 'blue' blocks
# Blue in RGB is often (0, 0, 1) or a similar variation.
# pdfplumber color formats can vary (RGB, CMYK, etc.)
If the calendar has horizontal lines representing 15-minute intervals:
page.lines where y0 == y1).top).top and bottom to specific times based on their position relative to the lines.基于 SOC 职业分类