원클릭으로
pptx
Extract text and tables from PowerPoint (.pptx) files. Use when a user uploads a PPTX presentation that needs to be processed.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Extract text and tables from PowerPoint (.pptx) files. Use when a user uploads a PPTX presentation that needs to be processed.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Plan page-level PPT outlines, audience adaptation, and visual-mode semantics for the deck.
PPT visual production rules, local enhancements, and vendor skill composition for theme, assets, and slide generation.
Extract text and tables from Word (.docx) files. Use when a user uploads a Word document that needs to be processed.
Extract text and tables from Markdown (.md) files. Use when a user uploads a Markdown document that needs to be processed.
Extract text and tables from PDF files. Use when a user uploads a PDF document that needs to be processed.
Summarize and structure raw document text (PDF/Word) into a format optimized for PPT generation. Use when the user uploads a document and wants to generate a presentation from it.
| name | pptx |
| description | Extract text and tables from PowerPoint (.pptx) files. Use when a user uploads a PPTX presentation that needs to be processed. |
Extract readable text and structured tables from PPTX files. This skill feeds extracted content into downstream processing (e.g., document summarization for PPT generation, content analysis).
.pptx filepython-pptx is the standard library for reading PPTX content.
from pptx import Presentation
prs = Presentation("presentation.pptx")
text_parts = []
tables = []
slide_count = len(prs.slides)
for slide in prs.slides:
for shape in slide.shapes:
# Extract text from shapes
if shape.has_text_frame:
for paragraph in shape.text_frame.paragraphs:
line = paragraph.text.strip()
if line:
text_parts.append(line)
# Extract tables
if shape.has_table:
table_data = []
for row in shape.table.rows:
cells = [cell.text.strip() for cell in row.cells]
table_data.append(cells)
if table_data:
tables.append(table_data)
Why python-pptx first:
markitdown converts PPTX to plain text/markdown for easy extraction.
python -m markitdown presentation.pptx
import subprocess
result = subprocess.run(
["python", "-m", "markitdown", "presentation.pptx"],
capture_output=True, text=True, timeout=60
)
if result.returncode == 0:
raw_text = result.stdout.strip()
Use markitdown when:
# pip install "markitdown[pptx]"
import markitdown
md_converter = markitdown.MarkItDown()
result = md_converter.convert("presentation.pptx")
raw_text = result.text_content
PPTX files are ZIP archives containing XML. Extract raw text if needed:
import zipfile
with zipfile.ZipFile("presentation.pptx", "r") as z:
slide_files = sorted([f for f in z.namelist() if f.startswith("ppt/slides/slide") and f.endswith(".xml")])
for slide_file in slide_files:
content = z.read(slide_file).decode("utf-8")
# Extract text between <a:t> tags
import re
texts = re.findall(r"<a:t>([^<]+)</a:t>", content)
Use XML extraction when:
| Problem | Solution |
|---|---|
| Empty text (encrypted PPTX) | File may be password-protected — prompt user to unlock first |
| Text in grouped shapes | python-pptx automatically flattens groups; XML fallback may be needed |
| Slide master text bleeding through | Filter out text from slide masters/layouts (use shape.is_placeholder check) |
| Tables not extracting | Verify shape.has_table is True; some tables are images, not shapes |
| Very large presentations | Process slide-by-slide, limit total text (e.g., 100,000 characters) |
| .ppt (legacy format, not .pptx) | Convert with LibreOffice: soffice --headless --convert-to pptx file.ppt |
| python-pptx not installed | Use markitdown fallback |
PPTX content follows a slide hierarchy:
# Preserve slide-by-slide structure
for slide_idx, slide in enumerate(prs.slides):
slide_title = None
slide_content = []
for shape in slide.shapes:
if shape.has_text_frame:
text = shape.text_frame.text.strip()
if not text:
continue
# First text shape with large font is likely the title
if shape.shape_type == 1: # TITLE placeholder
slide_title = text
else:
slide_content.append(text)
slide_text = f"Slide {slide_idx + 1}"
if slide_title:
slide_text += f": {slide_title}"
if slide_content:
slide_text += "\n" + "\n".join(slide_content)
all_slides.append(slide_text)
Speaker notes can contain additional context:
if slide.has_notes_slide:
notes_frame = slide.notes_slide.notes_text_frame
notes_text = notes_frame.text.strip()
For PPT generation pipeline:
"\n\n".join(text_parts)max(1, slide_count)Return (raw_text: str, tables: list[list[list[str]]], slide_count: int):
raw_text: all text parts joined with double newlinestables: list of tables, each table is list[list[str]] (rows → cells)slide_count: total number of slides (used as page count estimate)