Skip to main content

book-to-skill

Converts books and documents (PDF, EPUB, DOCX, HTML, Markdown, plain text, RTF, MOBI/AZW with Calibre) into structured agent skills, extracting frameworks, mental models, principles, techniques, and anti-patterns. Use when the user wants to study a document through GitHub Copilot CLI, Amp, or Claude Code, apply an author's frameworks while working, or build a reusable knowledge base from a file.

Ir para a instalação

Informações da origem

Repositório
LongLeo287/SEOSONA-OS
Última atividade na origem
4 de agosto de 2026 às 08:21
Idioma detectado do SKILL.md
inglês
Estrelas
2
Forks
1

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Explorador de arquivos
48 arquivos

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
book-to-skill
description
Converts books and documents (PDF, EPUB, DOCX, HTML, Markdown, plain text, RTF, MOBI/AZW with Calibre) into structured agent skills, extracting frameworks, mental models, principles, techniques, and anti-patterns. Use when the user wants to study a document through GitHub Copilot CLI, Amp, or Claude Code, apply an author's frameworks while working, or build a reusable knowledge base from a file.
<!-- Cross-agent notes (informational; ignored by host agents): - Compatible skill roots: GitHub Copilot CLI (~/.copilot/skills, ~/.agents/skills, .github/skills, .claude/skills, .agents/skills), Amp (.agents/skills, ~/.config/agents/skills, ~/.config/amp/skills), Claude Code (~/.claude/skills). - `allowed-tools` is intentionally omitted to stay agent-neutral: Copilot CLI uses `shell`/MCP-server names, Claude uses `Bash`/`Read`/`Write`/`Glob`/`Grep`, Amp adds `shell_command`. The skill needs shell (to run extract.py) and file read/write — each host will prompt for those on first use. - Argument hint: <path-to-document-folder-or-glob>... [skill-name-slug] --> # Book-to-Skill Converter Transform written knowledge into actionable agent skills by extracting structure — not producing summaries. ## Philosophy Books contain crystallized expertise: frameworks, principles, and techniques that took years to develop. This skill extracts that knowledge into a format GitHub Copilot CLI, Amp, Claude Code, or another compatible agent can leverage repeatedly. **Extract structure, not summaries.** A skill isn't a book report. It's a toolkit of: - Named frameworks (mental models with clear application) - Actionable principles (rules that guide decisions) - Techniques (step-by-step methods) - Anti-patterns (what to avoid and why) - Voice calibration (how the author thinks and communicates) **Preserve the author's precision.** Frameworks often have specific names for reasons. "The 5 Whys" isn't interchangeable with "ask why multiple times." Capture the exact formulation. **Layer depth appropriately.** Simple books → simple skills. Complex books with 10+ frameworks → skills with reference files and on-demand chapters. --- ## Modes of Operation Four paths available. Route based on what the user asks: ### 1. Full Conversion (Default) **Trigger:** User provides one or more document/directory/glob paths without special instructions **Action:** Run all steps below (Steps 0–9) **Output:** Complete skill with SKILL.md, chapters/, glossary, patterns, cheatsheet ### 2. Analyze Only **Trigger:** User says "analyze", "just extract", or "I want to review before generating" **Action:** Run Steps 0–3, then produce a structured extraction report (frameworks, principles, techniques found). Stop — do NOT generate skill files. **Output:** Analysis report for user review ### 3. Generate from Prior Analysis **Trigger:** User has existing analysis notes or previously ran analyze-only **Action:** Skip Steps 0–3, use the provided analysis as input, run Steps 4–9 **Output:** Skill files from the provided analysis ### 4. Update / Fold-in (Existing Skill) **Trigger:** User provides one or more new source paths and indicates they want to update an existing skill (either by pointing to the existing skill folder, providing a skill slug that already exists in `SKILLS_HOME`, or explicitly requesting an update). **Action:** Run Step 0 (out-of-scope check), Step 1 (validate inputs), Step 1.5 (identify book type), and Step 2 (extract new files). Then skip to Step 5 (identify/detect existing skill path) and run the **Update / Fold-in Workflow** to merge the new content into the existing skill files. **Output:** Updated existing skill with new/revised chapter summaries and merged indexes/glossaries. --- ## Skill Locations This converter can run from multiple skill systems. When looking for this converter's helper script or writing the generated book skill, prefer these locations in order: 1. GitHub Copilot CLI personal skills: `~/.copilot/skills/` 2. Cross-agent personal skills (Copilot + Amp): `~/.agents/skills/` 3. Claude Code personal skills: `~/.claude/skills/` 4. Project-local Copilot skills: `.github/skills/` 5. Project-local Claude skills: `.claude/skills/` 6. Project-local Amp / Copilot skills: `.agents/skills/` 7. Amp global skills: `~/.config/agents/skills/` 8. Amp legacy global skills: `~/.config/amp/skills/` For **generated** book skills, pick a destination that the user's host agent can actually discover (see Step 5). When more than one valid root exists, ask the user once and remember the answer for the session — do not silently default. --- ## Step 0 — Out-of-scope check If no arguments are provided, stop and respond: > "book-to-skill requires a supported document path, folder, or glob pattern. Usage: `book-to-skill <path-to-document-folder-or-glob>... [skill-name-slug]`" Throughout the workflow: - Identify the input paths and the optional skill slug. - If the last argument is not a file, folder, or glob that exists or matches any files, and it looks like a skill slug (e.g. lowercase hyphens, alphanumeric), treat it as `SKILL_NAME`. - Treat all other arguments as the list of `INPUT_PATHS`. - If any input path is an existing skill directory (contains `SKILL.md` and a `chapters/` sub-folder), or if `SKILL_NAME` matches an existing skill slug in `SKILLS_HOME`, flag this run as an **Update/Fold-in** operation (Mode 4). --- ## Step 1 — Validate input Verify that there is at least one supported file, directory, or glob pattern among the `INPUT_PATHS`. For directories and globs, expand them to find matching supported files (`.pdf`, `.epub`, `.docx`, `.txt`, `.md`, `.markdown`, `.rst`, `.adoc`, `.html`, `.htm`, `.rtf`, `.mobi`, `.azw`, `.azw3`). If no supported files are found, stop with a clear error message. --- ## Step 1.5 — Identify content type Before extracting, ask the user: > "What kind of content do these sources have? This helps me choose the best extraction method. > > 1. **Technical** — has code blocks, tables, formulas, diagrams (e.g. programming books, academic papers, architecture guides) > 2. **Text-heavy** — mostly prose, few or no tables/code (e.g. management, productivity, narrative non-fiction) > 3. **Not sure** — I'll use the fast method and warn you if quality seems limited" Store the answer as `BOOK_TYPE`: - Option 1 → `BOOK_TYPE=technical` - Option 2 → `BOOK_TYPE=text` - Option 3 → `BOOK_TYPE=text` **If `BOOK_TYPE=technical`**, inform the user before proceeding: > "📐 Technical mode selected — using Docling for structure-aware extraction (tables, code blocks, formulas preserved as markdown). This takes ~1.5s per page, so expect a few minutes for longer sources. Starting now…" **If `BOOK_TYPE=text`**, inform: > "📄 Text mode selected — using the fastest suitable extractor for each file type. Plain text/Markdown/HTML are usually ready in seconds; PDFs use pdftotext when available." --- ## Step 2 — Extract text from the source documents Run the extraction script, passing the input paths: ```bash SCRIPT_PATH="" for candidate in \ "$HOME/.copilot/skills/book-to-skill/scripts/extract.py" \ "$HOME/.agents/skills/book-to-skill/scripts/extract.py" \ "$HOME/.claude/skills/book-to-skill/scripts/extract.py" \ ".github/skills/book-to-skill/scripts/extract.py" \ ".claude/skills/book-to-skill/scripts/extract.py" \ ".agents/skills/book-to-skill/scripts/extract.py" \ "$HOME/.config/agents/skills/book-to-skill/scripts/extract.py" \ "$HOME/.config/amp/skills/book-to-skill/scripts/extract.py" do if [ -f "$candidate" ]; then SCRIPT_PATH="$candidate" break fi done if [ -z "$SCRIPT_PATH" ]; then echo "Could not find scripts/extract.py for book-to-skill" >&2 exit 1 fi PYTHON_BIN="${PYTHON_BIN:-python3}" if ! command -v "$PYTHON_BIN" >/dev/null 2>&1; then PYTHON_BIN="python" fi "$PYTHON_BIN" "$SCRIPT_PATH" $INPUT_PATHS --mode <BOOK_TYPE> --install-missing ask ``` Before extraction, the script checks optional Python packages needed for the detected format. If a better extractor is missing, it prompts the user with the available fallback. Non-interactive sessions default to fallback unless install mode is explicitly `yes`. **Tip — preflight the environment:** run `"$PYTHON_BIN" "$SCRIPT_PATH" --check` to print a per-format report of which extractors are installed and the exact command to install whatever is missing, without processing any file. Useful when a user reports a setup or quality problem. This creates: - `<tempdir>/book_skill_work/full_text.txt` — combined extracted text of all sources with clear visually demarcated boundaries. - `<tempdir>/book_skill_work/metadata.json` — overall combined size, words, pages, token counts, and a detailed list of individual processed `sources`. Read `<tempdir>/book_skill_work/metadata.json` to inspect the results. --- ## Step 2.5 — Pre-flight cost estimate Read `<tempdir>/book_skill_work/metadata.json` and present the user with an estimate **before doing any generation**: ``` 📖 Sources detected: <total_sources> source(s) <list each source filename and format from the sources metadata list> 📄 Combined Pages/Sections: ~<N> | Words: ~<N> | Total tokens: ~<N>K 💰 Estimated token cost (Full Conversion / Update): Input (reading + prompts): ~<N>K tokens Output (skill files generated/updated): ~<N>K tokens Total: ~<N>K tokens Cost: multiply the token counts above by your model's current input/output per-1M-token rates (prices and model names change often — do not hardcode them; quote today's rate and label it as an estimate). ⏱ Estimated time: ~<N> minutes 📁 Files to be generated/updated: SKILL.md + chapter files + glossary + patterns + cheatsheet ➡ Proceed with Full Conversion / Update? (or type "analyze only" to preview first) ``` **How to estimate:** - Input tokens ≈ `estimated_tokens` from metadata × 1.3 (prompts overhead per chapter pass) - Output tokens ≈ chapters × per-chapter budget + 4,000 (SKILL.md) + 4,500 (glossary + patterns + cheatsheet) - Per-chapter budget midpoint by `BOOK_TYPE` (DEPTH is decided later in Step 4 and can raise it): `text` ≈ 1,000, `technical` ≈ 1,800. If the user has already indicated reference-only vs deep study, use the matching row of the Step 7 matrix. - Cost: report the token counts and multiply by the user's current per-1M-token input/output rates. Do NOT hardcode dollar figures — model names and prices change; if you show one, label it an estimate and date it. Wait for the user to confirm before proceeding. If they say "analyze only", switch to Mode 2. --- ## Step 2.6 — REPL-style access for large books (> 50k tokens) Inspired by the Recursive Language Model (RLM) paradigm: treat `full_text.txt` as a queryable corpus, not a single read. Loading the whole file into context burns budget you will need later for generation. For books over ~50k tokens, prefer programmatic probes over `Read(full_text.txt)` without bounds: ```bash # Size check before any Read wc -w "$FULL_TEXT_PATH" # Find chapter offsets without loading the whole file grep -n -E "^\s*(Chapter|CHAPTER)\s+[0-9]+" "$FULL_TEXT_PATH" | head -40 # Pull only the chapter you need (lines start..end inclusive) sed -n '<start>,<end>p' "$FULL_TEXT_PATH" # Verify a framework is actually mentioned before claiming it in SKILL.md grep -c -i "westrum\|dora" "$FULL_TEXT_PATH" # Targeted Read with offset/limit avoids dumping the full file # Read(file_path=full_text.txt, offset=<line>, limit=<lines>) ``` Use this approach for Step 3 (structure analysis), Step 7 (per-chapter summaries), and Step 8 (glossary / patterns extraction). On books under 50k tokens, a single `Read` is fine. Why this matters: a 200-page book is ~75k tokens. Re-reading it once per chapter (28 passes) costs ~2M input tokens; using grep + sed to pull only relevant slices keeps generation cost proportional to the output, not the source. --- ## Step 3 — Analyze book structure Read the first 8,000 characters of the extracted `full_text.txt` to identify: - Book **title** and **author(s)** - **Chapter structure** (look for "Chapter N", "PART I", numbered headings, table of contents) - **Core themes** and subject domain - Approximate number of chapters Then read the Table of Contents section if present to map all chapters. **If mode is "Analyze Only":** produce the extraction report now and stop. Structure: ``` ## Extraction Report — <Title> ### Author's Core Frameworks - **<Framework Name>**: <what it is and when to apply> ### Key Principles - <Principle>: <actionable rule> ### Techniques & Methods - <Technique>: <step-by-step or how-to> ### Anti-patterns - <What to avoid>: <why> ### Suggested Skill Name `{author-lastname}-{core-concept}` — e.g. `cialdini-influence` ### Chapters Detected | # | Title | Main Frameworks | ``` --- ## Step 4 — Ask purpose (Full Conversion only) Before generating, ask the user: > "What should this skill help you do? (Pick one or more) > 1. Apply the author's frameworks while working > 2. Think with the author's mental models > 3. Reference specific chapters and concepts > 4. All of the above" Use the answer to weight what gets highlighted in the SKILL.md Core section. **Derive `DEPTH` from the answer (no extra prompt):** - Answer is **only** option 3 (reference) → `DEPTH=reference` — lean, fast-lookup chapters. - Answer includes option 1, 2, or 4 → `DEPTH=study` — deeper chapters with more worked detail, examples, and reasoning. `DEPTH` and `BOOK_TYPE` together set the per-chapter token budget in Step 7. Do **not** ask a separate "study vs reference" question — it is inferred here. (In Modes 2/3, where Step 4 is skipped, default `DEPTH=study`.) --- ## Step 5 — Determine skill name If `SKILL_NAME` was provided, use it as the skill slug. Otherwise, propose two options and let the user choose: - **By author-concept**: `{author-lastname}-{core-concept}` (e.g. `cialdini-influence`, `meadows-systems`) - **By title**: lowercase hyphens from book title (e.g. `designing-data-intensive-apps`) Default to author-concept format if the book has a strong methodological identity. Choose the destination skill root (`SKILLS_HOME`). Probe the user's filesystem for existing skill homes and pick by **the host the user is running in**: | Host agent | Personal skill root (probe in order) | Project-local root | |---|---|---| | **GitHub Copilot CLI** | `~/.copilot/skills` → `~/.agents/skills` | `.github/skills` → `.claude/skills` → `.agents/skills` | | **Amp** | `~/.agents/skills` → `~/.config/agents/skills` → `~/.config/amp/skills` | `.agents/skills` | | **Claude Code** | `~/.claude/skills` | `.claude/skills` | Selection rules: 1. If **exactly one** of the host's candidate roots exists on disk, use it without asking. 2. If **none** exist (fresh machine), ask the user which root to create — present the host-appropriate options and remember the choice for the session. Do not silently pick. 3. If the user explicitly asked for project-local output, prefer the project-local row. 4. If you cannot identify the host, ask: "Which agent are you running this in — GitHub Copilot CLI, Amp, or Claude Code?" Set `SKILLS_HOME` to the selected root and check if `$SKILLS_HOME/<skill_name>/` already exists. If it does, prompt the user to choose:
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub