Skip to main content

pp-substack-reader

Read any Substack publication as a local, full-text-searchable corpus — keyless for free posts, your own session for what you subscribe to. Trigger phrases: `archive this Substack`, `read this Substack post`, `search my Substack corpus`, `what's new in my newsletters`, `use substack-reader`, `run substack`.

Aller à l'installation

Informations de source

Dépôt
mvanhorn/printing-press-library
Dernière activité de la source
10 août 2026 à 02:50
Langue détectée de SKILL.md
anglais
Étoiles
1 918
Forks
572

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
pp-substack-reader
description
Read any Substack publication as a local, full-text-searchable corpus — keyless for free posts, your own session for what you subscribe to. Trigger phrases: `archive this Substack`, `read this Substack post`, `search my Substack corpus`, `what's new in my newsletters`, `use substack-reader`, `run substack`.
author
Maxime Delavergne
license
Apache-2.0
argument-hint
<command> [args] | install cli|mcp
allowed-tools
Read Bash
metadata
{"openclaw":{"requires":{"bins":"[Truncated]"},"install":["[Truncated]"]}}
<!-- GENERATED FILE — DO NOT EDIT. This file is a verbatim mirror of library/media-and-entertainment/substack-reader/SKILL.md, regenerated post-merge by tools/generate-skills/. Hand-edits here are silently overwritten on the next regen. Edit the library/ source instead. See the repository agent guide, section "Generated artifacts: registry.json, cli-skills/". --> # Substack Reader — Printing Press CLI ## Prerequisites: Install the CLI This skill drives the `substack-reader-pp-cli` binary. **You must verify the CLI is installed before invoking any command from this skill.** If it is missing, install it first: 1. Install via the Printing Press installer. It defaults binaries to `$HOME/.local/bin` on macOS/Linux and `%LOCALAPPDATA%\Programs\PrintingPress\bin` on Windows: ```bash npx -y @mvanhorn/printing-press-library install substack-reader --cli-only ``` 2. Verify: `substack-reader-pp-cli --version` 3. Ensure the reported install directory is on `$PATH` for the agent/runtime that will invoke this skill. If the `npx` install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.5 or newer). This installs into `$GOPATH/bin` (default `$HOME/go/bin`), so add that directory to `$PATH` instead: ```bash go install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/substack-reader/cmd/substack-reader-pp-cli@latest ``` If `--version` reports "command not found" after install, the runtime cannot see the binary directory on `$PATH`. Do not proceed with skill commands until verification succeeds. Substack Reader archives whole publications into a local SQLite mirror you can search, SQL-query, and read offline. Free posts need no login; paid posts you're entitled to unlock with your own session cookie — never redistributed, always opt-in. Unlike every other Substack tool it builds a corpus that compounds instead of fetching live per call. ## When to Use This CLI Use Substack Reader when you want a durable, searchable local copy of one or more Substack publications for reading, agent workflows, or analysis — especially reading a specific post's full text or searching across newsletters offline. It is the right tool when you value a corpus that compounds over live per-call fetching. ## Anti-triggers Do not use this CLI for: - Do not use it to publish, schedule, or manage a Substack you own (this is read-only) — use the Substack web app or a publishing tool. - Do not use it to bulk-scrape or redistribute paid content you are not entitled to — it reads only your own entitled content, on demand. - Do not use it to manage subscribers, payments, or analytics for your own publication. ## Unique Capabilities These capabilities aren't available in any other tool for this API. ### Local corpus that compounds - **`archive`** — Archive a whole Substack publication into a local SQLite mirror you can read, search, and query offline — no other Substack tool builds a persistent corpus. `--limit 0` walks the whole archive; when a run stops at `--limit` instead of exhausting the archive, the output says so (JSON carries `"exhausted": false`) — never treat a limit-stopped run as a complete archive. _Reach for this to turn a live newsletter into a durable, queryable knowledge base instead of re-fetching every time._ ```bash substack-reader-pp-cli archive astralcodexten --limit 0 ``` - **`sql`** — Run read-only SQL over your local Substack corpus for arbitrary analytics — post cadence, audience mix, longest posts — from data you've already archived. _Reach for this for ad-hoc analytics over what you've archived, without re-fetching or writing code._ ```bash substack-reader-pp-cli sql "SELECT json_extract(data,'$.audience') AS audience, COUNT(*) FROM resources WHERE resource_type='posts' GROUP BY audience" ``` ### Entitlement-aware reading - **`read`** — Read one or more posts' full text in a single call; free posts keyless, and paid posts you subscribe to via your own session cookie — with an honest 'preview only, you're not entitled' signal. With several posts, JSON mode emits an array of envelopes (a failed post becomes a `{"post", "error"}` entry and the rest still return); a single post keeps the single-object envelope. _Use to pull posts' full text into an agent workflow — several slugs in one call, no shell loop — respecting exactly what the user is entitled to._ ```bash substack-reader-pp-cli read astralcodexten/open-thread-441 substack-reader-pp-cli read astralcodexten/slug-one astralcodexten/slug-two --agent --select slug,subtitle ``` ### Topic & comparative intelligence - **`digest`** — A time-windowed digest across every publication in your local corpus — what's new since you last synced, ranked, in one view. _Use as a personal 'what did I miss across my newsletters' briefing._ ```bash substack-reader-pp-cli digest --since 7d ``` - **`author-compare`** — Compare two publications' cadence, topics, and free/paid mix from the local corpus. _Use to size up a newsletter before subscribing, or to study what a successful author publishes._ ```bash substack-reader-pp-cli author-compare astralcodexten blog.bytebytego.com ``` ## Command Reference **categories** — Browse Substack's publication categories - `substack-reader-pp-cli categories browse` — List publications in a category - `substack-reader-pp-cli categories list` — List all Substack categories **publications** — Discover Substack publications - `substack-reader-pp-cli publications <query>` — Search Substack publications by name (best-effort; may return few results anonymously) ### Finding the right command When you know what you want to do but not which command does it, ask the CLI directly: ```bash substack-reader-pp-cli which "<capability in your own words>" ``` `which` resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code `0` means at least one match; exit code `2` means no confident match — fall back to `--help` or use a narrower query. ## Recipes ### Build a searchable corpus ```bash substack-reader-pp-cli archive astralcodexten --limit 0 && substack-reader-pp-cli search "prediction markets" ``` Mirror a whole publication (`--limit 0` = no cap) then search it offline with FTS ranking. Compact/agent search hits carry key fields plus a bounded `snippet` — never full post bodies; use `read` (or `--select`) when you want a hit's full text. ### Narrow a large post to fields ```bash substack-reader-pp-cli read astralcodexten/open-thread-441 --agent --select title,post_date,audience,body_html ``` Pull only the fields an agent needs from a verbose post object. ### Read several posts in one call ```bash substack-reader-pp-cli read astralcodexten/slug-one astralcodexten/slug-two astralcodexten/slug-three --agent --select slug,subtitle,audience ``` Batch-read a list of slugs without a shell loop; JSON mode returns an array of post envelopes, and one bad slug doesn't sink the rest. ### Audience mix analytics ```bash substack-reader-pp-cli sql "SELECT json_extract(data,'$.audience') AS audience, COUNT(*) FROM resources WHERE resource_type='posts' GROUP BY audience" ``` Read-only SQL over the local corpus for arbitrary analytics. ## Auth Setup Free/public posts are keyless — zero setup. To read paid posts you already subscribe to, provide your own Substack session cookie (substack.sid); this reads only what you are already entitled to and is never required for free content. Run `substack-reader-pp-cli doctor` to verify setup. ## Agent Mode Add `--agent` to any command. Expands to: `--json --compact --no-input --no-color --yes`. - **Pipeable** — JSON on stdout, errors on stderr - **Filterable** — `--select` keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs: ```bash substack-reader-pp-cli categories list --agent --select id,name,status ``` - **Previewable** — `--dry-run` shows the request without sending - **Offline-friendly** — sync/search commands can use the local SQLite store when available - **Non-interactive** — never prompts, every input is a flag - **Read-only** — do not use this CLI for create, update, delete, publish, comment, upvote, invite, order, send, or other mutating requests ### Response envelope Commands that read from the local store or the API wrap output in a provenance envelope: ```json { "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."}, "results": <data> } ``` Parse `.results` for data and `.meta.source` to know whether it's live or local. A human-readable `N results (live)` summary is printed to stderr only when stdout is a terminal AND no machine-format flag (`--json`, `--csv`, `--compact`, `--quiet`, `--plain`, `--select`) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout. ## Paths and state Agents should treat the CLI's path resolver as part of the runtime contract: - Use `--home <dir>` for one invocation, or set `SUBSTACK_READER_HOME=<dir>` to relocate all four path kinds under one root. - Use per-kind env vars only when a specific kind must diverge: `SUBSTACK_READER_CONFIG_DIR`, `SUBSTACK_READER_DATA_DIR`, `SUBSTACK_READER_STATE_DIR`, `SUBSTACK_READER_CACHE_DIR`. - Resolution order is per-kind env var, `--home`, `SUBSTACK_READER_HOME`, XDG (`XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`), then platform defaults. - `config` contains settings like `config.toml` and profiles. `data` contains `credentials.toml`, `data.db`, cookies, and auth sidecars. `state` contains persisted queries, jobs, and `teach.log`. `cache` contains regenerable HTTP/cache files. - Stored secrets live in `credentials.toml` under the data dir. Existing legacy `config.toml` secrets are read for compatibility and leave `config.toml` on the first auth write. - Run `substack-reader-pp-cli doctor --fail-on warn` to surface path and credential-location warnings. `agent-context` exposes a schema v4 `paths` block for agents that need the resolved dirs. - For MCP, pass relocation through the MCP host config. The MCP binary does not inherit CLI flags: ```json { "mcpServers": { "substack-reader": { "command": "substack-reader-pp-mcp", "env": { "SUBSTACK_READER_HOME": "/srv/substack-reader" } } } } ``` Fleet precedence: an inherited per-kind env var overrides an explicit `--home` for that kind. Use `SUBSTACK_READER_HOME` or per-kind vars as durable fleet levers, and use `--home` only for a single invocation. Relocation is not reversible by unsetting env vars; move files manually before clearing `SUBSTACK_READER_HOME`, or `doctor` will not find credentials left under the former root. ## Automatic learning This CLI ships a self-capturing learning loop. The CLI does its own bookkeeping: every invocation is journaled locally, a failed flag followed by a corrected retry auto-derives a `flag_alias` candidate, and a `teach` on a query family without a playbook auto-synthesizes a `playbook_candidate` from the session's journal. Your job is judgment only: `recall` first, act on surfaced candidates, `teach` the final answer, `playbook amend` when you observe a correction. You never record failures by hand. ### Step 1: `recall` before any discovery Before list/search/drill commands on a new user question, run: ```bash substack-reader-pp-cli recall "<user's question>" --agent ``` The response envelope: ```json { "query": "...", "normalized": "<normalized form>", "query_entities": ["..."], "found": true | false, "match_score": 0.0, "results": [ { "resource_id": "...", "resource_type": "...", "venue": "...", "confidence": 2, "entity_match": "exact|partial|unknown", "source": "taught|preseed|pattern", "warnings": ["..."] } ], "mismatches": [ /* only when --debug-mismatches */ ], "warnings": [ /* top-level */ ], "candidates": [ { "id": 12, "class": "flag_alias | playbook_candidate", "summary": "...", "sightings": 3, "last_seen": "...", "rationale": "...", "next_action": ["<trial command>", "substack-reader-pp-cli learnings confirm 12"] } ], "playbook": { "query_family": "...", "playbook": { "steps": [ { "cmd": "<command with {slot} substitution>", "purpose": "..." } ], "entity_slots": ["$ENTITY"], "expected_tool_calls": 3 }, "slots_resolved": { "$ENTITY": { "token": "<live token>", "canonical": "<canonical>" } }, "notes": "<workarounds + gotchas for this query family>" }, "notes": "<duplicate surface for non-playbook callers>" } ``` Empty-store short-circuit: if the store has no learnings, playbooks, or candidates yet (recall finds nothing and `learnings list` and `learnings candidates` are both empty), skip recall for the rest of this session instead of taxing every query; resume recall-first once something has been taught. ### Step 2: decision tree Read `candidates`, `playbook`, `notes`, `results[0]`, and warnings in that order: ``` if Candidates present (warnings include "candidates_present"): -> candidates are try-then-confirm, never facts. Follow each candidate's two-step next_action verbatim: run the trial command first, then run `learnings confirm <id>` only after the trial verified the behavior. Reject a wrong candidate with `learnings reject <id>`. -> NEVER re-teach something recall surfaced as a candidate; confirm or reject that candidate instead of teaching a duplicate. -> candidates ride alongside playbooks and resource hits, not instead of them; continue with the branches below after acting on them. if Playbook present: -> READ Playbook.notes verbatim FIRST (workarounds + gotchas the CLI surface doesn't expose) -> replay Playbook.steps in order, substituting Playbook.slots_resolved entries for the entity slot tokens. If a step's slot is unresolved, fall back to discovery for that step only. -> the Playbook's expected_tool_calls is a budget; if you find yourself running materially more, record the divergence via `substack-reader-pp-cli playbook amend` at end-of-session. elif Notes present (no Playbook): -> read Notes verbatim before any discovery step; they carry known gotchas for this query family even when no structured choreography exists yet. elif Found AND Results[0].EntityMatch == "exact" AND Results[0].Confidence >= 2: -> skip discovery; fetch live data for Results[*].ResourceID in parallel elif Found AND Results[0].EntityMatch == "partial": -> candidate hint, NOT a hit; read the resource title to validate before trusting elif (any row in Mismatches[] when --debug-mismatches was passed): -> treat as cold start; the stored learning is for a different entity (different canonical resolved from query_entities) else: // Found == false, no playbook, no notes -> cold start; run discovery normally; teach the answer afterward (Step 4). If the family has no playbook yet, that teach auto-synthesizes a playbook candidate from this session's journal - you do not need to record one by hand. ``` Playbook and Notes are orthogonal to the per-resource path. A recall response can carry both a Playbook AND a `Results[]` hit - use both: the Playbook tells you which choreography to run; the resource hits short-circuit specific steps. Default to skipping `mismatches`; pass `--debug-mismatches` only when investigating cold-start surprises. Candidate judgment details: `learnings confirm <id>` prints the candidate's full payload before materializing it - check that the printed payload matches the behavior you verified. `learnings reject <id>` tombstones the derivation signature so the same candidate does not resurface. The envelope carries only the few candidates worth acting on now; `substack-reader-pp-cli learnings candidates` lists the full open set. Graceful degradation: if `learnings confirm` is an unknown command, you are driving an older binary - ignore the candidates guidance and follow the rest of the protocol. ### Step 3: always read `warnings` - `low_confidence`: row exists at `confidence<2`. Treat as a hint, not a skip-discovery hit. - `resource_not_in_store`: the local store doesn't have the resource the learning points at. The match validator couldn't classify entities — direct-fetch and re-evaluate.
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub