Skip to main content

pp-substack-reader

Read any Substack publication as a local, full-text-searchable corpus — keyless for free posts, your own session for what you subscribe to. Trigger phrases: `archive this Substack`, `read this Substack post`, `search my Substack corpus`, `what's new in my newsletters`, `use substack-reader`, `run substack`.

跳到安装

来源信息

仓库
mvanhorn/printing-press-library
最近来源活动
2026年8月10日 02:50
检测到的 SKILL.md 语言
英语
星标
1,918
分支
572

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
pp-substack-reader
description
Read any Substack publication as a local, full-text-searchable corpus — keyless for free posts, your own session for what you subscribe to. Trigger phrases: `archive this Substack`, `read this Substack post`, `search my Substack corpus`, `what's new in my newsletters`, `use substack-reader`, `run substack`.
author
Maxime Delavergne
license
Apache-2.0
argument-hint
<command> [args] | install cli|mcp
allowed-tools
Read Bash
metadata
{"openclaw":{"requires":{"bins":"[Truncated]"},"install":["[Truncated]"]}}
<!-- GENERATED FILE — DO NOT EDIT. This file is a verbatim mirror of library/media-and-entertainment/substack-reader/SKILL.md, regenerated post-merge by tools/generate-skills/. Hand-edits here are silently overwritten on the next regen. Edit the library/ source instead. See the repository agent guide, section "Generated artifacts: registry.json, cli-skills/". --> # Substack Reader — Printing Press CLI ## Prerequisites: Install the CLI This skill drives the `substack-reader-pp-cli` binary. **You must verify the CLI is installed before invoking any command from this skill.** If it is missing, install it first: 1. Install via the Printing Press installer. It defaults binaries to `$HOME/.local/bin` on macOS/Linux and `%LOCALAPPDATA%\Programs\PrintingPress\bin` on Windows: ```bash npx -y @mvanhorn/printing-press-library install substack-reader --cli-only ``` 2. Verify: `substack-reader-pp-cli --version` 3. Ensure the reported install directory is on `$PATH` for the agent/runtime that will invoke this skill. If the `npx` install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.5 or newer). This installs into `$GOPATH/bin` (default `$HOME/go/bin`), so add that directory to `$PATH` instead: ```bash go install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/substack-reader/cmd/substack-reader-pp-cli@latest ``` If `--version` reports "command not found" after install, the runtime cannot see the binary directory on `$PATH`. Do not proceed with skill commands until verification succeeds. Substack Reader archives whole publications into a local SQLite mirror you can search, SQL-query, and read offline. Free posts need no login; paid posts you're entitled to unlock with your own session cookie — never redistributed, always opt-in. Unlike every other Substack tool it builds a corpus that compounds instead of fetching live per call. ## When to Use This CLI Use Substack Reader when you want a durable, searchable local copy of one or more Substack publications for reading, agent workflows, or analysis — especially reading a specific post's full text or searching across newsletters offline. It is the right tool when you value a corpus that compounds over live per-call fetching. ## Anti-triggers Do not use this CLI for: - Do not use it to publish, schedule, or manage a Substack you own (this is read-only) — use the Substack web app or a publishing tool. - Do not use it to bulk-scrape or redistribute paid content you are not entitled to — it reads only your own entitled content, on demand. - Do not use it to manage subscribers, payments, or analytics for your own publication. ## Unique Capabilities These capabilities aren't available in any other tool for this API. ### Local corpus that compounds - **`archive`** — Archive a whole Substack publication into a local SQLite mirror you can read, search, and query offline — no other Substack tool builds a persistent corpus. `--limit 0` walks the whole archive; when a run stops at `--limit` instead of exhausting the archive, the output says so (JSON carries `"exhausted": false`) — never treat a limit-stopped run as a complete archive. _Reach for this to turn a live newsletter into a durable, queryable knowledge base instead of re-fetching every time._ ```bash substack-reader-pp-cli archive astralcodexten --limit 0 ``` - **`sql`** — Run read-only SQL over your local Substack corpus for arbitrary analytics — post cadence, audience mix, longest posts — from data you've already archived. _Reach for this for ad-hoc analytics over what you've archived, without re-fetching or writing code._ ```bash substack-reader-pp-cli sql "SELECT json_extract(data,'$.audience') AS audience, COUNT(*) FROM resources WHERE resource_type='posts' GROUP BY audience" ``` ### Entitlement-aware reading - **`read`** — Read one or more posts' full text in a single call; free posts keyless, and paid posts you subscribe to via your own session cookie — with an honest 'preview only, you're not entitled' signal. With several posts, JSON mode emits an array of envelopes (a failed post becomes a `{"post", "error"}` entry and the rest still return); a single post keeps the single-object envelope. _Use to pull posts' full text into an agent workflow — several slugs in one call, no shell loop — respecting exactly what the user is entitled to._ ```bash substack-reader-pp-cli read astralcodexten/open-thread-441 substack-reader-pp-cli read astralcodexten/slug-one astralcodexten/slug-two --agent --select slug,subtitle ``` ### Topic & comparative intelligence - **`digest`** — A time-windowed digest across every publication in your local corpus — what's new since you last synced, ranked, in one view. _Use as a personal 'what did I miss across my newsletters' briefing._ ```bash substack-reader-pp-cli digest --since 7d ``` - **`author-compare`** — Compare two publications' cadence, topics, and free/paid mix from the local corpus. _Use to size up a newsletter before subscribing, or to study what a successful author publishes._ ```bash substack-reader-pp-cli author-compare astralcodexten blog.bytebytego.com ``` ## Command Reference **categories** — Browse Substack's publication categories - `substack-reader-pp-cli categories browse` — List publications in a category - `substack-reader-pp-cli categories list` — List all Substack categories **publications** — Discover Substack publications - `substack-reader-pp-cli publications <query>` — Search Substack publications by name (best-effort; may return few results anonymously) ### Finding the right command When you know what you want to do but not which command does it, ask the CLI directly: ```bash substack-reader-pp-cli which "<capability in your own words>" ``` `which` resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code `0` means at least one match; exit code `2` means no confident match — fall back to `--help` or use a narrower query. ## Recipes ### Build a searchable corpus ```bash substack-reader-pp-cli archive astralcodexten --limit 0 && substack-reader-pp-cli search "prediction markets" ``` Mirror a whole publication (`--limit 0` = no cap) then search it offline with FTS ranking. Compact/agent search hits carry key fields plus a bounded `snippet` — never full post bodies; use `read` (or `--select`) when you want a hit's full text. ### Narrow a large post to fields ```bash substack-reader-pp-cli read astralcodexten/open-thread-441 --agent --select title,post_date,audience,body_html ``` Pull only the fields an agent needs from a verbose post object. ### Read several posts in one call ```bash substack-reader-pp-cli read astralcodexten/slug-one astralcodexten/slug-two astralcodexten/slug-three --agent --select slug,subtitle,audience ``` Batch-read a list of slugs without a shell loop; JSON mode returns an array of post envelopes, and one bad slug doesn't sink the rest. ### Audience mix analytics ```bash substack-reader-pp-cli sql "SELECT json_extract(data,'$.audience') AS audience, COUNT(*) FROM resources WHERE resource_type='posts' GROUP BY audience" ``` Read-only SQL over the local corpus for arbitrary analytics. ## Auth Setup Free/public posts are keyless — zero setup. To read paid posts you already subscribe to, provide your own Substack session cookie (substack.sid); this reads only what you are already entitled to and is never required for free content. Run `substack-reader-pp-cli doctor` to verify setup. ## Agent Mode Add `--agent` to any command. Expands to: `--json --compact --no-input --no-color --yes`. - **Pipeable** — JSON on stdout, errors on stderr - **Filterable** — `--select` keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs: ```bash substack-reader-pp-cli categories list --agent --select id,name,status ``` - **Previewable** — `--dry-run` shows the request without sending - **Offline-friendly** — sync/search commands can use the local SQLite store when available - **Non-interactive** — never prompts, every input is a flag - **Read-only** — do not use this CLI for create, update, delete, publish, comment, upvote, invite, order, send, or other mutating requests ### Response envelope Commands that read from the local store or the API wrap output in a provenance envelope: ```json { "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."}, "results": <data> } ``` Parse `.results` for data and `.meta.source` to know whether it's live or local. A human-readable `N results (live)` summary is printed to stderr only when stdout is a terminal AND no machine-format flag (`--json`, `--csv`, `--compact`, `--quiet`, `--plain`, `--select`) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout. ## Paths and state Agents should treat the CLI's path resolver as part of the runtime contract: - Use `--home <dir>` for one invocation, or set `SUBSTACK_READER_HOME=<dir>` to relocate all four path kinds under one root. - Use per-kind env vars only when a specific kind must diverge: `SUBSTACK_READER_CONFIG_DIR`, `SUBSTACK_READER_DATA_DIR`, `SUBSTACK_READER_STATE_DIR`, `SUBSTACK_READER_CACHE_DIR`. - Resolution order is per-kind env var, `--home`, `SUBSTACK_READER_HOME`, XDG (`XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`), then platform defaults. - `config` contains settings like `config.toml` and profiles. `data` contains `credentials.toml`, `data.db`, cookies, and auth sidecars. `state` contains persisted queries, jobs, and `teach.log`. `cache` contains regenerable HTTP/cache files. - Stored secrets live in `credentials.toml` under the data dir. Existing legacy `config.toml` secrets are read for compatibility and leave `config.toml` on the first auth write. - Run `substack-reader-pp-cli doctor --fail-on warn` to surface path and credential-location warnings. `agent-context` exposes a schema v4 `paths` block for agents that need the resolved dirs. - For MCP, pass relocation through the MCP host config. The MCP binary does not inherit CLI flags: ```json { "mcpServers": { "substack-reader": { "command": "substack-reader-pp-mcp", "env": { "SUBSTACK_READER_HOME": "/srv/substack-reader" } } } } ``` Fleet precedence: an inherited per-kind env var overrides an explicit `--home` for that kind. Use `SUBSTACK_READER_HOME` or per-kind vars as durable fleet levers, and use `--home` only for a single invocation. Relocation is not reversible by unsetting env vars; move files manually before clearing `SUBSTACK_READER_HOME`, or `doctor` will not find credentials left under the former root. ## Automatic learning This CLI ships a self-capturing learning loop. The CLI does its own bookkeeping: every invocation is journaled locally, a failed flag followed by a corrected retry auto-derives a `flag_alias` candidate, and a `teach` on a query family without a playbook auto-synthesizes a `playbook_candidate` from the session's journal. Your job is judgment only: `recall` first, act on surfaced candidates, `teach` the final answer, `playbook amend` when you observe a correction. You never record failures by hand. ### Step 1: `recall` before any discovery Before list/search/drill commands on a new user question, run: ```bash substack-reader-pp-cli recall "<user's question>" --agent ``` The response envelope: ```json { "query": "...", "normalized": "<normalized form>", "query_entities": ["..."], "found": true | false, "match_score": 0.0, "results": [ { "resource_id": "...", "resource_type": "...", "venue": "...", "confidence": 2, "entity_match": "exact|partial|unknown", "source": "taught|preseed|pattern", "warnings": ["..."] } ], "mismatches": [ /* only when --debug-mismatches */ ], "warnings": [ /* top-level */ ], "candidates": [ { "id": 12, "class": "flag_alias | playbook_candidate", "summary": "...", "sightings": 3, "last_seen": "...", "rationale": "...", "next_action": ["<trial command>", "substack-reader-pp-cli learnings confirm 12"] } ], "playbook": { "query_family": "...", "playbook": { "steps": [ { "cmd": "<command with {slot} substitution>", "purpose": "..." } ], "entity_slots": ["$ENTITY"], "expected_tool_calls": 3 }, "slots_resolved": { "$ENTITY": { "token": "<live token>", "canonical": "<canonical>" } }, "notes": "<workarounds + gotchas for this query family>" }, "notes": "<duplicate surface for non-playbook callers>" } ``` Empty-store short-circuit: if the store has no learnings, playbooks, or candidates yet (recall finds nothing and `learnings list` and `learnings candidates` are both empty), skip recall for the rest of this session instead of taxing every query; resume recall-first once something has been taught. ### Step 2: decision tree Read `candidates`, `playbook`, `notes`, `results[0]`, and warnings in that order: ``` if Candidates present (warnings include "candidates_present"): -> candidates are try-then-confirm, never facts. Follow each candidate's two-step next_action verbatim: run the trial command first, then run `learnings confirm <id>` only after the trial verified the behavior. Reject a wrong candidate with `learnings reject <id>`. -> NEVER re-teach something recall surfaced as a candidate; confirm or reject that candidate instead of teaching a duplicate. -> candidates ride alongside playbooks and resource hits, not instead of them; continue with the branches below after acting on them. if Playbook present: -> READ Playbook.notes verbatim FIRST (workarounds + gotchas the CLI surface doesn't expose) -> replay Playbook.steps in order, substituting Playbook.slots_resolved entries for the entity slot tokens. If a step's slot is unresolved, fall back to discovery for that step only. -> the Playbook's expected_tool_calls is a budget; if you find yourself running materially more, record the divergence via `substack-reader-pp-cli playbook amend` at end-of-session. elif Notes present (no Playbook): -> read Notes verbatim before any discovery step; they carry known gotchas for this query family even when no structured choreography exists yet. elif Found AND Results[0].EntityMatch == "exact" AND Results[0].Confidence >= 2: -> skip discovery; fetch live data for Results[*].ResourceID in parallel elif Found AND Results[0].EntityMatch == "partial": -> candidate hint, NOT a hit; read the resource title to validate before trusting elif (any row in Mismatches[] when --debug-mismatches was passed): -> treat as cold start; the stored learning is for a different entity (different canonical resolved from query_entities) else: // Found == false, no playbook, no notes -> cold start; run discovery normally; teach the answer afterward (Step 4). If the family has no playbook yet, that teach auto-synthesizes a playbook candidate from this session's journal - you do not need to record one by hand. ``` Playbook and Notes are orthogonal to the per-resource path. A recall response can carry both a Playbook AND a `Results[]` hit - use both: the Playbook tells you which choreography to run; the resource hits short-circuit specific steps. Default to skipping `mismatches`; pass `--debug-mismatches` only when investigating cold-start surprises. Candidate judgment details: `learnings confirm <id>` prints the candidate's full payload before materializing it - check that the printed payload matches the behavior you verified. `learnings reject <id>` tombstones the derivation signature so the same candidate does not resurface. The envelope carries only the few candidates worth acting on now; `substack-reader-pp-cli learnings candidates` lists the full open set. Graceful degradation: if `learnings confirm` is an unknown command, you are driving an older binary - ignore the candidates guidance and follow the rest of the protocol. ### Step 3: always read `warnings` - `low_confidence`: row exists at `confidence<2`. Treat as a hint, not a skip-discovery hit. - `resource_not_in_store`: the local store doesn't have the resource the learning points at. The match validator couldn't classify entities — direct-fetch and re-evaluate.
在 GitHub 查看
这个 SKILL.md 很大,SkillsMP 这里只预览前一段内容。 在 GitHub 查看