| name | summary |
| description | Create clear, source-grounded summaries and reading versions adapted to the source and the reader's goal. For videos and podcasts, default to a faithful Chinese paragraph-based edited dialogue or narration that preserves the source's sequence and voice; use shorter guided explanations when explicitly requested. Reconstruct nonfiction books as self-contained long-form articles organized by ideas rather than chapters; use dense structured reference summaries only when the source or user benefits from density. Use whenever the user asks to summarize, prune, condense, distill, explain, make something easier to understand, translate and organize a transcript, or create a reading version of a book, web article, YouTube video, podcast, transcript, local document, movie, TV series, or list of sources. |
Summary
Create a replacement reading experience that is easier to understand than the source. Preserve the source's reasoning, evidence, concrete details, narrative movement, and useful uncertainty while removing repetition and filler.
Comprehension comes before compression. A summary fails if it contains the right ideas but makes the reader work harder than the original. Let concrete scenes and examples carry abstract ideas; introduce terminology only after the reader has something tangible to attach it to.
Defaults
- Write in Chinese unless the user requests another language.
- Save results under
outputs/ unless the user provides a destination.
- Create one Markdown file per source.
- Select the output form from the source and the user's purpose instead of forcing every source into one template.
- For a video or podcast supplied without further format instructions, default to a substantial Chinese edited reading version: paragraph-based dialogue for conversations and paragraph-based narration for a single speaker. Preserve the original progression and voice rather than replacing the source with an abstract overview.
- Use a shorter story-driven guided explanation when the user explicitly asks for a summary, overview, key points, quick read, or explanation rather than a source-like reading experience.
- For technical references, research material, or an explicit request for a high-density version, use a dense structured reconstruction.
- Add deep Q&A only when it resolves real tensions or likely misunderstandings; do not include it by default merely to fill a template.
- Do not overwrite an existing output unless the user explicitly requests it.
Choose the reading experience
Infer the reader's goal from the request and source. Ask only when different choices would materially change the result and no safe default is available.
- Edited video dialogue or narration: Use by default for YouTube videos, podcasts, interviews, conversations, and spoken presentations when the user provides a source without asking for a shorter summary. Translate into natural Chinese, group speech into complete thought paragraphs, preserve the source's order, examples, uncertainty, and recognizable voice, and remove spoken noise. Use speaker labels only when the evidence supports them.
- Guided explanation: Use for interviews, podcasts, conversations, lectures with personal stories, philosophical exploration, or whenever the user asks to understand rather than archive. Follow the source's progression, keep a visible central question, and explain each major idea through a concrete story or example before naming the concept.
- Nonfiction article reconstruction: Use for nonfiction books. Write the article the author might have written if forced to make the book's most important case in one self-contained long-form essay. Preserve original insights, mechanisms, frameworks, strongest evidence, and necessary limits; remove chapter scaffolding and repeated persuasion.
- Structured explanatory summary: Use for essays, articles, documentaries, and lectures with a clear argument. Preserve the argument's sequence and explain claims, evidence, and implications in plain language.
- Dense reference version: Use when the user explicitly asks for high density, comprehensive notes, a reference document, or when the source is primarily technical or factual. Optimize for retrieval without turning every minor branch into an equal-weight section.
- Chronological retelling: Use for movies, TV, memoir, narrative nonfiction, or other sources where sequence and causality carry the meaning.
When in doubt, choose the form that a curious non-expert can read once and explain back in their own words.
Workflow
- Resolve the input type and source.
- For a local file, read it directly with the appropriate file tool.
- For a web article, retrieve the original page and prefer the article body over search snippets or secondary commentary.
- For YouTube, follow the mandatory script-first procedure below. Never make the browser or transcript panel the first attempt.
- For a title-only book request, follow the mandatory model-knowledge-only procedure below. Use supplied book text when the user actually provides it.
- For a movie or TV title, use supplied source material first. If none is available, model knowledge may be used only with a prominent provenance warning.
- For a text file containing multiple inputs, ignore blank lines and
# comments, then process each remaining line independently.
- Record the canonical title, source URL or local path, source language, and evidence quality. Choose the reading experience before outlining.
- Inspect the whole source structure before drafting. First outline it in its original order and identify:
- the central question or narrative drive;
- 3–7 main turns that answer or complicate it;
- the concrete stories, scenes, examples, or demonstrations that make each turn understandable;
- supporting branches that can be shortened or omitted.
For long sources, process coherent sections; do not repeatedly reload or resend the entire source when section-level evidence is sufficient. The inspection outline is not necessarily the output outline: for nonfiction books, rebuild the final order around causal, progressive, contrastive, and application relationships instead of mirroring the table of contents.
- Reconstruct the source for comprehension:
- retain claims, reasoning steps, facts, examples, causal links, key scenes, and meaningful changes of mind;
- remove greetings, sponsorships, navigation, repetition, and decorative filler;
- distinguish the source's claims from the summarizer's inference;
- preserve chronological order for narrative works and conversational progression for exploratory interviews;
- introduce an unfamiliar term only after a plain-language explanation or concrete example;
- separate the main thread from side branches instead of assigning every topic equal weight.
- for edited video dialogue, merge consecutive remarks into complete thought paragraphs rather than preserving caption-sized lines; retain a question-and-answer rhythm without keeping every acknowledgement or interruption;
- for edited video narration, preserve the speaker's explanatory or story sequence in connected Chinese paragraphs instead of turning every point into bullets;
- translate meaning and voice directly from the acquired transcript. Do not first invent a polished source-language transcript and then treat that invention as the source.
- Use short bridges that tell the reader why the conversation is moving to the next idea. Do not merely place compressed concepts next to each other.
- Add concise takeaways and excerpts only when they improve understanding and the exact source text is available. Generate deep questions only when they add a new layer of understanding rather than repeat the overview.
- Read
references/output-formats.md and select the matching format before writing the final file.
- Verify the finished artifact and report its path.
Plain-language discipline
- Prefer ordinary verbs and concrete nouns over stacked abstractions.
- Treat density as the ratio of useful ideas to total length, not the number of abstractions packed into each sentence. Cut repetition before compressing explanations.
- Keep one main idea per paragraph. Use enough explanation and breathing room that the reader does not need to decode every sentence.
- When the source uses a memorable metaphor, keep the metaphor as an anchor and explain it once. Do not repeatedly rename the same idea with new terminology.
- Preserve a speaker's uncertainty when they are thinking aloud. Do not upgrade tentative exploration into a confident universal theory.
- Avoid literary intensification that is not in the source, such as turning “a useful constraint” into “self-destruction” or an interesting result into a “masterpiece.”
- For guided explanations, a useful local pattern is: what happened → what the speaker means → why it matters to the reader. Use it naturally, not as a rigid heading repeated in every section.
- Prefer a small number of well-developed sections over exhaustive fragmentation. A long source may contain many topics while still making only a few central moves.
Videos and podcasts: default to a Chinese edited reading version
When the user supplies a video or podcast without asking for a short summary, create the reading experience they could use instead of watching or listening. This is more faithful and more substantial than a conventional overview, but more readable than raw captions.
- Preserve the source's conversational or explanatory order. Group the material into thematic sections only where the source itself makes a meaningful turn; do not reorder it into a taxonomy of concepts.
- Translate into natural Chinese while preserving each speaker's role, examples, metaphors, uncertainty, humor, disagreement, and changes of mind. The result should sound like an edited conversation or talk, not like a third-person report about it.
- Make each paragraph one complete thought. Combine caption fragments and consecutive remarks; remove greetings, sponsorships, repeated captions, filler words, stutters, empty acknowledgements, and interruptions that add no substance.
- Keep speaker turns at paragraph scale. A useful exchange may contain several paragraphs from one speaker before the other responds. Do not format one sentence per line and do not preserve every
yeah, mhm, or false start.
- Preserve all major turns and their strongest concrete examples. Shorten minor detours, repeated proof, and housekeeping. Do not compress a long, rich source into generic key points merely because the word “summary” appears in the skill name.
- Identify speakers conservatively. Prefer verified names; otherwise use reliable roles such as
主持人 and 嘉宾. If automatic captions do not support trustworthy diarization, do not alternate labels mechanically. Use a narrated reconstruction or state the limitation.
- Add an editorial note that distinguishes the result from a verbatim transcript and states the evidence quality. Never claim that translated or reconstructed wording is an exact quotation.
- Read
references/output-formats.md and use its edited video template. Let an explicit user request for brevity, density, commentary, bilingual text, or a verbatim transcript override this default where source quality and copyright constraints allow.
Nonfiction books: reconstruct a self-contained article
For a nonfiction book, aim to preserve the book's intellectual contribution while removing the machinery required to sustain a full-length book.
- Identify the real problem, central thesis, key mechanisms, original frameworks, strongest supporting evidence, and meaningful limits or counterarguments.
- Keep an anecdote, person, experiment, dataset, or case only when it is inseparable from the claim or is the clearest way to understand an abstraction. Prefer one strong example over several interchangeable examples.
- Merge repeated arguments and remove chapter previews, recaps, rhetorical detours, motivational padding, and stories that merely add length.
- Rebuild the logic instead of summarizing chapter by chapter. A useful progression is: problem → mechanism → framework → evidence → limits → practical meaning, but adapt it to the book.
- Preserve the author's distinctive terms, named frameworks, and important distinctions. Explain each on first use, then keep wording consistent rather than replacing it with generic synonyms.
- Make the result stand alone. Avoid summary-language such as “the author next says,” “the following chapter discusses,” or repeated references to the book's structure. State the ideas directly unless attribution is needed to separate the author's view from editorial judgment.
- Write primarily in connected prose. Use lists, tables, and quotations only when they make a relationship clearer than paragraphs would.
- End with
## 如果只记住三件事, containing three concise points that make sense without the rest of the article.
YouTube: mandatory script-first procedure
For every youtube.com or youtu.be URL:
-
Run the bundled helper from the repository root before opening any browser page or transcript panel. Prefer the virtual-environment interpreter when it exists:
.venv/bin/python .agents/skills/summary/scripts/fetch_youtube_transcript.py \
"<youtube-url>" \
-o ".summary-cache/<source-slug>/transcript.md"
Otherwise run the same command with python3.
-
If the helper reports an authentication or access error, retry the helper once with --cookies-from-browser chrome.
-
Use the generated transcript file as the primary source and continue the normal outline and reconstruction workflow.
-
Use a browser, transcript panel, or alternate transcript provider only after the helper has exited unsuccessfully. State the script failure when switching methods.
-
If yt-dlp is missing, report the missing dependency and use requirements.txt to offer setup. Do not silently bypass the script-first policy.
-
Unless the user asks for a shorter form, turn the acquired transcript directly into the Chinese edited dialogue or narration described above. Do not treat raw caption segmentation as meaningful paragraph or speaker boundaries.
Books: mandatory model-knowledge-only procedure
When the user supplies a book title without source text:
- Assess whether model knowledge is sufficient to distinguish this book's actual ideas from author background, related books, and general impressions.
- If sufficiently familiar, start directly from model knowledge. If not, say that the knowledge is insufficient and request the book text, table of contents, excerpts, or notes. Do not fill the gap by guessing.
- Do not browse the web, search catalogs or reviews, fetch bibliographic metadata, locate previews, or download a remote PDF/EPUB. Do not perform a “quick verification” unless the user explicitly asks for research or fact-checking.
- Label the output
基于模型知识,未核对原书全文.
- Build a logical topic outline from known content, but do not claim it is the book's exact published table of contents unless known with high confidence.
- Do not provide purported verbatim quotations, page numbers, edition-specific details, or precise studies and statistics that are not firmly known. Use
关键概念与案例 instead of 原文摘录.
- Use local or attached book text only when the user provides it or explicitly asks to summarize that source. A title alone is not permission to acquire a remote copy.
Source integrity
- Never fabricate a quotation, page number, timestamp, episode title, dialogue line, or bibliographic fact.
- Use blockquotes only for text verified in the acquired source.
- When relying on model knowledge, label the result
基于模型知识,未核对原文/剧本; replace 原文摘录 with 关键概念 or 关键场景, and paraphrase rather than quote.
- For books, do not reconstruct content from the title, author biography, related books, genre conventions, or scattered impressions. Omit uncertain details or label them explicitly.
- Keep the source's claim separate from editorial judgment. When the available evidence clearly shows a disputed claim, weak support, or narrow applicability, use a short
[注:……] only if it materially helps; do not turn the summary into a review.
- When a source is incomplete, paywalled, unavailable, or transcript quality is poor, state the limitation in the document instead of hiding it.
- Do not silently summarize search-result snippets as though they were the full source.
Large and resumable work
- For a long source, store temporary outline and section drafts under
.summary-cache/<source-slug>/.
- Reuse cached work only when it clearly belongs to the same source and version.
- Keep section numbering stable so interrupted work can resume.
- Remove the matching cache directory after the final output is complete and verified.
- Parallelize independent sections only when the active environment supports it and doing so will not weaken source consistency.
Quality check
Before delivery, confirm that:
- the title and source metadata are correct;
- the central question is clear near the beginning and remains visible;
- every major turn is represented, while minor branches do not crowd out the main thread;
- the result contains concrete information rather than generic commentary;
- important abstract claims are grounded in a story, scene, example, or plain-language explanation;
- a curious non-expert can read it once without maintaining a glossary of invented labels;
- section order helps the reader understand how one idea leads to the next;
- every quotation is traceable to acquired source text;
- uncertainty and model-knowledge sections are labeled;
- headings form valid Markdown and the output file is non-empty;
- any deep Q&A adds understanding instead of repeating the overview;
- the finished summary is demonstrably easier to read than the source, not merely shorter or more organized.
- for a default video or podcast request, the result is a substantial Chinese edited dialogue or narration, uses complete thought paragraphs, preserves the source's progression and voice, and clearly states that it is not a verbatim transcript;
- speaker labels are supported by the source or safely inferred from explicit roles, never created by blindly alternating ambiguous caption markers;
- for a nonfiction book, the result reads as a coherent article rather than a chapter-by-chapter report, preserves distinctive terminology, and ends with three independently understandable takeaways.