| name | reflect |
| description | Spawn three parallel review subagents over the active transcript, surface learnings, and route each to a concrete edit on an existing skill. Use when the user says reflect. |
Reflect
Mine the current conversation for durable learnings, then route them into skill edits.
When to invoke
- The user said "reflect" or "/reflect".
- A complex task (5+ tool calls) just landed cleanly and the recipe is worth keeping.
- The agent hit dead ends, found the working path, and the path generalizes.
- The user corrected the agent's approach mid-task.
- A non-trivial workflow emerged that isn't captured anywhere.
Skip when the conversation is trivial, off-topic, or already covered by an existing skill the parent followed correctly. One-offs are not learnings.
Process
1. Gather the active conversation
Use the active conversation already in context. When session_store_sql is available, query only the current session ID and the time range needed for this task. Do not search unrelated sessions. If the current session cannot be queried, write a tight digest from the visible conversation and pass that instead.
2. Spawn three reviewers in parallel
One message, three Task calls with agent_type: "general-purpose". Use configured models when present and omit model for auto. Each prompt forbids file writes; the parent applies edits.
| Lens | model | Prompt template |
|---|
| Judgment | configured reflect-judgment model or Task default | references/judgment-reviewer.md |
| Tooling | configured reflect-tooling model or Task default | references/tooling-reviewer.md |
| Divergent | configured reflect-judgment model or Task default | references/divergent-reviewer.md |
Pass each template verbatim, substituting the scoped conversation digest where marked. Use a transcript reference only when the current Copilot host explicitly provides one. Reviewers return findings in the Task response body.
3. Synthesize
One Task call with agent_type: "general-purpose" and the configured reflect-judgment model. Omit model when unconfigured or set to auto. The prompt forbids file writes but allows available integrations for citation checks. Use references/synthesizer.md verbatim, with each reviewer's full output inlined where marked. The synthesizer returns a structured Accepted / Rejected / Backlog list.
4. Structural enforcement check
Sanity-check the synthesizer's Accepted list. For any item that would be enforced more reliably by a lint rule, script, metadata flag, or runtime check, move it from Accepted to Backlog. The synthesizer already applies this criterion; this is a final pass before edits land. See the encode-lessons-in-structure principle skill.
5. Apply
Before applying any Accepted edit, present the synthesizer's full Accepted/Rejected/Backlog output to the user and wait for explicit approval. The user picks which subset to apply and may redirect routings. Skill changes affect every future agent in the org; do not auto-apply.
Backlog items file to whatever devex / backlog tracker your team uses automatically. Those are tracker submissions, not skill edits. Only the Accepted list waits for approval.
For each approved Accepted item, follow the Routing field exactly:
- Trivial existing-skill edit (a one-line bullet, a tightened sentence, a stale fact corrected): parent does directly.
- Substantive existing-skill edit (a new section, a new pattern table, more than ~10 lines): follow the poteto-mode Authoring a skill playbook and run focused task or trigger evals when an eval harness is available.
tune description: <skill path> (the skill exists but did not trigger when it should have): add realistic positive and negative trigger cases, then tighten the description against them.
new skill via authoring playbook: <kebab-name>: follow the Authoring a skill playbook. Do not invent the shape ad hoc.
If your environment ships a SKILL.md validator, run it on every touched skill before declaring done. Skip this step if it doesn't.
6. Summarize for the user
Short list, no preamble:
- Edits applied:
<skill path>. What changed, one line each.
- New skills created:
<skill path>. One line each (rare).
- Backlog filed to the devex tracker:
<issue title> (<tags>). One line each.
- Dropped: one line per rejected finding + reason from the synthesizer.