| name | reflect |
| description | Use for an occasional step-back, weekly, at a milestone, or after a big push, NOT every session. A reflection for both the agent and the human: what got done, the patterns worth noticing, what was learned that should change how the work is done, and what to prune or add. |
Reflect: step back, for the agent and the human
Most sessions only need the everyday habits, keep the docs current as you go, journal a few lines at the end. But every so often it's worth stepping back to a higher altitude and asking what the run has been, not just what today's task was. This is a reflection for both of you: the agent learns how to work better, the human sees what they've actually been doing.
Do it occasionally — not every session, and paced by THROUGHPUT, not the calendar. Weekly is the floor; a dense stretch pulls it earlier: a major arc just closed (big ship, incident, direction change), or the session log since the last reflect carries more lessons than a normal week's worth. A quiet agent's scheduled reflect that finds nothing should say so in one line and stop. A heavy nightly version is the thing nobody keeps up; the point is that it's light and real, not a ritual.
Surface, for both audiences:
- The shape of the period. What actually moved, and what didn't. Not a task log, the gist.
- The patterns. What did you keep reaching for, or keep fighting? A recurring friction is a signal worth naming.
- What was learned that should change how the work is done. A lesson is only worth capturing if it changes something, turn it into a doc or skill update there and then (keep it current), don't just note it.
- How the agent itself evolved. The persona, the loops, the way it works, the things that drift slowly and never get caught in a same-turn edit.
- What to prune or add. A rule that's aged, a skill that's earned its place, a habit that isn't serving. Reflection that changes nothing was just journaling. And test efficacy, not just existence: a durable rule should name the concrete failure it prevents, and each reflect re-checks a couple of old rules against that claim — measured result (AutoSaddler, arXiv:2608.23041): keeping every locally-plausible change with no re-check performed WORSE than never improving. When proposing a change, name its layer (prompt text vs tool vs config vs loop logic) — unconstrained self-editing collapses 91% onto prose while the valuable fixes are usually tool- or config-shaped.
- Verify a sample of stored memories. Writing is staffed by every agent; pruning by none — so each reflect takes ~10 memories from the agent's memory folder (biased oldest / least-recently-touched) and checks each one's claim against the live thing it names: the file, the flag, the tool, the person's role. Still true → leave. Drifted → fix in place. Wrong or superseded → delete (and its index line). Duplicated → merge. Log kept/fixed/deleted counts in the reflection. Same treatment for the project's learnings docs when one surfaces stale mid-week: fix at touch, don't wait for Sunday. (Generated surfaces — goanna indexes, dashboards — are excluded: their freshness is the generator's job, not a memory sweep's.)
Output: a short written reflection (in the journal, or a dated reflections/ note), plus any concrete change it triggers, applied the same session. The value isn't the reflection, it's what you and the agent do differently because of it.
This is the periodic step-back. The everyday is covered by keeping docs current as you go and a light end-of-session journal; reflect is the occasional higher-altitude look those can't give you.
What would have let you do this better? (load ask-for-what-you-need)
Part of every reflect's output, not an afterthought: the week's list of access, permissions,
resources, scope or tools that would have made the work more effective, each as one ask Jez
can grant in a word. Silence here is read as "nothing needed", which is almost never true.