retro
Aggregate the last N days of traces, MRs, and CI runs to surface patterns worth fixing.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Aggregate the last N days of traces, MRs, and CI runs to surface patterns worth fixing.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
SOC 職業分類に基づく
Walk a project from "no values provisioned" to "doctor --secrets is green" — eight idempotent steps with resume support via setup-state.toml. Wraps the secret framework (ADR-023 §3.8) for AI agents and headless onboarding.
Bootstrap devboy from scratch — install the CLI if missing, register the MCP server, run `devboy onboard` for the active agent, optionally bootstrap the secret framework, verify with `doctor`. First-run skill for both manual installs and the Claude Code / Codex plugin.
First-run wizard for the devboy secret framework — walk a fresh project from "no secret manifest, no router, no daemon" to "every required secret provisioned and verified". Idempotent eight-step flow per ADR-023 §3.8 with state at ~/.devboy/secrets/setup-state.toml so the user can resume or skip.
Analyse the user's Claude Code (or other agent) logs and auto-configure the layered-pipeline compression profiles for their tools, models, and workflow.
Diagnose and fix a broken devboy-tools setup — corrupt config, missing tokens, keychain trouble, wrong paths, plugin install failures.
Enumerate and introspect the active tool bundle — names, categories, schemas, how to invoke each tool from the CLI.
| name | retro |
| description | Aggregate the last N days of traces, MRs, and CI runs to surface patterns worth fixing. |
| category | self-feedback |
| version | 1 |
| compatibility | devboy-tools >= 0.18 |
| activation | ["weekly retro","what patterns are there","where should I improve"] |
| tools | ["trace","get_merge_requests","get_merge_request_discussions","get_pipeline","get_job_logs"] |
Looks back over the last N days of session traces, recent merge requests, and CI pipelines to surface recurring patterns: skills whose success rate slipped, review feedback that keeps coming back, flaky jobs. The output is a user-facing report with suggestions — the skill never files tickets, never edits other skills, and never opens MRs.
--days 7 (default). Collect traces from the last N calendar days
under <scope>/.devboy/sessions/<YYYY-MM-DD>/ where <scope> is
either the repo root (default) or ~/.devboy/ when --global is
passed.result=$(devboy trace begin --skill retro)
SESSION_DIR=$(echo "$result" | jq -r .session_dir)
SESSION_ID=$(echo "$result" | jq -r .session_id)
Emit a decision event recording the window and the scope.
Walk every <date>/<skill>/<session_id>/meta.json in the window —
the trace subsystem nests each session one level below <skill>/.
Per skill, aggregate across all its session directories:
total runs, success / failure / aborted counts, total tool_calls,
total errors, total duration, average duration, and the most
common summary strings for failing runs.
Additionally, read each failing session's trace.jsonl to find
retry loops — sequences of verify events with ok: false followed
by more tool_call attempts. A skill with many retry loops is a
skill that could benefit from a stronger precondition check; record
the ratio retried / total_failures per skill.
Emit one note event per skill containing the aggregate numbers so
future retros have a stable trail.
For every merged MR in the window:
devboy tools call get_merge_requests '{"state":"merged","limit":100}'
Filter the result to merge timestamps inside the window, then for the first ~20 call:
devboy tools call get_pipeline \
'{"mrKey":"mr#482","includeFailedLogs":true}'
Collect failing-job frequency keyed by job name. For the top three
failing jobs, call get_job_logs in search mode to pull the most
common error signature:
devboy tools call get_job_logs \
'{"jobId":"<id>","pattern":"error|fail|panic","context":2,"maxMatches":10}'
Keep only the error shapes that repeat across multiple runs — a single broken job is signal for the developer, not a pattern.
For the same merged MRs:
devboy tools call get_merge_request_discussions \
'{"key":"mr#482","limit":50}'
Group the discussion bodies by naïve keyword bucket (type-safety, error-handling, testing, naming, i18n, performance, security). Count how often each bucket appears across MRs. The top three buckets go into the report.
Markdown to stdout:
# Retro — last 7 days
## Skills with degraded success rate
- solve-issue — 6/10 success (was 9/10 the previous week);
60% of failures retry more than twice; top summary:
"gitlab returned 429".
## Frequent review feedback
- testing (mentioned in 9 MRs)
- error-handling (mentioned in 5 MRs)
- type-safety (mentioned in 4 MRs)
## Flaky CI signal
- integration::auth — 5/20 runs failed with "connection refused"
- clippy — 3/20 runs failed with "-D warnings" on a single rule
## Suggestions
- Add a 429 back-off to the get_issues call inside solve-issue.
- Update review-mr's checklist to call out type-safety explicitly.
- Investigate integration::auth — likely a race on the test fixture.
Omit sections with no entries. Keep the report tight; two screens of text at most.
devboy trace end \
--session-dir "$SESSION_DIR" --session-id "$SESSION_ID" \
--skill retro \
--outcome "$OUTCOME" \
--summary "<N> sessions, <M> MRs, <K> jobs analysed"
SKILL.md, never post a
comment. The suggestions are text for a human to read.<redacted:credential>, <redacted:token-pattern>)
are treated as opaque. Count them, do not try to un-redact them.get_pipeline or get_merge_request_discussions fails, note
the degradation in the report ("CI section omitted — pipeline
lookup failed") rather than pretending everything is fine.daily-report — that is a single-day
summary, this one is a multi-day pattern detector.