Prepare the next upstream migration PR for devops-bench (gke-labs -> kubernetes-sigs). Invoke when the user wants to "send the next migration PR", "what migration PR is next", "prep the <module> export", "cut the next forward PR", or otherwise export a module to kubernetes-sigs using the migration toolkit. Picks the next PR from the wave plan + frontier + upstream state, builds a scoped export branch with prep-export.sh, carries along any not-yet-migrated dependencies, and gates it submit-ready with lint + tests before it goes out.
2026-07-08
Keep the docs current after a code change — map changed code areas to the docs that describe them, update those docs in place, and prune known-issues rows the code has verifiably fixed. Invoke after editing the pipeline, adding a model/agent/task/metric, moving a directory, or whenever someone asks to "sync the docs", "update the docs for this change", or "do the docs still match the code?".
2026-07-06
cleanup-orphaned-resources
Discover and remove cloud or local resources leaked by aborted or failed eval runs — stale per-run state, leftover kind clusters and `-eval` containers, stuck runner/agent processes, and orphaned GKE clusters, node service accounts, secrets, VPCs, and Artifact Registry repos in the sandbox project. Invoke when a "fresh" run fails instantly, when someone reports leaked or orphaned resources, or asks to "clean up after a failed run", "sweep the sandbox project", or "why does re-running 409?".
2026-06-30
Use when the user asks for a CODE review of devops-bench changes — e.g. "review this PR", "review my changes", "review the working tree", "code-review this diff", "is this harness/deployer/metric change sound". Reviews a PR (number/URL) or the current working tree across seven code lenses — correctness, testability, maintainability, API hygiene, domain modeling, conventions, and security — and returns ranked, actionable findings with severity + file:line evidence + a concrete fix. Review-only: it analyzes statically and may run unit tests, ruff, and format checks, but it NEVER runs benchmark evals or provisions infra. For a NEW or CHANGED benchmark task (task.yaml + its stack), use the sibling `task-review` skill instead.
2026-06-30
Explain WHY a model scored low on an eval — read the judge's reasons and the agent's trajectory, compare against the task rubric, and produce a rationale (capability gap vs. task/rubric problem). Distinct from infrastructure failures. Invoke when someone asks "why did the model fail this task?", "why is OutcomeValidity low?", "explain this low score", or "did the agent actually do the work?".
2026-06-30