Build or materially redesign an agent skill from a rough idea or a workflow that just succeeded, with a spec-first design, risk-based evidence, real baseline and with-skill evaluation, blind cross-family forward tests, and a typed completion receipt. Use for “create/add a skill,” “turn what we just did into a skill,” or “redesign a skill's behavior, setup, scripts, or eval.” NOT for prose-only edits (use writing-great-skills), static scoring (use linting-and-scoring), prompt-to-script audits (use determinize-refactor), or installing an existing skill (use skill-installer).
Steward one substantial goal from a rough request to verified completion - clarify consequential decisions, compile a verifiable goal brief, walk the closed seven-route cascade to the smallest fit (four outcome routes total-tdd, wayfinder, direct, to-tickets; three interstitial outcomes grilling pass, human pause, no-work receipt), drive the dependency-ready ticket frontier through existing harness and tracker mechanisms, re-plan when evidence invalidates the graph, and finish with integration-level verification against the original rubric. Use when the user says "take this goal to done", "turn this into tickets and drive it through", "steward this project to completion", or asks to run a substantial goal, feature, plan, or goal brief through decomposition and execution. Not for one quick fix (just do it) and not a replacement for any composed mechanism.
Run a fast two-persona or deep four-persona adversarial council across live-verified configured seats. Use when asked to “review this plan,” “red-team this spec,” “challenge this decision,” or independently review an ADR, PRD, or diff. Uses ignored per-device profiles and prohibits Gemini and unplanned runtime substitution. NOT for visual UI critique (use visual-critique) or deterministic claim checking (use verify-this).
Systematic whole-app feature audit → test → fix loop, backed by tracker.py + render.py over one canonical CSV state machine. Inventory every feature into user stories with code-derived expected behavior, then loop: test every story, document errors, fix logic/UX bugs, re-test. Use for "/total-tdd", "auditing an entire app", "building a feature/user-story spec from the code", or a full test-and-fix sweep across all features. Not for a single feature or bug — use `tdd` (red-green-refactor); not for verifying one claim — use `verify-this`; not for reviewing a diff — use `review`.
Run three independent Codex vision inspections of a render, pose, screenshot, or UI, then synthesize a consensus report. Use to verify visible anatomy, geometry, alignment, or polish before proceeding; use design-critique for structured UX feedback and adversarial-review for text artifacts.
Turn a rough task into a launch-ready /goal brief — a verifiable spec with context-access, a verification plan, and a binary rubric — so a dispatched agent runs to completion unattended. Use when prepping a task for "/goal", or when the user says "spec this for goal", "write the rubric", "make this verifiable", "prep a dispatch", or any turn whose real job is assembling the spec+rubric for a later /goal launch. NOT for doing the task yourself (just do it) and NOT for scoring a finished skill (use "linting-and-scoring"). Pairs with "adversarial-review" and "grilling" (stress-test the brief); distinct from "writing-plans"/"to-prd", which produce implementation plans, not verifiable dispatch briefs.
Audit the layered instruction stack (in-conversation user → soul.md/global → project guide → skill → tool/system) for conflicting or ambiguous directives, and surface which layer should win. Use when the user says "check for instruction conflicts", "do my layers contradict", "audit soul.md vs project/skill", "why is the agent ignoring X", or before relying on a deep skill stack. NOT for scoring one skill's quality (use linting-and-scoring) or auditing whether a skill obeys itself (use the adherence audit). Built from ManyIH (arXiv:2604.09443) via apply-paper.
Ultra-compressed communication mode. Cuts token usage ~75% by dropping filler, articles, and pleasantries while keeping full technical accuracy. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman.