Skip to main content

review-work

Post-implementation review orchestrator. Launches 5 parallel background sub-agents: Oracle (goal/constraint verification), Oracle (code quality), Oracle (security), unspecified-high (hands-on QA execution), unspecified-high (context mining from GitHub/git/Slack/Notion). All must pass for review to pass. MUST USE after completing any significant implementation work. Triggers: 'review work', 'review my work', 'review changes', 'QA my work', 'verify implementation', 'check my work', 'validate changes', 'post-implementation review'. Use when this capability is needed.

Ir para a instalação

Informações da origem

Repositório
tomevault-io/tomes
Última atividade na origem
23 de julho de 2026 às 21:48
Idioma detectado do SKILL.md
inglês
Estrelas
1
Forks
0

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
review-work
description
Post-implementation review orchestrator. Launches 5 parallel background sub-agents: Oracle (goal/constraint verification), Oracle (code quality), Oracle (security), unspecified-high (hands-on QA execution), unspecified-high (context mining from GitHub/git/Slack/Notion). All must pass for review to pass. MUST USE after completing any significant implementation work. Triggers: 'review work', 'review my work', 'review changes', 'QA my work', 'verify implementation', 'check my work', 'validate changes', 'post-implementation review'. Use when this capability is needed.
## Codex Harness Tool Compatibility This skill may include examples copied from the OpenCode harness. In Codex, do not call OpenCode-only tools such as `call_omo_agent(...)`, `task(...)`, `background_output(...)`, or `team_*(...)` literally. Translate those examples to Codex native tools: | OpenCode example | Codex tool to use | | --- | --- | | `call_omo_agent(subagent_type="explore", ...)` | `multi_agent_v1.spawn_agent({"message":"TASK: act as an explorer. ...","agent_type":"explorer","fork_context":false})` | | `call_omo_agent(subagent_type="librarian", ...)` | `multi_agent_v1.spawn_agent({"message":"TASK: act as a librarian. ...","agent_type":"librarian","fork_context":false})` | | `task(subagent_type="plan", ...)` | `multi_agent_v1.spawn_agent({"message":"TASK: act as a planning agent. ...","agent_type":"plan","fork_context":false})` | | `task(subagent_type="oracle", ...)` for final verification | `multi_agent_v1.spawn_agent({"message":"TASK: act as a rigorous reviewer. ...","agent_type":"lazycodex-gate-reviewer","fork_context":false})` | | `task(category="...", ...)` for implementation or QA | `multi_agent_v1.spawn_agent({"message":"TASK: act as an implementation or QA worker. ...","fork_context":false})` | | `background_output(task_id="...")` | `multi_agent_v1.wait_agent(...)` for mailbox signals | | `team_*(...)` | Use Codex native subagents via `multi_agent_v1.spawn_agent`, `multi_agent_v1.send_input`, `multi_agent_v1.wait_agent`, and `multi_agent_v1.close_agent` | Role-specific behavior must be described in a self-contained `message`. Use `fork_context: false` to start the child with only the initial prompt (no parent history); use `fork_context: true` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's `message`. OMO installs these selectable agent roles into `~/.codex/agents/`: `explorer`, `librarian`, `plan`, `momus`, `metis`, `lazycodex-code-reviewer`, `lazycodex-qa-executor`, and `lazycodex-gate-reviewer` - pass the matching name as `agent_type` so the child gets that role's model and instructions. If the spawn tool exposes no `agent_type` parameter, omit it and describe the role inside `message`. If a code block below conflicts with this section, this section wins. On `multi_agent_v2` sessions the same `agent_type` applies (the OMO installer exposes it) with `fork_turns` instead of `fork_context`. If a code block below conflicts with this section, this section wins. When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`. If a code block below conflicts with this section, this section wins. For work likely to exceed one wait cycle, require the child to send `WORKING: <task> - <current phase>` before long passes and `BLOCKED: <reason>` only when progress stops. A `multi_agent_v1.wait_agent` timeout only means no new mailbox update arrived. Treat a running child as alive. Fallback only when the child is completed without the deliverable, ack-only after followup, explicitly `BLOCKED:`, or no longer running. ## Codex Subagent Reliability Every `multi_agent_v1.spawn_agent` message must be self-contained. Start with `TASK: <imperative assignment>`, then name `DELIVERABLE`, `SCOPE`, and `VERIFY`. State that it is an executable assignment, not a context handoff. Role or specialty instructions belong inside `message`. Use `fork_context: false` unless full history is truly required; paste only the review context that worker needs. Plan and reviewer agents may run for a long time; spawn them in the background, keep doing independent root work, and poll with short `multi_agent_v1.wait_agent` cycles sized to the work. Never use a single long blocking wait for them, and never spin on tiny timeouts as a failure budget. Treat child status as a progress signal, not a timeout counter. For work likely to exceed one wait cycle, require the child to send `WORKING: <task> - <current phase>` before long reading, testing, or review passes, and `BLOCKED: <reason>` only when it cannot progress. While any child is active, keep the parent visibly alive with active subagent count, agent names, latest `WORKING:` phase, and whether the parent is waiting for mailbox updates. Track spawned agent names locally. Use `multi_agent_v1.wait_agent` for mailbox signals, not proof of completion. A timeout only means no new mailbox update arrived. Treat a running child as alive. Fallback only when the child is completed without the deliverable, ack-only after followup, explicitly `BLOCKED:`, or no longer running. Then mark that review lane `INCONCLUSIVE`, do not count it as PASS or approval, close if safe, and respawn a smaller `fork_context: false` reviewer with the missing deliverable. Preserve completed lane results immediately. If the retry budget is exhausted, keep the lane `INCONCLUSIVE` and still emit a final aggregate result. # Review Work - 5-Agent Parallel Review Orchestrator Launch 5 specialized sub-agents in parallel to review completed implementation work from every angle. All 5 must pass for the review to pass. If even ONE fails, the review fails. When `review-work` is used as a final implementation, PR, or `$start-work` gate, it is blocking. A timeout, missing deliverable, ack-only response, explicit `BLOCKED:`, or inconclusive lane is not a pass. Treat that lane as failed, investigate the underlying uncertainty with the `debugging` skill when runtime behavior may be wrong, fix with evidence, and rerun the affected lane before claiming completion, creating or handing off a PR, or merging. When reviewing a PR or branch, collect diff, file contents, and verification results from a dedicated review worktree attached to that branch. Never checkout, test, or edit the review branch in the main worktree. Review evidence must be safe to share. Redact or mask secrets and sensitive user data before including evidence in logs, PR bodies, or handoffs. Never include raw tokens, credentials, auth headers, cookies, API keys, env dumps, private logs, or PII; summarize with lengths, hashes, and short non-sensitive prefixes when identity is needed. The 5 agents cover complementary concerns - together they form a comprehensive review that no single reviewer could match: | # | Agent | Type | Role | Focus Level | |---|-------|------|------|-------------| | 1 | Goal Verifier | Oracle | Did we build what was asked? | MAIN | | 2 | QA Executor | unspecified-high | Does it actually work? | MAIN | | 3 | Code Reviewer | Oracle | Is the code well-written? | MAIN | | 4 | Security Auditor | Oracle | Is it secure? | SUB | | 5 | Context Miner | unspecified-high | Did we miss any context? | MAIN | --- ## Phase 0: Gather Review Context Before launching agents, collect these inputs. Extract from conversation history first - the user's original request, constraints discussed, and decisions made are usually already in the thread. Only ask if truly missing. <required_inputs> - **GOAL**: The original objective. What was the user trying to achieve? Pull from the initial request in this conversation. - **CONSTRAINTS**: Rules, requirements, or limitations. Tech stack restrictions, performance targets, API contracts, design patterns to follow, backward compatibility needs. - **BACKGROUND**: Why this work was needed. Business context, user stories, related systems, prior decisions that informed the approach. - **CHANGED_FILES**: Auto-collect via `git diff --name-only HEAD~1` or against the appropriate base (branch point, specific commit). - **DIFF**: Auto-collect via `git diff HEAD~1` or against the appropriate base. - **FILE_CONTENTS**: Read the full content of each changed file (not just the diff). Oracle agents cannot read files - they need full context in the prompt. - **RUN_COMMAND**: How to start/run the application. Check `package.json` scripts, `Makefile`, `docker-compose.yml`, or ask the user. </required_inputs> Review PRs and branches from a dedicated review worktree only: create or attach one with `git worktree add <path> <branch>` before collecting changed files, diff, file contents, or running checks. The main worktree is read-only context; never checkout, test, or edit the review branch there. **Auto-collection sequence:** ```bash # 1. Get changed files git diff --name-only HEAD~1 # or: git diff --name-only main...HEAD # 2. Get diff git diff HEAD~1 # or: git diff main...HEAD # 3. Detect run command # Check package.json -> "scripts.dev" or "scripts.start" # Check Makefile -> default target # Check docker-compose.yml -> services ``` For GOAL, CONSTRAINTS, BACKGROUND - review the full conversation history. The user's original message almost always contains the goal. Constraints often emerge during discussion. If anything critical is ambiguous, ask ONE focused question - not a checklist. --- ## Phase 1: Launch 5 Agents Launch ALL 5 in a single turn. Every agent uses `run_in_background=true`. No sequential launches. No waiting between them. **Oracle agents receive everything in the prompt** (they cannot read files or run commands). Include DIFF + FILE_CONTENTS + all context directly in the prompt text. **unspecified-high agents are autonomous** - they can read files, run commands, and use tools. Give them goals and pointers, not raw content dumps. --- ### Agent 1: Goal & Constraint Verification (Oracle) - MAIN This agent answers: "Did we build exactly what was asked, within the rules we were given?" ``` task( subagent_type="oracle", run_in_background=true, load_skills=[], description="Verify implementation against original goal and constraints", prompt=""" <review_type>GOAL & CONSTRAINT VERIFICATION</review_type> <original_goal> {GOAL - paste the user's original request and any clarifications} </original_goal> <constraints> {CONSTRAINTS - every rule, requirement, or limitation discussed} </constraints> <background> {BACKGROUND - why this work was needed, broader context} </background> <changed_files> {CHANGED_FILES - list of modified file paths} </changed_files> <file_contents> {FILE_CONTENTS - full content of every changed file, clearly delimited per file} </file_contents> <diff> {DIFF - the actual git diff} </diff> Review whether this implementation correctly and completely achieves the stated goal within the given constraints. Be obsessively thorough - the point of this review is to catch what the implementer missed. REVIEW CHECKLIST: 1. **Goal Completeness**: Break the goal into every sub-requirement (explicit AND implied). For each, mark ACHIEVED / MISSED / PARTIAL. Missing even one implied requirement that a reasonable engineer would have addressed = PARTIAL at minimum. 2. **Constraint Compliance**: List every constraint. For each, verify compliance with specific code evidence. A constraint violated = automatic FAIL. 3. **Requirement Gaps**: Requirements the user clearly wanted but didn't spell out. Things implied by the goal or background that a thoughtful engineer would have included. 4. **Over-Engineering**: Anything added that wasn't requested - unnecessary abstractions, extra features, premature optimizations, speculative generality. Flag these as scope creep. 5. **Edge Cases**: Given the goal, what inputs or scenarios would break this? Trace through at least 5 edge cases mentally. 6. **Behavioral Correctness**: Walk through the code logic for 3+ representative scenarios. Does the code actually produce the expected behavior in each case? OUTPUT FORMAT: <verdict>PASS or FAIL</verdict> <confidence>HIGH / MEDIUM / LOW</confidence> <summary>1-3 sentence overall assessment</summary> <goal_breakdown> For each sub-requirement: - [ACHIEVED/MISSED/PARTIAL] Requirement description - Evidence: specific code reference or gap </goal_breakdown> <constraint_compliance> For each constraint: - [ACHIEVED/MISSED] Constraint description - evidence </constraint_compliance> <findings> - [PASS/FAIL/WARN] Category: Description - File: path (line range if applicable) - Evidence: specific code or logic reference </findings> <blocking_issues>Issues that MUST be fixed. Empty if PASS.</blocking_issues> """) ``` --- ### Agent 2: QA via App Execution (unspecified-high) - MAIN This agent answers: "Does it actually work when you run it?" The QA agent follows a structured process: brainstorm scenarios exhaustively first, then self-review and augment, then create a task list, then execute systematically. ``` task( category="unspecified-high", run_in_background=true, load_skills=["browser:control-in-app-browser", "playwright", "dev-browser"], description="QA by actually running and using the application", prompt=""" <review_type>QA - HANDS-ON APP EXECUTION</review_type> <original_goal> {GOAL} </original_goal> <constraints> {CONSTRAINTS} </constraints> <changed_files> {CHANGED_FILES} </changed_files> <run_command> {RUN_COMMAND - how to start the application, or "unknown" if not determined} </run_command> You are a QA engineer. Your job is to RUN the application and verify it works through hands-on testing. You do not review code - you test behavior. MANDATORY PROCESS (follow in order): ### Step 1: Scenario Brainstorm Before touching the app, write down EVERY test scenario you can think of. Be exhaustive. Think about: - **Happy paths**: The primary use cases this implementation enables. What's the main thing the user wanted to do? - **Boundary conditions**: Empty inputs, maximum-length inputs, zero values, negative numbers, special characters, unicode, very large datasets. - **Error paths**: Invalid inputs, network failures, missing files, permission denied, timeout conditions. - **Regression scenarios**: Existing features that touch the same code paths. Things that worked before and must still work. - **State transitions**: What happens when you do things out of order? Rapid repeated actions? Concurrent usage? - **UX scenarios** (if applicable): Layout on different sizes, keyboard navigation, screen reader compatibility, loading states, error messages. - **Integration points**: Does this feature interact with external services, databases, or other modules? Test those boundaries. Write each scenario as a one-liner with expected behavior. Aim for 15-30 scenarios minimum. ### Step 2: Scenario Augmentation Review your scenario list with fresh eyes. For each scenario, ask: - "What could go wrong here that I haven't considered?" - "What would a malicious or careless user do?" - "What environmental conditions could affect this?" (disk full, slow network, expired tokens) Add at least 5 more scenarios from this reflection. Group scenarios by priority: P0 (must pass), P1 (should pass), P2 (nice to pass). ### Step 3: Create Task List Convert your augmented scenario list into a structured task list (use TaskCreate/TaskUpdate or your todo system). Each task = one test scenario with: - Test name - Steps to execute - Expected result - Priority (P0/P1/P2) ### Step 4: Execute Systematically Work through the task list in priority order (P0 first). For each test: 1. Execute the test steps 2. Record actual result 3. Compare with expected result 4. Mark PASS or FAIL 5. If FAIL: capture evidence (screenshot, terminal output, error message) 6. Mark the task complete **Execution guidance by app type:** - **Web app**: In Codex, use `browser:control-in-app-browser` first for browser work that does not need an authenticated user session. Fall back to playwright/dev-browser when the Browser plugin is unavailable, lacks the needed action, or the test specifically needs a persistent/authenticated browser profile. Navigate, click, fill forms, and verify visual output through the chosen browser surface. - **CLI tool**: Run commands with various arguments, pipe inputs, check exit codes and output. - **Library/SDK**: Write and execute a test script that imports and exercises the public API. - **Backend API**: Use curl/httpie to hit endpoints with various payloads, verify response codes and bodies. - **Mobile/Desktop**: If not directly runnable, write integration tests and execute them. If the app cannot be started (build failure), that's an immediate FAIL - no need to continue. ### Step 5: Compile Results OUTPUT FORMAT: <verdict>PASS or FAIL</verdict> <confidence>HIGH / MEDIUM / LOW</confidence> <summary>1-3 sentence overall assessment</summary> <scenario_coverage> Total scenarios: N P0: X tested, Y passed P1: X tested, Y passed P2: X tested, Y passed </scenario_coverage> <test_results> For each test: - [PASS/FAIL] Test name (Priority) - Steps: What you did - Expected: What should happen - Actual: What actually happened - Evidence: Screenshot path or terminal output snippet (if FAIL) </test_results> <blocking_issues>P0 or P1 failures only. Empty if PASS.</blocking_issues> """) ``` ---
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub