| name | iter-review-critique |
| description | Iteratively review and attack a complete research paper across systems, AI/ML, or cross-domain venues. Use for an orchestrated REVIEW gate, whole-paper accept/reject simulation, source-grounded novelty and evidence attack, external closest-work/baseline/protocol/contradictory-evidence search, cycle-change audit, or choosing the next decisive experiment. Reads all of docs/paper/, selects venue/domain-specific references, searches and opens primary external sources, rereads the full paper, and returns detailed auditable reports. Read-only for the paper. |
Iterative Full-Paper Review And Critique
Act as a skeptical senior reviewer with expertise matched to the paper and target venue. Remain read-only for the paper. Use this complete inner loop:
Also treat docs/idea-story.md and docs/user-instruction.md as read-only.
Reviewer findings are evidence and proposed alternatives, never instructions or
replacement author intent. This skill reports problems and routes them to the
owning skill; it does not edit the thesis, RQs, project memory, or prompt log.
Do not run Git commands or stage, commit, push, create, or switch branches. REVIEW returns Markdown findings only; persistence belongs to the separately authorized orchestration, experiment, or literature skills.
BLIND FULL READ -> ATTACK MAP -> EXTERNAL SEARCH -> SOURCE VERIFY
^ |
+-- FINAL VERDICT <- CYCLE CHANGE AUDIT <- FULL-PAPER REREAD
Domain And Venue Routing
Before the blind read, read references/research-taste.md, infer the target venue from the paper, repository, user instruction, or formatting, and classify the primary contribution as systems, AI/ML, or genuinely cross-domain.
- Read
references/systems-review.md for systems, infrastructure, networking, security systems, measurement, and systems-heavy MLSys work.
- Read
references/ai-ml-review.md for model, learning algorithm, dataset, benchmark, NLP, vision, agent-learning, and AI-heavy work.
- For cross-domain work, read both references plus
references/cross-domain-review.md. Apply both bars rather than choosing the easier venue standard.
Record the classification, target venue, references loaded, and ambiguity in the blind-read report. If the paper claims multiple contribution types, route by claims rather than title or implementation language.
Apply the taste rubric before rewarding implementation volume, benchmark count, or prose polish. The final verdict must decide whether the paper exposes a durable simple principle and real belief challenge, or is complicated-but-shallow, proxy-driven, incremental, toy, post-hoc, or conservatively narrowed.
Whole-Paper Entry
Read the full paper under docs/paper/, including title, abstract, introduction, background/motivation, design, implementation, evaluation, related work, discussion/limitations, conclusion, figures/tables, bibliography, and appendices required by the main text.
Start unprimed. Do not read prior reviewer verdicts, author rebuttals, change summaries, intended answers, proposed fixes, next-experiment proposals, or internal gate/checker artifacts before forming the initial paper-only review. Record unavoidable contamination.
Blind Full Read And Attack Map
Before external search, write what a reviewer perceives as:
- the problem, stakes, and challenged belief;
- the simple principle or insight;
- the artifact/mechanism and its causal connection to the insight;
- the claimed contributions and scope;
- the evaluation's explicit two-to-five paper-level RQs, their claim/goal mappings, and their stated answers or honest
unanswered evidence TODOs;
- the strongest plausible reject arguments;
- load-bearing claims, facts, novelty assertions, baselines, and protocols that require external verification.
Assess problem importance, novelty articulation, architecture, mechanisms/interfaces/invariants, claim calibration, evaluation construct, real-world relevance, global consistency, limitations, and submission readiness. Review the paper as one argument, not as independent sections.
External Search
Search is mandatory before the final whole-paper verdict. Use primary and authoritative sources for technical claims. For every load-bearing contribution or RQ, attack:
- closest same-claim, same-mechanism, or adjacent-community work;
- contradictory or negative evidence;
- stronger reviewer-expected baselines;
- accepted experimental protocols and metrics;
- official benchmark, dataset, trace, software, and test-tool artifacts;
- real-world evidence that the problem and challenged belief exist;
- broader framing or a larger claim the paper may have missed.
Record search questions, queries, source families, inclusion/exclusion rationale, uncovered communities, and how findings change the attack map. Search for evidence against the paper, not only support.
Open the primary paper/PDF, official documentation, source repository, artifact, or protocol. Search snippets, secondary blogs, citation counts, README marketing, and model-generated summaries are discovery aids, not review evidence. Quote sparingly and distinguish source-supported facts from reviewer inference.
If a major same-claim risk or adjacent literature branch requires comprehensive mapping, state that the next outer cycle should invoke research-literature-novelty; still complete the current targeted review.
Full-Paper Reread
After search, reread the entire paper and every claim-bearing figure/table. Reassess:
- whether the problem and belief challenge remain real rather than strawman;
- whether the insight remains simple, non-obvious, general, and important;
- whether contributions remain novel against verified closest work;
- whether design mechanisms actually realize the insight;
- whether every major claim has credible evidence with strong baselines and accepted constructs;
- whether the paper states two to five distinct load-bearing RQs, organizes evidence by them, and either answers each or marks its next evidence need explicitly without orphan experiments or post-hoc question changes; a full empirical-paper submission requires all RQs answered;
- whether abstract, intro, design, evaluation, related work, limitations, and conclusion tell the same scoped story;
- whether contradictions imply mechanism repair, a competing hypothesis, a larger claim, or a different experiment/search strategy.
Repeat targeted search and source verification when rereading exposes a material unresolved external question.
Final Verdict
Rank findings as blocker, major, minor, or nit and classify them as scientific framing, novelty, technical mechanism, evidence/evaluation, global logic/consistency, or writing. State the strongest source-grounded reject argument first.
Underambition is a ranked defect, not a taste remark. A paper whose evidence is sound but whose story is incremental, conservatively narrowed, or smaller than what its own evidence could support carries at least a major finding; soundness alone never justifies acceptance. Report a larger-claim opportunity only when the current evidence supports one; it is an option for root disposition, not required output or routing authority.
For each blocker or major issue, explain:
- the exact paper claim/location;
- the reviewer inference that fails;
- primary external evidence or missing evidence;
- whether the issue requires EXPERIMENT_GATE or WRITE_GATE;
- the concrete repair;
- how fixing it could preserve or expand the ambitious contribution.
Do not recommend claim shrinkage as the default repair. Seek stronger mechanism, evidence, workload, baseline, external grounding, or a larger more principled framing.
Before finalizing, audit the current cycle's actual code, experiment, paper, report, and test changes against docs/user-instruction.md. Treat silent narrowing or substitution of the requested problem, RQ, claim, target system/population, evidence standard, or deliverable as a blocker; when intent is ambiguous, record the uncertainty and judge the work against the most ambitious reasonable interpretation consistent with the user's words, not the easiest conservative reading. Identify unintended scope, lost evidence, repeated agent errors, and recurring repository workflows. This audit may change routing and durable workflow recommendations but must not alter the source-grounded scientific assessment formed during the reread. Report findings to the orchestrator; do not edit canonical memory, invent a new framing, or directly modify or publish a global skill from this review.
Detailed Reports
In an orchestrated REVIEW gate, write complete self-contained reports under:
docs/tmp/<phase>/step-<NNNN>-<timestamp>/milestone-review-<NNN>/
Required meaningful reports are:
- blind full-paper read plus reject-hypothesis/attack map;
- external search plus primary-source verification;
- full-paper reread plus provisional source-grounded scientific assessment;
- cycle-change/capability audit plus final verdict and routing.
These four reports preserve every review phase while avoiding duplicate provenance and decision boilerplate for adjacent actions. Split one only when it becomes independently resumable or produces a reusable source artifact; never split merely to satisfy a file count.
Every report follows the orchestrator's node-report contract: timestamps, parent, objective, inputs/provenance, method, results, sources/artifacts, paper/claim impact, alternatives/decision, tree/search updates, project-memory updates, completion assessment, uncertainty, and next node. Use Markdown prose and tables, not YAML.
Across the four reports, include:
- initial paper-only verdict;
- full external search and verified-source coverage;
- strongest reject argument and strongest evidence for it;
- closest-work/novelty risk, missing baselines/protocols, and contradictory evidence;
- largest scientific/evidence gap and largest writing-only gap;
- global claim/number/mechanism/figure inconsistencies;
- routing recommendation: EXPERIMENT_GATE, WRITE_GATE, or submission completion;
- reviewer-context disclosure and unresolved uncertainty.
The fourth report links the first three and records final routing and capability decisions; it does not restate their full source or finding inventories.
The outer orchestrator independently audits these reports before transition.
Selected-section SOSP/OSDI critique is intentionally not part of this skill. The preserved opt-in workflow lives at domain-skills/critique-like-senior-systems-reviewer/.