| name | iter-refine-writing |
| description | Iteratively refine the WRITING of a systems paper through 12 serial review-fix rounds. Covers macro structure, micro structure, section conventions, explicit RQ-organized evaluation, logic flow, abstract/intro rebuild, consistency, language (3 passes), terminology/claim tone, citation verification, and a final meaning-preservation audit against the entry baseline. Does not handle claims, RQ meaning, research framing, or scope. Use when the user asks to iterate, polish, or improve paper writing, or when an orchestrated WRITE gate requires full writing refinement. Never invoke iter-refine-ideas; report scientific-contract defects to the caller. Works with LaTeX papers targeting any systems venue. |
Iterative Writing Refinement
Refine a systems paper's writing through 12 serial review-fix rounds. Claims
and framing are assumed stable because the user or root orchestrator owns the
scientific contract. If they are not stable, report the defect to the caller;
do not invoke iter-refine-ideas or alter scientific meaning.
Never stage, commit, push, create, or switch branches; read-only Git (git rev-parse, git diff, and git stash create for the Round 11 baseline) is allowed. Return paper and report changes to the caller; persistence belongs to the separately authorized orchestration, experiment, or literature skills.
Subagents review, the main agent (you) fixes. All rounds are serial — one subagent at a time.
First Step
Read the complete paper under docs/paper/ (main.tex or the build entrypoint the user specifies).
Before Round 0, record the entry baseline in the Round 0 report: git rev-parse HEAD, or if the tree is dirty the commit returned by git stash create (which touches neither the tree nor any ref). Round 11 diffs this baseline against the final paper; do not copy paper sources anywhere.
Locate the explicit paper-level RQ set before Round 0. It must contain two to
five numbered RQs already established by the user or root orchestrator. Treat
their scientific meaning as read-only. If the set is missing, outside the
two-to-five range, scientifically vague, or needs adding/removing/merging/
splitting, report the defect to the caller before continuing; do not invoke
iter-refine-ideas, invent RQs, or redefine them during writing refinement.
Always run the full cycle. NEVER stop, pause, or ask for permission between rounds. Launch Round 0 immediately after reading and proceed through all 12 rounds without pausing. Apply fixes directly when each subagent completes, then launch the next round.
Iteration Logs
In an orchestrated WRITE gate, every round writes a complete round report under the active step directory. For standalone use, create a timestamped docs/tmp/iter-refine-writing-<timestamp>/ run directory instead:
docs/tmp/<phase>/step-<NNNN>-<timestamp>/iter-refine-writing-<timestamp>/
round-0-macro-structure.md
round-1-micro-structure.md
round-2-section-conventions.md
round-3-logic-flow.md
round-4-abstract-intro-rebuild.md
round-5-consistency.md
round-6-language-sentence.md
round-7-language-word.md
round-8-terminology-claim-tone.md
round-9-language-flow.md
round-10-citations.md
round-11-meaning-preservation.md
docs/tmp/iter-refine-writing-<timestamp>/ # standalone
Each report is self-contained and records started/completed timestamps, step/parent node, objective, entry paper/evidence revision, files and sources read, method, raw subagent findings, applied and rejected fixes with before/after locations, claim/number preservation checks, compilation/page evidence, alternatives and decision, tree/memory changes, remaining concerns, and next node. Use Markdown prose and tables, not YAML. Do NOT resume a previous run by default: every invocation starts a fresh full cycle from Round 0 in a new run directory. Read old reports and resume only when the user explicitly asks or the outer orchestrator identifies an incomplete current node after context recovery.
Twelve Rounds (serial)
Each round: fork a read-only subagent → receive findings → apply fixes → compile → verify → next round.
Round 0 — Macro structure. Subagent invokes check-paper-structure-flow via the Skill tool with focus on Level 1 (macro). Checks: are all required sections present and correctly ordered? Is design separated from implementation? Is there an overview/architecture diagram? Does Evaluation explicitly list two to five RQs and organize its evidence blocks by them? Are subsections balanced? You apply fixes, compile, verify.
Round 1 — Micro structure. Subagent invokes check-paper-structure-flow via the Skill tool with focus on Levels 2-3 (micro). Checks: does each paragraph play the right role for its section? Does the abstract follow the correct paragraph structure? Does the intro cover the required paragraph roles for its format (per the check-paper-structure-flow references — required roles always present, optional root-cause/challenges paragraphs per their conditions)? Does each RQ evidence block open with the explicit RQ and close with its evidence-backed answer, or — during BOOTSTRAP only — an unanswered evidence placeholder? A submission-ready paper has no unanswered block. Topic sentences? One idea per paragraph? You apply fixes, compile, verify.
Round 2 — Section conventions. Subagent invokes check-paper-structure-flow via the Skill tool with focus on section-specific conventions. Checks: abstract word count and structure, intro paragraph roles, design goals paragraph, the explicit two-to-five-RQ Evaluation overview, one evidence block per RQ, no orphan experiment, related work grouping, and conclusion structure. You apply fixes, compile, verify.
Round 3 — Logic flow. A fresh subagent reads the complete paper, then checks whether the writing supports the claims, contains logic gaps, maintains a coherent argument across sections, and tells the same story from intro through conclusion. Do not invoke iter-review-critique inside WRITE_GATE; the full domain-routed external-search review runs after this writing loop. You apply fixes, compile, verify.
Round 4 — Abstract/intro rebuild. No fork for this round: you (the main agent) invoke rewrite-abstract-intro via the Skill tool and run its complete referenced procedure every time — role mapping, logic-chain diagnosis, paragraph-by-paragraph rebuild, and sentence-level abstract derivation last. Write the reorganization plan into the round log instead of asking the user. Compile, verify.
Round 5 — Consistency. Subagent invokes check-terminology-infoflow via the Skill tool in paper-consistency scope. Checks: architecture and workflow contradictions, claim/number drift, cross-refs, figure/table alignment, design goals ↔ contributions ↔ RQs ↔ evidence, stable RQ wording, and abstract ↔ intro ↔ body ↔ conclusion alignment. During an orchestrated BOOTSTRAP phase, treat the paper as the intended completed submission: do not flag present-tense descriptions of the system, method, implementation, or contributions merely because internal status files say implementation is unfinished, and never rewrite them as plans, targets, or future work. Missing result values remain explicit placeholders and must not be invented. You apply fixes, compile, verify.
Round 6 — Language: sentence structure. Subagent invokes paper-writing-style via the Skill tool with focus on sentence structure. Checks: semicolons joining independent clauses, fragment sentences, subject-verb distance (>7 words), weak openings (It is / There are), colons before unlabeled lists, dangling modifiers. You apply fixes, compile, verify.
Round 7 — Language: word choice. Subagent invokes paper-writing-style via the Skill tool with focus on word choice. Checks: jargon inflation (count compound terms), nominalizations, vague referents (this/it without antecedent), redundant hedging, verbose phrases (in order to, utilize, due to the fact that). You apply fixes, compile, verify.
Round 8 — Terminology and claim tone. Subagent invokes check-terminology-infoflow via the Skill tool for invented terms, definition order, synonym drift, and cross-section concept consistency, then invokes paper-writing-style for self-attacking sentences and claim-tone mechanics. Every non-standard term must be established usage or formally defined at first use. Delete any wording that weakens the contribution, self-attacks, apologizes, or sounds defensive; do not move it to Limitations. Delete apologies and defensive asides ("though neither is central to...", "a full soundness proof is beyond scope"). A negative, contradictory, or inconclusive result in the paper is an evidence defect, not a wording defect: report it to the caller as requiring a replacement experiment and leave its numbers in place until that replacement exists. A language round never deletes evidence to make the paper read positively. You apply fixes, compile, verify.
Round 9 — Language: flow and polish. Subagent invokes paper-writing-style via the Skill tool with focus on flow. Checks: topic position / stress position, old-to-new information thread, paragraph transitions, register consistency across sections. You apply fixes, compile, verify.
Round 10 — Citation gate. Check the .bib file for annotation blocks (VERIFIED, REAL, PDF, ABSTRACT, USED_FOR). If any entry is missing annotations or has REAL: unverified, invoke check-paper-citations via the Skill tool for a full verification run. If all annotations are present and verified, run Pass 3 only (missing citations in .tex). You apply fixes, compile, verify page count.
Round 11 — Meaning-preservation audit. A fresh subagent runs read-only git diff <entry-baseline> -- docs/paper/ and reports as Must-fix any cumulative drift the per-round diffs missed: a claim's technical meaning changed; technical content, citations, or numbers disappeared or changed without a round finding ordering it and its source; RQ wording drift; or terminology changed technical meaning. Every deletion must trace to redundancy or an explicit finding, never to silent compression. You restore what the audit flags, compile, rerun the audit once to confirm the restorations, and deliver the PDF. This round only restores; it introduces no new polish.
Subagent Rules
- All rounds are serial. One subagent at a time. No parallel launches.
- Use
subagent_type: "fork" for all review subagents.
- Subagents must invoke the target skill via the Skill tool. Tell the subagent: "Use the Skill tool to invoke
<skill-name> with args <file path and focus area>, then follow its instructions." This loads the full skill checklist.
- Every subagent prompt must end with: "Do NOT edit the file — only report findings as Must-fix / Should-fix / Consider with section, problem, and concrete fix suggestion."
- After each round, compile (
make or latexmk) and check for errors before starting the next round.
Fix Application Rules
When applying a round's findings, the main agent must follow these mechanical constraints:
- Apply beyond Must-fix. Must-fix items are mandatory. Should-fix items are applied by default — skipping one requires a stated reason in the round log. Consider items are individually evaluated and applied when they improve the paper; note each accepted/rejected Consider item with a one-line reason. Do not silently discard anything below Must-fix.
- One subsection at a time. Apply fixes subsection by subsection, never as one whole-file rewrite. If a round demands restructuring beyond one subsection, write the target outline into the round log first, then apply it in subsection-sized edits.
- No batch text replacement on .tex. Never use sed/regex bulk substitution on paper source; every edit is a reviewed Edit-tool change.
- Numbers are read-only. Language and structure rounds must not alter any quantitative value. A number changes only in the citation/consistency rounds with its source record shown.
- RQ meaning is read-only. Writing rounds may make the existing two-to-five RQs explicit and structurally consistent, but may not add, remove, merge, split, broaden, narrow, or scientifically reword them. Report required scientific changes to the caller; never invoke
iter-refine-ideas.
- Diff every round. Before reporting a round complete, diff the source and confirm: the sections the round claimed to fix actually changed ("rewrote the intro" with an untouched intro is a failed round), the
\cite{} count did not decrease, and no technical content disappeared without an explicit finding that ordered its removal.
Quality Gates
Stop iterating when:
- No Must-fix items remain
- Should-fix items are cosmetic, not structural
- The Round 11 meaning-preservation audit reports no unrestored loss
- Paper compiles cleanly with correct page count