| name | paper-scale-research-loop |
| description | Use this skill when the goal is to move AutoLabOS outputs from workflow completion toward evidence-backed, baseline-bearing, paper-scale research quality. |
| contract_version | 1 |
| contract_kind | codex_skill |
| runtime_contract | true |
| gate | paper_scale_evidence_ceiling |
| validation | paper_quality_bar_review |
Paper-Scale Research Loop
Purpose
Drive real experiments, hypothesis refinement, evidence review, and manuscript improvement until the strongest branch is honestly judged against paper-scale standards.
This skill exists to prevent:
- workflow completion from being misreported as research completion
- write_paper completion from being misreported as paper readiness
- polished text from being mistaken for evidence-backed claims
Use this skill when
Use this skill when the user asks to:
- assess research quality
- improve experiment quality
- judge manuscript readiness honestly
- push a topic toward paper scale
- decide whether outputs are only a memo, a draft, or a serious experimental paper candidate
Required research discipline
Keep the broad topic fixed, but allow hypotheses and experiment design to evolve.
Always evaluate:
- broad topic
- current hypothesis
- related-work grounding
- baseline or comparator presence
- executed experiments
- quantitative result tables
- claim-to-evidence linkage
- limitations and failure cases
- strongest-branch selection
- reviewer-grade evidence scale: sample size, one-example gains, seed replication, train-budget depth, and whether claimed interactions exceed the observed grid evidence
Hard gate
Downgrade the output if any of the following are missing:
- fixed broad topic
- explicit current hypothesis
- grounded related work
- baseline/comparator
- real executed experiment
- quantitative comparison table
- claim-to-evidence linkage
- limitations/failure cases
Also downgrade or block paper-scale progression when:
- the headline gain can be explained by a single changed evaluation example
- positive tuning claims have no repeated-seed support
- evaluation sets or optimizer steps are only smoke/preflight scale
- a method-centered paper misses canonical method references needed for related-work grounding
Review-gate continuation rule
When a review gate blocks paper-scale progression, treat the block as a governed outcome, not as a writing problem.
- Do not force
write_paper after a paper-quality reject unless the user explicitly asks for a downgraded memo/draft.
- Follow the recommended transition or backtrack target before attempting manuscript generation again.
- If the objective metric is not met, preserve the negative result and lower the claim ceiling instead of polishing around it.
- If implementation executes fewer conditions, seeds, baselines, or comparisons than the approved design required, treat that as an evidence-scope blocker until rerun or explicitly downgraded.
- Require the next hypothesis/design pass to state what evidence gap it is repairing.
- When review emits node-strengthening recommendations, feed them into the meta-harness or the named upstream node rather than trying to repair the problem only in
write_paper.
Output format
- Current artifact status
- Research-question status
- Hypothesis status
- Related-work depth
- Experimental evidence status
- Baseline/comparator status
- Quantitative result status
- Claim-to-evidence linkage status
- Limitations/failure-case status
- Honest output genre recommendation
- Strongest next action
Guardrails
- Do not reward polish without evidence.
- Do not over-credit abstract-only or plan-only runs.
- Do not call something paper-ready unless the evidence supports it.
- Prefer the strongest validated branch, not the most verbose one.