Fresh-context cross-model review protocol that prevents self-congratulatory score inflation - routes verdict-bearing reviews (idea critique, paper review, claim audits) to a different model or a zero-context fresh thread with a bias guard. Use when running any review/audit step of the pipeline, when the user asks for an unbiased second opinion, cross-model review, or mentions review score inflation.
Orchestrate the full research lifecycle (literature → ideation → novelty check → idea critique → experiments → paper writing → review simulation → rebuttal → thesis) with human decision cards between stages. Use when the user asks to start/resume a research project, run the research pipeline, or asks "what's the next step" in an ongoing research effort. Do not use for a single isolated task (use the stage-specific skill directly).
Check whether a proposed research novelty (given a problem statement and claimed idea/novelty) overlaps with existing published work. TRIGGER when the user asks to verify research novelty, check if an idea is new, compare a proposed contribution to prior art, do a literature/prior-art search for a specific claim, or assess whether a paper idea has been scooped. DO NOT TRIGGER for general literature reviews unrelated to a specific novelty claim, or for writing related-work sections of an already-validated idea.
Select and run the correct hypothesis test based on data properties. Covers parametric/non-parametric tests, effect sizes, and multiple comparison correction.
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission. Do not use for general experiment design (use experiment-design).
Analyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results" or needs to interpret experimental data. Do not use for run-vs-run alignment and tracking (use compare).
Autonomously improve a generated paper via GPT-5.6-Sol xhigh review → implement fixes → recompile, for 2 rounds. Use when user says "改论文", "improve paper", "论文润色循环", "auto improve", or wants to iteratively polish a generated paper.
Autonomous multi-round research review loop. Repeatedly reviews via external reviewer backend (Codex or manual), implements fixes, and re-reviews until positive assessment or max rounds reached. Use when user says "auto review loop", "review until it passes", or wants autonomous iterative improvement.