Experiment design, validation planning, baseline selection, ablation study planning, statistical testing, and reproducibility planning for STEM papers. Use when users need to design experiments for algorithms, models, systems, materials, instruments, simulations, biomedical studies, control systems, communication systems, or engineering prototypes. Trigger phrases include experiment design, validation plan, ablation study, baseline, benchmark, metrics, statistical test, reproducibility, and Chinese requests about shiyan sheji or xiaorong shiyan.
설치
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
Experiment design, validation planning, baseline selection, ablation study planning, statistical testing, and reproducibility planning for STEM papers. Use when users need to design experiments for algorithms, models, systems, materials, instruments, simulations, biomedical studies, control systems, communication systems, or engineering prototypes. Trigger phrases include experiment design, validation plan, ablation study, baseline, benchmark, metrics, statistical test, reproducibility, and Chinese requests about shiyan sheji or xiaorong shiyan.
STEM Experiment Designer
Purpose
Use this skill to design experiments that convincingly test a STEM paper's claims. It focuses on matching claims to evidence, selecting fair baselines, and making evaluation reproducible.
Inputs To Ask For
Main claims or hypotheses.
Method or system being evaluated.
Dataset, sample, simulation, apparatus, or benchmark.
Constraints such as budget, runtime, safety, ethics, equipment, or available data.
Target venue and expected standards.
Workflow
Convert claims into testable questions.
For each claim, define what evidence would support it and what result would weaken it.
Choose baselines.
Include classical, current strong, domain-standard, and simple sanity-check baselines.
Explain why each baseline is fair.
Define metrics.
Use primary metrics aligned with the claim.
Add secondary metrics for trade-offs: latency, memory, energy, cost, robustness, uncertainty, safety, interpretability, or sample efficiency.
Plan core experiments.
Main comparison.
Ablation study.
Sensitivity analysis.
Robustness or generalization test.
Failure-case analysis.
Runtime or resource analysis when relevant.
Add statistical and reproducibility controls.
Repeats, random seeds, confidence intervals, effect sizes, significance tests, calibration, train/test separation, blinding, or cross-validation as appropriate.
Define data splits and exclusion criteria before seeing results when possible.
Map experiments to paper figures and tables.
Decide which result belongs in the main paper and which belongs in appendix or supplementary material.