一键导入
bio-logic
Evaluate scientific claims, methods, biases, and evidence strength. Use when stress-testing a study design, paper, analysis, or causal interpretation.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Evaluate scientific claims, methods, biases, and evidence strength. Use when stress-testing a study design, paper, analysis, or causal interpretation.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Review, score, compare, and rank AI-generated biology or bioinformatics research artifacts. Use when auditing AI-scientist notebooks, code, figures, analyses, manuscripts, or reports for rigor, reproducibility, novelty, and task completion.
Search arXiv through its official API and save local Markdown summaries. Use when finding recent CS, math, physics, or quantitative-biology preprints or resolving arXiv IDs.
Create publication-quality static charts with matplotlib or seaborn. Use when scientific figures need readable axes, accessible palettes, tight layouts, and high data-ink design.
Annotate genes or proteins and infer taxonomy from sequence homology. Use when assigning functions, domains, or taxonomic labels to genomes, contigs, or protein sets.
Assemble genomes or metagenomes and assess assembly quality. Use when turning sequence reads into contigs and reporting completeness, continuity, and contamination evidence.
Bin and refine metagenomic contigs, then assess MAG quality. Use when recovering genomes with QuickBin and checking completeness, contamination, and bin consistency.
| name | bio-logic |
| description | Evaluate scientific claims, methods, biases, and evidence strength. Use when stress-testing a study design, paper, analysis, or causal interpretation. |
Use structured frameworks to evaluate scientific claims, methodology, and evidence strength.
/polars-dovmed, then revise hypotheses against the literature.| Profile | Main evidence checks | Rating language |
|---|---|---|
| Intervention | allocation, controls, adherence, attrition, estimand, harms | GRADE when the review question and evidence synthesis support it |
| Observational or quasi-experimental | identification assumptions, temporality, confounding, negative controls, sensitivity analyses | risk-of-bias plus causal-confidence statement |
| Computational or machine learning | data provenance, leakage, splits, baselines, calibration, external validation, reproducibility | computational evidence confidence |
| Evolutionary or comparative genomics | orthology, taxon/reference sampling, model fit, support, contamination, topology sensitivity | phylogenetic/comparative evidence confidence |
| Descriptive or exploratory omics | sampling, measurement, multiple testing, effect sizes, replication, alternative explanations | descriptive confidence; avoid causal labels |
Use this loop for omics projects, unexpected results, exploratory analyses, and any request that asks what results mean.
Use relevant sections based on the review scope. Skip items not applicable to the study type.
## Methodology
- [ ] Design identifies the causal estimand through randomization or a justified quasi-experimental or natural-experiment strategy
- [ ] Sample size justified (power analysis reported)
- [ ] Randomization/blinding implemented where feasible
- [ ] Confounders identified and controlled
- [ ] Measurements validated and reliable
## Statistics
- [ ] Tests appropriate for data type
- [ ] Assumptions checked
- [ ] Multiple comparisons corrected
- [ ] Effect sizes + CIs reported (not just p-values)
- [ ] Missing data handled appropriately
## Interpretation
- [ ] Conclusions match evidence strength
- [ ] Limitations acknowledged
- [ ] Causal claims require an identified causal design: randomized evidence or
a justified quasi-experimental/natural-experiment strategy with its
assumptions and sensitivity checks
- [ ] No cherry-picking or overgeneralization
## Red Flags
- [ ] P-values clustered just below .05
- [ ] Outcomes differ from registration
- [ ] Correlation presented as causation
- [ ] Subgroups analyzed without preregistration
Claim strength ladder:
| Language | Requires |
|---|---|
| "Demonstrates" | Strong evidence under a design that identifies the target claim |
| "Suggests" / "Indicates" | Observational with controlled confounds |
| "Associated with" | Observational, no causal claim |
| "May" / "Might" | Preliminary or hypothesis-generating |
## Summary
[1-2 sentences: What was studied and main finding]
## Strengths
- [Specific methodological strengths]
## Concerns
### Critical (threaten main conclusions)
- [Issue + why it matters]
### Important (affect interpretation)
- [Issue + why it matters]
### Minor (worth noting)
- [Issue]
## Evidence Rating
[Framework appropriate to the study profile, rating, and justification. Use
GRADE only when it applies to the review question.]
## Bottom Line
[What can/cannot be concluded from this evidence]
## Current Result
[Observed intermediate/final result and QC status]
## Hypothesis Register
| Rank | Hypothesis | Status | Evidence For | Evidence Against | Next Discriminating Check |
|------|------------|--------|--------------|------------------|---------------------------|
| 1 | [Hypothesis] | supported/weakened/ruled out/unresolved | [Evidence] | [Evidence] | [Test] |
| 2 | [Hypothesis] | ... | ... | ... | ... |
| 3 | [Hypothesis] | ... | ... | ... | ... |
| 4 | [Hypothesis] | ... | ... | ... | ... |
| 5 | [Hypothesis] | ... | ... | ... | ... |
## Literature Context
- [Paper or search result, DOI/PMCID if available]: supports/contradicts/narrows [hypothesis] because [reason].
## Analysis Playbook
- [Literature-derived analysis/comparison/outlier check and why it applies]
## Revision
[What changed since the previous register and why]
## Next Step
[Most informative analysis, control, or experiment]
| Task | Action |
|---|---|
| Review a paper | Use the Critique Checklist below |
| Evaluate a claim | Use Claim Assessment steps |
| Assess evidence strength | See references/evidence.md |
| Identify biases | See references/biases.md |
| Statistical errors | See references/stats.md |
| Logical fallacies | See references/fallacies.md |
| Study design review | See references/design.md |
Input: "Review this RCT on vitamin D and depression"
Output:
## Summary
RCT of 200 adults testing vitamin D supplementation for depression over 12 weeks.
## Strengths
- Randomized, double-blind, placebo-controlled
- Validated depression scale (PHQ-9)
- Adequate power analysis
## Concerns
### Critical
- 35% dropout, differential by group (attrition bias)
- ITT analysis not performed
### Important
- Single-site limits generalizability
## Evidence Rating
Moderate (downgraded from high due to attrition bias)
## Bottom Line
Suggestive but not conclusive due to differential attrition.
Input: "This study proves that coffee prevents Alzheimer's"
Assessment: If based on observational data, this is a correlation→causation error. Appropriate language: "Coffee consumption is associated with lower Alzheimer's risk."
Issue: Insufficient methods details Solution: Request missing design/statistics information before rating evidence.
Issue: Conflicting results across studies Solution: Report uncertainty and suggest stronger study designs for resolution.
/ai-scientist-evaluator — rubric-scored evaluation of AI scientist outputs/manuscript-review-council — multi-reviewer critique of peer-reviewed manuscripts/proposal-review — structured critique of funding proposals