| name | evidence-synthesis |
| description | Synthesize multiple issues, configs, seeds, metrics, and artifacts into conservative mechanism-level conclusions with caveats. |
| category | benchmark-evidence |
| kind | analysis |
| phase | analysis |
| requires_write | true |
| requires_slurm | false |
| requires_benchmark_artifacts | true |
| delegates_to | ["artifact-provenance","benchmark-row-status","paper-facing-docs"] |
| output_schema | evidence_synthesis_summary.v1 |
Evidence Synthesis
When to use
Use this skill when multiple issues, campaigns, configs, seeds, metrics, and artifacts need one conservative conclusion.
Workflow
- Inventory source issues, PRs, context notes, configs, seeds, commands, artifacts, and metric tables.
- Classify each evidence row with
benchmark-row-status when benchmark data is involved.
- Classify artifacts with
artifact-provenance.
- Build the required synthesis table.
- Separate observed evidence from hypothesis and state caveats before conclusions.
- For mixed or limited benchmark evidence, open the synthesis with the claim boundary: evidence
tier, fallback/degraded exclusions, major caveats, and uncertainty before result interpretation.
Use evidence status terms consistently:
diagnostic-only, smoke evidence,
nominal benchmark evidence, or paper-grade.
For ordinary exploratory runs, prefer per-experiment hypothesis notes over a central ledger. The
minimum reusable fields are hypothesis, variant/config, expected signal, result classification,
artifact pointer or snapshot, and next decision. Create a central hypothesis ledger only when a
research family has enough related runs that the synthesis question becomes "what do we believe
now?" instead of "what should we run next?".
Required table
| Mechanism | Source issue | Evidence tier | Config | Seeds | Artifacts | Metrics | Verdict | Caveats |
|---|
Guardrails
- Do not upgrade exploratory evidence into paper-facing claims.
- Do not average away blocked, fallback, or degraded rows.
- Do not require central hypothesis ledgers for every exploratory result; use them only for
cross-run belief management, repeated/conflicting results, claim-boundary movement, or paper
synthesis.
- Do not put rankings, success language, or mechanism conclusions before the claim boundary when
evidence is diagnostic-only, smoke-level, mixed, fallback-tainted, or below high confidence.
- Surface fallback/degraded rows as explicit caveats or exclusions before ranking claims.
- For uncertain comparisons below high confidence (about 95%), place numeric uncertainty and change
conditions near the claim boundary.
- Use
paper-facing-docs before manuscript-support language is published.
Output
Use evidence_synthesis_summary.v1.