| name | criterion-level-verification |
| description | Verify a scientific report and supporting artifacts against an AutoSciRub executable rubric criterion by criterion. Use after execution has produced a report, code, results, tables, figures, or other artifacts, when Codex needs strict satisfaction labels, evidence found, evidence gaps, and required actions for targeted revision. |
Criterion-Level Verification
Purpose
Check whether the current research artifact satisfies each executable-rubric criterion. This implements Verify(rho_{i,j}, A_i^{(t)}) from AutoSciRub.
Verification diagnoses omissions; it does not edit the report or run a revision.
All .autoscirub/ paths below refer to the state directory resolved by the controller. When invoked alone, use the user’s override, then AUTOSCIRUB_STATE_DIR, then project config state_dir, then .autoscirub/.
Inputs
Read:
.autoscirub/executable_rubric.json
- current report or generated artifact specified by the user
- supporting code, results, tables, figures, logs, and analysis files as needed
- optional prior verification report when comparing rounds
Do not inspect hidden benchmark answers, target reports, private grading checklists, or files excluded by a profile.
Output
Write .autoscirub/verification_report.json or a round-specific file under .autoscirub/revisions/round-XXX/verification.json:
{
"schema_version": "1.0",
"round": 1,
"all_satisfied": false,
"criteria": [
{
"criterion_id": "C1",
"satisfied": false,
"evidence_found": ["report section, figure, table, code path, result file, or quoted result"],
"evidence_gap": "Specific missing or insufficient evidence.",
"required_actions": ["action needed to satisfy the criterion"],
"artifact_references": ["relative/path or report anchor"]
}
],
"summary": {
"passed": 0,
"failed": 0,
"highest_priority_gaps": ["..."]
}
}
Verification Rules
- Check each criterion independently before judging the overall artifact.
- Mark a criterion satisfied only when the required analysis and supporting evidence are both present.
- Treat literature-only statements as insufficient when the criterion requires task-generated evidence.
- Check that conclusions are supported by artifacts, not merely asserted.
- Check that figures and tables have interpretable captions or report descriptions when the rubric expects them.
- Name the concrete missing experiment, comparison, metric, artifact, or explanation.
- Use stable criterion ids from the executable rubric; do not renumber during verification.
Pass Standard
Set all_satisfied to true only when every criterion in the agreed rubric is satisfied. Count all passed and failed criteria in the summary. Missing inputs and honest limitations remain unmet criteria, not passes. If the user changes the scope, record that change in the rubric before verifying; do not waive requirements unilaterally.