evaluate-protocol-outputs
Evaluate model outputs for biological and scientific procedural tasks, including question answering against reference answers, detecting whether protocols contain errors, restoring shuffled step order, comparing generated protocols with references, and aggregating precomputed judge consistency decisions. Use when an agent needs to score, compare, validate, or diagnose prediction files with accuracy, calibration, classification, ordering, lexical, or semantic metrics. Accept generic JSON field mappings and BioProBench-compatible files.
2026-07-19