Skip to main content
Jeden Skill in Manus ausführen
mit einem Klick

science-evals

Sterne0
Forks0
Aktualisiert4. Juli 2026 um 07:58

Prepare, record, grade, validate, and compare reproducible scientific-agent benchmark runs across Codex, Claude, or other systems. Use when measuring evidence traceability, uncertainty, protocol discipline, safety gates, reproducibility, loop decisions, regression behavior, or claims of scientific-agent parity on the same tasks and resource constraints.

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

Datei-Explorer
6 Dateien
SKILL.md
readonly