Guidance for agents working in codebases that depend on the `themis-eval` package or import `themis`. Use whenever a task mentions `evaluate(...)`, `Experiment(...)`, `Experiment.from_config(...)`, `themis` CLI commands, `themis.catalog`, shipped benchmarks, builtin ids such as `builtin/exact_match`, `run_id`, replay/resume, stores, adapters, prompt specs, or custom generators/parsers/reducers/metrics. Also use it when the user asks about benchmark wiring, `quick-eval benchmark`, reusable parser or metric components, or code-execution sandbox backends, even if they do not explicitly ask for "Themis docs."
2026-04-05