with one click
eval-pipeline
How the eval engine works: generate → grade → review → report
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
Menu
How the eval engine works: generate → grade → review → report
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
Based on SOC occupation classification
How to write comprehensive architectural proposals that drive alignment before code is written
Record final outcomes to history.md, not intermediate requests or reversed decisions
Tone enforcement patterns for external-facing community responses
Team initialization flow (Phase 1 proposal + Phase 2 creation)
Core conventions and patterns for this codebase
Expert guidance for authoring and maintaining .prompt.md evaluation files for hyoka. Covers frontmatter formats, file structure, filtering, and best practices.
| name | eval-pipeline |
| description | How the eval engine works: generate → grade → review → report |
| domain | architecture |
| confidence | high |
| source | hyoka/internal/eval/engine.go, architecture.md |
The evaluation pipeline is the core workflow of hyoka. It orchestrates AI agent code generation, multi-model grading, and report generation. Understanding the pipeline is essential for debugging eval failures and extending the engine.
--max-session-actions)reports/{run_id}/Workspace Isolation:
/tmp/hyoka-{run_id})Action Timeline Capture:
Async/Parallel Processing:
--workers flag)Error Handling:
hyoka/internal/eval/engine.gohyoka/internal/eval/copilot.gohyoka/internal/eval/action.gohyoka/internal/eval/workspace.gohyoka/internal/eval/proctracker.golog.Fatal in pipeline steps — return errors