Skip to main content
Run any Skill in Manus
with one click

eval-harness

Stars4
Forks1
UpdatedJuly 9, 2026 at 09:49

Automated evaluation harness for measuring agent and LLM performance, preventing regressions, and enabling eval-driven development. Use when setting up systematic quality gates for AI-assisted workflows or when you need measurable pass/fail criteria for non-deterministic outputs. Distinguishes itself through pass@k/pass^k metrics, LLM-as-judge patterns, cost tracking per eval run, and CI/CD integration with coverage mapping.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly