Skip to main content

agent-eval

Stars14
Forks1
UpdatedMay 10, 2026 at 22:13

Evaluate agentic systems at component and end-to-end level, producing a structured report with metrics, datasets, and regression strategy. Use this skill whenever the user wants to evaluate, benchmark, or test any agentic or AI pipeline system, including RAG pipelines, multi-agent systems, LLM chains, retrieval systems, or any system with modelling components. Triggers include: "evaluate my agent", "benchmark my pipeline", "write evals for", "test my RAG", "how is my system performing", and "set up evals". Always use this skill when the user shares a codebase and asks how well it works.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly