Skip to main content

agent-eval

Sterne14
Forks1
Aktualisiert10. Mai 2026 um 22:13

Evaluate agentic systems at component and end-to-end level, producing a structured report with metrics, datasets, and regression strategy. Use this skill whenever the user wants to evaluate, benchmark, or test any agentic or AI pipeline system, including RAG pipelines, multi-agent systems, LLM chains, retrieval systems, or any system with modelling components. Triggers include: "evaluate my agent", "benchmark my pipeline", "write evals for", "test my RAG", "how is my system performing", and "set up evals". Always use this skill when the user shares a codebase and asks how well it works.

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

SKILL.md
readonly