Skip to main content

agent-evals

Étoiles10
Forks1
Mis à jour30 juillet 2026 à 08:19

Builds evaluations for your own Copilot skills, prompts, and agents so changes are measured, not vibed. Starts from the native VS Code tooling (Chat Customizations Evaluations analysis and the Waza eval runner) and falls back to a bundled run-evals.ps1 harness. Covers capability vs regression eval sets, grader types (deterministic / LLM-as-judge / human), pass@k vs pass^k, and eval-driven development. Rule of thumb: start from 20-50 real failures, not synthetic prompts. USE FOR: evaluate agent, evaluate skill, evaluate prompt, eval harness, build evals, Waza, analyze-prompt, Chat Customizations Evaluations, LLM-as-judge, grader, capability eval, regression eval, pass@k, pass^k, eval-driven development, does my skill work, test a prompt, measure agent quality, run-evals. DO NOT USE FOR: authoring the skill itself (use skill-creator), MCP server eval questions (use mcp-builder Phase 4), unit-testing PowerShell code (use pester-patterns), security review of an agent (use agent-security-review).

Installation

Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.

Explorateur de fichiers
3 fichiers
SKILL.md
readonly