Skip to main content

agent-evals

النجوم١٠
التفرعات١
آخر تحديث٣٠ يوليو ٢٠٢٦ في ٠٨:١٩

Builds evaluations for your own Copilot skills, prompts, and agents so changes are measured, not vibed. Starts from the native VS Code tooling (Chat Customizations Evaluations analysis and the Waza eval runner) and falls back to a bundled run-evals.ps1 harness. Covers capability vs regression eval sets, grader types (deterministic / LLM-as-judge / human), pass@k vs pass^k, and eval-driven development. Rule of thumb: start from 20-50 real failures, not synthetic prompts. USE FOR: evaluate agent, evaluate skill, evaluate prompt, eval harness, build evals, Waza, analyze-prompt, Chat Customizations Evaluations, LLM-as-judge, grader, capability eval, regression eval, pass@k, pass^k, eval-driven development, does my skill work, test a prompt, measure agent quality, run-evals. DO NOT USE FOR: authoring the skill itself (use skill-creator), MCP server eval questions (use mcp-builder Phase 4), unit-testing PowerShell code (use pester-patterns), security review of an agent (use agent-security-review).

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
3 ملفات
SKILL.md
readonly