Skip to main content
Run any Skill in Manus
with one click

eval-driven-dev

Stars8
Forks0
UpdatedMarch 30, 2026 at 22:28

Evaluation-driven development for Python LLM applications using the Microsoft Evaluations SDK (`azure-ai-evaluation`) and Microsoft Foundry. Use when: - Setting up evals, QA, or testing for the KB Agent or any LLM-calling code - Running local evaluations with LLM-as-judge evaluators backed by Foundry models - Publishing evaluation results to Microsoft Foundry for tracking and comparison - Building golden datasets for regression testing - Investigating LLM response quality failures - Benchmarking prompt changes

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly