Skip to main content

eval-harness

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.

Informations de source

Dépôt
a5c-ai/babysitter
Dernière activité de la source
1 juin 2026 à 07:46
Langue détectée de SKILL.md
anglais
Étoiles
1 813
Forks
111

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
2 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
eval-harness
description
Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.
allowed-tools
Read, Write, Edit, Bash, Grep, Glob
graph
{"domains":["domain:software-engineering"],"skillAreas":["skill-area:agentic-loops","skill-area:orchestration-loop"],"workflows":["workflow:feature-development"],"topics":["topic:developer-experience"],"roles":["role:tech-lead","role:backend-engineer"]}
- Define test cases with known-correct outputs - Run agent against each test case - Score: accuracy, completeness, relevance - Compare against baseline performance - Track performance over time ### 2. Skill Quality Testing - Verify skill instructions produce expected outcomes - Test edge cases and boundary conditions - Measure consistency across multiple runs - Check for harmful or incorrect outputs - Validate against ground truth ### 3. Regression Suite - Collection of previously-passing test cases - Run after any agent/skill modification - Flag regressions with before/after comparison - Maintain pass rate threshold (>= 95%) ### 4. Process Verification - End-to-end process execution with known inputs - Verify each phase produces expected outputs - Check task ordering and dependency satisfaction - Measure total execution time ## Quality Scoring ### Accuracy Score (0-100) - Correctness of output vs expected - Partial credit for partially correct outputs - Penalty for hallucinated or fabricated content ### Completeness Score (0-100) - Coverage of required output elements - Missing sections flagged and scored - Bonus for useful additional context ### Consistency Score (0-100) - Run same input 3 times - Compare outputs for semantic similarity - Flag inconsistencies ### Composite Score - (accuracy * 0.4 + completeness * 0.3 + consistency * 0.3) - Threshold: 80 to pass ## When to Use - After creating new agents or skills - After modifying existing agents or skills - Periodic quality audits - Before promoting skills to production ## Agents Used - Used by process-level evaluation orchestrators - No specific agent dependency (evaluates other agents)
Voir sur GitHub