Skip to main content

ai-evaluation-harness

MANUAL-ONLY; never auto-invoke. Design and run the evaluation harness for an LLM feature — a versioned dataset (representative, adversarial/red-team, and regression cases), graders per dimension (task quality, schema adherence, safety/refusal, groundedness/hallucination, injection resistance, latency, cost), pass/fail thresholds, and a CI gate that blocks a prompt/model/retrieval/provider change on regression. Absorbs the AI security test harness: injection, jailbreak, data-exfiltration, and tool-misuse suites are first-class dimensions. Running it spends tokens and money, so it is manual-only. Use when building AI evals, a golden dataset, regression gates for AI changes, or encoding red-team cases from ai-threat-modeler / prompt-injection-defender / agent-tool-safety-guard. Do NOT use for the code-change test pyramid (regression-suite-curator / qa-automation-architect), the threat model itself (ai-threat-modeler), or live production monitoring (observability-operator).

Aller à l'installation

Informations de source

Dépôt
ModernNomad-98/Project-Aegis
Dernière activité de la source
18 juillet 2026 à 12:46
Langue détectée de SKILL.md
anglais
Étoiles
3
Forks
0

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.