Skip to main content

ai-evaluation-harness

MANUAL-ONLY; never auto-invoke. Design and run the evaluation harness for an LLM feature — a versioned dataset (representative, adversarial/red-team, and regression cases), graders per dimension (task quality, schema adherence, safety/refusal, groundedness/hallucination, injection resistance, latency, cost), pass/fail thresholds, and a CI gate that blocks a prompt/model/retrieval/provider change on regression. Absorbs the AI security test harness: injection, jailbreak, data-exfiltration, and tool-misuse suites are first-class dimensions. Running it spends tokens and money, so it is manual-only. Use when building AI evals, a golden dataset, regression gates for AI changes, or encoding red-team cases from ai-threat-modeler / prompt-injection-defender / agent-tool-safety-guard. Do NOT use for the code-change test pyramid (regression-suite-curator / qa-automation-architect), the threat model itself (ai-threat-modeler), or live production monitoring (observability-operator).

الانتقال إلى التثبيت

معلومات المصدر

المستودع
ModernNomad-98/Project-Aegis
آخر نشاط في المصدر
١٨ يوليو ٢٠٢٦ في ١٢:٤٦
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.