Skip to main content

run-assert-eval

Run an ASSERT evaluation against a described risk. Use when the user wants to evaluate, test, or check an AI agent, LLM app, or model against requirements/policies (e.g. "evaluate my agent for budget violations", "test that the support bot never gives legal advice"). Risks come either from Clarity — recommended, driving the real Clarity MCP tools (run_clarity) in-IDE to discover failure modes the user has not considered — or directly from the user as a description, PRD, design doc, threat model, red-team finding, or risk assessment. Then researches how that risk has been evaluated in the literature and turns it into an evidence-backed, cited config per selected risk at examples/<domain>/<risk>/eval_config.yaml, gets it approved, runs the pipeline, and reports pass/violation rates with trace-cited failure examples.

Zur Installation springen

Quellinformationen

Repository
responsibleai/ASSERT
Letzte Quellaktivität
2. September 2026 um 17:52
Erkannte Sprache von SKILL.md
Englisch
Sterne
235
Forks
36

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.