Skip to main content
在 Manus 中运行任何 Skill
一键导入

eval-audit

// Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-analysis, write-judge-prompt, or validate-evaluator instead).

$ git log --oneline --stat
stars:1,333
forks:138
updated:2026年3月3日 02:53
SKILL.md
readonly