Skip to main content
在 Manus 中运行任何 Skill
一键导入

agent-eval-no-ground-truth

星标1
分支0
更新时间2026年5月11日 12:12

Design and reason about evaluation frameworks for production AI agents when no labeled dataset or ground truth exists. Use this skill whenever someone asks how to evaluate an agent, measure agent quality, build an eval signal stack, assess agent correctness, or trust agent outputs in production — especially when benchmarks don't exist for the domain, outputs are non-deterministic, or multiple correct answers are possible. Trigger on phrases like "how do I know if my agent is right", "eval for my agent", "measure agent accuracy", "agent quality signals", "no ground truth", "production agent metrics", or any request to evaluate/grade an agentic system's outputs in the real world.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly