Skip to main content
在 Manus 中运行任何 Skill
一键导入

evals

星标3
分支0
更新时间2026年7月8日 02:33

Comprehensive AI agent evaluation framework with three grader types (code-based: deterministic; model-based: LLM rubric; human: gold standard) and pass@k / pass^k scoring. Evaluates agent transcripts, tool-call sequences, and multi-turn conversations — not just single outputs. Capability evals (~70% target) and regression evals (~99% target). Workflows: RunEval, CompareModels, ComparePrompts, CreateJudge, CreateUseCase, RunScenario, CreateScenario, ViewResults. Integrates with ALGORITHM ISC rows for automated verification. Domain patterns pre-configured for coding, conversational, research, computer-use agents. USE WHEN eval, evaluate, benchmark, regression test, compare models, create judge, test agent, pass@k, scenario simulation. NOT FOR scientific method framing (use Science).

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

文件资源管理器
44 个文件
SKILL.md
readonly