Skip to main content
Manusで任意のスキルを実行
ワンクリックで

evals

スター3
フォーク0
更新日2026年7月8日 02:33

Comprehensive AI agent evaluation framework with three grader types (code-based: deterministic; model-based: LLM rubric; human: gold standard) and pass@k / pass^k scoring. Evaluates agent transcripts, tool-call sequences, and multi-turn conversations — not just single outputs. Capability evals (~70% target) and regression evals (~99% target). Workflows: RunEval, CompareModels, ComparePrompts, CreateJudge, CreateUseCase, RunScenario, CreateScenario, ViewResults. Integrates with ALGORITHM ISC rows for automated verification. Domain patterns pre-configured for coding, conversational, research, computer-use agents. USE WHEN eval, evaluate, benchmark, regression test, compare models, create judge, test agent, pass@k, scenario simulation. NOT FOR scientific method framing (use Science).

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
44 ファイル
SKILL.md
readonly