Skip to main content
Manusで任意のスキルを実行
ワンクリックで

evals-design

スター25
フォーク0
更新日2026年4月13日 07:22

Design and review evaluation suites for LLM and agent systems. Use when the task involves creating, extending, or reviewing eval datasets, rubrics, judge configs, benchmark scripts, or shipping-readiness scorecards. Triggers: "eval suite", "golden dataset", "rubric", "judge model", "benchmark", "scorecard". Negative triggers: generic unit/integration testing with no LLM or agent component.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
3 ファイル
SKILL.md
readonly