Skip to main content
Run any Skill in Manus
with one click

evals-design

Stars25
Forks0
UpdatedApril 13, 2026 at 07:22

Design and review evaluation suites for LLM and agent systems. Use when the task involves creating, extending, or reviewing eval datasets, rubrics, judge configs, benchmark scripts, or shipping-readiness scorecards. Triggers: "eval suite", "golden dataset", "rubric", "judge model", "benchmark", "scorecard". Negative triggers: generic unit/integration testing with no LLM or agent component.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
3 files
SKILL.md
readonly