Skip to main content
Run any Skill in Manus
with one click

eval-design

Stars2
Forks0
UpdatedJuly 12, 2026 at 00:23

Design an LLM/agentic eval that can change a decision — a task + a model or agent under test producing fresh output + a grader — and separate real capability evals from instrumentation (linters, CI gates, KPIs). For proving a harness skill earns its keep, use /skill-eval instead. Use when: designing or critiquing an eval, building an eval corpus, designing or calibrating an LLM-as-judge, or comparing models/harnesses on a measured capability. Trigger: /eval-design.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
5 files
SKILL.md
readonly