Skip to main content
Run any Skill in Manus
with one click

dev-eval

Stars1
Forks0
UpdatedMay 4, 2026 at 03:21

Run comparative development evaluations: assign the same task to agent-worker team AND claude-code (team subagents), then have an evaluator agent score both and a reviewer agent verify the evaluation. Fully orchestrated โ€” you coordinate all phases, agents notify via channel/callback when done. Use this skill when the user wants to benchmark agent-worker vs claude-code, run a dev eval, compare agent performance, or prove that the workspace team outperforms solo agents. Trigger on phrases like 'dev eval', 'run eval', 'compare performance', 'benchmark', 'ๅฏนๆฏ”่ฏ„ไผฐ', 'ๅผ€ๅ‘่ฏ„ไผฐ', '่ท‘่ฏ„ไผฐ', 'eval iteration', '/dev-eval'.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly