Skip to main content
Run any Skill in Manus
with one click

skill-ab-eval

Stars0
Forks0
UpdatedMay 31, 2026 at 02:09

Empirically measure two things for any task or domain: (1) does loading a SKILL.md actually change behavior — with_skill vs without_skill (baseline) — and (2) which CLI agent harness does the task best — claude vs codex vs gemini vs agy vs an OpenAI-compatible API. Runs the same prompt across the matrix in fresh, isolated contexts (subagents or separate CLI processes), a judge grades every output, repeated over trials, and you get a skill-lift table plus a harness leaderboard. Use to validate a new skill before shipping, catch skills that do nothing (or hurt), compare CLIs on a task you give, or pick the best harness for a domain. Triggers: evaluate skill, test skill, does my skill work, skill A/B, measure skill lift, compare CLIs, which agent is best, claude vs codex vs gemini, harness eval, benchmark agents on a task.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly