Skip to main content
Exécutez n'importe quel Skill dans Manus
en un clic

skill-ab-eval

Étoiles0
Forks0
Mis à jour31 mai 2026 à 02:09

Empirically measure two things for any task or domain: (1) does loading a SKILL.md actually change behavior — with_skill vs without_skill (baseline) — and (2) which CLI agent harness does the task best — claude vs codex vs gemini vs agy vs an OpenAI-compatible API. Runs the same prompt across the matrix in fresh, isolated contexts (subagents or separate CLI processes), a judge grades every output, repeated over trials, and you get a skill-lift table plus a harness leaderboard. Use to validate a new skill before shipping, catch skills that do nothing (or hurt), compare CLIs on a task you give, or pick the best harness for a domain. Triggers: evaluate skill, test skill, does my skill work, skill A/B, measure skill lift, compare CLIs, which agent is best, claude vs codex vs gemini, harness eval, benchmark agents on a task.

Installation

Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.

SKILL.md
readonly