Skip to main content
Manus에서 모든 스킬 실행
원클릭으로

skill-ab-eval

스타0
포크0
업데이트2026년 5월 31일 02:09

Empirically measure two things for any task or domain: (1) does loading a SKILL.md actually change behavior — with_skill vs without_skill (baseline) — and (2) which CLI agent harness does the task best — claude vs codex vs gemini vs agy vs an OpenAI-compatible API. Runs the same prompt across the matrix in fresh, isolated contexts (subagents or separate CLI processes), a judge grades every output, repeated over trials, and you get a skill-lift table plus a harness leaderboard. Use to validate a new skill before shipping, catch skills that do nothing (or hurt), compare CLIs on a task you give, or pick the best harness for a domain. Triggers: evaluate skill, test skill, does my skill work, skill A/B, measure skill lift, compare CLIs, which agent is best, claude vs codex vs gemini, harness eval, benchmark agents on a task.

설치

Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.

SKILL.md
readonly