Compare two eval runs and report what changed. Reads both runs' events, transcripts, and produced artifacts. Writes a short markdown summary classifying differences as regression, improvement, or neutral.
원문 언어: 영어
메뉴
SkillsMP는 adam-s/agent-spec에서 8개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.
수집된 skill 8개 중 8개를 표시합니다.
Compare two eval runs and report what changed. Reads both runs' events, transcripts, and produced artifacts. Writes a short markdown summary classifying differences as regression, improvement, or neutral.
원문 언어: 영어
Generalized recursive iteration loop. Runs parallel sub-agents against a target, scores deterministically, diagnoses instruction gaps, applies fixes, and recurses until the stop condition is met or max depth is reached.
원문 언어: 영어
Run an evaluation against an eval with a specific config
원문 언어: 영어
Write a handoff document so a new chat can continue the work
원문 언어: 영어
Show evaluation results and comparisons
원문 언어: 영어
Test-driven development of a Hono/Bun WebSocket application. Read requirements, read tests, build server, verify, iterate until all tests pass.
원문 언어: 영어
Scaffold a new evaluation
원문 언어: 영어
Stop all agent-spec processes, clear ports, remove sandboxes, verify clean state.
원문 언어: 영어