Compare two eval runs and report what changed. Reads both runs' events, transcripts, and produced artifacts. Writes a short markdown summary classifying differences as regression, improvement, or neutral.
لغة النص الأصلي: الإنجليزية
القائمة
جمع SkillsMP عدد ٨ من skills من adam-s/agent-spec. افتح أي skill لمراجعة مصدره وتفاصيله.
عرض ٨ من أصل ٨ skills مجمعة.
Compare two eval runs and report what changed. Reads both runs' events, transcripts, and produced artifacts. Writes a short markdown summary classifying differences as regression, improvement, or neutral.
لغة النص الأصلي: الإنجليزية
Generalized recursive iteration loop. Runs parallel sub-agents against a target, scores deterministically, diagnoses instruction gaps, applies fixes, and recurses until the stop condition is met or max depth is reached.
لغة النص الأصلي: الإنجليزية
Run an evaluation against an eval with a specific config
لغة النص الأصلي: الإنجليزية
Write a handoff document so a new chat can continue the work
لغة النص الأصلي: الإنجليزية
Show evaluation results and comparisons
لغة النص الأصلي: الإنجليزية
Test-driven development of a Hono/Bun WebSocket application. Read requirements, read tests, build server, verify, iterate until all tests pass.
لغة النص الأصلي: الإنجليزية
Scaffold a new evaluation
لغة النص الأصلي: الإنجليزية
Stop all agent-spec processes, clear ports, remove sandboxes, verify clean state.
لغة النص الأصلي: الإنجليزية