一键导入
eval
eval 收录了来自 eforge-build 的 2 个 skills,并提供仓库级职业覆盖和站内 skill 详情页。
这个仓库中的 skills
Compare eforge eval runs of the same scenario across two or more backends. Reads eforge.log, result.json, and workspace artifacts to judge decision quality stage-by-stage — pipeline composer, planner, builder, tester, reviewer, review-fixer, evaluator — and writes a ranked analysis to the results directory. Use when the user asks to compare backends, rank backends, analyze which backend made better decisions, or review a specific eval run.
Compile a directory of per-scenario eval analyses into a top-level README.md suitable for publishing. Reads analyses/<dated-dir>/*.md, optionally re-aggregates from results/<timestamp>/ result.json files, and writes a factual, non-sensational summary following analyses/_TEMPLATE/. Use when the user asks to publish, compile, or finalize an eval analysis set, or points at an analyses dir containing per-scenario files but no aggregated README.md.