Skip to main content
在 Manus 中运行任何 Skill
一键导入

coding-agent-bench-runner

星标0
分支0
更新时间2026年6月28日 20:06

Benchmark coding agents against open-source models with CodingAgentBench to find which agent plus which model actually codes best for a given task profile. Takes a task suite (or helps define one) and a set of agent/model combinations, runs them, scores results on pass rate, quality, cost, and latency, and emits a ranked leaderboard with a recommendation. Use whenever the user wants to compare coding agents or models, asks 'which model should I use for coding', 'benchmark these agents', 'which open model is best with Claude Code / OpenClaw', 'is model X good enough for this task', or wants evidence-based model selection rather than a guess. Trigger for any agent-times-model evaluation.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

文件资源管理器
2 个文件
SKILL.md
readonly