Skip to main content
Manusで任意のスキルを実行
ワンクリックで

agent-benchmark

スター10
フォーク2
更新日2026年6月4日 21:12

Run agent tasks across multiple LLMs in parallel tmux sessions, score results with a weighted rubric, track improvements over time in JSONL, and generate harness improvement recommendations. Use when comparing model performance, evaluating harness changes, or identifying agent failure patterns at scale.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
6 ファイル
SKILL.md
readonly