Skip to main content
Manusで任意のスキルを実行
ワンクリックで

model-run-error-analysis

スター6
フォーク1
更新日2026年3月19日 04:55

Analyze benchmark runs to identify dominant error modes per model, shared hard turns, grader or benchmark issues, and representative failing examples. Prefer judged runs, but fall back to transcript + benchmark-contract review when judged artifacts are missing or unreliable.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
2 ファイル
SKILL.md
readonly