Skip to main content

ai-eval-efficiency-audit

スター4
フォーク0
更新日2026年7月18日 15:25

Audit AI evaluation infrastructure for wasted compute — full benchmark suites re-run on unchanged cases, LLM-as-judge grading without caching, oversized judge models, redundant eval passes per commit, and missing result reuse. Use this skill whenever the user shares eval harness configs or CI eval steps, complains that evals are slow or expensive, mentions LLM-as-judge costs, or runs benchmark suites on every change. Part of Lean Agentic AI Skills; emits lean-findings.json.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
2 ファイル
SKILL.md
readonly