Skip to main content

ai-eval-efficiency-audit

Stars4
Forks0
UpdatedJuly 18, 2026 at 15:25

Audit AI evaluation infrastructure for wasted compute โ€” full benchmark suites re-run on unchanged cases, LLM-as-judge grading without caching, oversized judge models, redundant eval passes per commit, and missing result reuse. Use this skill whenever the user shares eval harness configs or CI eval steps, complains that evals are slow or expensive, mentions LLM-as-judge costs, or runs benchmark suites on every change. Part of Lean Agentic AI Skills; emits lean-findings.json.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
2 files
SKILL.md
readonly