with one click
audio-arena-bench
audio-arena-bench contains 2 collected skills from Design-Arena, with repository-level occupation coverage and site-owned skill detail pages.
Skills in this repository
Run repeated benchmark consistency studies, find turn-level pass/fail flips across runs of the same model, identify suspected data-spec or grader issues versus real model failures, and generate a polished HTML review. Use when asked to run N repeated evals, inspect unstable datapoints, separate likely benchmark/judge problems from model problems, or produce a consistency review report.
Analyze benchmark runs to identify dominant error modes per model, shared hard turns, grader or benchmark issues, and representative failing examples. Prefer judged runs, but fall back to transcript + benchmark-contract review when judged artifacts are missing or unreliable.