| name | llm-benchmark-analyst |
| description | search and analyze llm benchmark results within a fixed benchmark universe, then produce evidence-based model strength and weakness reports or domain-leader summaries. use when comparing a model across benchmarks, ranking the best models by domain, explaining what a benchmark measures, checking predecessor-vs-current progress, or writing benchmark reports that must prioritize exact model version, evaluation date, benchmark variant, score semantics, sub-scores, and benchmark defect warnings. works with browser, web, and multimodal extraction for text, table, canvas, or image-only leaderboards. |
LLM Benchmark Analyst
Overview
Use this skill to research benchmark evidence and write structured reports about:
- a single model's strengths and weaknesses
- best models in a capability domain
- what a benchmark measures and how trustworthy it is
- predecessor vs current-model progress
Default to the user's language. Never invent scores, ranks, dates, benchmark variants, or missing table values.
Core constraints
- Restrict the benchmark universe to
references/benchmark-source.md. If a benchmark is not in that file, exclude it.
- Use
references/core-dimensions.md to collapse scattered benchmarks into a small set of report dimensions.
- Follow
references/search-playbook.md for routing, overlap expansion, evidence gathering, and comparison anchors.
- Follow
references/report-template.md for output structure.
- Apply
references/data-defect-warnings.md benchmark by benchmark, inline and again in the limitations section.
- Prefer official benchmark or benchmark-author pages. Use aggregators mainly to discover links and context.
- Record the evaluation mode exactly: benchmark version, split, difficulty, public/private, verified/original, with-tools/without-tools, pass@k, and any visible sub-score names.