Skip to main content

danielrosehill/Claude-Eval-Runner-Plugin

SkillsMP has collected 6 skills from danielrosehill/Claude-Eval-Runner-Plugin. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
6
GitHub stars
0
GitHub forks
0

Skills in this repository

Showing 6 of 6 collected skills.

occupation
Software Quality Assurance Analysts & Testers
description

Design a custom eval from scratch, or remix an existing benchmark. Use when the user wants to define the eval itself — task framing, dataset composition, scoring rubric, and reporting format — rather than simply wiring up a framework. Produces a fully…

updated
occupation
Software Developers
description

Provision a new eval-runner workspace on disk. Use when the user wants to start a new evaluation project — scaffolds evals/, datasets/, results/, and docs/ directories, personalises CLAUDE.md, and (by default) creates a GitHub repo.

updated
occupation
Data Scientists
description

Publish an eval dataset to Hugging Face Hub (or GitHub as a fallback). Use when the user wants to share the inputs/labels used by an eval — with a dataset card, licensing, splits, and a content hash so downstream runs can verify integrity.

updated
occupation
Software Developers
description

Publish an eval (definition + results) so others can reproduce it. Use when the user wants to share an eval publicly — as a GitHub repo, Hugging Face space, or a standalone writeup. Produces a clean, self-contained bundle with README, task spec, rubric,…

updated
occupation
Software Quality Assurance Analysts & Testers
description

Execute an eval defined in the current workspace and capture results with full metadata. Use when the user wants to actually run an eval (one or many SUTs), collect scored outputs under results/, and produce a run manifest so findings are reproducible and…

updated
occupation
Software Developers
description

Set up an evaluation in the current workspace. Use when the user wants to scaffold a single eval — choosing an existing framework (DeepEval, Inspect AI, OpenAI Evals, lm-evaluation-harness, LightEval, OLMES, Promptfoo, etc.), adapting an existing benchmark,…

updated
Showing 6 of 6 collected skills.