Skip to main content

danielrosehill/Claude-Eval-Runner-Plugin

SkillsMP는 danielrosehill/Claude-Eval-Runner-Plugin에서 6개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.

최근 기록된 소스 활동
SkillsMP 카탈로그 업데이트
수집된 skills
6
GitHub 스타
0
GitHub 포크
0

이 저장소의 skills

수집된 skill 6개 중 6개를 표시합니다.

직업 분류
소프트웨어 품질 보증 분석가·테스터
설명

Design a custom eval from scratch, or remix an existing benchmark. Use when the user wants to define the eval itself — task framing, dataset composition, scoring rubric, and reporting format — rather than simply wiring up a framework. Produces a fully…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Provision a new eval-runner workspace on disk. Use when the user wants to start a new evaluation project — scaffolds evals/, datasets/, results/, and docs/ directories, personalises CLAUDE.md, and (by default) creates a GitHub repo.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Publish an eval dataset to Hugging Face Hub (or GitHub as a fallback). Use when the user wants to share the inputs/labels used by an eval — with a dataset card, licensing, splits, and a content hash so downstream runs can verify integrity.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Publish an eval (definition + results) so others can reproduce it. Use when the user wants to share an eval publicly — as a GitHub repo, Hugging Face space, or a standalone writeup. Produces a clean, self-contained bundle with README, task spec, rubric,…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 품질 보증 분석가·테스터
설명

Execute an eval defined in the current workspace and capture results with full metadata. Use when the user wants to actually run an eval (one or many SUTs), collect scored outputs under results/, and produce a run manifest so findings are reproducible and…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Set up an evaluation in the current workspace. Use when the user wants to scaffold a single eval — choosing an existing framework (DeepEval, Inspect AI, OpenAI Evals, lm-evaluation-harness, LightEval, OLMES, Promptfoo, etc.), adapting an existing benchmark,…

원문 언어: 영어

업데이트
수집된 skill 6개 중 6개를 표시합니다.