Skip to main content

ai-evals-course/evals-skills

SkillsMP has collected 4 skills from ai-evals-course/evals-skills. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
4
GitHub stars
451
GitHub forks
36

Skills in this repository

Showing 4 of 4 collected skills.

occupation
Software Quality Assurance Analysts & Testers
description

Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for something no other skill in this plugin matches. Do NOT use when a more specific skill in this plugin already matches; load that skill directly.

updated
occupation
Data Scientists
description

Run error analysis on a dataset. Build a review UI, select diverse samples, monitor annotations, and organize failure modes.

updated
occupation
Web Developers
description

Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to build an annotation tool, review traces, or collect human labels.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT…

updated
Showing 4 of 4 collected skills.