Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
shihongDev/evalyn
SkillsMP has collected 5 skills from shihongDev/evalyn. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 5
- GitHub stars
- 257
- GitHub forks
- 26
Skills in this repository
1 occupation categories · 100% classified
Showing 5 of 5 collected skills.
skill
occupation
description
updated
occupation
Software Quality Assurance Analysts & Testers
description
updated
occupation
Software Quality Assurance Analysts & Testers
description
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
updated
occupation
Software Quality Assurance Analysts & Testers
description
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
updated
occupation
Software Quality Assurance Analysts & Testers
description
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
updated
occupation
Software Quality Assurance Analysts & Testers
description
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
updated
Showing 5 of 5 collected skills.