Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
Langue du texte source : anglais
Menu
SkillsMP a collecté 5 skills depuis shihongDev/evalyn. Ouvrez un skill pour examiner sa source et ses détails.
Affichage de 5 skills collectés sur 5.
Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
Langue du texte source : anglais
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
Langue du texte source : anglais
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
Langue du texte source : anglais
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
Langue du texte source : anglais
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
Langue du texte source : anglais