Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
Idioma do texto original: inglês
Menu
O SkillsMP coletou 5 skills de shihongDev/evalyn. Abra uma skill para revisar a origem e os detalhes.
Mostrando 5 de 5 skills coletadas.
Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
Idioma do texto original: inglês
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
Idioma do texto original: inglês
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
Idioma do texto original: inglês
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
Idioma do texto original: inglês
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
Idioma do texto original: inglês