Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
原文语言:英语
菜单
SkillsMP 已收集 shihongDev/evalyn 中的 5 个 Skill。打开任一 Skill 可查看来源和详情。
已展示 5 / 5 个已收集 Skill。
Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
原文语言:英语
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
原文语言:英语
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
原文语言:英语
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
原文语言:英语
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
原文语言:英语