Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
원문 언어: 영어
메뉴
SkillsMP는 shihongDev/evalyn에서 5개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.
수집된 skill 5개 중 5개를 표시합니다.
Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
원문 언어: 영어
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
원문 언어: 영어
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
원문 언어: 영어
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
원문 언어: 영어
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
원문 언어: 영어