Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
原文の言語: 英語
メニュー
SkillsMP は shihongDev/evalyn から 5 件の skill を収集しています。skill を開くとソースと詳細を確認できます。
収集済み skill 5 件中 5 件を表示しています。
Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
原文の言語: 英語
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
原文の言語: 英語
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
原文の言語: 英語
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
原文の言語: 英語
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
原文の言語: 英語