Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
لغة النص الأصلي: الإنجليزية
القائمة
جمع SkillsMP عدد ٥ من skills من shihongDev/evalyn. افتح أي skill لمراجعة مصدره وتفاصيله.
عرض ٥ من أصل ٥ skills مجمعة.
Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed
لغة النص الأصلي: الإنجليزية
Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn
لغة النص الأصلي: الإنجليزية
Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.
لغة النص الأصلي: الإنجليزية
Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator
لغة النص الأصلي: الإنجليزية
Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance
لغة النص الأصلي: الإنجليزية