Skip to main content

llm-evaluation

Implement comprehensive evaluation strategies for LLM applications — automated metrics (BLEU, ROUGE, BERTScore, RAG metrics), A/B testing with statistical rigor, regression detection, and benchmarking. Use when measuring agent quality, comparing models or prompts, or building eval pipelines for LangGraph or Google ADK agents.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
kumaran-is/claude-code-onboarding
آخر نشاط في المصدر
١٥ مارس ٢٠٢٦ في ٢٣:١٨
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣٤
التفرعات
٢١

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.