Skip to main content

llm-evaluation

Implement comprehensive evaluation strategies for LLM applications — automated metrics (BLEU, ROUGE, BERTScore, RAG metrics), A/B testing with statistical rigor, regression detection, and benchmarking. Use when measuring agent quality, comparing models or prompts, or building eval pipelines for LangGraph or Google ADK agents.

Jump to install

Source facts

Repository
kumaran-is/claude-code-onboarding
Last source activity
March 15, 2026 at 23:18
Detected SKILL.md language
English
Stars
34
Forks
21

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.