Skip to main content
Manusで任意のスキルを実行
ワンクリックで

llm-eval

スター0
フォーク0
更新日2026年7月7日 19:51

Author LLM/RAG/agent evaluation suites in DeepEval that prove a feature is correct with gated numbers, not vibes. Use when asked to "eval an LLM", "test a prompt", "measure RAG quality", "check for hallucination", "score answer relevancy", "verify tool calls", or gate a release on model output quality. Ships the metric definitions baked in — faithfulness, answer relevancy, contextual precision/recall/relevancy, hallucination, tool correctness, G-Eval, bias, toxicity — with the inputs each needs, the score formula, and which direction passes. Enforces guardrails against unpinned judge models, exact-matching non-deterministic output, contaminated goldens, single-run scores, and metric-picking that ignores the real failure mode. See `reference.md` for runnable DeepEval suites and CI wiring, `tooling.md` for install/versions and alternatives.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

ファイルエクスプローラー
3 ファイル
SKILL.md
readonly