Skip to main content

design-llm-evaluation-suite

Design a regression-oriented evaluation suite for an LLM, RAG pipeline, tool-using agent, or multi-agent workflow by defining behavioral cases, datasets, deterministic and model-graded oracles, trace or receipt requirements, stochastic thresholds, framework selection, CI tiers, and retained evidence. Use when deciding what LLM or agent evals to add, choosing an evaluation harness for Python, Rust, TypeScript, or an HTTP service, preparing a prompt, model, tool, or orchestration migration, or explicitly implementing an eval harness; remain read-only unless implementation is requested.

インストールへ移動

ソース情報

リポジトリ
wcygan/agent-skills
ソースの最終更新活動
2026年8月15日 23:13
検出された SKILL.md の言語
英語
スター
0
フォーク
0

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。