Skip to main content

agent-eval-workflow

This skill should be used when the user wants to evaluate an AI agent end-to-end: scaffold an evaluation, design metrics that test a real hypothesis, make an agent measurable, audit generated eval config, read evaluation results, or run an improvement ("hill climbing") loop. Covers evaluation methodology, metric design, dataset coverage, reading deterministic vs LLM-judged metrics, and the traps that make eval runs silently measure nothing. Use alongside the tool-specific skills (agents-cli-eval, adk-eval-guide) — those cover commands and schemas, this covers the process and judgement.

설치로 이동

소스 정보

저장소
GoogleCloudPlatform/professional-services
최근 소스 활동
2026년 8월 11일 15:01
감지된 SKILL.md 언어
영어
스타
3,070
포크
1,470

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.