Skip to main content

google-agents-cli-eval

5-stage Agent Quality Flywheel evaluation skill for ADK agents using the agents-cli toolchain.

소스 정보

저장소
palladius/adk-sre-benjamin
최근 소스 활동
2026년 8월 25일 08:58
감지된 SKILL.md 언어
영어
스타
2
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
google-agents-cli-eval
description
5-stage Agent Quality Flywheel evaluation skill for ADK agents using the agents-cli toolchain.
metadata
{"version":"1.0.0"}
# Google Agents CLI Eval Skill This skill encodes the 5-stage Agent Quality Flywheel for ADK agents built with the `agents-cli` toolchain, as detailed in the Google Developers Blog ("Driving the Agent Quality Flywheel from Your Coding Agent"). ## The 5 Flywheel Stages 1. **Prepare Data**: - Synthesize synthetic scenarios via `agents-cli eval dataset synthesize` or compile traces from OTel logs / hand-crafted test cases. - Example: ```bash agents-cli eval dataset synthesize -n 5 --max-turns 8 --model gemini-2.5-flash \ --instruction "$(cat eval_instruction.txt)" \ -o traces_dataset.json ``` 2. **Run Inference**: - Execute agent models over the evaluation dataset to produce conversation traces (skip if traces already exist). 3. **Grade**: - Score traces using adaptive AutoRaters (`multi_turn_task_success`, `multi_turn_trajectory_quality`) or custom categorical rubrics. - Example: ```bash agents-cli eval grade --traces traces_merged.json --config eval_config.yaml ``` 4. **Analyze Failures**: - Read rubric verdicts to determine root causes. For $\ge 10$ failures, trigger Automatic Loss Analysis. 5. **Optimize & Iterate**: - Propose targeted prompt or tool routing fixes, re-run grading, and compare against baseline metrics.
GitHub에서 보기