google-agents-cli-eval
5-stage Agent Quality Flywheel evaluation skill for ADK agents using the agents-cli toolchain.
소스 정보
- 저장소
- palladius/adk-sre-benjamin
- 최근 소스 활동
- 2026년 8월 25일 08:58
- 감지된 SKILL.md 언어
- 영어
- 스타
- 2
- 포크
- 0
설치 방법
기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.
소스 파일 검토
설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.
SKILL.md 표시 중
SKILL.md
소스 지침 · 읽기 전용 미리보기- name
- google-agents-cli-eval
- description
- 5-stage Agent Quality Flywheel evaluation skill for ADK agents using the agents-cli toolchain.
- metadata
- {"version":"1.0.0"}
# Google Agents CLI Eval Skill
This skill encodes the 5-stage Agent Quality Flywheel for ADK agents built with the `agents-cli` toolchain, as detailed in the Google Developers Blog ("Driving the Agent Quality Flywheel from Your Coding Agent").
## The 5 Flywheel Stages
1. **Prepare Data**:
- Synthesize synthetic scenarios via `agents-cli eval dataset synthesize` or compile traces from OTel logs / hand-crafted test cases.
- Example:
```bash
agents-cli eval dataset synthesize -n 5 --max-turns 8 --model gemini-2.5-flash \
--instruction "$(cat eval_instruction.txt)" \
-o traces_dataset.json
```
2. **Run Inference**:
- Execute agent models over the evaluation dataset to produce conversation traces (skip if traces already exist).
3. **Grade**:
- Score traces using adaptive AutoRaters (`multi_turn_task_success`, `multi_turn_trajectory_quality`) or custom categorical rubrics.
- Example:
```bash
agents-cli eval grade --traces traces_merged.json --config eval_config.yaml
```
4. **Analyze Failures**:
- Read rubric verdicts to determine root causes. For $\ge 10$ failures, trigger Automatic Loss Analysis.
5. **Optimize & Iterate**:
- Propose targeted prompt or tool routing fixes, re-run grading, and compare against baseline metrics.
GitHub에서 보기