用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/jb-612/mortgage_concierge --skill eval-baseline命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Scaffold a new ADK function tool with TDD lifecycle and eval case addition
Adversarial testing loop inspired by autoresearch. Generate adversarial cases, run against agent, evaluate, accumulate findings.
Three-pass code review checking prerequisites, code quality, and test/eval coverage
基于 SOC 职业分类
正在显示 SKILL.md
| name | eval-baseline |
| description | Run ADK evaluations and capture current scores as a baseline before making changes |
| argument-hint | [workitem-path] |
| allowed-tools | Read, Write, Bash, Glob |
Capture eval scores for $ARGUMENTS.
ls tests/eval/data/*.evalset.jsonadk eval mortgage_concierge tests/eval/data/<file> \
--config_file_path=tests/eval/data/test_config.json \
--print_detailed_results
tool_trajectory_avg_scoreresponse_match_score.workitems/<current>/eval-baseline.json:
{
"timestamp": "<ISO-8601>",
"evalsets": {
"<evalset-name>": {
"tool_trajectory_avg_score": 0.85,
"response_match_score": 0.72
}
},
"thresholds": {
"tool_trajectory_avg_score": 0.8,
"response_match_score": 0.7
}
}
When run AFTER changes, compare against existing baseline: