职业分类
软件质量保证分析师与测试员
描述
LLM judge agent for grading AI agent eval transcripts. Checks deterministic assertions (transcript_contains, tool_called) and uses LLM reasoning for behavioral assertions (llm_judge). Returns structured JSON grades.
原文语言:英语
更新
菜单
SkillsMP 已收集 awslabs/agent-builder-toolkit-aws-transform 中的 2 个 Skill。打开任一 Skill 可查看来源和详情。
已展示 2 / 2 个已收集 Skill。
LLM judge agent for grading AI agent eval transcripts. Checks deterministic assertions (transcript_contains, tool_called) and uses LLM reasoning for behavioral assertions (llm_judge). Returns structured JSON grades.
原文语言:英语
Simulated human agent for eval scenarios. Interacts with the agent under test via the ACP bridge, following the scenario goal and guidance to respond to agent questions, approve tool calls, and drive the multi-turn flow to completion.
原文语言:英语