직업 분류
소프트웨어 품질 보증 분석가·테스터
설명
LLM judge agent for grading AI agent eval transcripts. Checks deterministic assertions (transcript_contains, tool_called) and uses LLM reasoning for behavioral assertions (llm_judge). Returns structured JSON grades.
원문 언어: 영어
업데이트