المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف
LLM judge agent for grading AI agent eval transcripts. Checks deterministic assertions (transcript_contains, tool_called) and uses LLM reasoning for behavioral assertions (llm_judge). Returns structured JSON grades.
لغة النص الأصلي: الإنجليزية
آخر تحديث