Use when the user wants to build an LLM-as-a-judge, evaluator, classifier, or guardrail for their AI application. Triggers on: evaluating AI outputs, judging response quality, classifying text, building guardrails, detecting hallucinations, checking grounding, prompt adherence, safety checks, content moderation, or any task where they need to automatically score/label/classify LLM outputs. Do NOT trigger for: general code questions, debugging, or tasks unrelated to evaluation/classification.
2026-06-17