Skip to main content
Run any Skill in Manus
with one click

toadapt-tutor-response-evaluation

Stars0
Forks0
UpdatedJuly 9, 2026 at 04:35

Pädagogische Qualitätsbewertung der LLM-Tutor-Antworten in ToAdapt nach der NAACL-2025-Taxonomie (Maurya et al., "Unifying AI Tutor Evaluation") - acht Dimensionen, LLM-as-Judge, Human-Validierung, Desirability-Quoten. Lade diese Skill, wenn du (a) messen willst, wie gut die Chat-Agenten oder Denkanstöße pädagogisch antworten ("verrät der Agent Lösungen?", "sind Antworten actionable?"), (b) scripts/evaluate_tutor_responses.py bedienen willst (Erhebung, Annotation-Workbook, Aggregation), (c) einen Agent-/Formative-Prompt geändert hast und den Pflicht-Regressionsnachweis brauchst, (d) ein TUTOR-Modell auswählen oder wechseln willst (OPENROUTER_MODEL) und Kandidaten vergleichen musst, oder (e) die Judge-Verlässlichkeit gegen menschliche Annotation validieren willst. Keywords - Tutor-Evaluation, NAACL, MRBench, Mistake Identification, Revealing of the Answer, Providing Guidance, Actionability, Tutor Tone, Humanlikeness, Desirability, Pedagogical Ability, LLM-as-Judge, Modellwahl.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly