multi-turn-preference-evaluation
Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.
معلومات المصدر
- المستودع
- Dingxingdi/paper_fast_search_backup
- آخر نشاط في المصدر
- ٨ أبريل ٢٠٢٦ في ١٥:١٤
- لغة SKILL.md المكتشفة
- الإنجليزية
- النجوم
- ٠
- التفرعات
- ٠
خيارات التثبيت
يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.
مراجعة ملفات المصدر
اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.
عرض SKILL.md
- name
- multi-turn-preference-evaluation
- description
- Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.