multi-turn-preference-evaluation
Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.
Datos de origen
- Repositorio
- Dingxingdi/paper_fast_search_backup
- Última actividad en el origen
- 8 de abril de 2026 a las 15:14
- Idioma detectado de SKILL.md
- inglés
- Estrellas
- 0
- Forks
- 0
Opciones de instalación
De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.
Revisa los archivos de origen
Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.
Mostrando SKILL.md
- name
- multi-turn-preference-evaluation
- description
- Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.