multi-turn-preference-evaluation
Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.
Source facts
- Repository
- Dingxingdi/paper_fast_search_backup
- Last source activity
- April 8, 2026 at 15:14
- Detected SKILL.md language
- English
- Stars
- 0
- Forks
- 0
Install options
The review-first prompt is selected by default. You can switch to a direct command or download a local copy.
Review the source files
Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.
Showing SKILL.md
- name
- multi-turn-preference-evaluation
- description
- Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.