nlp-preference-optimization
Set up and run NLP preference optimization experiments (DPO, SimPO, CPO). Use this skill when configuring trainer environments, installing dependencies (torch, trl, transformers, peft), or running preference optimization unit tests. Triggers on: DPO, SimPO, preference optimization, trl trainer, RLHF, alignment training.
Source facts
- Repository
- cxcscmu/SkillLearnBench
- Last source activity
- April 24, 2026 at 05:14
- Detected SKILL.md language
- English
- Stars
- 77
- Forks
- 4
Install options
The review-first prompt is selected by default. You can switch to a direct command or download a local copy.
Review the source files
Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.