Skip to main content

dpo

Direct Preference Optimization for learning from preference pairs. Covers DPOTrainer, preference dataset preparation, implicit reward modeling, and beta tuning for stable preference learning without explicit reward models. Includes thinking quality patterns.

Jump to install

Source facts

Repository
majiayu000/claude-skill-registry
Last source activity
June 23, 2026 at 12:15
Detected SKILL.md language
English
Stars
543
Forks
85

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.