Skip to main content
在 Manus 中运行任何 Skill
一键导入

simpo-training

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

概览

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

安装命令
npx skills add https://github.com/NousResearch/hermes-agent --skill simpo-training

复制此命令并粘贴到 Claude Code 中以安装该技能

星标178,912
分支30,651
更新时间2026年5月8日 21:27
文件资源管理器
4 个文件
SKILL.md
readonly