用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/cxcscmu/SkillLearnBench --skill simpo-loss命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
基于 SOC 职业分类
正在显示 SKILL.md
| name | simpo-loss |
| description | SimPO (Simple Preference Optimization) loss computation for LLM alignment without a reference model. |
SimPO is a reference-free preference optimization algorithm. Its key innovation is using the average log probability of a sequence as the implicit reward, plus a target reward margin γ.
L_SimPO = -E log σ(β/|yw| · log πθ(yw|x) - β/|yl| · log πθ(yl|x) - γ)
Since the log probabilities passed to simpo_loss are already length-normalized (average log prob), the loss simplifies to:
logits = β * policy_chosen_logps - β * policy_rejected_logps - γ
where γ = gamma_beta_ratio * beta.
losses = -log σ(logits) * (1 - label_smoothing) - log σ(-logits) * label_smoothinglosses = relu(1 - logits)chosen_rewards = β * policy_chosen_logpsrejected_rewards = β * policy_rejected_logps