Skip to main content
Manus에서 모든 스킬 실행
원클릭으로

rl-reward

// Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win rate across group rollouts); generating preference pairs for DPO/RLAIF; and normalizing scores for training stability. Use when building reward models, scoring rollouts for GRPO/REINFORCE, generating preference data for DPO, or doing Best-of-N selection.

$ git log --oneline --stat
stars:619
forks:52
updated:2026년 3월 16일 04:23
파일 탐색기
3 개 파일
SKILL.md
readonly