Skip to main content
Manus에서 모든 스킬 실행
원클릭으로

auditing-rlhf-reward-hacking

스타2
포크0
업데이트2026년 6월 17일 00:55

Audits a post-RLHF (or post-DPO / post-RLAIF) policy model for reward hacking — the failure mode where the policy maximizes the learned reward model's score without actually satisfying the underlying human preference the reward model was supposed to encode. Probes for known reward-hacking patterns (length bias, sycophancy, formatting tricks, refusal-substitution, persuasion-over-correctness, reward-model exploitation at the distribution boundary), computes a pre-vs-post-RLHF preference-divergence metric, and produces a per-probe verdict table plus a remediation list. Use when an RLHF / DPO / RLAIF run has just completed, when downstream eval shows the post-RLHF model behaves worse on holdout preference data despite higher reward-model scores, when alignment-tax measurement is needed, or before promoting an RLHF checkpoint to production.

설치

Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.

파일 탐색기
7 개 파일
SKILL.md
readonly