Skip to main content
Run any Skill in Manus
with one click

auditing-rlhf-reward-hacking

Stars2
Forks0
UpdatedJune 17, 2026 at 00:55

Audits a post-RLHF (or post-DPO / post-RLAIF) policy model for reward hacking — the failure mode where the policy maximizes the learned reward model's score without actually satisfying the underlying human preference the reward model was supposed to encode. Probes for known reward-hacking patterns (length bias, sycophancy, formatting tricks, refusal-substitution, persuasion-over-correctness, reward-model exploitation at the distribution boundary), computes a pre-vs-post-RLHF preference-divergence metric, and produces a per-probe verdict table plus a remediation list. Use when an RLHF / DPO / RLAIF run has just completed, when downstream eval shows the post-RLHF model behaves worse on holdout preference data despite higher reward-model scores, when alignment-tax measurement is needed, or before promoting an RLHF checkpoint to production.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
7 files
SKILL.md
readonly