Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

auditing-rlhf-reward-hacking

النجوم٢
التفرعات٠
آخر تحديث١٧ يونيو ٢٠٢٦ في ٠٠:٥٥

Audits a post-RLHF (or post-DPO / post-RLAIF) policy model for reward hacking — the failure mode where the policy maximizes the learned reward model's score without actually satisfying the underlying human preference the reward model was supposed to encode. Probes for known reward-hacking patterns (length bias, sycophancy, formatting tricks, refusal-substitution, persuasion-over-correctness, reward-model exploitation at the distribution boundary), computes a pre-vs-post-RLHF preference-divergence metric, and produces a per-probe verdict table plus a remediation list. Use when an RLHF / DPO / RLAIF run has just completed, when downstream eval shows the post-RLHF model behaves worse on holdout preference data despite higher reward-model scores, when alignment-tax measurement is needed, or before promoting an RLHF checkpoint to production.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
7 ملفات
SKILL.md
readonly