Skip to main content

grpo-troubleshooting

النجوم١٢
التفرعات٠
آخر تحديث١٠ يونيو ٢٠٢٦ في ١٨:٢٣

Get Stage-3 GRPO running and converging on a small GPU. Use when GRPO "won't run", crashes on launch, OOMs, is impossibly slow, throws bitsandbytes/libnvJitLink errors, "Attempting to unscale FP16 gradients", TRL "unexpected keyword argument 'max_prompt_length'", or when training runs but reward-hacks / doesn't converge (completions pinned at max length, grad_norm nan). Trigger on "GRPO 안 돌아가", "GRPO 터진다", "reward hacking", "max_prompt_length 에러", "fp16 에러", "4bit 너무 느려".

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly