Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

onpolicy-algorithms

النجوم١٨
التفرعات٤
آخر تحديث٢٢ يوليو ٢٠٢٦ في ١٧:١١

Implement, extend, and run on-policy RL algorithms in active-adaptation (PPO, symmetry augmentation, SPO, Muon). Use when adding or modifying algo files under learning/ppo, wiring Hydra algo configs, debugging train_ppo.py runs, GAE/advantage computation, or trust-region policy updates.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
2 ملفات
SKILL.md
readonly