Skip to main content
在 Manus 中运行任何 Skill
一键导入

reward-engineering

星标25
分支4
更新时间2026年7月17日 02:09

Use when designing or reviewing a reward function for a new so101-nexus environment (or auditing an existing one), especially any task with a multi-phase or dwelling-prone completion condition (grasp-then-release, reach-then-hold, multi-object sequencing). Covers dwelling-vs-potential-based shaping, keeping the potential monotone along the ideal trajectory, how to spot a reward-hacking trap before it costs a training run, and the primitives/citations to use.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly