Use when designing or reviewing a reward function for a new so101-nexus environment (or auditing an existing one), especially any task with a multi-phase or dwelling-prone completion condition (grasp-then-release, reach-then-hold, multi-object sequencing). Covers dwelling-vs-potential-based shaping, keeping the potential monotone along the ideal trajectory, how to spot a reward-hacking trap before it costs a training run, and the primitives/citations to use.
2026-07-17