Skip to main content

capopt-reward-preference-modeling-engineer

Capability/optimization role: **Reward & preference modeling engineer** (Human engineering role (AI/robotics)) — builds the reward, preference, and constitutional signals that shape behavior (RLHF, RLAIF, rule-based rewards). Part of the layer that decides *how* robot and machine capabilities are built — across model tiers (LLM, SLM, tiny LM, deterministic) and many training methods (imitation, model-based/offline RL, RLHF/RLAIF, sim-to-real, distillation, classical control, formal methods). Use this skill when choosing or building how a capability is trained, optimized, or run on-device, even if the user only describes the underlying need. Works under a alignment lead.

インストールへ移動

ソース情報

リポジトリ
TuringWorks/civstack
ソースの最終更新活動
2026年6月16日 07:03
検出された SKILL.md の言語
英語
スター
1
フォーク
0

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。