Skip to main content

odysseyarena-benchmarking-long-horizon-active

Design and run inductive agent benchmarks where LLMs must discover hidden rules through long-horizon interaction loops rather than following explicit instructions. Use when the user mentions 'inductive agent evaluation', 'long-horizon benchmarking', 'hidden rule discovery', 'active exploration benchmark', 'OdysseyArena', or 'agent world-model induction'.

インストールへ移動

ソース情報

リポジトリ
ndpvt-web/arxiv-claude-skills
ソースの最終更新活動
2026年2月13日 09:38
検出された SKILL.md の言語
英語
スター
14
フォーク
3

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。