Skip to main content

latent-revise-zero-hit-reasoning

LatentRevise — First-order latent revision method that recovers training signal from zero-hit prompts in RLVR. Optimizes input embeddings of failed reasoning prefixes under dual gradients (away from failed continuation, toward gold answer), constrained to vocabulary embedding convex hull.

インストールへ移動

ソース情報

リポジトリ
hiyenwong/ai_collection
ソースの最終更新活動
2026年7月13日 02:00
検出された SKILL.md の言語
英語
スター
2
フォーク
0

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。