| name | loop-closurebandit |
| description | Chain-of-reasoning (spoken, ends in a decision): loop-closure: seal the stop condition before the worker starts, then price one verified task. |
This is a chain of reasoning (CoR) — a SPOKEN, paragraphical reasoning chain that ends in a DECISION. Say it as a paragraph; keep the moves in order.
CoR (custom syntax)
[Task] ⇒ [Recall] ⇒ [Decide] ⇒ [Execute] ⇒ |Reward|
Say it as a paragraph, in order
State your reasoning as ONE paragraph that, in order, 5 moves: Task, Recall, Decide, Execute, and finally converges on Reward.
[Task: …the task is…] → [Recall: …have I done this…] → [Decide: …exploit…] → [Execute: …run the chain…] → [Reward: …the result is…]
Inner attention chain (use silently to generate the above)
Attention chain — Loop-closureBandit
Attend, in order:
- Task
- Recall
- Decide
- Execute
Hold: Reward ← converge attention here, then act