| name | founderquest-rl |
| description | Run only the synthetic, offline FounderQuest-RL environment and preserve protected rewards, blind heldout separation, and human/professional authority boundaries. |
FounderQuest-RL offline research policy
This lab includes deterministic behavior cloning and an evaluation harness. It is not reinforcement learning or an autonomous founder, legal,
banking, healthcare, or regulatory agent.
Allowed:
- Read synthetic task fixtures.
- Create a proposal for the local replay session.
- Evaluate action, target, evidence, and authority against the protected
synthetic fixture.
- Persist a local session and export a secret-free receipt.
- Fit the declared local baseline from train labels and use validation only for
candidate selection.
- Export standardized step trajectories without protected labels in state.
Never:
- Call a model, browser, API, portal, or MCP server.
- Train weights, run online RL, expose heldout labels to candidate inference,
or write to a real external environment.
- Submit an application, accept terms, activate payments, make a credit
decision, publish, or represent an external approval.
- Treat fixture labels as real legal, financial, clinical, or regulatory advice.
Any externally consequential action must receive zero reward and be reverted.