| name | until-it-holds |
| description | Explicitly invoked evidence-led improvement loop for building, inspecting, criticizing, repairing, and rechecking an artifact or outcome until it meets a real bar or reaches an honest stop condition. Use only when the user invokes $until-it-holds or /until-it-holds for code, design, writing, research, plans, data, releases, or mixed work. Separates builders from critics, preserves authority boundaries, rejects evaluator gaming, and stops before iteration becomes theatre. |
Until It Holds
Define → build → fresh critic → repair → verify → stop.
Improve the artifact, not its story about itself.
Invocation and authority
Run only when the user explicitly invokes $until-it-holds or
/until-it-holds. Invocation authorizes the improvement loop, not new external
side effects. Keep normal authority gates for publishing, deploying, merging,
sending, buying, submitting, deleting, exposing private data, or altering an
unrelated system.
Use relevant medium-specific skills and tools for the work itself. This skill
owns the loop, critic separation, evidence, regressions, and stop decision.
1. Define what must hold
Freeze a compact contract before editing:
- outcome and intended audience effect;
- exact artifact or external state in scope;
- observable bar and authoritative evidence;
- hard constraints and protected material;
- allowed actions and actions that still require confirmation;
- baseline identity and known failures;
- total-attempt cap.
Default to at most three total builder–critic rounds. Use one for bounded,
low-risk work. Use up to five only when the user explicitly asks for a deep pass.
Failed attempts count; report accepted attempts separately.
Do not invent an objective merely to complete the contract. Ask at most one
strategic question only when no responsible default exists. When Actual Goal is
available, it may repair a suspect prompt or bar first; otherwise resolve the
object-level outcome here. Freeze the bar before seeing results and never move it
to flatter the latest artifact.
Read references/bars-by-medium.md when the bar is
unclear or the artifact crosses media.
2. Build one bounded improvement
Inspect the real baseline and protect user work. Checkpoint the exact delta owned
by this round. Choose the largest material gap, then make the smallest causal
change likely to close it. Preserve what already works. The builder never
certifies its own work.
3. Give the artifact to a fresh critic
Give a separate agent or fresh context the artifact, frozen bar, constraints,
and relevant evidence—not the builder's effort story or preferred verdict. The
critic inspects the actual output, tries to falsify success, and names the
evidence, consequence, and smallest useful repair for any material gap.
When fresh review is unavailable, label the critique non-independent. A
same-context critic may award holds only when deterministic or authoritative
evidence completely decides the checked bar. Subjective, creative, and
whole-artifact criteria remain partial until fresh review.
Read references/critic-protocol.md for verdicts,
multiple critic lenses, disagreement, and detailed inspection order.
4. Repair or discard only this round
Keep a change only when direct evidence shows net improvement without a protected
regression. Repair the largest remaining material gap, not the easiest visible
flaw.
Discard only the exact checkpointed delta owned by this round. Never use broad
reset or checkout. If the delta overlaps pre-existing work and cannot be
separated safely, repair in place or stop blocked. Preserve conflicting
evidence instead of averaging it into a pass.
5. Verify the whole
After every accepted change, rerun relevant regressions on the changed artifact.
Then inspect the complete artifact once for seams, duplication, drift,
cross-boundary failure, accidental complexity, and audience effect. Individually
improved pieces can still fail together.
For release or operational claims, use Release Steward when available. Otherwise
record each surface's action receipt and independent readback separately.
Iterate external work in drafts, staging, previews, test accounts, or sandboxes.
Perform an irreversible or user-facing action once, only after the pre-action
artifact holds and authority is current. Read the target back afterward. Any
corrective second send, submission, purchase, production deploy, or publication
requires fresh authority.
6. Stop honestly
Stop at the first true condition:
- the hard constraints and material bar are directly verified;
- two consecutive attempts produce no material accepted gain;
- the next gain costs more complexity, risk, or distortion than it is worth;
- required evidence, access, authority, or a user decision is unavailable;
- the total-attempt cap is exhausted;
- the objective or bar itself must be reframed;
- continuing would game an evaluator, conceal uncertainty, or imitate protected
work.
Do not lower the bar to manufacture closure or keep iterating to perform effort.
For medical, legal, financial, safety, security, or other consequential work,
fresh AI critics are not qualified independent authorities. Without appropriate
authoritative evidence and qualified human review, the artifact may hold as
research or decision support, but the real-world recommendation or action stays
partial or blocked.
Report
Lead with holds, partial, or blocked. State material changes kept,
decisive evidence, regressions checked, attempts made and accepted, each external
action actually performed, remaining gaps, and checks not performed. Never claim
“perfect,” “live,” or “done everywhere” beyond direct evidence.
Use references/run-ledger.md for a durable multi-round
receipt. Read references/method-and-boundaries.md
for lineage, adjacent skills, edge cases, and the method's limits.