Skip to main content

phoenix-ralph

Persistence loop that keeps working a backlog until the goal is OBJECTIVELY proven done — not until the agent thinks it's done. Wraps Geoffrey Huntley's Ralph loop (fresh context per iteration, filesystem as memory) with Phoenix's failure-first gate ledger, so completion is derived from the tamper-evident trace, never self-reported. Use when a task must run to completion across many iterations, when the user says /phoenix-ralph, "ralph", "don't stop", "keep going until it's done", or "run until the tests pass". Not for one-shot fixes (use phoenix-build) or scoping a vague idea (use phoenix-goal).

Aller à l'installation

Informations de source

Dépôt
All-The-Vibes/ATV-Phoenix
Dernière activité de la source
18 juin 2026 à 02:25
Langue détectée de SKILL.md
anglais
Étoiles
6
Forks
4

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
type
Phoenix Skill
name
phoenix-ralph
description
Persistence loop that keeps working a backlog until the goal is OBJECTIVELY proven done — not until the agent thinks it's done. Wraps Geoffrey Huntley's Ralph loop (fresh context per iteration, filesystem as memory) with Phoenix's failure-first gate ledger, so completion is derived from the tamper-evident trace, never self-reported. Use when a task must run to completion across many iterations, when the user says /phoenix-ralph, "ralph", "don't stop", "keep going until it's done", or "run until the tests pass". Not for one-shot fixes (use phoenix-build) or scoping a vague idea (use phoenix-goal).
license
MIT
# phoenix-ralph — the persistence loop, gated by objective proof Ralph (Geoffrey Huntley, [ghuntley.com/ralph](https://ghuntley.com/ralph)) is a dead-simple, powerful idea: run the agent in a loop with a fixed prompt and a **fresh context every iteration**, with the **filesystem as the only memory**. One task per loop. Huntley's original is literally `while :; do cat PROMPT.md | agent; done`. Phoenix adds the one thing Ralph (and Claude Code's ralph, and every "autonomous" loop) lacks: an **objective, tamper-evident completion proof**. The loop does not stop when the agent *says* it's done — it stops when the driver *proves* the top-level acceptance check is **failure-first satisfied** (`phoenix-mcp accept`): the trace shows the check went **red → green** for the same check, the chain is intact, and it is **green right now**. > Every other persistence loop ends in an opinion ("the reviewer approved", "the model thinks it's > done"). phoenix-ralph ends in evidence. ## Two ways to run it **A. Interactive (inside the Copilot CLI) — the common case.** You're in a live `copilot` session and invoke `/phoenix-ralph`. There is **no external script**. The loop runs *in-session*: within your agentic turn you iterate edit → `phoenix_sense` → heal → `phoenix_sense`, one backlog item at a time, and you do not say "done" until **`phoenix_accept`** (the gate ledger, an MCP tool) returns ok=true for the done-check. The agent's own tool-use loop *is* the loop; the skill's rules are what keep it honest. If the job outgrows one turn, the user just says "continue" and you re-read the state files and resume. **B. Unattended / large (`copilot -p`) — fresh context per iteration.** For overnight jobs, CI, or work too big for one context window, the external driver `dist/ralph/phoenix-ralph.ps1` (+ bash twin) re-invokes the agent with a **fresh context every loop** (Huntley's key trick against the ~150k-token degradation). Same gate, same state files — the driver calls `phoenix-mcp accept` instead of the in-session tool. Use this when nobody's watching or the backlog is long. Both modes share the identical law: **completion is proven, never self-reported.** In-session you call the `phoenix_accept` tool; unattended the driver calls `phoenix-mcp accept`. Either way "done" requires the trace to show the done-check went **red → green** on an intact chain, green now. ## The loop (mode B shown; mode A is the same minus the external re-invocation) ``` driver: accept(done-check)? ── proven green ──▶ DONE (driver writes completed.json + git tag) │ not yet ▼ copilot -p PROMPT.md (fresh context; agent reads backlog.json + progress.md, does ONE item) │ verify-trace intact? ── broken ──▶ STOP (tamper/corruption) │ ok state changed? ── no, N× ──▶ STOP (stuck = planning problem) └──────────────────── loop ◀──────────── ``` ## State (the brain, on disk — `.phoenix-ralph/`) - `PROMPT.md` — fixed per-loop instructions (scaffold from `dist/ralph/PROMPT.template.md`). - `backlog.json` — `[{id, title, priority, done, check}]`. **Each item's `check` is an objective Phoenix check**, not prose. This is the differentiator over Huntley's `fix_plan.md` (a bullet list) and Claude Code's PRD (LLM-reviewed criteria). - `progress.md` — append-only learnings; the working memory that survives the fresh context. - `done-check.json` — the single top-level acceptance check (a real `command_exit`). - `completed.json` — the proof bundle, written by the **driver** on success (mode B). In interactive mode you instead report the `phoenix_accept` result as the proof. Never hand-write a fake one. ## What you (the agent) do each iteration 1. **Re-read** `backlog.json` + `progress.md` first — you have amnesia between loops. 2. **Pick one** highest-priority item with `done:false`. 3. **Search before building** — code search is non-deterministic; "not found" ≠ "not done". Use `phoenix-context` (the graph), not blind grep. 4. **Reproduce the failure first**: `phoenix_sense` the item's `check`. It must be RED now. *This is what makes the gate real* — a check never seen failing is rejected by `accept`. 5. **Implement** the smallest full change (no placeholders). `phoenix_snapshot` before risky edits. 6. **Verify**: `phoenix_sense` again. Red → `phoenix_heal` (≤3) or fix; never proceed on red. 7. **Record**: set `done:true` only after red→green; append what you did + learnings to `progress.md`. 8. **Prove the goal**: call **`phoenix_accept`** on the done-check. - *Interactive (mode A):* this is your objective stop signal — when it returns ok=true, the goal is proven done; report success with the evidence. Until then, loop back to step 2. - *Unattended (mode B):* the external driver runs `accept` and owns `completed.json` + the git tag, so just append a final note to `progress.md`; don't write those yourself. ## Common Rationalizations | Rationalization | Reality | |---|---| | "I'll mark the item done; I'm sure it works." | The driver derives done from the trace. An unproven `done:true` is ignored — and dishonest edits break the hash chain. | | "The check passes already, I'll skip reproducing the failure." | A check never seen red is vacuous; `accept` rejects it. Watch it fail first, or it proves nothing. | | "I'll write `completed.json` / tag it to finish faster." | Only the driver writes those, only on a proven `accept`. Faking them changes nothing and corrupts state. | | "I'll do three items this loop to save iterations." | One task per loop keeps the context fresh and the trace legible. Batching is how loops go off the rails. | | "It's been failing the same way for 5 loops, I'll keep trying." | A stuck item is a planning problem. Record the blocker and stop; grinding burns tokens. | ## Red Flags — stop - You're about to set `done:true` without having watched the item's check go red → green. → Reproduce first. - You're writing `completed.json` or `git tag`. → That's the driver's job; you can't fake completion. - The backlog item's `check` is something like `test -f file` or matches text already present. → Vacuous gate; make it a real `command_exit` that exercises the behavior. - You're re-reading the 4th file to recall what you did last loop. → It belongs in `progress.md`; write it there. - The same check has been red for 3 iterations. → Stop; re-plan (`phoenix-plan`), don't grind. ## Relationship to the other skills phoenix-ralph is the **persistence wrapper**. It drives `phoenix-build`/`phoenix-test`/`phoenix-debug` one item at a time. To turn a *vague goal* into a backlog + a real acceptance check first, use **phoenix-goal**, which scopes the work and then hands off to this loop.
Voir sur GitHub