Skip to main content

phoenix-ralph

Persistence loop that keeps working a backlog until the goal is OBJECTIVELY proven done — not until the agent thinks it's done. Wraps Geoffrey Huntley's Ralph loop (fresh context per iteration, filesystem as memory) with Phoenix's failure-first gate ledger, so completion is derived from the tamper-evident trace, never self-reported. Use when a task must run to completion across many iterations, when the user says /phoenix-ralph, "ralph", "don't stop", "keep going until it's done", or "run until the tests pass". Not for one-shot fixes (use phoenix-build) or scoping a vague idea (use phoenix-goal).

Jump to install

Source facts

Repository
All-The-Vibes/ATV-Phoenix
Last source activity
June 18, 2026 at 02:25
Detected SKILL.md language
English
Stars
6
Forks
4

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
type
Phoenix Skill
name
phoenix-ralph
description
Persistence loop that keeps working a backlog until the goal is OBJECTIVELY proven done — not until the agent thinks it's done. Wraps Geoffrey Huntley's Ralph loop (fresh context per iteration, filesystem as memory) with Phoenix's failure-first gate ledger, so completion is derived from the tamper-evident trace, never self-reported. Use when a task must run to completion across many iterations, when the user says /phoenix-ralph, "ralph", "don't stop", "keep going until it's done", or "run until the tests pass". Not for one-shot fixes (use phoenix-build) or scoping a vague idea (use phoenix-goal).
license
MIT
# phoenix-ralph — the persistence loop, gated by objective proof Ralph (Geoffrey Huntley, [ghuntley.com/ralph](https://ghuntley.com/ralph)) is a dead-simple, powerful idea: run the agent in a loop with a fixed prompt and a **fresh context every iteration**, with the **filesystem as the only memory**. One task per loop. Huntley's original is literally `while :; do cat PROMPT.md | agent; done`. Phoenix adds the one thing Ralph (and Claude Code's ralph, and every "autonomous" loop) lacks: an **objective, tamper-evident completion proof**. The loop does not stop when the agent *says* it's done — it stops when the driver *proves* the top-level acceptance check is **failure-first satisfied** (`phoenix-mcp accept`): the trace shows the check went **red → green** for the same check, the chain is intact, and it is **green right now**. > Every other persistence loop ends in an opinion ("the reviewer approved", "the model thinks it's > done"). phoenix-ralph ends in evidence. ## Two ways to run it **A. Interactive (inside the Copilot CLI) — the common case.** You're in a live `copilot` session and invoke `/phoenix-ralph`. There is **no external script**. The loop runs *in-session*: within your agentic turn you iterate edit → `phoenix_sense` → heal → `phoenix_sense`, one backlog item at a time, and you do not say "done" until **`phoenix_accept`** (the gate ledger, an MCP tool) returns ok=true for the done-check. The agent's own tool-use loop *is* the loop; the skill's rules are what keep it honest. If the job outgrows one turn, the user just says "continue" and you re-read the state files and resume. **B. Unattended / large (`copilot -p`) — fresh context per iteration.** For overnight jobs, CI, or work too big for one context window, the external driver `dist/ralph/phoenix-ralph.ps1` (+ bash twin) re-invokes the agent with a **fresh context every loop** (Huntley's key trick against the ~150k-token degradation). Same gate, same state files — the driver calls `phoenix-mcp accept` instead of the in-session tool. Use this when nobody's watching or the backlog is long. Both modes share the identical law: **completion is proven, never self-reported.** In-session you call the `phoenix_accept` tool; unattended the driver calls `phoenix-mcp accept`. Either way "done" requires the trace to show the done-check went **red → green** on an intact chain, green now. ## The loop (mode B shown; mode A is the same minus the external re-invocation) ``` driver: accept(done-check)? ── proven green ──▶ DONE (driver writes completed.json + git tag) │ not yet ▼ copilot -p PROMPT.md (fresh context; agent reads backlog.json + progress.md, does ONE item) │ verify-trace intact? ── broken ──▶ STOP (tamper/corruption) │ ok state changed? ── no, N× ──▶ STOP (stuck = planning problem) └──────────────────── loop ◀──────────── ``` ## State (the brain, on disk — `.phoenix-ralph/`) - `PROMPT.md` — fixed per-loop instructions (scaffold from `dist/ralph/PROMPT.template.md`). - `backlog.json` — `[{id, title, priority, done, check}]`. **Each item's `check` is an objective Phoenix check**, not prose. This is the differentiator over Huntley's `fix_plan.md` (a bullet list) and Claude Code's PRD (LLM-reviewed criteria). - `progress.md` — append-only learnings; the working memory that survives the fresh context. - `done-check.json` — the single top-level acceptance check (a real `command_exit`). - `completed.json` — the proof bundle, written by the **driver** on success (mode B). In interactive mode you instead report the `phoenix_accept` result as the proof. Never hand-write a fake one. ## What you (the agent) do each iteration 1. **Re-read** `backlog.json` + `progress.md` first — you have amnesia between loops. 2. **Pick one** highest-priority item with `done:false`. 3. **Search before building** — code search is non-deterministic; "not found" ≠ "not done". Use `phoenix-context` (the graph), not blind grep. 4. **Reproduce the failure first**: `phoenix_sense` the item's `check`. It must be RED now. *This is what makes the gate real* — a check never seen failing is rejected by `accept`. 5. **Implement** the smallest full change (no placeholders). `phoenix_snapshot` before risky edits. 6. **Verify**: `phoenix_sense` again. Red → `phoenix_heal` (≤3) or fix; never proceed on red. 7. **Record**: set `done:true` only after red→green; append what you did + learnings to `progress.md`. 8. **Prove the goal**: call **`phoenix_accept`** on the done-check. - *Interactive (mode A):* this is your objective stop signal — when it returns ok=true, the goal is proven done; report success with the evidence. Until then, loop back to step 2. - *Unattended (mode B):* the external driver runs `accept` and owns `completed.json` + the git tag, so just append a final note to `progress.md`; don't write those yourself. ## Common Rationalizations | Rationalization | Reality | |---|---| | "I'll mark the item done; I'm sure it works." | The driver derives done from the trace. An unproven `done:true` is ignored — and dishonest edits break the hash chain. | | "The check passes already, I'll skip reproducing the failure." | A check never seen red is vacuous; `accept` rejects it. Watch it fail first, or it proves nothing. | | "I'll write `completed.json` / tag it to finish faster." | Only the driver writes those, only on a proven `accept`. Faking them changes nothing and corrupts state. | | "I'll do three items this loop to save iterations." | One task per loop keeps the context fresh and the trace legible. Batching is how loops go off the rails. | | "It's been failing the same way for 5 loops, I'll keep trying." | A stuck item is a planning problem. Record the blocker and stop; grinding burns tokens. | ## Red Flags — stop - You're about to set `done:true` without having watched the item's check go red → green. → Reproduce first. - You're writing `completed.json` or `git tag`. → That's the driver's job; you can't fake completion. - The backlog item's `check` is something like `test -f file` or matches text already present. → Vacuous gate; make it a real `command_exit` that exercises the behavior. - You're re-reading the 4th file to recall what you did last loop. → It belongs in `progress.md`; write it there. - The same check has been red for 3 iterations. → Stop; re-plan (`phoenix-plan`), don't grind. ## Relationship to the other skills phoenix-ralph is the **persistence wrapper**. It drives `phoenix-build`/`phoenix-test`/`phoenix-debug` one item at a time. To turn a *vague goal* into a backlog + a real acceptance check first, use **phoenix-goal**, which scopes the work and then hands off to this loop.
View on GitHub