Skip to main content

debugger

Force the project-agent to use a real debugger instead of guessing: set breakpoints where the problem might be, stop execution at those breakpoints, inspect live variable state, and analyze the observed runtime state before patching. Use when a project agent is stuck, sees confusing or repeated failures, suspects state mutation, routing, async, serialization, cache, closure, test fixture, UI/backend mismatch, or any bug where logs/static reading would lead to speculation. Use before further patching after two failed attempts or whenever the user asks for debugger, breakpoints, debug mode, variable state, inspect locals, step through, VS Code debugger, or prove runtime behavior.

소스 정보

저장소
grahama1970/agent-stack-public
최근 소스 활동
2026년 9월 24일 15:51
감지된 SKILL.md 언어
영어
스타
0
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
82 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
debugger
description
Force the project-agent to use a real debugger instead of guessing: set breakpoints where the problem might be, stop execution at those breakpoints, inspect live variable state, and analyze the observed runtime state before patching. Use when a project agent is stuck, sees confusing or repeated failures, suspects state mutation, routing, async, serialization, cache, closure, test fixture, UI/backend mismatch, or any bug where logs/static reading would lead to speculation. Use before further patching after two failed attempts or whenever the user asks for debugger, breakpoints, debug mode, variable state, inspect locals, step through, VS Code debugger, or prove runtime behavior.
triggers
["debugger","use the debugger","project-agent debugger","debug mode","VS Code debugger","set breakpoints","stop at a breakpoint","hit a breakpoint","inspect variable state","inspect runtime state","inspect locals","step through the bug","stop guessing","prove runtime behavior","analyze variable states before patching"]
provides
["runtime-state-inspection","breakpoint-debugging","debugger-proof"]
composes
["brave-search","dogpile","agentic-evals"]
taxonomy
["validation","debugging","evidence","resilience"]
disciplines
["developer-tooling"]
# Debugger Use this skill to replace LLM inference with observed runtime state. It is primarily for the project-agent, not for the human: the agent must invoke it on itself when it is stuck, seeing confusing state, or at risk of patching from guesses. ## Front door (drive it in one line) Other skills, project agents, and the human drive the debugger through `skills/debugger/run.sh` — the same pattern as `memory/run.sh`. Every subcommand owns its own env plumbing (venv, workspace detection, extension-host kind); a caller never exports `UV_PROJECT_ENVIRONMENT` or assembles `uv run` commands. ```bash ./run.sh break <file:line> [--local NAME ...] -- <python-cmd> # headless breakpoint proof -> debugger.proof.v1 ./run.sh stop <file:line> [--local NAME ...] [--expand N[:D]] # live VS Code stop; prints STATUS_PATH + settled STATUS ./run.sh open <file> --json-field FIELD [--bridge] # reveal/select a file range; preserves user focus by default ./run.sh windows list|close [--workspace NAME] [--execute] # list/close VS Code windows; close is explicit and scoped ./run.sh walkthrough <spec.json> [--speak] [--voice] # narrated breakpoint tour (review/blocked) ./run.sh session [--wait-seconds N] # collaborative live session with explained pauses ./run.sh validate <proof.json> [--expect-valid] [--repo-root P] # independent proof validation ./run.sh matrix [--suite NAME ...] # capability-gated eval matrix + receipts ./run.sh recall <query> # stored debugger lessons ./run.sh verify # deterministic self-check -> DEBUGGER-VERIFY-OK ``` `break` adds the caller's cwd to `PYTHONPATH`, so a multi-module scenario runs from its own directory. Live subcommands default the workspace to the git toplevel of `$PWD` (override with `DEBUGGER_VSCODE_WORKSPACE`) and fail closed with `BRIDGE_BLOCKED` when no open, trusted VS Code bridge answers. `open --bridge` is the safe human handoff path for files, JSON fields, and exact selected ranges. It uses VS Code `preserveFocus` and leaves the user's active window and geometry unchanged by default, so a human typing in another app does not lose keystrokes to VS Code. Only move or focus VS Code when the human asks: ```bash ./run.sh open report.json --json-field cases[].trials[].stderr --bridge ./run.sh open report.json --json-field cases[].trials[].stderr --bridge \ --place-window --frontmost --monitor right --window-layout half-vertical ./run.sh open report.json --json-field cases[].trials[].stderr --bridge \ --place-window --frontmost --monitor left --window-layout quarter ``` VS Code window hygiene: ```bash ./run.sh windows list ./run.sh windows close --workspace agent-skills # dry-run plan only ./run.sh windows close --workspace agent-skills --execute ``` `windows close` refuses unfiltered closes and skips Remote SSH windows unless `--include-remote` is explicitly supplied. Use this after a debugger handoff if VS Code windows accumulated, but do not close a human's unrelated project window. When `--frontmost` is explicitly requested, `$debugger` first reads the current virtual desktop with `xdotool get_desktop`, moves the VS Code window to that same desktop with `wmctrl -t`, and only then activates it. `wmctrl` and `xdotool` use zero-based desktop indexes: desktop `6` is the visible "Desktop 7". Prompt examples for humans: - "Use `$debugger` to open the failing receipt at `cases[].trials[].stderr`, but do not steal focus." - "Use `$debugger` to show me the selected field on the right monitor, half width, full height, frontmost." - "Use `$debugger` to pause at `src/server.py:184`, inspect `request` and `selected_handler`, then tell me what changed." Gates: `fixtures/front-door.json` for the front door and `fixtures/vscode-selection.json` for selected-range reveal, focus preservation, explicit window placement, and fail-closed missing-bridge behavior. ### Debug it, then explain what happened `./run.sh spec-from-proof <proof.json> --out spec.json [--workspace P] [--narrate <handler>]` turns any real captured session (`run.sh break`) into a runnable `debugger.walkthrough.v1` spec: stops in true session order (first hit per file:line), `expect` pinned to the first observed locals, launch module + PYTHONPATH derived from the capture, repeat counts narrated ("this line ran 5 times"). No hand-authored JSON. Fail-closed: zero-hit or validation-failing proofs refuse to generate. `--narrate` optionally rewrites the say-lines naturally through /ask (fail-soft to the deterministic template). Gate: `fixtures/spec-from-proof.json`, including a live replay of a freshly generated spec to WALKTHROUGH-COMPLETE. The core functions of this skill are: 1. The project-agent must set breakpoints where the problem might be. 2. The project-agent must stop at a breakpoint and analyze live variable state before deciding what to patch. 3. When reporting debugger work to the human, the project-agent must provide a concrete breakpoint location the human can examine, with the expected relevant variable state at that pause. If these functions are not demonstrated, the skill has not been used. ## When To Use And Why Use `$debugger` when the next correct action depends on live runtime state, not on what the code appears to do. The purpose is to prevent the project-agent from patching by inference when a debugger can show the actual value, branch, frame, request, response, object mutation, or adapter payload. When in doubt, ask this gate question: ```text Would seeing the actual paused variable/frame/request state change the patch I am about to make? ``` If yes, use `$debugger`. Mandatory triggers: - The human asks for `$debugger`, debugger, VS Code debugger, breakpoints, debug mode, stepping, locals, variable state, or proof of runtime behavior. - The same defect, failed test, bad UI state, or confusing behavior survives two focused fix/verification attempts. - A prior success claim is disproved by the human, a screenshot, a runtime artifact, or a visible UI state. - The suspected bug involves state that changes at runtime: async order, routing, request parsing, branch selection, cache, mutation, serialization, fixture setup, closure state, subprocess output, model payloads, browser state, or UI/backend adapter state. - The project-agent is about to patch code based on a guess about what a variable contains, which branch runs, which handler receives a request, or which object is passed across a boundary. - Logs, static reading, DOM assertions, or test pass/fail status do not explain why the observed behavior is wrong. - The human wants a collaborative breakpoint review: the agent pauses execution, reports the relevant variables, asks whether the state is semantically correct, then continues to the next breakpoint. Use `$debugger` before patching in these cases because it gives positive evidence: - the breakpoint was set and verified - execution stopped at the expected source line - the paused frame and thread are known - relevant variables and watches were inspected while execution was stopped - the agent can say which state is already wrong, which state is still correct, and what transition should be inspected next This positive evidence matters because it prevents the agent from confidently patching a wrong hypothesis. If the observed state contradicts the planned edit, the edit must change or stop. Do not use `$debugger` as busywork for problems already explained by deterministic evidence, such as syntax errors, formatter failures, missing imports, dependency resolution, environment setup, type-checker diagnostics, formatter output, or a test assertion that directly names the incorrect literal value and requires no hidden state. Fix those directly, then test. ## Route To The Right Evidence Layer First (operator 2026-08-04) A breakpoint is one evidence source, not the only one, and it is the WRONG one for most stuck states in this repo. Before setting a breakpoint, run this triage. It is a lookup, not a judgement call: the project agent does not get to decide which layer to read. | The stuck state | Read this FIRST | Not this | | --- | --- | --- | | A `/tau` DAG was rejected before dispatch | The `tau.dag_error.v1` payload: `verdict`, `failure_code`, `severity`, every `evidence.errors[]` entry, and `recommended_action{type,next_agent,reason}` | A breakpoint in the compiler; Tau already named the cause and the next step | | A `/tau` node was blocked at runtime | `receipt.alerts[]` — each carries `code`, `message`, and an `evidence` object naming node and handler | Inferring the failure from exit codes | | An `/ask` browser lane failed | `lane-diagnostics.json` in the lane artifact dir: the fixed check series plus its derived `diagnosis` | Guessing whether the tab died, drifted, or was rate-limited | | A browser page's live state is in question | `surf js --tab-id <id> --no-activate` | Attaching a breakpoint debugger to a lane blocked on Chrome — it gets zero hits | | A seam artifact is malformed | The `SeamViolation` error list, which carries the pydantic errors verbatim | Reading the producer's source to imagine what it emitted | | Your own Python does something you cannot explain from its receipts | **This skill.** Set the breakpoint | — | Only the last row is a debugger problem. The rows above it are already answered in writing by a tool that fails closed; a breakpoint there is slower, and it replaces an authoritative answer with a reconstruction of one. The rule this encodes: **`/debugger` is for state nothing wrote down.** When a contract validator, a receipt, or a diagnostic probe has already recorded the answer, reading it is the debugging step. Reach for a breakpoint when the failing transition happens inside your own process and left no artifact behind. ### The escalation ladder The table above is a DISPATCH, not a sequence: exactly one row owns any given symptom, and the project agent reads that row's evidence first. Running `surf js` against a Tau contract rejection is wasted motion — Tau already wrote the cause and the next step into the payload. Escalate only when the owning layer did not resolve it. ```text 0. DISPATCH - the symptom selects ONE row above. Read that evidence. Most stuck states end here: the artifact names the cause. 1. BREAKPOINT - the artifact did not explain it, or no artifact exists. $debugger: break at the failing transition, inspect the frame. 2. RESEARCH - the observed state is real but its MEANING is unknown (a provider changed, an API contract is unfamiliar, an error string is undocumented). $brave-search or $dogpile, then retry ONCE with what the search returned. 3. STOP - report NEEDS_ATTENTION with the evidence from every rung run. ``` Rung 2 is mandatory, not optional, after two failed focused attempts — the same bar `$tau` enforces on its own subagents. Do not take a third attempt from the same stale context: a retry with no new input is spray-and-pray, and the search exists to supply the new input. Each rung must produce an artifact before the next one starts. "I looked at the receipt" without quoting the field, or "I searched" without the query and what it returned, does not advance the ladder — it just relabels a guess. ### The ladder is enforced in code, not by this document Write a `debugger.ladder.v1` receipt as you climb, and validate it: ```bash skills/debugger/scripts/validate_debugger_ladder.py ladder-receipt.json --expect-valid ``` The validator refuses the receipt when a rung is skipped or reordered, when a non-final rung claims to have resolved the problem, when a cited artifact does not exist on disk, when a dispatch rung names no field it read, when a breakpoint rung cites no `debugger.proof.v1`, when a research rung records no query or no result URL, or when `attempts >= 2` with no research rung. The existence check is the load-bearing one: an agent can assert it read a receipt, but it cannot conjure the file it claims to have read. Schema: `schemas/debugger.ladder.v1.schema.json`. Gate: `./sanity-ladder.sh`. The validator has no third-party dependencies and needs no Tau checkout, which is what keeps this skill self-contained. ### Running the ladder as a Tau DAG `templates/debugger-ladder.dag.yaml` expresses the ladder as a `tau.dag_contract.v1` (validated against Tau's own `validate_dag_contract`). Use it when the stuck work is ALREADY running under Tau and you want the rungs scheduled, receipted, and resumable, with `ladder-gate` on every path to a terminal node. The DAG is orchestration, not enforcement — it calls the validator above rather than reimplementing it. Note that Tau already ships `brave_search_required_after_two_attempts` as a `fail_closed_on` invariant, which is Tau's own name for the research rung; the ladder-specific invariants stay in the validator, because Tau rejects invented invariant codes before dispatch. ## Self-Contained Scope This skill is project-agnostic and must remain self-contained in the `agent-skills` repo. Any project agent can use it against the current project by choosing breakpoints in that project's code and running the local reproduction command under the bundled harness or an equivalent platform debugger. Do not move the reusable debugger workflow into a project repo. Project-specific debugger UI, adapters, or debug-session APIs may live in that project, but they are consumers of this skill. The skill remains the cross-project contract: stop, hypothesize, break, run, inspect real variables, then patch. All command examples assume: ```bash export SKILL_DIR="${SKILL_DIR:-/path/to/agent-skills/skills/debugger}" export UV_PROJECT_ENVIRONMENT="${UV_PROJECT_ENVIRONMENT:-/mnt/storage12tb/skills/debugger/.venv}" ``` ## Required Loop 1. Stop coding and state the bug or uncertainty as a runtime-state question. 2. Identify where the relevant state enters, changes, branches, or exits. 3. Set breakpoints at the smallest useful source locations around that transition. 4. Run the real failing command, request, test, UI action, or reproduction under a debugger. 5. Stop at the breakpoint; do not replace this with logs or a static explanation. 6. Inspect locals, selected globals, watched expressions, request/response objects, and return/error state from the paused frame. 7. Analyze the variable state: what is already wrong, what is still correct, and which next branch or mutation follows. 8. Escalate to the human only when the agent is blocked, the observed state requires human/domain judgment, or the agent cannot honestly determine whether the state is correct. When escalating, ask the human to examine the specific paused values, similar to `$interview`: "Does this paused variable state look correct?" or "Which value is wrong?" 9. Continue or step only as needed to observe the next state transition. 10. Give the human a concrete breakpoint to examine: file, line, source statement, and expected relevant variable state at that pause. 11. Report the exact breakpoint locations, inspected values, human confirmation or correction when requested, and what conclusion follows. 12. Patch only after the runtime state explains the failure. Do not satisfy this skill with a source-code explanation, log skim, print-only trace, or generic test rerun. Those can support the investigation, but the core proof is paused runtime state. ## VS Code Debugger Requirement When the user asks for the VS Code debugger, use VS Code's debugger path or Debug Adapter Protocol path, not only `pdb`, print statements, or the bundled Python harness. The project-agent must automatically create or update the target project's `.vscode/launch.json` with a runnable configuration for the reproduction. Do not leave VS Code activation as prose instructions only. The human should be able to open the project in VS Code, select the generated configuration, set or inspect the listed breakpoint, and press Start Debugging. ### Visible VS Code GUI Boundary Be precise about what is being controlled: - A standalone terminal CLI can generate `.vscode/launch.json`, open files or workspaces in VS Code, and drive an external DAP/debugpy session, but it is not a supported remote-control API for the visible VS Code workbench debugger. - A VS Code extension can start a visible VS Code debug session with VS Code's `vscode.debug.startDebugging(...)` API, add/remove breakpoints through `vscode.debug.addBreakpoints(...)`, observe debug lifecycle events, and send requests to the active debug adapter. - Neither a standalone CLI nor a VS Code extension should claim to read the Variables pane UI directly. Variable state should be captured through DAP requests such as `threads`, `stackTrace`, `scopes`, `variables`, and `evaluate`, or through a DAP tracker/proxy. Therefore, this skill has two honest modes: 1. **Automated DAP proof mode:** create/update `launch.json`, run the reproduction under a DAP/debugpy controller, stop at breakpoints, inspect variables, and write a proof artifact. This is self-contained in `agent-skills`. 2. **Visible VS Code GUI mode:** create/update `launch.json` and use the bundled companion VS Code extension bridge or a DAP proxy to start/control the visible VS Code session. Without that bridge, report that visible GUI control is not available instead of claiming it. ### Bundled VS Code Extension Bridge The bundled bridge lives at `$SKILL_DIR/vscode-bridge`. It runs inside the VS Code extension host and provides the missing boundary that a terminal process cannot cross directly. Plainly: the terminal writer creates `.vscode/debugger-bridge/request.json`, the extension reads that request inside the trusted workspace, starts or continues the visible debug session, queries the stopped adapter for selected locals/watches, and writes a status/proof artifact back. The bridge is session-oriented. A start/restart/process request returns `debugger.session.v1` state with the VS Code debug session ID, selected thread/frame, current stop sequence, requested/verified breakpoints, and an event log reference. Follow-up `inspect`, `continue`, `stepOver`, `stepIn`, `stepOut`, `pause`, `runTo`, `removeBreakpoints`, `selectFrame`, `selectThread`, and `terminate` requests must bind `sessionId` and `expectedStopSequence` so a stale agent command cannot control a newer pause. `inspect` reads the already-paused selected frame without continuing execution. The bridge is intentionally fail-closed: - it does not auto-process stale request files on VS Code startup - every request must include a fresh `createdAt` and unique request id - duplicate or stale request ids are rejected - the workspace must be trusted before the bridge starts, restarts, continues, or evaluates debugger state - the request writer records `status: pending` and a request hash before atomically replacing `request.json` - the bridge captures only explicitly requested locals - watch expressions are opt-in per request and should be used only when the expression is known to be side-effect safe - visible bridge status must not be treated as adapter breakpoint verification unless it includes adapter proof; bridge status can prove the session stopped at the requested source line, while direct DAP proof can additionally show the adapter `setBreakpoints` response Detailed bridge commands, watcher behavior, status ownership rules, and known limitations live in `references/vscode-bridge.md`. Install or update it with: ```bash "$SKILL_DIR/scripts/install_vscode_bridge.sh" ``` If VS Code was already open before installation, reload the VS Code window so the extension activates. To request a visible VS Code debug session from a VS Code integrated terminal, write the bridge request file: ```bash uv run --project "$SKILL_DIR" \ python "$SKILL_DIR/scripts/request_vscode_bridge.py" \ --workspace /path/to/project \ --launch-config-name "Debug failing pytest with $debugger" \ --break path/to/file.py:123 \ --local some_var \ --watch 'some_obj.field' \ --allow-watch-eval ``` The extension writes status/proof to: ```text .vscode/debugger-bridge/status.json ``` The bridge does not scrape the Variables pane UI. It captures the same class of runtime state through DAP while the visible VS Code debug session is stopped. Use the bundled writer: ```bash uv run --project "$SKILL_DIR" \ python "$SKILL_DIR/scripts/write_vscode_launch.py" \ --workspace /path/to/project \ --name "Debug failing pytest with $debugger" \ --python '${workspaceFolder}/backend/.venv/bin/python3' \ --module pytest \ --arg -q \ --arg path/to/test.py::test_name \ --env 'PYTHONPATH=${workspaceFolder}/backend/src' ``` A valid VS Code debugger proof includes: - the VS Code debug adapter or extension used - the generated `.vscode/launch.json` configuration path and name - the `setBreakpoints` request or visible breakpoint configuration - proof the breakpoint was verified when the adapter exposes it, or an explicit `adapterBreakpointVerification: unavailable-vscode-api` limitation plus proof the stopped frame matches the requested source line - for Remote SSH workspaces, proof that the bridge extension ran in the workspace extension host with the expected `remoteName` and that request, status, and session artifacts were written in the remote workspace - when a breakpoint is requested on a declaration line, receipt evidence for requested path/line, VS Code breakpoint state, actual stopped frame, and the current source hash/symbol range that justifies any relocated executable line - proof execution stopped with reason `breakpoint` - the paused source file, line, and frame - inspected variables from the paused frame - analysis of what those variables prove Run the Remote SSH bridge authority gate when debugging from Graham's normal local-client/remote-Ubuntu workflow: ```bash bash "$SKILL_DIR/sanity-bridge-remote-ssh.sh" --allow-live --out /tmp/debugger-remote-ssh-proof ``` If the shell is not inside a Remote SSH workspace, that command must write a typed blocked receipt instead of treating local VS Code as equivalent. The project-agent may drive the VS Code debugger through DAP in a terminal when a GUI is not required. The proof still must show that the breakpoint was set, hit, and used to inspect live runtime state. ## Language-Neutral Debugging Contract The project-agent and human should not need a different debugging workflow for Python, TypeScript, Rust, or any other implementation language. The useful object is always the paused variable state at a particular breakpoint. Language only determines the adapter used to stop execution and read frame state: - Python uses the bundled Python harness, VS Code debugpy, or another Python
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기