Skip to main content

debug-task

Debug agent task execution infrastructure issues — wrong cwd, wrong files, wrong venv, browser startup failures, agent retry loops, skill loading failures

Jump to install

Source facts

Repository
zpoint/vibe-seller
Last source activity
July 27, 2026 at 15:42
Detected SKILL.md language
English
Stars
83
Forks
19

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
debug-task
description
Debug agent task execution infrastructure issues — wrong cwd, wrong files, wrong venv, browser startup failures, agent retry loops, skill loading failures
# Debug Task Infra Debug agent task execution infrastructure issues by analyzing agent logs. ## Core Philosophy **Every agent failure, detour, or workaround is an infra/platform bug — not an agent problem.** When an agent: - Tries a wrong file path then searches for the right one → **catalog or path resolution bug** - Uses wrong CLI syntax then self-corrects → **skill docs not loaded or incomplete** - Takes a screenshot but can't find/view it → **tool integration bug** - Retries a command with different args → **missing docs or unclear API** - Reflects "no failures" when there were failures → **reflection prompt bug** Even if the agent eventually succeeds, each detour: 1. Wastes tokens and time 2. Signals a gap in the platform that WILL hit other tasks 3. Must be traced to a specific infra root cause and filed as a bug **Do NOT dismiss detours as "the agent will self-correct."** Your job is to find every single failure step, trace it to a code/config/prompt root cause, and report it. ## When to use When a task execution shows symptoms like: - Agent sees unexpected files or can't find expected ones - Agent runs in wrong working directory - Browser (Ziniao/Chrome) fails to start or takes multiple retries - Task uses wrong Python environment - Store context (bookmarks, browser config) missing from agent prompt - Agent repeats same bash command 3+ times (infra issue, not agent fault) - Agent repeats browser-open attempts (browser startup failure) - Agent can't load skills (wrong skill path or skill not synced) - Agent passes wrong arguments to tools (schema mismatch) - Agent uses workarounds for things that should "just work" - Agent's post-task reflection misses failures that clearly happened ## Environment first: are you in WSL debugging a native-Windows server? **Before running ANY command below, figure out where the server actually runs.** vibe-seller ships as a native-Windows install *and* as a Linux/WSL install, and the two do not share a data directory. Debugging the wrong one is the single biggest time-sink — you'll read a stale DB, grep empty logs, and "fix" a store that isn't the one serving traffic. **Detect the topology:** ```bash # The API answers, but is there a Python server process on the WSL side? curl -s -m3 http://127.0.0.1:7777/api/auth/me # 200 with a user → a server is up ps aux | grep -Ei 'uvicorn|app\.main' | grep -v grep # empty → NOT served from WSL # If empty, the server is the native-Windows process (WSL2 mirrored # networking makes 127.0.0.1:7777 reach the Windows listener): powershell.exe -NoProfile -Command "Get-Process python,pythonw -ErrorAction SilentlyContinue | Select Id,Path" # → paths under C:\Users\<WinUser>\AppData\Local\Programs\VibeSeller\ confirm it ``` **Need to call the HTTP API from this skill (not just read the DB)?** Get a session cookie via the JWT-auth workaround described in `debug-store/SKILL.md` § "Skipping JWT-cookie auth" — it uses a single stable `taskbot_debug` account, rotates the password per session, and deactivates on exit. **When the live server is native-Windows, everything relocates.** The real runtime data root is **`/mnt/c/Users/<WinUser>/.vibe-seller/`** (commonly `<WinUser>=Administrator`), NOT `~/.vibe-seller`. Re-point every path in this doc: | This doc says | On a WSL→Windows box, use | |---|---| | `~/.vibe-seller/data/vibe_seller.db` | `/mnt/c/Users/<WinUser>/.vibe-seller/data/vibe_seller.db` | | `logs/backend_7777.log` | `/mnt/c/Users/<WinUser>/.vibe-seller/logs/` | | `~/.vibe-seller/tasks/{id}/` | `/mnt/c/Users/<WinUser>/.vibe-seller/tasks/{id}/` | | `~/.vibe-seller/bin/<slug>/browser-use` | `/mnt/c/Users/<WinUser>/.vibe-seller/bin/<slug>/browser-use` | | `~/.vibe-seller/.claude/skills/` | `/mnt/c/Users/<WinUser>/.vibe-seller/.claude/skills/` | | app code (`app/skills/…`) | `…/Programs/VibeSeller/.venv/Lib/site-packages/app/…` | `sqlite3` reads the Windows DB fine over `/mnt/c`. There is often ALSO a stale `~/.vibe-seller` sandbox in WSL from earlier dev — it has a *different* `stores` table (decoy store names, wrong ids). If a store you expect is missing, or the ids don't match the live `/api/stores` response, you're reading the wrong DB. ### Driving a store's `browser-use` wrapper from WSL The per-store wrapper is generated by the Windows install with **CRLF line endings**, and it `exec`s a Windows `browser-use.EXE`. Two failures bite immediately: 1. **CRLF breaks WSL bash** — `exec`ing it directly gives `/usr/bin/env: 'bash\r': No such file or directory`; running it via WSL `bash` gives `set: pipefail: invalid option name`. Do NOT `dos2unix` the file (it's regenerated on every server boot, and a sanitized copy then can't find the Windows `.EXE`). Instead run it through the install's **bundled Git-for-Windows bash**, which tolerates CRLF and resolves the Windows exec path: ```bash BASHEXE='/mnt/c/Users/<WinUser>/AppData/Local/Programs/VibeSeller/git/bin/bash.exe' WU_WIN='C:\Users\<WinUser>\.vibe-seller\bin\<slug>\browser-use' # native path, backslashes TASK_ID="$(uuidgen | tr '[:upper:]' '[:lower:]')" run(){ "$BASHEXE" -c "export VIBE_TASK_ID='$TASK_ID'; export PYTHONIOENCODING=utf-8 PYTHONUTF8=1; '$WU_WIN' $1"; } run "open https://sellercentral.amazon.<tld>/home" run "state" ``` 2. **GBK console → UnicodeEncodeError** — the Windows Python console codec is GBK, so any page with Arabic/CJK text crashes `state`/`eval` with `'gbk' codec can't encode character`. Always export `PYTHONIOENCODING=utf-8` and `PYTHONUTF8=1` (shown above) before the wrapper call. Downloads triggered through the CDP-proxied browser land in `/mnt/c/Users/<WinUser>/.vibe-seller/downloads/<slug>/` (the proxy overrides Chrome's download dir per `ziniao.py`). Screenshots: pass a **native Windows path** (`C:\...\downloads\<slug>\shot.png`) to `screenshot`, then Read it from the `/mnt/c` equivalent. For the browser-driving specifics (session rotation, wedged-daemon recovery, aux sessions) see the `debug-store` skill — the same WSL→Windows path/CRLF/UTF-8 rules apply there. ### Restarting the native-Windows server `./restart.sh` is WSL-only. To restart the Windows server, stop its `pythonw`/tray process and relaunch the installed app (see `docs/windows-setup.md`). Skills deploy by syncing repo `app/skills/` into `…/Programs/VibeSeller/.venv/Lib/site-packages/app/skills/`; the server's `skills_sync.fetch()` copies them to `…/.vibe-seller/.claude/skills/` at boot. ## Debug methodology ### RULE: Read agent logs BEFORE any conclusion **Every time this skill is invoked, you MUST read the actual agent debug logs before making any claim about what happened.** Do not assume from file existence, code reading, or task_messages alone. The agent debug logs are the source of truth. ```bash # 1. Find which run(s) happened — tasks can restart with different profiles! grep "Starting agent.*{task_id}" logs/backend_7777.log # 2. Get agent debug logs for the CORRECT run (match the date/time) grep "AGENT_DEBUG.*{task_id}" logs/backend_7777.log | grep "{date}" # 3. Find the transcript file (full session log, most detailed) ls ~/.claude/projects/*{task_id}*/*.jsonl # 4. Check if a skill body was actually loaded into context # (use strings unique to that skill's SKILL.md body) grep -c "unique string from SKILL.md" <transcript.jsonl> ``` **Common mistakes to avoid:** - Reading logs from the wrong run (task restarted with different profile) - Checking transcript before the task finishes — skill loading may happen mid-task, not at start. Wait for task to complete before concluding. - Assuming a skill is "loaded" because the file exists in the workspace - Assuming `slash_commands` listing = skill content in LLM context (it doesn't — only metadata loads on discovery, body loads on trigger) - Making claims about what the agent "saw" without transcript evidence - Skipping the transcript file — it's the only way to confirm skill body loading ### Step 0: Pick the task to debug If the user didn't specify a task ID, find the most recent task from the last hour: ```bash sqlite3 ~/.vibe-seller/data/vibe_seller.db \ "SELECT id, status, substr(result,1,100), substr(error,1,100) \ FROM tasks WHERE created_at > datetime('now', '-1 hour') \ ORDER BY created_at DESC LIMIT 5;" ``` If exactly one task was active in the last hour, debug that one. If multiple, ask the user which one. If none, ask the user for a task ID. ### Step 1: Pull ALL task messages Read every message — not just errors. Detours hide in `thinking` messages where the agent says things like "let me try another approach" or "file doesn't exist, let me search." ```bash # Full task log — read ALL of it, not just tail sqlite3 ~/.vibe-seller/data/vibe_seller.db \ "SELECT seq, role, content FROM task_messages \ WHERE task_id='{task_id}' ORDER BY seq;" ``` ### Step 2: Walk through sequentially and flag EVERY detour For each step, ask: "Did this step succeed on the first try?" If not, classify the failure: | Failure Type | Signal in Logs | Root Cause Category | |---|---|---| | Wrong file path → search/glob → find correct path | `Read` error → `Glob` → `Read` success | **Catalog/path resolution bug** | | Hallucinated CLI command → error → correct syntax | `Bash` error with usage/invalid choice message | **Skill docs not effective enough** | | Tool output not usable → workaround | Screenshot returns bytes, agent searches for .png | **Tool integration gap** | | File doesn't exist but agent expected it | `Read` error "File does not exist" | **Catalog lists nonexistent file, or agent guessed** | | Duplicate table headers in catalog | Agent reads garbled catalog | **Catalog generation bug** | | Browser open retried (even once) | >1 `browser-use open` to same URL | **CDP proxy / Ziniao / wrapper script bug** | | Agent says "no failures" in reflection | Contradicts actual log | **Reflection prompt doesn't force log review** | ### Step 3: For each detour, trace to root cause Every detour must map to ONE of these: 1. **Catalog bug** — wrong path, missing `project/` prefix, duplicate headers, unlisted file 2. **Skill docs bug** — missing command, wrong syntax example, undocumented limitation 3. **Tool integration bug** — output not consumable, file not saved where expected 4. **System prompt bug** — unclear instructions, missing "ONLY read catalog files" 5. **Workspace bug** — missing symlinks, wrong isolation, files not synced 6. **Reflection prompt bug** — agent misses failures, doesn't save learnings ### Step 4: Report findings For each bug, provide: - **Seq range**: which message steps were wasted - **What happened**: agent action → error → workaround - **Steps wasted**: count of unnecessary steps - **Root cause**: specific code/file/prompt that caused it - **Fix location**: exact file and what to change ## Known bug patterns (from past investigations) ### Catalog path resolution The L1 catalog (`app/knowledge/CATALOG.md`) lists files like `common/amazon-sites.md`. These are synced to `~/.vibe-seller/knowledge/project/common/amazon-sites.md`. But the store CATALOG.md copies L1 entries verbatim without adding the `project/` prefix. From the agent's CWD, `knowledge/` symlinks to `~/.vibe-seller/knowledge/`, so the correct relative path is `knowledge/project/common/amazon-sites.md` — but the catalog says `common/amazon-sites.md`. **Fix**: `_filter_l1_for_store()` in `app/workspace/knowledge_sync.py` should prepend `project/` to L1 file paths in the store catalog, or the system prompt should tell the agent the base path. ### Duplicate table header in store CATALOG `_filter_l1_for_store()` returns a string starting with `| File | Relevance | Summary |` header. If `_build_store_catalog()` also adds a header, or if the function is called twice, the catalog gets a duplicate header row. **Check**: `app/workspace/knowledge_sync.py` lines around `_build_store_catalog` and `_filter_l1_for_store`. ### Agent hallucinates CLI commands (browser-use and others) Agent invents CLI syntax that doesn't exist instead of using the commands documented in the loaded skill. Examples seen in the wild: - `browser-use scroll 10` (correct: `browser-use scroll down --amount 500`) - `browser-use get text` without index (correct: `browser-use get text <index>`) - `browser-use screenshot` without path then searching for .png files **How skill loading works** (important for diagnosis): Skills go through 3 stages — **discovery ≠ loading ≠ in context**: 1. **File exists** — `.claude/skills/` is **copied** (not symlinked) into the task workspace from `~/.vibe-seller/.claude/skills/` (`app/workspace/manager.py` ~line 710, `shutil.copytree`) 2. **Discovered** — Claude Code finds the SKILL.md via `--add-dir` and lists it in the init event's `slash_commands` array 3. **Content in LLM context** — the SKILL.md content is actually sent to the LLM as part of the prompt. **This is the step that matters and the step that's hardest to verify.** **CRITICAL**: A skill appearing in `slash_commands` in the init event does NOT prove its content is in the LLM's context window. Claude Code may defer loading skill content until the skill is invoked or until a matching `allowed-tools` pattern fires. **IMPORTANT**: browser-use SKILL.md is an **upstream/official** doc from the browser-use project. Do NOT modify it to fix agent behavior issues. **Diagnosis steps** (in order of evidence strength): 1. Check the init event for the correct run (watch for task restarts!): ```bash # Find ALL starts — tasks can restart with different profiles grep "Starting agent.*{task_id}" logs/backend_7777.log # Then filter debug logs by the correct date/time grep "AGENT_DEBUG.*{task_id}.*system.*init" logs/backend_7777.log ``` Look for `slash_commands` — does the skill name appear? If NO → file missing or `--add-dir` wrong → **infra bug**. If YES → skill was **discovered** but this does NOT mean body was loaded. 2. Check the **transcript file** (definitive evidence): Claude Code saves full transcripts at (replace `<user>` with your macOS / Linux username and `<repo>` with the repo dir): `~/.claude/projects/-Users-<user>-Desktop-<repo>-tasks-{task_id}/{session_id}.jsonl` Search for unique strings from the SKILL.md body: ```bash # Find the transcript file ls ~/.claude/projects/*{task_id}*/*.jsonl # Search for SKILL.md body content (use strings unique to the skill) # For browser-use: "browser-use doctor", "browser-use tunnel", # "Browser Automation with browser-use CLI" grep -c "browser-use doctor" <transcript.jsonl> grep -c "Browser Automation with browser-use CLI" <transcript.jsonl> ``` If count is 0 → **skill body was NEVER loaded into context**. If count > 0 → skill body was loaded. 3. Check if agent ever explicitly Read the skill file: ```bash sqlite3 ~/.vibe-seller/data/vibe_seller.db \ "SELECT content FROM task_messages \
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub