Skip to main content

debug-task

Debug agent task execution infrastructure issues โ€” wrong cwd, wrong files, wrong venv, browser startup failures, agent retry loops, skill loading failures

Jump to install

Source facts

Repository
zpoint/vibe-seller
Last source activity
July 27, 2026 at 15:42
Detected SKILL.md language
English
Stars
68
Forks
14

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions ยท Read-only preview
name
debug-task
description
Debug agent task execution infrastructure issues โ€” wrong cwd, wrong files, wrong venv, browser startup failures, agent retry loops, skill loading failures
# Debug Task Infra Debug agent task execution infrastructure issues by analyzing agent logs. ## Core Philosophy **Every agent failure, detour, or workaround is an infra/platform bug โ€” not an agent problem.** When an agent: - Tries a wrong file path then searches for the right one โ†’ **catalog or path resolution bug** - Uses wrong CLI syntax then self-corrects โ†’ **skill docs not loaded or incomplete** - Takes a screenshot but can't find/view it โ†’ **tool integration bug** - Retries a command with different args โ†’ **missing docs or unclear API** - Reflects "no failures" when there were failures โ†’ **reflection prompt bug** Even if the agent eventually succeeds, each detour: 1. Wastes tokens and time 2. Signals a gap in the platform that WILL hit other tasks 3. Must be traced to a specific infra root cause and filed as a bug **Do NOT dismiss detours as "the agent will self-correct."** Your job is to find every single failure step, trace it to a code/config/prompt root cause, and report it. ## When to use When a task execution shows symptoms like: - Agent sees unexpected files or can't find expected ones - Agent runs in wrong working directory - Browser (Ziniao/Chrome) fails to start or takes multiple retries - Task uses wrong Python environment - Store context (bookmarks, browser config) missing from agent prompt - Agent repeats same bash command 3+ times (infra issue, not agent fault) - Agent repeats browser-open attempts (browser startup failure) - Agent can't load skills (wrong skill path or skill not synced) - Agent passes wrong arguments to tools (schema mismatch) - Agent uses workarounds for things that should "just work" - Agent's post-task reflection misses failures that clearly happened ## Environment first: are you in WSL debugging a native-Windows server? **Before running ANY command below, figure out where the server actually runs.** vibe-seller ships as a native-Windows install *and* as a Linux/WSL install, and the two do not share a data directory. Debugging the wrong one is the single biggest time-sink โ€” you'll read a stale DB, grep empty logs, and "fix" a store that isn't the one serving traffic. **Detect the topology:** ```bash # The API answers, but is there a Python server process on the WSL side? curl -s -m3 http://127.0.0.1:7777/api/auth/me # 200 with a user โ†’ a server is up ps aux | grep -Ei 'uvicorn|app\.main' | grep -v grep # empty โ†’ NOT served from WSL # If empty, the server is the native-Windows process (WSL2 mirrored # networking makes 127.0.0.1:7777 reach the Windows listener): powershell.exe -NoProfile -Command "Get-Process python,pythonw -ErrorAction SilentlyContinue | Select Id,Path" # โ†’ paths under C:\Users\<WinUser>\AppData\Local\Programs\VibeSeller\ confirm it ``` **Need to call the HTTP API from this skill (not just read the DB)?** Get a session cookie via the JWT-auth workaround described in `debug-store/SKILL.md` ยง "Skipping JWT-cookie auth" โ€” it uses a single stable `taskbot_debug` account, rotates the password per session, and deactivates on exit. **When the live server is native-Windows, everything relocates.** The real runtime data root is **`/mnt/c/Users/<WinUser>/.vibe-seller/`** (commonly `<WinUser>=Administrator`), NOT `~/.vibe-seller`. Re-point every path in this doc: | This doc says | On a WSLโ†’Windows box, use | |---|---| | `~/.vibe-seller/data/vibe_seller.db` | `/mnt/c/Users/<WinUser>/.vibe-seller/data/vibe_seller.db` | | `logs/backend_7777.log` | `/mnt/c/Users/<WinUser>/.vibe-seller/logs/` | | `~/.vibe-seller/tasks/{id}/` | `/mnt/c/Users/<WinUser>/.vibe-seller/tasks/{id}/` | | `~/.vibe-seller/bin/<slug>/browser-use` | `/mnt/c/Users/<WinUser>/.vibe-seller/bin/<slug>/browser-use` | | `~/.vibe-seller/.claude/skills/` | `/mnt/c/Users/<WinUser>/.vibe-seller/.claude/skills/` | | app code (`app/skills/โ€ฆ`) | `โ€ฆ/Programs/VibeSeller/.venv/Lib/site-packages/app/โ€ฆ` | `sqlite3` reads the Windows DB fine over `/mnt/c`. There is often ALSO a stale `~/.vibe-seller` sandbox in WSL from earlier dev โ€” it has a *different* `stores` table (decoy store names, wrong ids). If a store you expect is missing, or the ids don't match the live `/api/stores` response, you're reading the wrong DB. ### Driving a store's `browser-use` wrapper from WSL The per-store wrapper is generated by the Windows install with **CRLF line endings**, and it `exec`s a Windows `browser-use.EXE`. Two failures bite immediately: 1. **CRLF breaks WSL bash** โ€” `exec`ing it directly gives `/usr/bin/env: 'bash\r': No such file or directory`; running it via WSL `bash` gives `set: pipefail: invalid option name`. Do NOT `dos2unix` the file (it's regenerated on every server boot, and a sanitized copy then can't find the Windows `.EXE`). Instead run it through the install's **bundled Git-for-Windows bash**, which tolerates CRLF and resolves the Windows exec path: ```bash BASHEXE='/mnt/c/Users/<WinUser>/AppData/Local/Programs/VibeSeller/git/bin/bash.exe' WU_WIN='C:\Users\<WinUser>\.vibe-seller\bin\<slug>\browser-use' # native path, backslashes TASK_ID="$(uuidgen | tr '[:upper:]' '[:lower:]')" run(){ "$BASHEXE" -c "export VIBE_TASK_ID='$TASK_ID'; export PYTHONIOENCODING=utf-8 PYTHONUTF8=1; '$WU_WIN' $1"; } run "open https://sellercentral.amazon.<tld>/home" run "state" ``` 2. **GBK console โ†’ UnicodeEncodeError** โ€” the Windows Python console codec is GBK, so any page with Arabic/CJK text crashes `state`/`eval` with `'gbk' codec can't encode character`. Always export `PYTHONIOENCODING=utf-8` and `PYTHONUTF8=1` (shown above) before the wrapper call. Downloads triggered through the CDP-proxied browser land in `/mnt/c/Users/<WinUser>/.vibe-seller/downloads/<slug>/` (the proxy overrides Chrome's download dir per `ziniao.py`). Screenshots: pass a **native Windows path** (`C:\...\downloads\<slug>\shot.png`) to `screenshot`, then Read it from the `/mnt/c` equivalent. For the browser-driving specifics (session rotation, wedged-daemon recovery, aux sessions) see the `debug-store` skill โ€” the same WSLโ†’Windows path/CRLF/UTF-8 rules apply there. ### Restarting the native-Windows server `./restart.sh` is WSL-only. To restart the Windows server, stop its `pythonw`/tray process and relaunch the installed app (see `docs/windows-setup.md`). Skills deploy by syncing repo `app/skills/` into `โ€ฆ/Programs/VibeSeller/.venv/Lib/site-packages/app/skills/`; the server's `skills_sync.fetch()` copies them to `โ€ฆ/.vibe-seller/.claude/skills/` at boot. ## Debug methodology ### RULE: Read agent logs BEFORE any conclusion **Every time this skill is invoked, you MUST read the actual agent debug logs before making any claim about what happened.** Do not assume from file existence, code reading, or task_messages alone. The agent debug logs are the source of truth. ```bash # 1. Find which run(s) happened โ€” tasks can restart with different profiles! grep "Starting agent.*{task_id}" logs/backend_7777.log # 2. Get agent debug logs for the CORRECT run (match the date/time) grep "AGENT_DEBUG.*{task_id}" logs/backend_7777.log | grep "{date}" # 3. Find the transcript file (full session log, most detailed) ls ~/.claude/projects/*{task_id}*/*.jsonl # 4. Check if a skill body was actually loaded into context # (use strings unique to that skill's SKILL.md body) grep -c "unique string from SKILL.md" <transcript.jsonl> ``` **Common mistakes to avoid:** - Reading logs from the wrong run (task restarted with different profile) - Checking transcript before the task finishes โ€” skill loading may happen mid-task, not at start. Wait for task to complete before concluding. - Assuming a skill is "loaded" because the file exists in the workspace - Assuming `slash_commands` listing = skill content in LLM context (it doesn't โ€” only metadata loads on discovery, body loads on trigger) - Making claims about what the agent "saw" without transcript evidence - Skipping the transcript file โ€” it's the only way to confirm skill body loading ### Step 0: Pick the task to debug If the user didn't specify a task ID, find the most recent task from the last hour: ```bash sqlite3 ~/.vibe-seller/data/vibe_seller.db \ "SELECT id, status, substr(result,1,100), substr(error,1,100) \ FROM tasks WHERE created_at > datetime('now', '-1 hour') \ ORDER BY created_at DESC LIMIT 5;" ``` If exactly one task was active in the last hour, debug that one. If multiple, ask the user which one. If none, ask the user for a task ID. ### Step 1: Pull ALL task messages Read every message โ€” not just errors. Detours hide in `thinking` messages where the agent says things like "let me try another approach" or "file doesn't exist, let me search." ```bash # Full task log โ€” read ALL of it, not just tail sqlite3 ~/.vibe-seller/data/vibe_seller.db \ "SELECT seq, role, content FROM task_messages \ WHERE task_id='{task_id}' ORDER BY seq;" ``` ### Step 2: Walk through sequentially and flag EVERY detour For each step, ask: "Did this step succeed on the first try?" If not, classify the failure: | Failure Type | Signal in Logs | Root Cause Category | |---|---|---| | Wrong file path โ†’ search/glob โ†’ find correct path | `Read` error โ†’ `Glob` โ†’ `Read` success | **Catalog/path resolution bug** | | Hallucinated CLI command โ†’ error โ†’ correct syntax | `Bash` error with usage/invalid choice message | **Skill docs not effective enough** | | Tool output not usable โ†’ workaround | Screenshot returns bytes, agent searches for .png | **Tool integration gap** | | File doesn't exist but agent expected it | `Read` error "File does not exist" | **Catalog lists nonexistent file, or agent guessed** | | Duplicate table headers in catalog | Agent reads garbled catalog | **Catalog generation bug** | | Browser open retried (even once) | >1 `browser-use open` to same URL | **CDP proxy / Ziniao / wrapper script bug** | | Agent says "no failures" in reflection | Contradicts actual log | **Reflection prompt doesn't force log review** | ### Step 3: For each detour, trace to root cause Every detour must map to ONE of these: 1. **Catalog bug** โ€” wrong path, missing `project/` prefix, duplicate headers, unlisted file 2. **Skill docs bug** โ€” missing command, wrong syntax example, undocumented limitation 3. **Tool integration bug** โ€” output not consumable, file not saved where expected 4. **System prompt bug** โ€” unclear instructions, missing "ONLY read catalog files" 5. **Workspace bug** โ€” missing symlinks, wrong isolation, files not synced 6. **Reflection prompt bug** โ€” agent misses failures, doesn't save learnings ### Step 4: Report findings For each bug, provide: - **Seq range**: which message steps were wasted - **What happened**: agent action โ†’ error โ†’ workaround - **Steps wasted**: count of unnecessary steps - **Root cause**: specific code/file/prompt that caused it - **Fix location**: exact file and what to change ## Known bug patterns (from past investigations) ### Catalog path resolution The L1 catalog (`app/knowledge/CATALOG.md`) lists files like `common/amazon-sites.md`. These are synced to `~/.vibe-seller/knowledge/project/common/amazon-sites.md`. But the store CATALOG.md copies L1 entries verbatim without adding the `project/` prefix. From the agent's CWD, `knowledge/` symlinks to `~/.vibe-seller/knowledge/`, so the correct relative path is `knowledge/project/common/amazon-sites.md` โ€” but the catalog says `common/amazon-sites.md`. **Fix**: `_filter_l1_for_store()` in `app/workspace/knowledge_sync.py` should prepend `project/` to L1 file paths in the store catalog, or the system prompt should tell the agent the base path. ### Duplicate table header in store CATALOG `_filter_l1_for_store()` returns a string starting with `| File | Relevance | Summary |` header. If `_build_store_catalog()` also adds a header, or if the function is called twice, the catalog gets a duplicate header row. **Check**: `app/workspace/knowledge_sync.py` lines around `_build_store_catalog` and `_filter_l1_for_store`. ### Agent hallucinates CLI commands (browser-use and others) Agent invents CLI syntax that doesn't exist instead of using the commands documented in the loaded skill. Examples seen in the wild: - `browser-use scroll 10` (correct: `browser-use scroll down --amount 500`) - `browser-use get text` without index (correct: `browser-use get text <index>`) - `browser-use screenshot` without path then searching for .png files **How skill loading works** (important for diagnosis): Skills go through 3 stages โ€” **discovery โ‰  loading โ‰  in context**: 1. **File exists** โ€” `.claude/skills/` is **copied** (not symlinked) into the task workspace from `~/.vibe-seller/.claude/skills/` (`app/workspace/manager.py` ~line 710, `shutil.copytree`) 2. **Discovered** โ€” Claude Code finds the SKILL.md via `--add-dir` and lists it in the init event's `slash_commands` array 3. **Content in LLM context** โ€” the SKILL.md content is actually sent to the LLM as part of the prompt. **This is the step that matters and the step that's hardest to verify.** **CRITICAL**: A skill appearing in `slash_commands` in the init event does NOT prove its content is in the LLM's context window. Claude Code may defer loading skill content until the skill is invoked or until a matching `allowed-tools` pattern fires. **IMPORTANT**: browser-use SKILL.md is an **upstream/official** doc from the browser-use project. Do NOT modify it to fix agent behavior issues. **Diagnosis steps** (in order of evidence strength): 1. Check the init event for the correct run (watch for task restarts!): ```bash # Find ALL starts โ€” tasks can restart with different profiles grep "Starting agent.*{task_id}" logs/backend_7777.log # Then filter debug logs by the correct date/time grep "AGENT_DEBUG.*{task_id}.*system.*init" logs/backend_7777.log ``` Look for `slash_commands` โ€” does the skill name appear? If NO โ†’ file missing or `--add-dir` wrong โ†’ **infra bug**. If YES โ†’ skill was **discovered** but this does NOT mean body was loaded. 2. Check the **transcript file** (definitive evidence): Claude Code saves full transcripts at (replace `<user>` with your macOS / Linux username and `<repo>` with the repo dir): `~/.claude/projects/-Users-<user>-Desktop-<repo>-tasks-{task_id}/{session_id}.jsonl` Search for unique strings from the SKILL.md body: ```bash # Find the transcript file ls ~/.claude/projects/*{task_id}*/*.jsonl # Search for SKILL.md body content (use strings unique to the skill) # For browser-use: "browser-use doctor", "browser-use tunnel", # "Browser Automation with browser-use CLI" grep -c "browser-use doctor" <transcript.jsonl> grep -c "Browser Automation with browser-use CLI" <transcript.jsonl> ``` If count is 0 โ†’ **skill body was NEVER loaded into context**. If count > 0 โ†’ skill body was loaded. 3. Check if agent ever explicitly Read the skill file: ```bash sqlite3 ~/.vibe-seller/data/vibe_seller.db \ "SELECT content FROM task_messages \
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub