Skip to main content

agent-browser

This skill should be used when automating browser interactions via Vercel's agent-browser CLI. It handles web page navigation, form filling, screenshots, and data scraping using ref-based element selection.

Source facts

Repository
jikig-ai/soleur
Last source activity
October 5, 2026 at 12:33
Detected SKILL.md language
English
Stars
16
Forks
5

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
55 files

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
agent-browser
description
This skill should be used when automating browser interactions via Vercel's agent-browser CLI. It handles web page navigation, form filling, screenshots, and data scraping using ref-based element selection.
<!-- grok-harness-invoke:start --> **Grok Build (`plugins/soleur/lib/harness.ts` `invokeSkill()`):** Read this SKILL.md in this process and run it to completion. A one-segment `soleur:<name>` in this document names a SKILL — on Grok Build, Read `plugins/soleur/skills/<name>/SKILL.md` in this process; it is not a nested tool_use. A multi-segment id such as `soleur:<domain>:<name>` names an AGENT: spawn it, never Read it, and on Grok Build spawn_subagent takes the id with its colons replaced by hyphens (`agentIdToGrokSubagentType`). **Claude Code:** Skill tool for a skill (`soleur:<name>`), Task tool with `subagent_type` for an agent. Forbidden is executing a subset, not the Read. <!-- grok-harness-invoke:end --> # agent-browser: CLI Browser Automation Vercel's headless browser automation CLI designed for AI agents. Uses ref-based selection (@e1, @e2) from accessibility snapshots. ## Setup Check ```bash # Check installation command -v agent-browser >/dev/null 2>&1 && echo "Installed" || echo "NOT INSTALLED - run: npm install --prefix ~/.local -g agent-browser@0.22.3 && agent-browser install" ``` ### Install if needed ```bash npm install --prefix ~/.local -g agent-browser@0.22.3 agent-browser install # Downloads Chrome for Testing (~300MB) # On Linux if system deps missing: # agent-browser install --with-deps ``` ### Required launch flag on Linux: `--no-sandbox` On Ubuntu 23.10+, containers, and VMs, the host's AppArmor policy restricts unprivileged user namespaces, so Chrome for Testing cannot initialize its zygote sandbox and the browser fails to launch. Export the no-sandbox flag once per session **before the first `agent-browser` command** — the daemon reads it at launch and every later command in the session inherits it: ```bash export AGENT_BROWSER_ARGS="--no-sandbox" ``` Inline alternative (first `open` only): `agent-browser open <url> --args "--no-sandbox"`. This confines `--no-sandbox` to agent-browser's own ephemeral automation Chrome — the same posture Playwright already runs here; it does not touch your real browser. ### Troubleshooting: Chrome fails to launch / `open` hangs (no usable sandbox) Symptom on pinned 0.22.3: `agent-browser open <url>` **hangs indefinitely** with zero stdout/stderr (even `--debug` prints nothing). On newer versions it fails fast with `No usable sandbox! ... unprivileged user namespaces ... AppArmor` and a `--args "--no-sandbox"` hint. Both are the same cause. 1. Set the launch flag: `export AGENT_BROWSER_ARGS="--no-sandbox"` (see above). 2. If it still fails, a stale/wedged daemon may be holding the socket. Clear it: `pkill -f agent-browser-linux-x64; rm -rf /tmp/agent-browser/* "/run/user/$(id -u)/agent-browser/"*` then retry. (Never kill `playwright-mcp` processes — those are a separate stack.) The commonest cause is a daemon whose worktree was REAPED: resolve each match's `/proc/<pid>/cwd` and expect one ending `(deleted)`, often weeks old and inherited by every later session on the machine. Such a daemon does not answer and **survives SIGTERM** — it needs `kill -9`. Nothing in the CLI's error names any of this; the only symptom is `Resource temporarily unavailable (os error 11)` (#7947). 3. Verify: `AGENT_BROWSER_ARGS="--no-sandbox" timeout 45 agent-browser open https://example.com --headless` → exit 0 + a `✓` line. ### Troubleshooting: Playwright MCP backend closed between calls This is the **other** browser-automation symptom #6605 reported (the "MCP tools de-register" half) — distinct from the agent-browser CLI hang above, and covering the Playwright **MCP** stack. If a `mcp__plugin_soleur_playwright__browser_*` call (or the same call on a host's own `mcp__playwright__*` registration) returns `browserBackend.callTool: Target page, context or browser has been closed`, the browser backend dropped while the MCP server itself stayed registered (a lifecycle event, not a dead tool). Two known causes, check in this order. **(1) A hook killing the browser at the end of a turn.** Older plugin versions registered `browser-cleanup-hook.sh` on `Stop`, which fires every turn and SIGTERMed every Playwright Chrome of the user (ADR-271; method in `knowledge-base/project/learnings/bug-fixes/2026-10-04-stop-hook-killed-live-playwright-chrome-heartbeat-theory-refuted.md`). Tell-tale: the next call after a turn boundary fails, and the session transcript (`~/.claude/projects/<slug>/<session>.jsonl`) carries a `hook_success` attachment whose `hookEvent` is `Stop` and whose `command` matches `browser-cleanup`; the kill text `Browser cleanup: killed N orphaned Playwright Chrome process(es)` is in `attachment.stderr`, not `stdout`. The hook is removed from the repository, but a session may run the installed plugin copy, which keeps it until the plugin is updated: 1. Check (read-only): take `installPath` for `soleur@soleur` from `~/.claude/plugins/installed_plugins.json`, then run `jq -c '.hooks.Stop[].hooks[].command' <installPath>/hooks/hooks.json`. A line naming `browser-cleanup-hook.sh` means the installed copy still kills the browser. 2. Update: `claude plugin update soleur@soleur` (or the id `claude plugin list` prints). The installed version is a git sha, so the update carries the fix only once the release that removes the hook has been published; before that it changes nothing. 3. Restart: the USER restarts Claude Code (an agent cannot restart its own host). A `/mcp` reconnect reuses the cached `.mcp.json` command and does not reload hooks. 4. Until then, complete a whole browser flow inside ONE turn. Interim and untested: the user can `chmod -x` the installed copy's `hooks/browser-cleanup-hook.sh`, which should make that hook fail without blocking until the next update overwrites it. **A second killer survives the update: a session in ANOTHER checkout still on the old `.mcp.json` launch string.** That string contains `pkill -9 -f` and, on every start or `/mcp` reconnect of that session, kills the proxy, server and Chrome of EVERY slot of a session on the new string. Read-only check: `for w in $(git worktree list --porcelain | sed -n 's/^worktree //p'); do grep -aq pkill "$w/.mcp.json" && echo "$w"; done` prints each checkout still carrying it. Bring those checkouts to this `.mcp.json` and restart their sessions (the string a running session loaded cannot be changed from here). **A login that should persist is missing in the repo registration:** that session was given a numbered slot or a `-p<pid>` directory. Run `grep -a 'was skipped:' "$(ls -t ~/.cache/claude-cli-nodejs/"$(pwd | sed 's/[^A-Za-z0-9]/-/g')"/mcp-logs-playwright/*.jsonl | head -1)"` from the project directory: the repo registration logs under `mcp-logs-playwright` in the directory named after the working directory with every character that is not a letter or digit replaced by a dash (`/.worktrees/` becomes `--worktrees-`). The plugin registration's `mcp-logs-plugin-soleur-playwright` never runs the slot script and never prints that line. The line names why slot 0 (the persistent profile) was not used, and the remedy depends on the reason: - "its lock is held by another live launch": the lease is held by the other session's proxy process, and `browser_close` does not release it (it closes Chrome, not the proxy). End that session, or disconnect its Playwright MCP server, then reconnect. - "its SingletonLock names a live Chrome": that Chrome must exit, for example `browser_close` in the session that owns it, then reconnect. - A reason naming another HOST (a renamed host or a copied home directory) is cleared by hand: after confirming no Chrome of yours uses that profile, `rm <slot-0 dir>/Singleton*`, which removes only the stale lock files and no login. A search for another killer must cover every enabled plugin's `hooks.json` (the installed copy too), not only `settings.json`. The server's ping heartbeat is not a cause on stdio (ADR-271). **(2) A Wayland/Vulkan GPU crash** — already diagnosed and remediated in `.claude/playwright-mcp.config.json` (forces the X11/XWayland backend and disables the GPU); see `knowledge-base/project/learnings/workflow-patterns/2026-06-17-playwright-mcp-wayland-vulkan-launch-crash.md`. **Whatever the cause**, recycle the context and re-navigate (the pattern in `plugins/soleur/skills/qa/SKILL.md`: `browser_close` — safe even if already closed — then `browser_navigate`); the backend restarts. Note that snapshot `ref=` handles do **not** survive the restart; target elements by name/selector (`button:has-text("Save")`, `input[aria-label="..."]`) across it. A separate `"these deferred tools are no longer available"` notice means the MCP server disconnected (reload via `ToolSearch`) — a different failure from the backend-close. ### Troubleshooting: version mismatch If you see "Version mismatch between agent-browser (expects 1200) and installed Playwright (1208)": 1. Check which binary is running: `which agent-browser && agent-browser --version` 2. If it resolves to `/usr/bin/agent-browser` (version 0.5.0), a stale system install is shadowing the correct version 3. Fix: `sudo npm uninstall -g agent-browser` to remove the system binary 4. Verify: `which agent-browser` should now resolve to `~/.local/bin/agent-browser` (0.22.3) ## Core Workflow **The snapshot + ref pattern is optimal for LLMs:** 1. **Navigate** to URL 2. **Snapshot** to get interactive elements with refs 3. **Interact** using refs (@e1, @e2, etc.) 4. **Re-snapshot** after navigation or DOM changes **Playwright MCP flow: when the flow is finished, call `browser_close`.** Nothing reaps the browser at turn end (ADR-271): it lives as long as the session that launched it, and a logged-in window left open after a credential hand-off stays logged in. For the `agent-browser` CLI the same duty is `agent-browser close`. ### Preflight: verify the plugin install before any snapshot The redactor is reached through `${CLAUDE_PLUGIN_ROOT}`. An ambient value pointing at a directory that is not a Soleur install would resolve to a path that does not exist — or, worse, to one an attacker chose. Verify plugin IDENTITY and halt if it does not hold (ADR-179 decision 2); a `test -f` on the script alone is a shape check and was measured bypassable. ```bash [ -f "${CLAUDE_PLUGIN_ROOT}/.claude-plugin/plugin.json" ] \ && grep -q '"name"[[:space:]]*:[[:space:]]*"soleur"' "${CLAUDE_PLUGIN_ROOT}/.claude-plugin/plugin.json" \ || { echo "SOLEUR_SNAPSHOT_HALT reason=plugin-root-unverified root=[${CLAUDE_PLUGIN_ROOT}]" >&2 echo " Cannot locate the snapshot redactor, so no accessibility snapshot may be taken here." >&2 echo " Root EMPTY: no Soleur plugin is loaded in this session. Install it and start a NEW session." >&2 echo " Root set but wrong: a repo checkout is not an install. Run 'claude plugin update soleur@soleur-marketplace' (or the id 'claude plugin list' prints, if you added the repository directly), then RESTART Claude Code." >&2 echo " Nothing has been captured yet, so nothing has leaked." >&2 exit 2; } ``` ```bash # Step 1: Open URL agent-browser open https://example.com # Step 2: Get interactive elements with refs agent-browser snapshot -i --json 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" # Step 3: Interact using refs agent-browser click @e1 agent-browser fill @e2 "search query" # Step 4: Re-snapshot after changes agent-browser snapshot -i 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" ``` ## Key Commands ### Navigation ```bash agent-browser open <url> # Navigate to URL agent-browser back # Go back agent-browser forward # Go forward agent-browser reload # Reload page agent-browser close # Close browser ``` ### Snapshots (Essential for AI) ```bash agent-browser snapshot 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" # Full accessibility tree agent-browser snapshot -i 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" # Interactive elements only (recommended) agent-browser snapshot -i --json 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" # JSON output for parsing agent-browser snapshot -c 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" # Compact (remove empty elements) agent-browser snapshot -d 3 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" # Limit depth ``` ### Interactions ```bash agent-browser click @e1 # Click element agent-browser dblclick @e1 # Double-click agent-browser fill @e1 "text" # Clear and fill input agent-browser type @e1 "text" # Type without clearing agent-browser press Enter # Press key agent-browser hover @e1 # Hover element agent-browser check @e1 # Check checkbox agent-browser uncheck @e1 # Uncheck checkbox agent-browser select @e1 "option" # Select dropdown option agent-browser scroll down 500 # Scroll (up/down/left/right) agent-browser scrollintoview @e1 # Scroll element into view ``` ### Get Information ```bash agent-browser get text @e1 # Get element text agent-browser get html @e1 # Get element HTML agent-browser get value @e1 # Get input value agent-browser get attr href @e1 # Get attribute agent-browser get title # Get page title agent-browser get url # Get current URL agent-browser get count "button" # Count matching elements ``` ### Screenshots & PDFs ```bash agent-browser screenshot # Viewport screenshot agent-browser screenshot --full # Full page agent-browser screenshot output.png # Save to file agent-browser screenshot --full output.png # Full page to file agent-browser pdf output.pdf # Save as PDF ``` ### Wait ```bash agent-browser wait @e1 # Wait for element agent-browser wait 2000 # Wait milliseconds agent-browser wait "text" # Wait for text to appear ``` ## Semantic Locators (Alternative to Refs) ```bash agent-browser find role button click --name "Submit" agent-browser find text "Sign up" click agent-browser find label "Email" fill "user@example.com" agent-browser find placeholder "Search..." fill "query" ``` ## Sessions (Parallel Browsers) ```bash # Run multiple independent browser sessions agent-browser --session-name browser1 open https://site1.com agent-browser --session-name browser2 open https://site2.com # List saved states agent-browser state list ``` ## Examples ### Credential safety on a login or credential page An accessibility snapshot serializes the **value** of input fields. A value the agent never typed — a password manager's autofill, a static `value=`, a JS assignment, or a freshly-minted credential shown in a panel — is rendered into the transcript and into any snapshot file written to disk. Nothing about the call looks credential-adjacent, which is why this is a gate and not advice: the PreToolUse hook `browser-snapshot-credential-guard.sh` denies an unrouted snapshot before it runs (#7947). **Route every snapshot on a credential-bearing page through the redactor**, [redact-a11y-snapshot.py](./scripts/redact-a11y-snapshot.py): ```bash agent-browser snapshot -i 2>&1 | python3 "${CLAUDE_PLUGIN_ROOT}/skills/agent-browser/scripts/redact-a11y-snapshot.py" ``` Stating the ceiling first, because the rule is otherwise read as "piping makes a snapshot safe": the redactor is **defense-in-depth on one sink**, not a control that makes snapshotting a credential page safe. It keys on the node's accessible name, because neither surface serializes the input's `type` — measured, see the
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub