- name
- agent-browser
- description
- Web browser automation CLI for AI agents. Use when the user needs to initialize or verify browser automation, interact with websites, navigate pages, fill forms, click buttons, take screenshots, extract data, test web apps, log in to sites, or automate any browser task.
- allowed-tools
- Bash(core/bin/agent-browser:*)
# Browser Automation with agent-browser
The CLI uses Chrome/Chromium via CDP directly through the local wrapper script. Use `core/bin/agent-browser` for all commands in this skill.
## Init Action
Use the init action before browser automation when `core/bin/agent-browser` may not be ready, Chrome may be missing, the environment connection method is unknown, or the user asks to initialize or verify browser automation.
Follow [references/init.md](references/init.md) for environment detection, Chrome connection setup, verification, and workflow output requirements. Follow [references/troubleshooting.md](references/troubleshooting.md) for known browser setup and connection failures.
## Core Workflow
Every browser automation follows this pattern:
1. **Navigate**: `core/bin/agent-browser open <url>`
2. **Snapshot**: `core/bin/agent-browser snapshot -i` (get element refs like `@e1`, `@e2`)
3. **Interact**: Use refs to click, fill, select
4. **Re-snapshot**: After navigation or DOM changes, get fresh refs
```bash
core/bin/agent-browser open https://example.com/form
core/bin/agent-browser snapshot -i
# Output: @e1 [input type="email"], @e2 [input type="password"], @e3 [button] "Submit"
core/bin/agent-browser fill @e1 "user@example.com"
core/bin/agent-browser fill @e2 "password123"
core/bin/agent-browser click @e3
core/bin/agent-browser wait --load networkidle
core/bin/agent-browser snapshot -i # Check result
```
## Command Chaining
Commands can be chained with `&&` in a single shell invocation. The browser persists between commands via a background daemon, so chaining is safe and more efficient than separate calls.
```bash
# Chain open + wait + snapshot in one call
core/bin/agent-browser open https://example.com && core/bin/agent-browser wait --load networkidle && core/bin/agent-browser snapshot -i
# Chain multiple interactions
core/bin/agent-browser fill @e1 "user@example.com" && core/bin/agent-browser fill @e2 "password123" && core/bin/agent-browser click @e3
# Navigate and capture
core/bin/agent-browser open https://example.com && core/bin/agent-browser wait --load networkidle && core/bin/agent-browser screenshot page.png
```
**When to chain:** Use `&&` when you don't need to read the output of an intermediate command before proceeding (e.g., open + wait + screenshot). Run commands separately when you need to parse the output first (e.g., snapshot to discover refs, then interact using those refs).
## Demo Mode (Showing the User How to Use a Website)
When the task is to demo a website or show the user how to use one, decide the browser mode based on whether the target site requires a password:
**Local development or any site that does NOT need a password** — just run in headed mode with a fresh session. Do NOT use the user's daily browser.
```bash
core/bin/agent-browser --headed open <url>
```
**Third-party sites that require login** (AWS, GitHub, Google, banks, SaaS dashboards, etc.) — **ask the user first** which browser to use:
1. **Open a new browser session** — clean profile, no saved logins. The user will need to enter credentials manually.
2. **Use the user's daily Chrome browser** — reuses existing logins/saved passwords, so the user does not need to re-enter credentials.
If the user chooses option 2, instruct them: "Please make sure your daily Chrome is open and signed in with the Google account you normally use. Let me know once it's ready." Then connect with `--auto-connect`:
```bash
core/bin/agent-browser --auto-connect open <url>
core/bin/agent-browser --auto-connect snapshot -i
```
If `--auto-connect` fails, Chrome may not have remote debugging enabled. See [references/init.md](references/init.md) for setup and [references/troubleshooting.md](references/troubleshooting.md) for known connection failures.
## Handling Authentication
When automating a site that requires login, choose the approach that fits:
**Option 1: Import auth from the user's browser (fastest for one-off tasks)**
```bash
# Connect to the user's running Chrome (they're already logged in)
core/bin/agent-browser --auto-connect state save ./auth.json
# Use that auth state
core/bin/agent-browser --state ./auth.json open https://app.example.com/dashboard
```
State files contain session tokens in plaintext -- add to `.gitignore` and delete when no longer needed. Set `AGENT_BROWSER_ENCRYPTION_KEY` for encryption at rest.
**Option 2: Persistent profile (simplest for recurring tasks)**
```bash
# First run: login manually or via automation
core/bin/agent-browser --profile ~/.myapp open https://app.example.com/login
# ... fill credentials, submit ...
# All future runs: already authenticated
core/bin/agent-browser --profile ~/.myapp open https://app.example.com/dashboard
```
**Option 3: Session name (auto-save/restore cookies + localStorage)**
```bash
core/bin/agent-browser --session-name myapp open https://app.example.com/login
# ... login flow ...
core/bin/agent-browser close # State auto-saved
# Next time: state auto-restored
core/bin/agent-browser --session-name myapp open https://app.example.com/dashboard
```
**Option 4: Auth vault (credentials stored encrypted, login by name)**
```bash
echo "$PASSWORD" | core/bin/agent-browser auth save myapp --url https://app.example.com/login --username user --password-stdin
core/bin/agent-browser auth login myapp
```
`auth login` navigates with `load` and then waits for login form selectors to appear before filling/clicking, which is more reliable on delayed SPA login screens.
**Option 5: State file (manual save/load)**
```bash
# After logging in:
core/bin/agent-browser state save ./auth.json
# In a future session:
core/bin/agent-browser state load ./auth.json
core/bin/agent-browser open https://app.example.com/dashboard
```
See [references/authentication.md](references/authentication.md) for OAuth, 2FA, cookie-based auth, and token refresh patterns.
## Essential Commands
```bash
# Navigation
core/bin/agent-browser open <url> # Navigate (aliases: goto, navigate)
core/bin/agent-browser close # Close browser
# Snapshot
core/bin/agent-browser snapshot -i # Interactive elements with refs (recommended)
core/bin/agent-browser snapshot -i -C # Include cursor-interactive elements (divs with onclick, cursor:pointer)
core/bin/agent-browser snapshot -s "#selector" # Scope to CSS selector
# Interaction (use @refs from snapshot)
core/bin/agent-browser click @e1 # Click element
core/bin/agent-browser click @e1 --new-tab # Click and open in new tab
core/bin/agent-browser fill @e2 "text" # Clear and type text
core/bin/agent-browser type @e2 "text" # Type without clearing
core/bin/agent-browser select @e1 "option" # Select dropdown option
core/bin/agent-browser check @e1 # Check checkbox
core/bin/agent-browser press Enter # Press key
core/bin/agent-browser keyboard type "text" # Type at current focus (no selector)
core/bin/agent-browser keyboard inserttext "text" # Insert without key events
core/bin/agent-browser scroll down 500 # Scroll page
core/bin/agent-browser scroll down 500 --selector "div.content" # Scroll within a specific container
# Get information
core/bin/agent-browser get text @e1 # Get element text
core/bin/agent-browser get url # Get current URL
core/bin/agent-browser get title # Get page title
core/bin/agent-browser get cdp-url # Get CDP WebSocket URL
# Wait
core/bin/agent-browser wait @e1 # Wait for element
core/bin/agent-browser wait --load networkidle # Wait for network idle
core/bin/agent-browser wait --url "**/page" # Wait for URL pattern
core/bin/agent-browser wait 2000 # Wait milliseconds
core/bin/agent-browser wait --text "Welcome" # Wait for text to appear (substring match)
core/bin/agent-browser wait --fn "!document.body.innerText.includes('Loading...')" # Wait for text to disappear
core/bin/agent-browser wait "#spinner" --state hidden # Wait for element to disappear
# Downloads
core/bin/agent-browser download @e1 ./file.pdf # Click element to trigger download
core/bin/agent-browser wait --download ./output.zip # Wait for any download to complete
core/bin/agent-browser --download-path ./downloads open <url> # Set default download directory
# Network
core/bin/agent-browser network requests # Inspect tracked requests
core/bin/agent-browser network route "**/api/*" --abort # Block matching requests
core/bin/agent-browser network har start # Start HAR recording
core/bin/agent-browser network har stop ./capture.har # Stop and save HAR file
# Viewport & Device Emulation
core/bin/agent-browser set viewport 1920 1080 # Set viewport size (default: 1280x720)
core/bin/agent-browser set viewport 1920 1080 2 # 2x retina (same CSS size, higher res screenshots)
core/bin/agent-browser set device "iPhone 14" # Emulate device (viewport + user agent)
# Capture
core/bin/agent-browser screenshot # Screenshot to temp dir
core/bin/agent-browser screenshot --full # Full page screenshot
core/bin/agent-browser screenshot --annotate # Annotated screenshot with numbered element labels
core/bin/agent-browser screenshot --screenshot-dir ./shots # Save to custom directory
core/bin/agent-browser screenshot --screenshot-format jpeg --screenshot-quality 80
core/bin/agent-browser pdf output.pdf # Save as PDF
# Clipboard
core/bin/agent-browser clipboard read # Read text from clipboard
core/bin/agent-browser clipboard write "Hello, World!" # Write text to clipboard
core/bin/agent-browser clipboard copy # Copy current selection
core/bin/agent-browser clipboard paste # Paste from clipboard
# Diff (compare page states)
core/bin/agent-browser diff snapshot # Compare current vs last snapshot
core/bin/agent-browser diff snapshot --baseline before.txt # Compare current vs saved file
core/bin/agent-browser diff screenshot --baseline before.png # Visual pixel diff
core/bin/agent-browser diff url <url1> <url2> # Compare two pages
core/bin/agent-browser diff url <url1> <url2> --wait-until networkidle # Custom wait strategy
core/bin/agent-browser diff url <url1> <url2> --selector "#main" # Scope to element
```
## Batch Execution
Execute multiple commands in a single invocation by piping a JSON array of string arrays to `batch`. This avoids per-command process startup overhead when running multi-step workflows.
```bash
echo '[
["open", "https://example.com"],
["snapshot", "-i"],
["click", "@e1"],
["screenshot", "result.png"]
]' | core/bin/agent-browser batch --json
# Stop on first error
core/bin/agent-browser batch --bail < commands.json
```
Use `batch` when you have a known sequence of commands that don't depend on intermediate output. Use separate commands or `&&` chaining when you need to parse output between steps (e.g., snapshot to discover refs, then interact).
## Common Patterns
### Form Submission
```bash
core/bin/agent-browser open https://example.com/signup
core/bin/agent-browser snapshot -i
core/bin/agent-browser fill @e1 "Jane Doe"
core/bin/agent-browser fill @e2 "jane@example.com"
core/bin/agent-browser select @e3 "California"
core/bin/agent-browser check @e4
core/bin/agent-browser click @e5
core/bin/agent-browser wait --load networkidle
```
### Authentication with Auth Vault (Recommended)
```bash
# Save credentials once (encrypted with AGENT_BROWSER_ENCRYPTION_KEY)
# Recommended: pipe password via stdin to avoid shell history exposure
echo "pass" | core/bin/agent-browser auth save github --url https://github.com/login --username user --password-stdin
# Login using saved profile (LLM never sees password)
core/bin/agent-browser auth login github
# List/show/delete profiles
core/bin/agent-browser auth list
core/bin/agent-browser auth show github
core/bin/agent-browser auth delete github
```
`auth login` waits for username/password/submit selectors before interacting, with a timeout tied to the default action timeout.
### Authentication with State Persistence
```bash
# Login once and save state
core/bin/agent-browser open https://app.example.com/login
core/bin/agent-browser snapshot -i
core/bin/agent-browser fill @e1 "$USERNAME"
core/bin/agent-browser fill @e2 "$PASSWORD"
core/bin/agent-browser click @e3
core/bin/agent-browser wait --url "**/dashboard"
core/bin/agent-browser state save auth.json
# Reuse in future sessions
core/bin/agent-browser state load auth.json
core/bin/agent-browser open https://app.example.com/dashboard
```
### Session Persistence
```bash
# Auto-save/restore cookies and localStorage across browser restarts
core/bin/agent-browser --session-name myapp open https://app.example.com/login
# ... login flow ...
core/bin/agent-browser close # State auto-saved to ~/.agent-browser/sessions/
# Next time, state is auto-loaded
core/bin/agent-browser --session-name myapp open https://app.example.com/dashboard
# Encrypt state at rest
export AGENT_BROWSER_ENCRYPTION_KEY=$(openssl rand -hex 32)
core/bin/agent-browser --session-name secure open https://app.example.com
# Manage saved states
core/bin/agent-browser state list
core/bin/agent-browser state show myapp-default.json
core/bin/agent-browser state clear myapp
core/bin/agent-browser state clean --older-than 7
```
### Working with Iframes
Iframe content is automatically inlined in snapshots. Refs inside iframes carry frame context, so you can interact with them directly.
```bash
core/bin/agent-browser open https://example.com/checkout
core/bin/agent-browser snapshot -i
# @e1 [heading] "Checkout"
# @e2 [Iframe] "payment-frame"
# @e3 [input] "Card number"
# @e4 [input] "Expiry"
# @e5 [button] "Pay"
# Interact directly — no frame switch needed
core/bin/agent-browser fill @e3 "4111111111111111"
core/bin/agent-browser fill @e4 "12/28"
core/bin/agent-browser click @e5
# To scope a snapshot to one iframe:
core/bin/agent-browser frame @e2
core/bin/agent-browser snapshot -i # Only iframe content
core/bin/agent-browser frame main # Return to main frame
```
### Data Extraction
```bash
core/bin/agent-browser open https://example.com/products
core/bin/agent-browser snapshot -i
core/bin/agent-browser get text @e5 # Get specific element text
core/bin/agent-browser get text body > page.txt # Get all page text
# JSON output for parsing
core/bin/agent-browser snapshot -i --json
core/bin/agent-browser get text @e1 --json
```
### Parallel Sessions
```bash
core/bin/agent-browser --session site1 open https://site-a.com
core/bin/agent-browser --session site2 open https://site-b.com
core/bin/agent-browser --session site1 snapshot -i
core/bin/agent-browser --session site2 snapshot -i
core/bin/agent-browser session list
```
### Connect to Existing Chrome
```bash
# Auto-discover running Chrome with remote debugging enabled
core/bin/agent-browser --auto-connect open https://example.com
core/bin/agent-browser --auto-connect snapshot
# Or with explicit CDP port
core/bin/agent-browser --cdp 9222 snapshot
```
在 GitHub 查看