Skip to main content

agent-browser

Web browser automation CLI for AI agents. Use when the user needs to initialize or verify browser automation, interact with websites, navigate pages, fill forms, click buttons, take screenshots, extract data, test web apps, log in to sites, or automate any browser task.

来源信息

仓库
X-School-Academy/skill-pilot
最近来源活动
2026年5月24日 02:31
检测到的 SKILL.md 语言
英语
星标
20
分支
13

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
18 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
agent-browser
description
Web browser automation CLI for AI agents. Use when the user needs to initialize or verify browser automation, interact with websites, navigate pages, fill forms, click buttons, take screenshots, extract data, test web apps, log in to sites, or automate any browser task.
allowed-tools
Bash(core/bin/agent-browser:*)
# Browser Automation with agent-browser The CLI uses Chrome/Chromium via CDP directly through the local wrapper script. Use `core/bin/agent-browser` for all commands in this skill. ## Init Action Use the init action before browser automation when `core/bin/agent-browser` may not be ready, Chrome may be missing, the environment connection method is unknown, or the user asks to initialize or verify browser automation. Follow [references/init.md](references/init.md) for environment detection, Chrome connection setup, verification, and workflow output requirements. Follow [references/troubleshooting.md](references/troubleshooting.md) for known browser setup and connection failures. ## Core Workflow Every browser automation follows this pattern: 1. **Navigate**: `core/bin/agent-browser open <url>` 2. **Snapshot**: `core/bin/agent-browser snapshot -i` (get element refs like `@e1`, `@e2`) 3. **Interact**: Use refs to click, fill, select 4. **Re-snapshot**: After navigation or DOM changes, get fresh refs ```bash core/bin/agent-browser open https://example.com/form core/bin/agent-browser snapshot -i # Output: @e1 [input type="email"], @e2 [input type="password"], @e3 [button] "Submit" core/bin/agent-browser fill @e1 "user@example.com" core/bin/agent-browser fill @e2 "password123" core/bin/agent-browser click @e3 core/bin/agent-browser wait --load networkidle core/bin/agent-browser snapshot -i # Check result ``` ## Command Chaining Commands can be chained with `&&` in a single shell invocation. The browser persists between commands via a background daemon, so chaining is safe and more efficient than separate calls. ```bash # Chain open + wait + snapshot in one call core/bin/agent-browser open https://example.com && core/bin/agent-browser wait --load networkidle && core/bin/agent-browser snapshot -i # Chain multiple interactions core/bin/agent-browser fill @e1 "user@example.com" && core/bin/agent-browser fill @e2 "password123" && core/bin/agent-browser click @e3 # Navigate and capture core/bin/agent-browser open https://example.com && core/bin/agent-browser wait --load networkidle && core/bin/agent-browser screenshot page.png ``` **When to chain:** Use `&&` when you don't need to read the output of an intermediate command before proceeding (e.g., open + wait + screenshot). Run commands separately when you need to parse the output first (e.g., snapshot to discover refs, then interact using those refs). ## Demo Mode (Showing the User How to Use a Website) When the task is to demo a website or show the user how to use one, decide the browser mode based on whether the target site requires a password: **Local development or any site that does NOT need a password** — just run in headed mode with a fresh session. Do NOT use the user's daily browser. ```bash core/bin/agent-browser --headed open <url> ``` **Third-party sites that require login** (AWS, GitHub, Google, banks, SaaS dashboards, etc.) — **ask the user first** which browser to use: 1. **Open a new browser session** — clean profile, no saved logins. The user will need to enter credentials manually. 2. **Use the user's daily Chrome browser** — reuses existing logins/saved passwords, so the user does not need to re-enter credentials. If the user chooses option 2, instruct them: "Please make sure your daily Chrome is open and signed in with the Google account you normally use. Let me know once it's ready." Then connect with `--auto-connect`: ```bash core/bin/agent-browser --auto-connect open <url> core/bin/agent-browser --auto-connect snapshot -i ``` If `--auto-connect` fails, Chrome may not have remote debugging enabled. See [references/init.md](references/init.md) for setup and [references/troubleshooting.md](references/troubleshooting.md) for known connection failures. ## Handling Authentication When automating a site that requires login, choose the approach that fits: **Option 1: Import auth from the user's browser (fastest for one-off tasks)** ```bash # Connect to the user's running Chrome (they're already logged in) core/bin/agent-browser --auto-connect state save ./auth.json # Use that auth state core/bin/agent-browser --state ./auth.json open https://app.example.com/dashboard ``` State files contain session tokens in plaintext -- add to `.gitignore` and delete when no longer needed. Set `AGENT_BROWSER_ENCRYPTION_KEY` for encryption at rest. **Option 2: Persistent profile (simplest for recurring tasks)** ```bash # First run: login manually or via automation core/bin/agent-browser --profile ~/.myapp open https://app.example.com/login # ... fill credentials, submit ... # All future runs: already authenticated core/bin/agent-browser --profile ~/.myapp open https://app.example.com/dashboard ``` **Option 3: Session name (auto-save/restore cookies + localStorage)** ```bash core/bin/agent-browser --session-name myapp open https://app.example.com/login # ... login flow ... core/bin/agent-browser close # State auto-saved # Next time: state auto-restored core/bin/agent-browser --session-name myapp open https://app.example.com/dashboard ``` **Option 4: Auth vault (credentials stored encrypted, login by name)** ```bash echo "$PASSWORD" | core/bin/agent-browser auth save myapp --url https://app.example.com/login --username user --password-stdin core/bin/agent-browser auth login myapp ``` `auth login` navigates with `load` and then waits for login form selectors to appear before filling/clicking, which is more reliable on delayed SPA login screens. **Option 5: State file (manual save/load)** ```bash # After logging in: core/bin/agent-browser state save ./auth.json # In a future session: core/bin/agent-browser state load ./auth.json core/bin/agent-browser open https://app.example.com/dashboard ``` See [references/authentication.md](references/authentication.md) for OAuth, 2FA, cookie-based auth, and token refresh patterns. ## Essential Commands ```bash # Navigation core/bin/agent-browser open <url> # Navigate (aliases: goto, navigate) core/bin/agent-browser close # Close browser # Snapshot core/bin/agent-browser snapshot -i # Interactive elements with refs (recommended) core/bin/agent-browser snapshot -i -C # Include cursor-interactive elements (divs with onclick, cursor:pointer) core/bin/agent-browser snapshot -s "#selector" # Scope to CSS selector # Interaction (use @refs from snapshot) core/bin/agent-browser click @e1 # Click element core/bin/agent-browser click @e1 --new-tab # Click and open in new tab core/bin/agent-browser fill @e2 "text" # Clear and type text core/bin/agent-browser type @e2 "text" # Type without clearing core/bin/agent-browser select @e1 "option" # Select dropdown option core/bin/agent-browser check @e1 # Check checkbox core/bin/agent-browser press Enter # Press key core/bin/agent-browser keyboard type "text" # Type at current focus (no selector) core/bin/agent-browser keyboard inserttext "text" # Insert without key events core/bin/agent-browser scroll down 500 # Scroll page core/bin/agent-browser scroll down 500 --selector "div.content" # Scroll within a specific container # Get information core/bin/agent-browser get text @e1 # Get element text core/bin/agent-browser get url # Get current URL core/bin/agent-browser get title # Get page title core/bin/agent-browser get cdp-url # Get CDP WebSocket URL # Wait core/bin/agent-browser wait @e1 # Wait for element core/bin/agent-browser wait --load networkidle # Wait for network idle core/bin/agent-browser wait --url "**/page" # Wait for URL pattern core/bin/agent-browser wait 2000 # Wait milliseconds core/bin/agent-browser wait --text "Welcome" # Wait for text to appear (substring match) core/bin/agent-browser wait --fn "!document.body.innerText.includes('Loading...')" # Wait for text to disappear core/bin/agent-browser wait "#spinner" --state hidden # Wait for element to disappear # Downloads core/bin/agent-browser download @e1 ./file.pdf # Click element to trigger download core/bin/agent-browser wait --download ./output.zip # Wait for any download to complete core/bin/agent-browser --download-path ./downloads open <url> # Set default download directory # Network core/bin/agent-browser network requests # Inspect tracked requests core/bin/agent-browser network route "**/api/*" --abort # Block matching requests core/bin/agent-browser network har start # Start HAR recording core/bin/agent-browser network har stop ./capture.har # Stop and save HAR file # Viewport & Device Emulation core/bin/agent-browser set viewport 1920 1080 # Set viewport size (default: 1280x720) core/bin/agent-browser set viewport 1920 1080 2 # 2x retina (same CSS size, higher res screenshots) core/bin/agent-browser set device "iPhone 14" # Emulate device (viewport + user agent) # Capture core/bin/agent-browser screenshot # Screenshot to temp dir core/bin/agent-browser screenshot --full # Full page screenshot core/bin/agent-browser screenshot --annotate # Annotated screenshot with numbered element labels core/bin/agent-browser screenshot --screenshot-dir ./shots # Save to custom directory core/bin/agent-browser screenshot --screenshot-format jpeg --screenshot-quality 80 core/bin/agent-browser pdf output.pdf # Save as PDF # Clipboard core/bin/agent-browser clipboard read # Read text from clipboard core/bin/agent-browser clipboard write "Hello, World!" # Write text to clipboard core/bin/agent-browser clipboard copy # Copy current selection core/bin/agent-browser clipboard paste # Paste from clipboard # Diff (compare page states) core/bin/agent-browser diff snapshot # Compare current vs last snapshot core/bin/agent-browser diff snapshot --baseline before.txt # Compare current vs saved file core/bin/agent-browser diff screenshot --baseline before.png # Visual pixel diff core/bin/agent-browser diff url <url1> <url2> # Compare two pages core/bin/agent-browser diff url <url1> <url2> --wait-until networkidle # Custom wait strategy core/bin/agent-browser diff url <url1> <url2> --selector "#main" # Scope to element ``` ## Batch Execution Execute multiple commands in a single invocation by piping a JSON array of string arrays to `batch`. This avoids per-command process startup overhead when running multi-step workflows. ```bash echo '[ ["open", "https://example.com"], ["snapshot", "-i"], ["click", "@e1"], ["screenshot", "result.png"] ]' | core/bin/agent-browser batch --json # Stop on first error core/bin/agent-browser batch --bail < commands.json ``` Use `batch` when you have a known sequence of commands that don't depend on intermediate output. Use separate commands or `&&` chaining when you need to parse output between steps (e.g., snapshot to discover refs, then interact). ## Common Patterns ### Form Submission ```bash core/bin/agent-browser open https://example.com/signup core/bin/agent-browser snapshot -i core/bin/agent-browser fill @e1 "Jane Doe" core/bin/agent-browser fill @e2 "jane@example.com" core/bin/agent-browser select @e3 "California" core/bin/agent-browser check @e4 core/bin/agent-browser click @e5 core/bin/agent-browser wait --load networkidle ``` ### Authentication with Auth Vault (Recommended) ```bash # Save credentials once (encrypted with AGENT_BROWSER_ENCRYPTION_KEY) # Recommended: pipe password via stdin to avoid shell history exposure echo "pass" | core/bin/agent-browser auth save github --url https://github.com/login --username user --password-stdin # Login using saved profile (LLM never sees password) core/bin/agent-browser auth login github # List/show/delete profiles core/bin/agent-browser auth list core/bin/agent-browser auth show github core/bin/agent-browser auth delete github ``` `auth login` waits for username/password/submit selectors before interacting, with a timeout tied to the default action timeout. ### Authentication with State Persistence ```bash # Login once and save state core/bin/agent-browser open https://app.example.com/login core/bin/agent-browser snapshot -i core/bin/agent-browser fill @e1 "$USERNAME" core/bin/agent-browser fill @e2 "$PASSWORD" core/bin/agent-browser click @e3 core/bin/agent-browser wait --url "**/dashboard" core/bin/agent-browser state save auth.json # Reuse in future sessions core/bin/agent-browser state load auth.json core/bin/agent-browser open https://app.example.com/dashboard ``` ### Session Persistence ```bash # Auto-save/restore cookies and localStorage across browser restarts core/bin/agent-browser --session-name myapp open https://app.example.com/login # ... login flow ... core/bin/agent-browser close # State auto-saved to ~/.agent-browser/sessions/ # Next time, state is auto-loaded core/bin/agent-browser --session-name myapp open https://app.example.com/dashboard # Encrypt state at rest export AGENT_BROWSER_ENCRYPTION_KEY=$(openssl rand -hex 32) core/bin/agent-browser --session-name secure open https://app.example.com # Manage saved states core/bin/agent-browser state list core/bin/agent-browser state show myapp-default.json core/bin/agent-browser state clear myapp core/bin/agent-browser state clean --older-than 7 ``` ### Working with Iframes Iframe content is automatically inlined in snapshots. Refs inside iframes carry frame context, so you can interact with them directly. ```bash core/bin/agent-browser open https://example.com/checkout core/bin/agent-browser snapshot -i # @e1 [heading] "Checkout" # @e2 [Iframe] "payment-frame" # @e3 [input] "Card number" # @e4 [input] "Expiry" # @e5 [button] "Pay" # Interact directly — no frame switch needed core/bin/agent-browser fill @e3 "4111111111111111" core/bin/agent-browser fill @e4 "12/28" core/bin/agent-browser click @e5 # To scope a snapshot to one iframe: core/bin/agent-browser frame @e2 core/bin/agent-browser snapshot -i # Only iframe content core/bin/agent-browser frame main # Return to main frame ``` ### Data Extraction ```bash core/bin/agent-browser open https://example.com/products core/bin/agent-browser snapshot -i core/bin/agent-browser get text @e5 # Get specific element text core/bin/agent-browser get text body > page.txt # Get all page text # JSON output for parsing core/bin/agent-browser snapshot -i --json core/bin/agent-browser get text @e1 --json ``` ### Parallel Sessions ```bash core/bin/agent-browser --session site1 open https://site-a.com core/bin/agent-browser --session site2 open https://site-b.com core/bin/agent-browser --session site1 snapshot -i core/bin/agent-browser --session site2 snapshot -i core/bin/agent-browser session list ``` ### Connect to Existing Chrome ```bash # Auto-discover running Chrome with remote debugging enabled core/bin/agent-browser --auto-connect open https://example.com core/bin/agent-browser --auto-connect snapshot # Or with explicit CDP port core/bin/agent-browser --cdp 9222 snapshot ```
在 GitHub 查看
这个 SKILL.md 很大,SkillsMP 这里只预览前一段内容。 在 GitHub 查看