| name | browser-tools |
| description | Interactive browser automation via Chrome DevTools Protocol. Use when you need to interact with web pages, test frontends, or when user interaction with a visible browser is required. |
Browser Tools
Chrome DevTools Protocol tools for agent-assisted web automation. These tools connect to Chrome running on :9222 with remote debugging enabled.
Setup
Run once before first use:
cd <base-dir>/scripts
npm install
Start Chrome
node <base-dir>/scripts/browser-start.js
node <base-dir>/scripts/browser-start.js --profile
Launch Chrome with remote debugging on :9222. Use --profile to preserve user's authentication state.
Navigate
node <base-dir>/scripts/browser-nav.js https://example.com
node <base-dir>/scripts/browser-nav.js https://example.com --new
Navigate to URLs. Use --new flag to open in a new tab instead of reusing current tab.
Evaluate JavaScript
node <base-dir>/scripts/browser-eval.js 'document.title'
node <base-dir>/scripts/browser-eval.js 'document.querySelectorAll("a").length'
Execute JavaScript in the active tab. Code runs in async context. Use this to extract data, inspect page state, or perform DOM operations programmatically.
Screenshot
node <base-dir>/scripts/browser-screenshot.js
Capture current viewport and return temporary file path. Use this to visually inspect page state or verify UI changes.
Pick Elements
node <base-dir>/scripts/browser-pick.js "Click the submit button"
IMPORTANT: Use this tool when the user wants to select specific DOM elements on the page. This launches an interactive picker that lets the user click elements to select them. The user can select multiple elements (Cmd/Ctrl+Click) and press Enter when done. The tool returns CSS selectors for the selected elements.
Common use cases:
- User says "I want to click that button" → Use this tool to let them select it
- User says "extract data from these items" → Use this tool to let them select the elements
- When you need specific selectors but the page structure is complex or ambiguous
Cookies
node <base-dir>/scripts/browser-cookies.js
Display all cookies for the current tab including domain, path, httpOnly, and secure flags. Use this to debug authentication issues or inspect session state.
Extract Page Content
node <base-dir>/scripts/browser-content.js https://example.com
Navigate to a URL and extract readable content as markdown. Uses Mozilla Readability for article extraction and Turndown for HTML-to-markdown conversion. Works on pages with JavaScript content (waits for page to load).
When to Use
- Testing frontend code in a real browser
- Interacting with pages that require JavaScript
- When user needs to visually see or interact with a page
- Debugging authentication or session issues
- Scraping dynamic content that requires JS execution
Efficiency Guide
DOM Inspection Over Screenshots
Don't take screenshots to see page state. Do parse the DOM directly:
document.body.innerHTML.slice(0, 5000)
Array.from(document.querySelectorAll('button, input, [role="button"]')).map(e => ({
id: e.id,
text: e.textContent.trim(),
class: e.className
}))
Complex Scripts in Single Calls
Wrap everything in an IIFE to run multi-statement code:
(function() {
const data = document.querySelector('#target').textContent;
const buttons = document.querySelectorAll('button');
buttons[0].click();
return JSON.stringify({ data, buttonCount: buttons.length });
})()
Batch Interactions
Don't make separate calls for each click. Do batch them:
(function() {
const actions = ["btn1", "btn2", "btn3"];
actions.forEach(id => document.getElementById(id).click());
return "Done";
})()
Typing/Input Sequences
(function() {
const text = "HELLO";
for (const char of text) {
document.getElementById("key-" + char).click();
}
document.getElementById("submit").click();
return "Submitted: " + text;
})()
Reading App/Game State
Extract structured state in one call:
(function() {
const state = {
score: document.querySelector('.score')?.textContent,
status: document.querySelector('.status')?.className,
items: Array.from(document.querySelectorAll('.item')).map(el => ({
text: el.textContent,
active: el.classList.contains('active')
}))
};
return JSON.stringify(state, null, 2);
})()
Waiting for Updates
If DOM updates after actions, add a small delay with bash:
sleep 0.5 && node <base-dir>/scripts/browser-eval.js '...'
Investigate Before Interacting
Always start by understanding the page structure:
(function() {
return {
title: document.title,
forms: document.forms.length,
buttons: document.querySelectorAll('button').length,
inputs: document.querySelectorAll('input').length,
mainContent: document.body.innerHTML.slice(0, 3000)
};
})()
Then target specific elements based on what you find.
Gotchas
{baseDir} is not substituted by OpenCode. Construct the full script path using the "Base directory" value injected at the bottom of this skill.
- Chrome must be running with
--remote-debugging-port=9222 before any scripts work. If scripts fail silently or with connection errors, check Chrome is running with remote debugging enabled.
browser-pick.js is interactive — it requires a human at the keyboard to click elements. Never invoke it autonomously without the user present.
browser-content.js waits for page load before extracting. Very slow or JS-heavy pages may time out. Check the page loaded before trusting empty output.
- Run
npm install in the skill directory before first use. Missing node_modules causes silent failures with unhelpful "module not found" errors.