Skip to main content

google-surf-mcp-search

Google search MCP server with academic PDF extraction, no API key required, CAPTCHA recovery, and parallel search capabilities

Jump to install

Source facts

Repository
reason-machines/mcp-skills
Last source activity
May 17, 2026 at 04:50
Detected SKILL.md language
English
Stars
7
Forks
2

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
google-surf-mcp-search
description
Google search MCP server with academic PDF extraction, no API key required, CAPTCHA recovery, and parallel search capabilities
triggers
["search google without an api key","set up google search mcp server","extract content from academic papers","search and extract web content in parallel","configure google-surf-mcp for claude","handle captcha in mcp search","fetch and extract article content","search multiple queries in parallel"]
# google-surf-mcp-search > Skill by [ara.so](https://ara.so) — MCP Skills collection. ## What It Does google-surf-mcp is an MCP server that provides Google search functionality without requiring API keys. It combines three capabilities in one: - Google search with ad/spam filtering - URL content extraction (HTML + PDF) - Academic paper extraction (arXiv, Nature, PubMed, etc.) Key features: - Works with actual Google search (not an API wrapper) - Automatic CAPTCHA recovery with persistent browser profiles - Parallel search and extraction - Token-efficient abstract mode for triage - Built-in rate limiting and caching - Geometric verification to drop sponsored ads and knowledge panels ## Installation ### Quick Install (npx) Add to your MCP client config (e.g., `~/.claude.json` for Claude Code): ```json { "mcpServers": { "google-surf": { "command": "npx", "args": ["-y", "google-surf-mcp"] } } } ``` ### Local Clone Installation ```bash git clone https://github.com/HarimxChoi/google-surf-mcp cd google-surf-mcp npm install npm run build ``` Config for local installation: ```json { "mcpServers": { "google-surf": { "command": "node", "args": ["/absolute/path/to/google-surf-mcp/build/index.js"] } } } ``` ### Manual Bootstrap (if auto-bootstrap fails) ```bash npm run bootstrap ``` With custom paths: ```bash CHROME_PATH=/usr/bin/google-chrome SURF_TZ=America/New_York npm run bootstrap ``` ## Available Tools ### 1. `search` - Single Google Search Performs a single Google search, returns filtered results (ads removed). **Parameters:** - `query` (string, required): Search query - `limit` (number, optional): Max results, default 10 **Returns:** - `results[]`: Array of `{ title, url, snippet }` - `dropped`: Count of filtered results (ads, knowledge panels) - `dropped_reasons[]`: Why items were dropped - `cache_hit`: Boolean indicating cache use **Example Usage:** ```typescript // Via MCP tool call { "query": "typescript async patterns", "limit": 5 } ``` **Response:** ```json { "results": [ { "title": "Async/Await in TypeScript", "url": "https://example.com/typescript-async", "snippet": "Learn how to use async/await patterns..." } ], "dropped": 2, "dropped_reasons": ["sponsored", "knowledge_panel"], "cache_hit": false } ``` ### 2. `search_parallel` - Parallel Multi-Query Search Execute multiple searches in parallel using a worker pool (max 10 queries). **Parameters:** - `queries` (string[], required): Array of search queries - `limit` (number, optional): Max results per query, default 10 **Returns:** - Array of search results (same format as `search`) **Example Usage:** ```typescript { "queries": [ "mcp server best practices", "playwright stealth techniques", "typescript pdf extraction", "google search scraping 2026" ], "limit": 3 } ``` ### 3. `extract` - Fetch and Extract Content Extract text content from a URL (HTML or PDF). **Parameters:** - `url` (string, required): URL to extract - `max_chars` (number, optional): Character limit, default 100k - `mode` (string, optional): `"full"` | `"abstract"` | `"metadata"` **Modes:** - `full`: Complete article text (HTML via Readability, PDF via unpdf) - `abstract`: ~1500 chars for triage (PDF page 1 or HTML meta description) - `metadata`: PDF page count only **Returns:** - `content`: Extracted text (markdown for HTML) - `title`: Document title - `excerpt`: Short summary - `length`: Character count - `is_pdf`: Boolean - `page_count`: Number (PDFs only) - `extraction_quality`: `"high"` | `"medium"` | `"low"` **Example Usage:** ```typescript // Extract full academic paper { "url": "https://arxiv.org/pdf/2301.12345.pdf", "mode": "full" } // Quick abstract for triage { "url": "https://nature.com/articles/s41586-023-12345-6", "mode": "abstract", "max_chars": 2000 } ``` **Response:** ```json { "content": "# Paper Title\n\nAbstract: This paper presents...", "title": "Novel Approach to AI Safety", "excerpt": "This paper presents a novel approach...", "length": 45678, "is_pdf": true, "page_count": 12, "extraction_quality": "high" } ``` ### 4. `search_extract` - Combined Search + Extract Search and extract content in one call. Efficiently parallelizes extraction. **Parameters:** - `query` (string, required): Search query - `limit` (number, optional): Max results to extract, default 5 - `max_chars` (number, optional): Per-result char limit - `mode` (string, optional): `"abstract"` (default) | `"full"` **Best Practices:** - Use `mode="abstract"` (default) for cheap triage with ~1500-char summaries - Use `mode="full"` only when you need complete article text (slower, more tokens) **Returns:** - `results[]`: Search results enriched with `extracted_content` **Example Usage:** ```typescript // Triage mode (default, token-efficient) { "query": "claude mcp server tutorials", "limit": 5, "mode": "abstract" } // Full extraction (when you need complete content) { "query": "machine learning interpretability survey", "limit": 3, "mode": "full", "max_chars": 50000 } ``` **Response:** ```json { "results": [ { "title": "Building MCP Servers", "url": "https://example.com/mcp-tutorial", "snippet": "Complete guide to MCP servers...", "extracted_content": { "content": "# Building MCP Servers\n\nMCP (Model Context Protocol)...", "title": "Building MCP Servers", "length": 1523, "is_pdf": false, "extraction_quality": "high" } } ] } ``` ### 5. `health` - Server Status Check server health and configuration. **Returns:** - `status`: `"healthy"` | `"degraded"` - `cascade_mode`: Current stealth mode - `rate_limiter`: Request counts and limits - `cache_stats`: Cache size and hit rates - `config`: Active configuration values **Example Usage:** ```typescript // No parameters {} ``` ## Configuration All configuration via environment variables: ### Essential Variables ```bash # Chrome binary path (auto-detected if not set) CHROME_PATH=/usr/bin/google-chrome # Profile storage (default: ~/.google-surf-mcp) SURF_PROFILE_ROOT=/custom/path/profiles # Browser locale and timezone SURF_LOCALE=en-US SURF_TZ=America/New_York ``` ### Headless & CAPTCHA Recovery ```bash # Run Chrome visibly (for demos/debugging) SURF_HEADLESS=false # Remote debugging mode (headless servers) SURF_REMOTE_DEBUG=true # Cloud/serverless mode (fail-fast on CAPTCHA) SURF_CLOUD_MODE=true ``` ### Performance Tuning ```bash # Idle close timeout (ms), 0 disables SURF_IDLE_CLOSE_MS=30000 # Rate limit (requests per minute) SURF_RATE_LIMIT_PER_MIN=10 # Search cache TTL (ms), 0 disables SURF_CACHE_TTL_SEARCH_MS=86400000 # Cache LRU size SURF_CACHE_MAX_ENTRIES=1000 ``` ### Security ```bash # Allow private IPs in extract (default: false) SURF_ALLOW_PRIVATE=true # Ignore TLS errors (auto-on in cloud mode) SURF_INSECURE_TLS=false # Disable sandbox (auto-on in cloud mode) SURF_NO_SANDBOX=false ``` ### Advanced ```bash # Disable cascade fallback (pin single mode) SURF_CASCADE_DISABLED=true SURF_USE_STEALTH=true # Humanlike browsing (off | background | inline) SURF_HUMANLIKE_MODE=background ``` ## Common Patterns ### Pattern 1: Research Assistant Search academic papers and extract abstracts for quick review: ```typescript // Step 1: Search and triage with abstracts const triage = await use_mcp_tool("google-surf", "search_extract", { query: "transformer architecture improvements 2026", limit: 10, mode: "abstract" }); // Step 2: Extract full text for promising papers const topPapers = triage.results.slice(0, 3); const fullTexts = await Promise.all( topPapers.map(paper => use_mcp_tool("google-surf", "extract", { url: paper.url, mode: "full", max_chars: 100000 }) ) ); ``` ### Pattern 2: Parallel Research Search multiple related topics simultaneously: ```typescript const relatedTopics = await use_mcp_tool("google-surf", "search_parallel", { queries: [ "MCP server authentication patterns", "MCP server error handling", "MCP server rate limiting", "MCP server caching strategies" ], limit: 5 }); // Process results by topic relatedTopics.forEach((topicResults, index) => { console.log(`Topic ${index + 1}:`, topicResults.results.length, "results"); }); ``` ### Pattern 3: Content Aggregation Build a comprehensive knowledge base: ```typescript // 1. Find relevant sources const sources = await use_mcp_tool("google-surf", "search", { query: "typescript best practices 2026", limit: 20 }); // 2. Extract abstracts to filter quality const abstracts = await Promise.all( sources.results.map(result => use_mcp_tool("google-surf", "extract", { url: result.url, mode: "abstract" }) ) ); // 3. Full extraction for high-quality sources const highQuality = abstracts .filter(a => a.extraction_quality === "high") .slice(0, 5); const fullContent = await Promise.all( highQuality.map(a => use_mcp_tool("google-surf", "extract", { url: a.url, mode: "full" }) ) ); ``` ### Pattern 4: Health Check Before Heavy Operations ```typescript // Check server health before batch operations const health = await use_mcp_tool("google-surf", "health", {}); if (health.status !== "healthy") { console.warn("Server degraded, reducing concurrency"); } const rateLimit = health.rate_limiter.requests_per_minute; if (rateLimit > 8) { // Wait before starting batch await sleep(60000); } ``` ## CAPTCHA Recovery Modes The server handles CAPTCHAs automatically based on environment: ### Mode 1: Local Desktop (default) ```bash # No config needed - default behavior ``` When CAPTCHA appears: 1. OS notification fires 2. Headed Chrome window opens 3. Human solves CAPTCHA 4. Call automatically retries 5. Profile reputation preserved ### Mode 2: Visible Chrome (demos/debugging) ```bash SURF_HEADLESS=false ``` - Chrome runs visibly at all times - CAPTCHA recovery skips notification (user is watching)
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub