| name | wigolo |
| description | Local-first web intelligence MCP server for AI coding agents. Ten tools for search, fetch, crawl, cache, extract, find similar, research, agent-driven data gathering, page diffing, and change watching. No API keys. Results cached in a local knowledge store. |
wigolo
Local-first web intelligence MCP server for AI coding agents. Ships ten tools over stdio. All network results land in a local knowledge cache.
Host-LLM synthesis (read me first)
Wigolo has no internal LLM. It returns structured evidence so the calling model (you) writes the final answer. Fold structure into your reply rather than collapsing it away:
search with format: "highlights" — ML-scored passages + citations. Quote and cite [N].
research — when MCP sampling is unavailable (common), the output carries a brief with topics, highlights, key_findings. Use it as the scaffold for the report you write.
find_similar — may return a cold_start string. Pass it to the user; it explains why results came from the web and how to warm the cache.
extract with mode: "structured" — one call for tables + <dl> definitions + JSON-LD + chart hints + key-value pairs.
fetch metadata — surfaces og_type, canonical_url, and og_image; use canonical_url to dedupe tracked/canonical URLs.
Quick Setup
Claude Code:
claude mcp add wigolo -- npx wigolo
Cursor / VS Code / any MCP client:
{
"mcpServers": {
"wigolo": {
"command": "npx",
"args": ["wigolo"]
}
}
}
Warmup (recommended, one-time):
npx wigolo warmup
npx wigolo warmup --all
npx wigolo warmup --force
Warmup flags: --force, --all, --reranker, --firefox, --webkit, --embeddings, --no-searxng, --verify.
Tools
fetch
Fetch a single URL and return clean markdown. Use when you already have a specific URL.
Parameters:
url (string, required)
render_js: "auto" (default) | "always" | "never"
use_auth: boolean (default false) — reuses the user's browser session
max_chars: number
section: string — return only the content under a heading
section_index: number (default 0) — which heading match when multiple hit
screenshot: boolean (default false)
headers: object
force_refresh: boolean — bypass cache
actions: array of {type, selector, text, ms, timeout, direction, amount} — click, type, wait, wait_for, scroll, screenshot. Forces browser rendering when present.
Example:
{ "url": "https://react.dev/reference/react/useState", "section": "Parameters" }
Tip: section is much cheaper than reading the full page. Repeat fetches of the same URL are free from cache unless force_refresh: true.
search
Search the web and return extracted markdown per result. Use when you don't have a URL yet.
Parameters:
query (string OR string[], required) — array runs variants in parallel and dedupes
max_results: number (default 5, cap 20)
include_content: boolean (default true)
content_max_chars: number (default 30000)
max_total_chars: number (default 50000)
time_range: "day" | "week" | "month" | "year"
include_domains / exclude_domains: string[]
from_date / to_date: ISO YYYY-MM-DD
category: "general" | "news" | "code" | "docs" | "papers" | "images"
language: string
search_engines: string[] — override engine selection
format: "full" (default) | "context" (token-budgeted string) | "highlights" (ML-scored passages + citations) | "answer" (synthesized via MCP sampling; falls back to highlights when unsupported) | "stream_answer" (answer + phase progress notifications)
max_highlights: number (default 10) — cap when format: "highlights"
force_refresh: boolean
Example:
{ "query": ["react server components patterns", "RSC data fetching", "react server components streaming"], "category": "docs", "include_domains": ["react.dev"], "max_results": 5 }
Tip: keyword queries beat natural-language questions. A 3–5 item query array usually finds more unique sources than one longer query.
crawl
Crawl a site starting from a seed URL.
Parameters:
url (string, required)
strategy: "bfs" (default) | "dfs" | "sitemap" | "map" (URL-only discovery, no content)
max_depth: number (default 2)
max_pages: number (default 20)
include_patterns / exclude_patterns: regex string[]
use_auth: boolean (default false)
extract_links: boolean (default false) — returns inter-page link graph
max_total_chars: number (default 100000)
Example:
{ "url": "https://docs.python.org/3/library/", "strategy": "sitemap", "max_pages": 30, "include_patterns": ["^https://docs\\.python\\.org/3/library/asyncio"] }
Tip: strategy: "sitemap" is faster and more complete than BFS on doc sites. strategy: "map" returns URLs only — cheap way to scope before targeted fetches.
cache
Search previously fetched content without hitting the network.
Parameters:
query: full-text search — supports AND, OR, NOT, "exact phrase"
url_pattern: glob (e.g. "*react.dev*")
since: ISO date
stats: boolean — returns total URLs, size, date range
clear: boolean — deletes matching entries (requires one of query, url_pattern, since)
check_changes: boolean — re-fetches matching URLs, reports changed/unchanged with diff summaries
Example:
{ "query": "useState OR useReducer", "url_pattern": "*react.dev*" }
Tip: cache hits are instant and cross-session. Run this before search or fetch when you suspect the content is already on disk.
extract
Structured extraction from URL or raw HTML.
Parameters:
url OR html (one required; url wins if both provided)
mode: "metadata" (default) | "selector" | "tables" | "schema" | "structured" (tables + <dl> definitions + JSON-LD + chart hints + microdata/data-attr/grid key-value pairs in one call)
css_selector: string — required for mode: "selector"
multiple: boolean (default false) — return all matches, selector mode only
schema: JSON Schema object with properties — required for mode: "schema"
Example:
{ "url": "https://example.com/product", "mode": "schema", "schema": { "type": "object", "properties": { "price": { "type": "string" }, "name": { "type": "string" }, "sku": { "type": "string" } } } }
Tip: mode: "schema" does heuristic matching over CSS classes, ARIA labels, microdata, and JSON-LD — no LLM call required. mode: "structured" returns every structured pattern on the page (tables, definitions, jsonld, chart_hints, key_value_pairs) in one response — prefer it over chaining multiple extract calls.
find_similar
Find pages related to a URL or a free-text concept.
Parameters:
url OR concept (one required)
max_results: number (default 10, cap 50)
include_domains / exclude_domains: string[]
include_cache: boolean (default true)
include_web: boolean (default true)
Example:
{ "url": "https://react.dev/reference/react/useState", "max_results": 8, "include_domains": ["react.dev", "developer.mozilla.org"] }
Tip: uses hybrid 3-way RRF fusion — keyword search + semantic embeddings + live web search. Each result carries match_signals with embedding_rank, fts5_rank, and fused_score. If the cache is empty or embeddings aren't set up, the response includes a cold_start string — pass it to the user to explain why results came from the web.
research
Multi-step research pipeline with decomposition, parallel search, and cited synthesis.
Parameters:
question (string, required)
depth: "quick" (~15s, 2 sub-queries, 5–8 sources) | "standard" (~40s, default) | "comprehensive" (~80s, 7 sub-queries, 20–25 sources)
max_sources: number (cap 50) — overrides depth default
include_domains / exclude_domains: string[]
schema: JSON Schema — if present, report is structured to fill these fields
stream: boolean — emit progress notifications per phase
Example:
{ "question": "How do modern JS bundlers tree-shake ESM vs CJS?", "depth": "standard", "include_domains": ["webpack.js.org", "rollupjs.org", "esbuild.github.io", "vitejs.dev"] }
Tip: research checks cache internally — no need to pre-probe. With MCP sampling, the tool synthesizes the report directly. Without sampling (the common case), the output ships a brief with topics, highlights, and key_findings, plus the raw sources — the host LLM writes the final report from the brief.
agent
Natural-language data gathering. Plans queries and URLs from a prompt, runs them in parallel within budget, optionally applies a schema.
Parameters:
prompt (string, required)
urls: string[] — seed URLs to include
schema: JSON Schema — extract structured fields per page and merge
max_pages: number (default 10, cap 100)
max_time_ms: number (default 60000, cap 600000)
stream: boolean
Example:
{ "prompt": "Compare pricing tiers for Supabase, Firebase, and Clerk", "schema": { "type": "object", "properties": { "provider": { "type": "string" }, "free_tier": { "type": "string" }, "paid_start": { "type": "string" } } }, "max_pages": 12 }
Tip: output includes a steps array showing every action (plan, search, fetch, extract, synthesize) with timings. Use this to debug why an agent run produced a weak result.
diff
Compare two versions of a page. Point it at a live URL and its cached copy, two URLs, or two markdown blobs.
Parameters:
old (object, required): one of { url, markdown, content_hash }
new (object, required): one of { url, markdown }
output: "unified" (default), "hunks", or "summary"
granularity: "line" (default), "word", or "section"
Example:
{ "old": { "url": "https://docs.example.com/api" }, "new": { "url": "https://docs.example.com/api" }, "output": "hunks", "granularity": "section" }
Tip: same URL on both sides diffs the cached copy against a fresh fetch. Use output: "summary" for counts only.
watch
Monitor a page for changes over time. Lazy execution — checks run when you call check or when an overdue job is picked up during another tool run.
Parameters:
action (string, required): "create", "list", "check", "pause", "resume", "delete"
url or urls: single-URL or batch create (mutually exclusive)
interval_seconds: required for create (min 60)
selector: create-only CSS selector to scope the diff
notification: "inline" (default) or an SSRF-guarded webhook URL
job_id: required for check/pause/resume/delete
Example: