Fetch web pages, PDFs, and documents with automatic fallbacks and content extraction. Use when user says "fetch this URL", "download this page", "crawl this website", "extract content from", "get the PDF", or provides URLs needing retrieval.
Fetch web pages, PDFs, and documents with automatic fallbacks and content extraction. Use when user says "fetch this URL", "download this page", "crawl this website", "extract content from", "get the PDF", or provides URLs needing retrieval.
allowed-tools
Bash, Read
triggers
["fetch this URL","download page","crawl website","extract content from","get the PDF","scrape this site","retrieve document"]
metadata
{"short-description":"Web crawling and document fetching CLI"}
provides
["web-fetch"]
composes
["extractor","memory","task-monitor"]
taxonomy
["web","ingestion"]
Fetcher - Web Crawling
Fetch web pages and documents with automatic fallbacks, proxy rotation, and content extraction.
Self-contained skill - auto-installs via uvx from git (no pre-installation needed).
Fully automatic - Playwright browsers are installed on first run for SPA/JS page support.
Simplest Usage
# Via wrapper (recommended - auto-installs)
.pi/skills/fetcher/run.sh get https://example.com
# Or directly if fetcher is installed
fetcher get https://example.com
Common Commands
./run.sh get https://example.com # Fetch single URL
./run.sh get-manifest urls.txt # Fetch list of URLs
./run.sh get-manifest - < urls.txt # Fetch from stdin
Common Patterns
Fetch a single URL
fetcher get https://www.nasa.gov --out run/nasa
Outputs to run/nasa/:
consumer_summary.json - structured result
Walkthrough.md - human-readable summary
downloads/ - raw content files
Fetch multiple URLs
# From file (one URL per line)
fetcher get-manifest urls.txt --out run/batch
# From stdinecho -e "https://example.com\nhttps://nasa.gov" | fetcher get-manifest -
Playwright auto-fallback should trigger; check used_playwright in summary
Stale cached results
Set FETCHER_HTTP_CACHE_DISABLE=1 for fresh fetch
Rate limited
Configure proxy rotation or reduce concurrency
Paywall detected
Check content_verdict and use alternates
Empty content
Check junk_results.jsonl for diagnosis
Run fetcher doctor to check environment and dependencies.
SPA/JavaScript Page Support
Fetcher automatically falls back to Playwright for known SPA domains. If a page returns thin/empty content:
Check if used_playwright: 1 in consumer_summary.json
If not, the domain may need to be added to SPA_FALLBACK_DOMAINS in fetcher source
Force fresh fetch with FETCHER_HTTP_CACHE_DISABLE=1
Managing SPA_FALLBACK_DOMAINS
To add a new domain that requires Playwright:
# In fetcher source: src/fetcher/workflows/web_fetch.py
SPA_FALLBACK_DOMAINS = {
"attack.mitre.org", # React SPA"csf.tools", # NIST CSF - JS-heavy"cwe.mitre.org", # CWE definitions - JS-heavy# Add new domain here
}
Seed to memory for future agents:
# Use /memory skill to record successful strategy
.pi/skills/memory/run.sh learn \
--problem "Fetching example.com returns empty JS shell" \
--solution "Use Playwright for example.com (SPA). Add to SPA_FALLBACK_DOMAINS." \
--tag "fetch" --tag "playwright" --tag "example.com"
Content Extraction Pipeline
Understanding the extraction flow helps diagnose "empty content" bugs:
URL fetch
↓
raw HTML → result.text (initially)
↓
evaluate_result_content()
↓
trafilatura/readability extraction
↓
result.text = extracted_text (clean text, not HTML)
↓
downloads/ → raw HTML preserved
extracted_text/ → clean text output
Critical: result.text should contain extracted text, not raw HTML. If downstream sees <!DOCTYPE html>, the extraction step failed.
Diagnostic Checklist
When content appears empty or wrong, check these fields in consumer_summary.json:
Field
Expected
Problem If
status
200
Non-200 = fetch failed
method
aiohttp or playwright
playwright expected for SPA domains
verdict
ok
thin/empty = extraction issue
used_playwright
0 or 1
0 when SPA domain = missing from fallback list
Quick verification commands:
# Check if Playwright was used
jq '.counts.used_playwright' consumer_summary.json
# Compare raw HTML vs extracted texthead -c 200 downloads/*.html # Should see <!DOCTYPE html>head -c 200 extracted_text/*.txt # Should see clean text# Verify content is meaningful
grep -c "expected-keyword" extracted_text/*.txt
Common False Positives
Not all "empty content" issues need Playwright:
Domain
Reality
Why
cwe.mitre.org
Server-side rendered
Returns full HTML without JS
capec.mitre.org
Server-side rendered
Same as CWE
Most .gov sites
Static HTML
JS is progressive enhancement
Test before adding to SPA_FALLBACK_DOMAINS:
# Fetch without Playwright
curl -s "https://example.com" | grep -c "<article>"# If content exists in raw HTML, Playwright isn't needed# The issue is likely in the extraction step, not rendering