Skip to main content

firecrawl

Scrape web pages to clean markdown using Firecrawl v2 — handles JS-heavy pages, site crawls, URL mapping, document parsing (PDF/DOCX/XLSX), LLM-powered extraction, autonomous agent scraping, and post-scrape browser interaction (Interact API). Free keyless mode available for search/scrape/parse/interact (no API key needed). Prefer over WebFetch for quality and completeness. Triggers on scrape URL, fetch page, crawl site, extract content, parse document, web to markdown, DeepWiki, Firecrawl.

Datos de origen

Repositorio
tdimino/claude-code-minoan
Última actividad en el origen
28 de agosto de 2026 a las 18:56
Idioma detectado de SKILL.md
inglés
Estrellas
41
Forks
4

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
13 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
firecrawl
description
Scrape web pages to clean markdown using Firecrawl v2 — handles JS-heavy pages, site crawls, URL mapping, document parsing (PDF/DOCX/XLSX), LLM-powered extraction, autonomous agent scraping, and post-scrape browser interaction (Interact API). Free keyless mode available for search/scrape/parse/interact (no API key needed). Prefer over WebFetch for quality and completeness. Triggers on scrape URL, fetch page, crawl site, extract content, parse document, web to markdown, DeepWiki, Firecrawl.
# Firecrawl & Jina Web Scraping ## Firecrawl vs WebFetch Prefer `firecrawl scrape URL --only-main-content` over the WebFetch tool—it produces cleaner markdown, handles JavaScript-heavy pages, and avoids content truncation (>80% benchmark coverage). WebFetch is acceptable as a fallback when Firecrawl is unavailable. ```bash # Preferred approach: firecrawl scrape https://docs.example.com/api --only-main-content ``` ## Keyless Mode (v2.11.0+) Firecrawl's core endpoints work **without an API key**, rate-limited per IP per day. No signup required. **Keyless-eligible commands:** Search, Scrape, Parse, Interact **Requires API key:** Crawl, Map, Batch Scrape, Agent, Extract, Monitor ```bash # Quick start — no API key needed: firecrawl scrape https://example.com --only-main-content firecrawl search "query" --limit 10 # Python script also works keyless for eligible commands: python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py scrape https://example.com ``` **Rate limits:** Per-IP daily cap on requests and credits (exact numbers unpublished, below 1,000 credits/day). Sign up for a free API key to get 1,000 credits and higher rate limits. **Keyless MCP endpoint:** `https://mcp.firecrawl.dev/v2/mcp` — exposes Search, Scrape, Parse without any key. Use for MCP-native integrations. **Full setup (recommended for heavier use):** ```bash npx -y firecrawl-cli@latest init --all --browser # Installs CLI + skill segments into coding agents + opens browser for OAuth auth ``` **Search accuracy:** 94.7% on SimpleQA — custom relevance model scores paragraphs against your query, returning the most relevant excerpts with 10x fewer tokens than full-page processing. ## Token-Efficient Scraping Inspired by Anthropic's [dynamic filtering](https://claude.com/blog/improved-web-search-with-dynamic-filtering)—always filter before reasoning. This reduced input tokens by ~24% and improved accuracy by ~11% in their benchmarks. ### The Principle: Search → Filter → Scrape → Filter → Reason **DO:** ``` Search (titles/URLs only) → Evaluate relevance → Scrape top hits → Filter by section → Reason ``` **DON'T:** ``` Search → Scrape everything → Reason over all of it ``` ### Step-by-Step Efficient Workflow ```bash # Step 1: Search — get titles/URLs only (cheap) firecrawl search "query" --limit 20 # Step 2: Evaluate results, pick 3-5 best URLs # Step 3: Scrape only those, filter to relevant sections firecrawl scrape URL1 --only-main-content | \ python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py \ --sections "API,Authentication" --max-chars 5000 ``` ### Post-Processing with filter_web_results.py Pipe any Firecrawl or Exa output through this script to reduce context before reasoning: ```bash # Extract only matching sections from scraped page firecrawl scrape URL --only-main-content | \ python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py --sections "Pricing,Plans" # Keep only paragraphs with keywords firecrawl search "query" --scrape --pretty | \ python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py --keywords "pricing,cost" --max-chars 5000 # Extract specific JSON fields from API output python3 ~/.claude/skills/exa-search/scripts/exa_search.py "query" --json | \ python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py --fields "title,url,text" --max-chars 3000 # Combine filters with stats firecrawl scrape URL --only-main-content | \ python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py --sections "API" --keywords "endpoint" --compact --stats ``` **Full path:** `python3 ~/.claude/skills/firecrawl/scripts/filter_web_results.py` **Flags:** `--sections`, `--keywords`, `--max-chars`, `--max-lines`, `--fields` (JSON), `--strip-links`, `--strip-images`, `--compact`, `--stats` ### Other Token-Saving Patterns - **Use `--only-main-content`** to strip navigation and footer boilerplate, reducing token consumption. Omit only when nav/footer content is specifically needed. - **Use `--only-clean-content`** (Python API script) for aggressive cleaning—strips nav, ads, and cookie banners. Stronger than `--only-main-content`; use when the page is still noisy after main-content filtering. - **Use `firecrawl map URL --search "topic"` first** to find relevant subpages before scraping - **Use `--format links` first** to get URL list, evaluate, then scrape selectively - **Use `--max-chars`** with `exa_contents.py` to cap extraction length - **Use `--formats summary`** (Python API script) over full text when you need the gist, not raw content ### Claude API Native Tools (for API Agent Builders) Anthropic's API now offers built-in dynamic filtering tools: ``` web_search_20260209 / web_fetch_20260209 Header: anthropic-beta: code-execution-web-tools-2026-02-09 ``` These have built-in dynamic filtering via code execution. Use them when building Claude API agents directly. Use Firecrawl/Exa when you need: autonomous agents, batch scraping, structured extraction, domain-specific crawling, or when not on the Claude API. --- ## Available Tools ### 1. Official Firecrawl CLI (`firecrawl`) — Primary **Keyless setup:** `npm install -g firecrawl-cli@latest` (search/scrape/parse/interact work immediately) **Full setup:** `npm install -g firecrawl-cli@latest && firecrawl login --api-key $FIRECRAWL_API_KEY` **Agent setup:** `npx -y firecrawl-cli@latest init --all --browser` (installs CLI + skills + OAuth) | Command | Purpose | Quick Example | |---------|---------|---------------| | `scrape` | Single page → markdown | `firecrawl scrape URL --only-main-content` | | `crawl` | Entire site with progress | `firecrawl crawl URL --wait --progress --limit 50` | | `map` | Discover all URLs on a site | `firecrawl map URL --search "API"` | | `search` | Web search, 94.7% SimpleQA accuracy (+ optional scrape) | `firecrawl search "query" --limit 10` | **Full CLI reference:** `references/cli-reference.md` ### 2. Auto-Save Alias (`fc-save`) — Shell Alias Requires shell alias setup (not bundled with this skill). ```bash fc-save URL # → Saves to ~/Desktop/Screencaps & Chats/Web-Scrapes/docs-example-com-api.md ``` ### 3. Python API Script (`firecrawl_api.py`) — Advanced Features **Command:** `python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py <command>` **Requires:** `pip install firecrawl-py requests`. `FIRECRAWL_API_KEY` optional for keyless commands (search, scrape, parse, interact); required for all others. | Command | Purpose | Quick Example | |---------|---------|---------------| | `search` | Web search with scraping | `firecrawl_api.py search "query" -n 10` | | `scrape` | Single URL with page actions | `firecrawl_api.py scrape URL --formats markdown summary` | | `batch-scrape` | Multiple URLs concurrently | `firecrawl_api.py batch-scrape URL1 URL2 URL3` | | `crawl` | Website crawling | `firecrawl_api.py crawl URL --limit 20` | | `map` | URL discovery | `firecrawl_api.py map URL --search "query"` | | `parse` | Parse local documents (PDF, DOCX, XLSX) | `firecrawl_api.py parse report.pdf` | | `extract` | LLM-powered structured extraction | `firecrawl_api.py extract URL --prompt "Find pricing"` | | `agent` | Autonomous extraction (no URLs needed) | `firecrawl_api.py agent "Find YC W24 AI startups"` | | `parallel-agent` | Bulk agent queries (v2.8.0+) | `firecrawl_api.py parallel-agent "Q1" "Q2" "Q3"` | | `interact` | Post-scrape browser interaction | `firecrawl_api.py interact SCRAPE_ID --prompt "Click pricing"` | | `interact-stop` | Stop an interact session | `firecrawl_api.py interact-stop SCRAPE_ID` | **Agent models:** `spark-1-fast` (10 credits, simple), `spark-1-mini` (default), `spark-1-pro` (thorough) **Full Python API reference:** `references/python-api-reference.md` ### 4. DeepWiki — GitHub Repo Documentation ```bash ~/.claude/skills/firecrawl/scripts/deepwiki.sh <owner/repo> [section] [options] ``` AI-generated wiki for any public GitHub repo. No API key required. ```bash # Overview ~/.claude/skills/firecrawl/scripts/deepwiki.sh karpathy/nanochat # Browse sections ~/.claude/skills/firecrawl/scripts/deepwiki.sh langchain-ai/langchain --toc # Specific section ~/.claude/skills/firecrawl/scripts/deepwiki.sh karpathy/nanochat 4.1-gpt-transformer-implementation # Full dump for RAG ~/.claude/skills/firecrawl/scripts/deepwiki.sh openai/openai-python --all --save ``` ### 5. Jina Reader (`jina`) — Fallback Use when Firecrawl fails or for **Twitter/X URLs** (Firecrawl blocks Twitter, Jina works). ```bash jina https://x.com/username/status/123456 ``` --- ## Firecrawl vs Exa vs Native Claude Tools | Need | Best Tool | Why | |------|-----------|-----| | Free scrape, no API key | `firecrawl scrape` (keyless) | Core commands work without auth | | Free search, no API key | `firecrawl search` (keyless) | 94.7% SimpleQA accuracy, keyless | | Single page → markdown | `firecrawl scrape --only-main-content` | Cleanest output | | Search + scrape in one shot | `firecrawl search --scrape` | Combined operation | | Crawl entire site | `firecrawl crawl --wait --progress` | Link following + progress | | Local file → markdown | `firecrawl_api.py parse FILE` | Direct upload, no URL needed | | Autonomous data finding | `firecrawl_api.py agent` | No URLs needed | | Semantic/neural search | Exa `exa_search.py` | AI-powered relevance | | Find research papers | Exa `--category "research paper"` | Academic index | | Quick research answer | Exa `exa_research.py` | Citations + synthesis | | Find similar pages | Exa `exa_similar.py` | Competitive analysis | | Claude API agent building | Native `web_search_20260209` | Built-in dynamic filtering | | Twitter/X content | `jina URL` | Only tool that works | | GitHub repo docs | `deepwiki.sh owner/repo` | AI-generated wiki | | Anti-bot / Cloudflare bypass | `scrapling` stealth fetch | Local Turnstile solver | | Element-level extraction | `scrapling` + CSS selectors | Precision targeting, adaptive tracking | | No API key scraping | `scrapling` HTTP fetch | 100% local, no credentials | | Site redesign resilience | `scrapling` adaptive mode | SQLite similarity matching | | Budget JS-rendered scrape | `cf_browser.py markdown URL` | CF free tier: 10 min/day, $0.09/hr paid | | Free static page fetch | `cf_browser.py markdown URL --no-render` | FREE during beta (no JS) | | Budget multi-page crawl | `cf_browser.py crawl URL` | 5 free crawls/day, 100 pages each | | Incremental re-crawl | `cf_browser.py crawl --modified-since` | Built-in, Firecrawl lacks this | | Page screenshot/PDF | `cf_browser.py screenshot/pdf URL` | Built-in CF endpoints, cheaper | | AI structured extraction | `cf_browser.py json URL --prompt "..."` | Workers AI included free | --- ## Common Workflows ### Single Page Scraping ```bash firecrawl scrape https://example.com/page --only-main-content # Or auto-save: fc-save URL # Or to file: firecrawl scrape URL --only-main-content -o page.md ``` ### Documentation Crawling ```bash # Map first, then crawl relevant paths firecrawl map https://docs.example.com --search "API" firecrawl crawl https://docs.example.com --include-paths /api,/guides --wait --progress ``` ### Research Workflow ```bash firecrawl search "machine learning best practices 2026" --scrape --scrape-formats markdown ``` ### Document Parsing (Local Files) Parse local documents into clean Markdown. Use `parse` for local or non-public files; use `scrape` for public URLs pointing to documents—both use the same Rust-based parser. ```bash # PDF to markdown python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py parse report.pdf # Excel spreadsheet with main content only python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py parse data.xlsx --only-main-content # Word doc with zero data retention, save to file python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py parse contract.docx --zero-data-retention -o contract.md # Raw JSON output for programmatic use python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py parse invoice.pdf --json ``` Supported formats: PDF, DOCX, DOC, XLSX, XLS, HTML, HTM, ODT, RTF (up to 50 MB). ### PDF Parsing (Fire-PDF v2.9) Fire-PDF is now the default parsing pipeline for all PDF scrapes. Three modes: ```bash # Auto mode (default) — detects text layer vs scanned, chooses best method python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py scrape "https://example.com/report.pdf" # Fast mode — text layer only, skip OCR (use for PDFs with selectable text) python3 ~/.claude/skills/firecrawl/scripts/firecrawl_api.py scrape URL --pdf-mode fast # OCR mode — force full OCR (use for scanned docs or image-only PDFs)
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub