Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server. Use when the user wants to fetch a page, follow links across a domain, enumerate URLs, or drive a real browser. Covers installation, the subcommands (scrape, crawl, map, interact, batch-scrape, batch-crawl, download, citations, version, mcp, serve), output formats (JSON + Markdown), browser fallback, and when to prefer the MCP server over shelling out.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server. Use when the user wants to fetch a page, follow links across a domain, enumerate URLs, or drive a real browser. Covers installation, the subcommands (scrape, crawl, map, interact, batch-scrape, batch-crawl, download, citations, version, mcp, serve), output formats (JSON + Markdown), browser fallback, and when to prefer the MCP server over shelling out.
Crawlberg is a Rust-native web crawler and scraper. It fetches static HTML
with reqwest, falls back to headless Chrome when a page needs JS or trips a
WAF, and converts every result to clean Markdown via the built-in
HTMLโMarkdown engine.
Use this skill when the user wants to:
Scrape a single URL to Markdown plus structured metadata.
Crawl a site following links bounded by depth, page count, and concurrency.
Enumerate URLs from sitemaps without paying for rendering.
Drive a real browser (click, type, scroll) and capture the resulting DOM.
Run the same operations from another agent harness via MCP tools.
Installation
The plugin shells out to a crawlberg binary on PATH. Install one of:
# or run without a persistent install (the CLI proxy package self-installs the binary):
help
help
# or build from source:
The serve and mcp subcommands are gated behind non-default cargo features
(api and mcp). The Homebrew tap is built with all features, so both
subcommands work out of the box. A from-source build must pass
--features mcp (and --features api for serve), or --features all, to
include them.
Verify:
crawlberg --version
Headless fallback needs Chrome/Chromium reachable locally (chromiumoxide
launches it on demand). Skip the install if you only plan to use
--browser-mode never.
Command map
crawlberg scrape <url> # single page โ JSON or Markdown
crawlberg crawl <url...> # follow links, BFS, depth-bounded
crawlberg map <url> # enumerate URLs via sitemaps + link extraction
crawlberg interact <url> # browser actions: click, type, scroll
crawlberg batch-scrape <url...> # scrape many URLs concurrently
crawlberg batch-crawl <url...> # crawl many seed URLs concurrently
crawlberg download <url> # download a document, report file metadata
crawlberg citations <input> # markdown links โ numbered citations (text or @file.md)
crawlberg version # print the crawlberg version as JSON
crawlberg mcp # MCP server (stdio) โ auto-registered (`mcp` feature)
crawlberg serve # REST API server (`api` feature)
crawl also handles batching implicitly: pass multiple seed URLs and it fans
out via batch_crawl internally. The explicit batch-scrape and batch-crawl
subcommands expose the same concurrency for many independent URLs.
The --config flag accepts the full CrawlConfig schema. Anything you set
explicitly on the CLI overrides the corresponding JSON field.
These shared flags apply to the crawl/scrape-family subcommands. --format,
--browser-mode, --browser-endpoint, and --config cover scrape, crawl,
map, interact, batch-scrape, and batch-crawl; download takes
--timeout, --browser-mode, --browser-endpoint, --max-size, and
--config (no --format). --respect-robots-txt applies to scrape,
crawl, map, batch-scrape, and batch-crawl. citations and version
take no shared flags.
JSON output (default) carries the rendered Markdown, page metadata
(PageMetadata), links by category, images, feeds, JSON-LD blocks, and
HTTP response metadata. Use Markdown output when piping into a file the user
will read.
See the scraping-html-to-markdown skill for the full flag surface.
Crawling is BFS by default, bounded by --depth, --max-pages, and
--concurrent. Per-domain politeness is enforced by --rate-limit
(milliseconds between requests to the same origin).
See the crawling-a-site skill for the recommended defaults and the full
flag surface.
map reads sitemap.xml (and nested sitemaps), then falls back to link
extraction from the seed page. It does not render pages โ use it to plan a
crawl or to feed URLs into another tool.
Action types are click, type, press, scroll, wait, screenshot,
executeJs, and scrape (to wait for an element, use wait with a
selector field). The result wraps the final HTML under
interaction.final_html. See the automating-the-browser skill for the full
action schema and limits.
MCP server
When this plugin is installed in a Claude Code / Codex / Cursor / Gemini /
opencode harness, the MCP server is auto-registered:
crawlberg mcp
mcp is a stdio-transport server and takes no arguments. It requires a binary
built with the mcp feature (see Installation).
The server registers nine tools (the same set is served over the Streamable
HTTP transport when running crawlberg serve):
Tool
Purpose
Parameters
scrape
Scrape one URL to Markdown or JSON (content, metadata, links).
url (required), format (markdown|json), use_browser (bool โ force browser)
crawl
Follow links from a URL, bounded by depth/page count.
url (required), max_depth, max_pages, format, stay_on_domain
map
Discover all URLs via links and sitemaps.
url (required), limit, search, respect_robots_txt
batch_scrape
Scrape multiple URLs concurrently.
urls (required array), format, concurrency
batch_crawl
Crawl multiple seed URLs concurrently.
urls (required array), max_depth, max_pages, format, stay_on_domain, concurrency
download
Download a document and return file metadata.
url (required), max_size
interact
Execute browser actions on a page (mutating/destructive).
url (required), actions (required array of action objects)
generate_citations
Rewrite markdown links as numbered citations + reference list.
markdown (required)
get_version
Return the crawlberg library version.
none
Prefer MCP tools over shelling out when both are available:
Typed schemas surface argument errors before the call.
Results stream back as structured tool output instead of stdout text.
No --format juggling โ the harness pulls whatever shape it needs.
Fall back to the CLI when you need to script a pipeline, capture stderr, or
chain with shell tools.
Headless fallback
In --browser-mode auto (default), the engine:
Fetches statically via reqwest.
Detects WAF blocks (8 vendors) and JS-only shells.
Re-fetches through headless Chrome with a real fingerprint when needed.
Force the browser path with --browser-mode always when you already know
the page needs JS. Use --browser-mode never for hot loops where the cost
of a stray Chrome launch is unacceptable.
Point --browser-endpoint ws://host:9222/devtools/browser/<id> at an
already-running Chrome to skip the local launch.
See the headless-fallback skill for symptoms, costs, and external-CDP
patterns.
Output formats
Mode
Use when
json
Downstream consumer needs metadata, links, images, etc.
markdown
Human reader or LLM-context payload.
Markdown output skips metadata. If you need both, run with --format json
and read result.markdown.content.
Robots, rate limits, ethics
--respect-robots-txt is off by default; pass it for any crawl on a host
you do not own.
The default --rate-limit 200 already produces a polite cadence; raise it
for shared hosts.
Identify the crawler honestly via --user-agent. Do not impersonate a
browser unless the operator has approved it.
Cross-references
skills/crawling-a-site/SKILL.md โ multi-page crawl with depth, page
caps, concurrency, rate limits, and domain scoping.
skills/scraping-html-to-markdown/SKILL.md โ single-page rendering, the
Markdown output shape, and common pitfalls.
skills/mapping-urls/SKILL.md โ map: sitemap + link URL discovery,
filtering, and seeding a crawl.
skills/automating-the-browser/SKILL.md โ interact: the full scripted
action schema, limits, and result shape.
skills/serving-the-api/SKILL.md โ serve: the Firecrawl-v1-compatible
REST API server and its endpoints.
skills/headless-fallback/SKILL.md โ when and how to force the browser
backend.