| name | anysite-mcp |
| description | How to use the anysite MCP server effectively - the six meta-tools (discover, execute, get_page, query_cache, export_data, search_requests), the source map for GTM signals (funding, hiring, tech stack, reviews, news, launches), email finding cascades, domain->company resolution, and cost-aware calling patterns. Consult this before any anysite data work. Use when unsure which source or endpoint covers a data need, how much a call costs / how many credits, why an endpoint is 'not found', how to reuse a cache_key, how to paginate or re-filter cached results, or how to combine sources into a signal chain. |
Anysite MCP — usage guide
The anysite MCP exposes hundreds of data sources through six universal meta-tools (plus the
crm_* family, see Working with CRM). This skill is the map: how to call them, which sources
cover which GTM need, and how to not waste credits.
The six meta-tools
| Tool | Purpose | Credits |
|---|
discover(source, category) | List endpoints + exact params for a source/category | free |
execute(source, category, endpoint, params) | Run an endpoint; returns first 10 items + cache_key | paid |
get_page(cache_key, offset, limit) | Page through a cached result | free |
query_cache(cache_key, conditions, sort_by, sort_order, aggregate, group_by, limit, offset) | Filter/sort/aggregate cached data with SQL-like ops | free |
export_data(cache_key, output_format, list_unpack) | Export cached data — output_format json (default) / csv / jsonl; list_unpack = how many nested-array elements to expand into CSV columns (default 1) | free |
search_requests(source, category, endpoint, query, since, until, limit, offset) | Find past execute() calls and their cache_keys — 7-day history, works across sessions | free |
Rules that prevent 90% of failures
- Always
discover before execute. Endpoint names and params are not guessable, and a
wrong source name returns the full source list — a wrong guess self-corrects for free.
execute takes the endpoint NAME exactly as discover returns it (products_reviews),
never a REST path segment (reviews) — resolution is an exact-match lookup.
- Never guess identifiers. LinkedIn aliases, URNs, Crunchbase aliases, Greenhouse board
tokens are unpredictable. Resolve them through the search endpoint of the same source first.
- Re-use the cache — it outlives the session.
execute returns a cache_key; further
filtering, sorting, counting and paging of that result is free, and the cache lives for
7 days across sessions. Before any paid execute, check search_requests (free) for
a recent identical call — same endpoint, matching params — and reuse its cache_key via
query_cache/get_page instead of refetching (verified live: a two-day-old cache_key
from another session served in full). Freshness rule: reuse when the data's age is fine
for the task (enrichment firmographics — usually yes; "what's new today" — no).
- Cheap-first cascade. When several endpoints can answer, call the cached/DB one first
(
*/db/*, *sql* endpoints, ~1 credit) and the live one only for the remainder.
- Estimate volume before bulk runs — plan-aware. First know the user's plan (the CRM
profile stores it after setup; if unknown, ask once: MCP Unlimited or credit-based?).
- Credit-based plan: before anything above ~100 calls, state the estimate
(
N targets × credits-per-call) and get a nod. Prefer cheap DB endpoints, batch hard.
- MCP Unlimited: credit warnings off, but keep batch sizes sane anyway — the real
limits are latency and upstream rate limits, so cap sweeps the same way and say
"this will take ~N minutes" instead of a price.
- Live LinkedIn search fails as an empty list, not an error.
search_users and
search_companies return {"results":[]} on queries that just don't hit ("stripe",
"databar" both came back empty live, while "microsoft" worked) — it is not a broken key.
On empty, switch to the search_sql_* DB endpoints; do NOT retry with broader keywords.
(This is why reverse-lookup via live is best-effort, not "usually one
match".)
GTM source map
Company discovery (bulk):
linkedin/search/search_sql_companies — the workhorse. Up to 1000 companies per call with
DSL filters (keywords, industry_name, employee_count_min/max, country_hq, founded_on_min/max,
has_website) and a sort param (relevance — default for filtered queries — or
last_modified for freshness/monitoring). Also batch lookup by urn and search by website.
Query craft (naive keywords return wrong-country token soup — measured 1/5 relevant vs
5/5 structured): the anysite-company-sourcing skill.
⚠️ website search is SUBSTRING match, ordered by last_modified. Verification is
MANDATORY on every resolve — position in the results means nothing. Verified live:
{website: "stripe.com", count: 1} → Soundstripe; {website: "stlabs.com", count: 5} →
five *labs.com companies, none of them stlabs.com (common tokens flood the result even in
a single-domain call). The domain-resolve rules:
- Verify exact
website match (normalize both sides: lowercase, strip
protocol/www./path) on EVERY resolve, single or batched. No exact match =
unresolved — never write anything to the CRM for it; wrong-company data lands in
blank fields where nobody will catch it.
- Never
count: 1 on a website resolve — the substring flood means the one row you
get is very likely the wrong company. Default one domain per call at a small count (5+).
OR-DSL batching
({website: "a.com|b.io", count: 10× domains}) is an optimization with a verification
tax: a domain with a common token can be flooded out of the batch entirely — every
domain that didn't come back exact-matched must be re-queried individually.
query_cache filters over the WHOLE cached set (verified) but returns at most limit
rows (default 10) — pass an explicit limit when you expect more matches back. Sanity
rule: aggregate {op: "count"} should equal the total from execute; if not, page
with get_page before concluding anything.
- A whole class of domains never appears in its own substring results (common-token
domains like stlabs.com) — so the website's own page is the STANDARD second step, not
an emergency:
webparser/parse {url: "https://<domain>", extract_minimal: true} →
top-level title says who they are, links[] usually carries their own
linkedin.com/company/... URL → for the exact URN (verified, ~1cr).
Live shape on stlabs.com: at the TOP
level, while came back and empty — read , and treat
as a fallback only, not the primary location.
Secondary fallback: by name → . Name search
alone is never a source of truth.
Bonus from a successful resolve: the row already carries
(free crunchbase alias — skip the live 20cr search) and
( — the numeric id goes straight into ).
⚠️ For company SIZE use , never — the two fields
can contradict each other in the same record (verified: Clay returns alongside ). The range field looks like the natural
key for size segmentation and would misfile that company by ~3x, silently. Fall back to
the range only when the exact count is empty, and say that you did.
Company detail: crunchbase/company — get its alias for free from the crunchbase_link
that search_sql_companies already returned (live crunchbase/search is the fallback, not
the first step). ⚠️ The alias is CASE-SENSITIVE ('Google' ≠ 'google') — take it verbatim
from the URL slug. Also linkedin/company. One crunchbase/company call carries free extras:
bombora_surges[] (B2B intent topics — but they show what THAT company's staff researches,
i.e. what they BUY; treat as a signal only when a topic matches what the user sells),
related.competitors[], predictions.funding_score, awards[]. Coverage caveat:
leadership_hires[] is often EMPTY for smaller companies — absence of the field is not
absence of hires. Normalize contacts.email (trailing dots observed: "x@y.ai.").
A third resolve path when crunchbase is already fetched: contacts.linkedin_url →
linkedin/company → exact URN (verified; bypasses both fuzzy searches).
Note: owler endpoints need an owler alias and its search has no name/keyword parameter —
not usable for looking up a named account.
Engagement graph (who interacted with content): linkedin/post/post_comments,
post_reactions, post_reposts and linkedin/company/company_posts (~1cr/10) — answers
"who paid attention to this content", incl. people outside your title filters. Identifiers:
comments/reposts carry a vanity alias; reactions give only an obfuscated /in/ACoAA... URL
plus internal_id — user_email accepts the internal_id, never the obfuscated URL.
Honest scaling: volume follows the SEED's audience, not the target's importance (large brand
post → dozens of engagers; 80-person company → 0–2 per post), and on a small account those
few are mostly the company's OWN staff plus engagement farmers (verified: 3 of 4 commenters
were employees) — filter by the author's company first and expect nothing left. Use this on
seeds with a real audience (a competitor's page), not on SMB target lists. For small accounts
the reliable nugget is company_posts → mentioned[]: hiring announcements name new people
with their vanity aliases. linkedin/company/company_employee_stats (1cr) gives
function/skill/location breakdown — cross-check totals against employee_count from
linkedin/company before trusting absolutes, and never sum the locations array: its
buckets are nested (US ⊃ California ⊃ SF Bay Area), so summing double-counts badly. Its
llm_hint promises seniority and growth trends that the response does not contain.
Tech stack & adoption signals: stackshare/companies (a company's declared stack by
slug — the forward direction wappalyzer can't do); producthunt/products/products_customers
(reverse stack: who uses a product, with a testimonial quote — a budget/intent tell).
There is NO ad-transparency source in the catalog (verified: not among the 591 sources,
and linkedin has no ads category) — do not reach for ad-library data, it isn't here.
People:
linkedin/search/search_sql_users — the 856M-profile DB, the bulk workhorse: derived
seniority/function filters, company domain/id (incl. past employers = alumni),
career-shape (months_in_role, tenure, promotions), lookalike graph (similar_to),
deterministic buckets for >1000. Craft guide: the anysite-people-sourcing skill.
Key semantics: over-count result is an unbiased SAMPLE (repeat = same people; walk
bucket_total/bucket_index instead), and has_* flags make coverage narrowing
explicit — set them when filtering by fields not every profile states.
linkedin/search/search_users (live) — one-off lookups and namesake disambiguation
(job_title + current_company/company_keywords; never bare keywords alone).
linkedin/user (full profile, needs alias/URL/URN — never guess the alias),
linkedin/user/user_posts, user_experience, user_comments.
Email finding (cascade, cheap → expensive):
linkedin/user/user_email — batch up to 10 profiles, cheap, low yield. Truths from live
testing: it returns a MIX of personal and work addresses (roughly half and half), one row
per EMAIL — not per profile — and a single person can come back with several rows,
including emails at PAST employers (measured: one alias → 4 rows spanning current and
former company domains). Its found field is always true (useless as a check). So: group
by alias/internal_id, then match the domain against the person's CURRENT company; if
more than one work address survives, treat it as unverified and pass to step 2. Personal
addresses are not outreach-ready.
linkedin/user/user_find_email_by_url {url} — high yield but expensive (50cr), run only
on the remainder after step 1. Takes a VANITY profile URL (/in/satyanadella/);
URN-style URLs (/in/ACoA...) are rejected — get the vanity URL from linkedin/user
first. Response includes email_status and valid_email — check them and pass only
valid work emails onward; an address with a bad status is a bounce, not a find.
- No work email found → keep the lead anyway; CRM contact upserts match by
linkedin_url
too (but note: creating a NEW contact requires an email — no email means update-only).
Reverse lookup (email → person), reliability order:
linkedin/email/email_sql_user (cached DB) → email_user (live) — cheap, but verified
to return empty even for people who are definitely on LinkedIn. Try, don't rely.
- The cascade that works when you know the name (a CRM does): email domain → resolve the
company (verified, see above) →
organizational_urn → search_users {first_name, last_name, current_company: [{"type": "company", "value": "<id>"}]} → usually exactly
one match, delivered WITH the fsd_profile URN needed for user_posts. The company
filter is mandatory — a bare name returns namesakes.
Hiring signals:
linkedin/search/search_jobs — by company; works for any company. The company param
takes [{"type": "company", "value": "<numeric id>"}]. linkedin/search/search_companies
returns urn ALREADY in that object form — pass it through as-is. Only search_sql_companies
returns string URNs (fsd_company:<id>) — there, extract the numeric id yourself. And
verify the company before using its URN: the first search hit is often a namesake
(verified: "Notion" → NOTION Media Production first, the real notionhq second) — check
name + industry + alias.
greenhouse/jobs/jobs_search {board_token, count} — full descriptions via content=true;
ashby/jobs/jobs_search {board_name, count} — descriptions always included. Both need the
company slug; 412 = wrong token, fall back to linkedin jobs.
glassdoor (resolve employer id via companies_search first), builtin, adzuna —
supplements; blind/layoffs/layoffs_search for layoffs.
Tech stack: wappalyzer/technologies — technology slug → who uses it (top_websites
sample), category alternatives (alternatives[]), country/language breakdown. Note: it is a
sample, not an exhaustive site list.
Software reviews: g2/products/products_search (search only),
capterra/products/products_reviews (includes switched_from[] and switching_reason —
direct competitor-switch evidence), trustradius and getapp products_reviews,
gartner/products — competitor review mining. Employer sentiment:
glassdoor/companies/companies_ratings (employer id via companies_search), kununu
(DACH only — country ∈ de/at/ch), comparably, blind/companies/companies_reviews
(+ companies_salaries comp percentiles, companies_posts anonymous chatter).
News & mentions: techmeme/stories/stories_search {keyword, count} (archive) and
stories_front_page; google/news/news_articles_search;
linkedin/search/search_posts (keyword or mentioned company URN; date_posted accepts
only past-24h / past-week / past-month); reddit, hackernews, twitter, bluesky for
community chatter; substack/medium for content signals.
Launches & products: producthunt/launches/launches_search,
producthunt/products/products_alternatives, products_reviews; indiehackers,
kickstarter/indiegogo for niche ICPs.
Web fallback: webparser/parse (static pages) → webparser/render (JS-rendered).
Covers any URL when no named source fits. Web search: duckduckgo/search, brave/search.
Combining into signal chains
The standard pattern for account signals (used by the crm-signals skill):
company domain
→ search_sql_companies {website} + exact verify (firmographics
↳ crunchbase_link → alias FREE ↳ organizational_urn)
→ crunchbase/company {alias} → funding_rounds, leadership_hires, news, layoffs, bombora
→ search_jobs {company: [{type, value from organizational_urn}]} → what they hire for
→ search_posts (company name, past-month) → mentions
Four paid calls per account instead of five — the live crunchbase/search drops out (the
alias comes free from crunchbase_link); five if the domain doesn't resolve and webparser
is needed.
Stack signals: one signal is a guess, 2–3 signals within ~30 days is a pattern worth acting on.
Working with CRM
CRM read/write goes through the crm_* tools, NOT through execute. Before any CRM write,
consult the anysite-crm-profile skill (field mapping law) and the Writing rules in
anysite-crm-setup. The server enforces fill-blank policy, protected fields and write logging
regardless of what you pass.