| name | data-enrichment |
| description | Design data enrichment pipelines that augment first-party data with external sources. Outputs enrichment strategy, provider evaluation, pipeline design, match rate optimisation, and quality controls. |
| argument-hint | ["data types to enrich","use cases","budget","privacy requirements","existing data assets"] |
| allowed-tools | Read, Write |
Data Enrichment
Data enrichment adds context to your first-party data using external sources. A customer record enriched with firmographic data (company size, industry) enables better segmentation, scoring, and personalisation. The challenge is match rates, data quality, freshness, and cost.
Enrichment Use Cases
FIRMOGRAPHIC (B2B)
Sources: Clearbit, ZoomInfo, Apollo, LinkedIn, Crunchbase
Adds: company_size, industry, revenue_range, funding_stage, tech_stack
Use: Lead scoring, ICP matching, tier assignment
DEMOGRAPHIC (B2C)
Sources: Experian, Acxiom, first-party surveys
Adds: age_range, income_bracket, household_size, location_type
Use: Personalisation, product recommendations
BEHAVIOURAL ENRICHMENT
Sources: Intent data providers (Bombora, G2), review sites
Adds: in_market_signals, competitor_usage, buying_intent
Use: Sales prioritisation, timing of outreach
GEOGRAPHIC
Sources: Google Maps API, MaxMind, IP geolocation
Adds: timezone, metro_area, country, region, lat/lng
Use: Regional pricing, localisation, compliance
TECHNOGRAPHIC
Sources: BuiltWith, Wappalyzer, SimilarTech
Adds: tech_stack, cms, ecommerce_platform, analytics_tools
Use: Integration prioritisation, competitive intelligence
Enrichment Pipeline
import httpx
from pydantic import BaseModel
from typing import Optional
import asyncio
class ClearbitEnrichment(BaseModel):
company_name: Optional[str] = None
company_domain: Optional[str] = None
company_size: Optional[str] = None
industry: Optional[str] = None
country: Optional[str] = None
funding_stage: Optional[str] = None
annual_revenue_range: Optional[str] = None
linkedin_url: Optional[str] = None
enriched_at: Optional[str] = None
match_confidence: Optional[float] = None
class DataEnricher:
def __init__(self, clearbit_api_key: str):
self.client = httpx.AsyncClient(
base_url="https://company.clearbit.com/v2",
headers={"Authorization": },
timeout=,
)
() -> ClearbitEnrichment:
:
resp = .client.get(
,
params={: email},
)
resp.status_code == :
data = resp.json()
company = data.get(, {})
ClearbitEnrichment(
company_name=company.get(),
company_domain=company.get(),
company_size=company.get(, {}).get(),
industry=company.get(, {}).get(),
country=company.get(, {}).get(),
funding_stage=company.get(, {}).get(),
enriched_at=datetime.utcnow().isoformat(),
match_confidence=,
)
resp.status_code == :
ClearbitEnrichment(match_confidence=)
resp.status_code == :
ClearbitEnrichment(match_confidence=)
httpx.TimeoutException:
ClearbitEnrichment(match_confidence=)
() -> [, ClearbitEnrichment]:
semaphore = asyncio.Semaphore()
():
semaphore:
result = .enrich_by_email(email)
asyncio.sleep()
email, result
results = asyncio.gather(*[enrich_one(e) e emails])
(results)
Match Rate Optimisation
SELECT
DATE_TRUNC('week', enriched_at) AS week,
COUNT(*) AS total_records,
SUM(CASE WHEN match_confidence > 0 THEN 1 ELSE 0 END) AS matched,
SUM(CASE WHEN match_confidence IS NULL THEN 1 ELSE 0 END) AS pending,
ROUND(100.0 * SUM(CASE WHEN match_confidence > 0 THEN 1 ELSE 0 END)
/ COUNT(*), 1) AS match_rate,
AVG(CASE WHEN match_confidence > 0 THEN match_confidence END) AS avg_confidence
FROM customer_enrichment
GROUP
;
email, created_at
customers
enriched_at
email
email
email
created_at
LIMIT ;
Anti-Patterns to Avoid
| Anti-Pattern | Problem | Fix |
|---|
| Enriching without consent | GDPR violation in EU | Review legal basis; document legitimate interest |
| Storing enriched PII without TTL | Data minimisation violation | Enrich at point of use; or set retention policy |
| Single provider dependency | Provider outage or price increase | Multi-provider strategy with fallback |
| No match confidence tracking | Low-quality matches corrupt downstream models | Track confidence; use threshold for scoring models |
| Enriching all records | Cost waste on inactive accounts | Prioritise high-value or recently active accounts |
10 Rules
- Legal basis for enrichment must be documented — GDPR legitimate interest or consent.
- Match confidence is a first-class metric — low-confidence enrichment degrades model quality.
- Enrich on demand or at point of use — not for all records by default.
- Prioritise enrichment budget on high-value segments — not uniform across all customers.
- Multi-provider waterfall: try primary, fall back to secondary for misses.
- Freshness matters — firmographic data changes; re-enrich key accounts quarterly.
- Store raw enriched data separate from derived attributes — enables reprocessing.
- Track enrichment coverage by segment — "we have company_size for 70% of enterprise accounts".
- Personal email domains (gmail, yahoo) match poorly — focus on business emails.
- Enrichment is a complement to first-party data — never more trusted than your own signals.