Deduplicate HubSpot contacts at production scale — surviving import storms, wrong-winner
merges, fuzzy-match blind spots, association orphans, rate-limit exhaustion, and silent
merge failures on conflicting lifecycle or opt-out status. Use when cleaning a CRM after
a bulk import, running a nightly dedup pipeline on millions of records, recovering from
a merge that destroyed the wrong timeline, or building fuzzy matching beyond HubSpot's
native email-uniqueness. Trigger with "hubspot dedup", "hubspot merge contacts",
"hubspot duplicate contacts", "hubspot contact cleanup", "hubspot import duplicates",
"hubspot fuzzy match contacts".
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Deduplicate HubSpot contacts at production scale — surviving import storms, wrong-winner
merges, fuzzy-match blind spots, association orphans, rate-limit exhaustion, and silent
merge failures on conflicting lifecycle or opt-out status. Use when cleaning a CRM after
a bulk import, running a nightly dedup pipeline on millions of records, recovering from
a merge that destroyed the wrong timeline, or building fuzzy matching beyond HubSpot's
native email-uniqueness. Trigger with "hubspot dedup", "hubspot merge contacts",
"hubspot duplicate contacts", "hubspot contact cleanup", "hubspot import duplicates",
"hubspot fuzzy match contacts".
allowed-tools
Read, Bash(curl:*), Bash(jq:*), Bash(python3:*)
version
2.9.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
compatibility
Designed for Claude Code
tags
["hubspot","crm","deduplication","data-quality"]
HubSpot Contact Deduplication
Overview
Merge duplicate contacts in HubSpot and operate that process in production, at scale, without data loss. This is not a one-click cleanup guide — it is the logic your pipeline runs when a sales ops team imports 80,000 leads from a tradeshow CSV that already exist in the CRM, when a merge destroys the "winner" contact's email history, when a fuzzy match on "Jon" vs "John" leaves a six-figure deal associated to a ghost record, and when on-call discovers that 40,000 contacts were merged without checking opt-out flags.
The six production failures this skill prevents:
Import storms creating thousands of exact duplicates — HubSpot enforces email uniqueness only at the property level; the merge API has no dedup-all-at-once endpoint. A 100K-row CSV import where 60% of rows already exist creates 60,000 duplicates that must be found and merged one pair at a time within a 100 req/10s rate envelope.
Merge destroying the wrong timeline — POST /crm/v3/objects/contacts/merge requires a primaryObjectId. Picking the wrong one demotes the older contact's full activity timeline — calls, emails, form submissions — to the discarded record's history.
Property-based dedup missing fuzzy matches — Email-exact dedup leaves "john@gmail.com" and "jon.smith@googlemail.com" as separate records. Phone dedup leaves "+1 (512) 867-5309" and "5128675309" as separate records. Without normalization your CRM accumulates a shadow population of semantically identical but technically distinct contacts.
Post-merge association orphans — When a secondary contact has deals, tickets, or company associations, HubSpot re-parents most automatically — but not all. Custom object associations and some third-party-integration links may not follow.
Rate-limit exhaustion on large catalogs — A 1-million-contact dedup scan requires 10,000 batch reads (2.7 hours at full throughput, before merge calls). Naive single-threaded loops exhaust the 500K daily quota before the search phase finishes.
Silent merge failures on conflicting lifecycle or opt-out status — The merge API returns 200 even when the resulting contact has hs_email_optout=true overriding the primary's opted-in status. HubSpot's "most recently updated value wins" rule is wrong for compliance flags.
Auth
Authenticate with a private app token (pat-na1-*) or OAuth access token. Pass it on every request:
Authorization: Bearer {your-token}
Required scopes: crm.objects.contacts.read, , , . See the for token caching, OAuth refresh, and scope-drift detection.
Python 3.10+ (requests, phonenumbers, rapidfuzz) for the full pipeline
HubSpot Professional or Enterprise account (batch merge at scale)
Private app token with required scopes (above)
jq for shell examples
For catalogs >500K contacts: confirm daily quota with HubSpot support
Instructions
Step 1. Discover duplicates with search
Find exact duplicates by email using the search API. Never pull all contacts into memory for comparison — use the search endpoint with specific filter values.
For full-portal scans across millions of contacts use the four-stage Python pipeline in implementation-guide.md. The pipeline writes a local SQLite checkpoint so rate-limit interruptions do not require starting over.
Step 2. Select the primary (winner) contact
The oldest contact by createdate is the primary — its timeline is most historically complete. Two overrides apply:
If the oldest contact has hs_email_optout=true and the newer one does not, prefer the opted-in record as primary to avoid propagating unsubscribe status.
If the oldest contact has a test-domain email (@mailinator.com, @example.com, @test.com), always make the real-address contact the primary.
from datetime import datetime
defpick_primary(contacts: list[dict]) -> tuple[dict, list[dict]]:
"""Return (primary, secondaries). contacts is a list of HubSpot result dicts."""
TEST_DOMAINS = {"mailinator.com","example.com","test.com","yopmail.com"}
defis_test(email: str) -> bool:
return (email or"").split("@")[-1].lower() in TEST_DOMAINS
# Sort oldest first (default primary)
sorted_c = sorted(contacts, key=lambda c: c["properties"]["createdate"])
primary = sorted_c[0]
# Opt-out overrideif primary["properties"].get("hs_email_optout") == "true":
opted_in = next((c for c in sorted_c[1:] if c["properties"].get("hs_email_optout") != "true"), None)
if opted_in:
primary = opted_in
# Test email overrideif is_test(primary["properties"].get("email", "")):
real = next((c for c in sorted_c ifnot is_test(c["properties"].get("email", ""))), None)
if real:
primary = real
secondaries = [c for c in contacts if c["id"] != primary["id"]]
return primary, secondaries
Step 3. Normalize emails and phones for fuzzy matching
Exact-email dedup leaves a shadow population. Normalize before comparing:
DAILY_STOP_AT = 480_000# Stop at 96% of 500K quotadefcheck_quota(resp: requests.Response) -> None:
remaining = int(resp.headers.get("X-HubSpot-RateLimit-Daily-Remaining", 500_000))
if (500_000 - remaining) >= DAILY_STOP_AT:
raise SystemExit("Daily quota near limit — stopping. Resume after midnight UTC reset.")
Step 6. Post-merge verification and association repair
After merging, verify that the surviving contact's hs_email_optout matches the expected value (Step 4) and patch it if it drifted. Then audit associations that may not have transferred automatically:
# Check associations on surviving contact (replace 12345 with actual primary contact ID)
curl -s "https://api.hubapi.com/crm/v4/objects/contacts/12345/associations/deals" \
-H "Authorization: Bearer {your-token}" | jq '[.results[].toObjectId]'# Manually create a missing association (replace 12345 with primary ID, 67890 with deal ID)
curl -s -X PUT \
"https://api.hubapi.com/crm/v4/objects/contacts/12345/associations/deals/67890" \
-H "Authorization: Bearer {your-token}" \
-H "Content-Type: application/json" \
-d '[{"associationCategory":"HUBSPOT_DEFINED","associationTypeId":3}]'
The full four-stage Python pipeline (scan → pair → qualify → execute) with automatic association repair is in implementation-guide.md.
Error Handling
HTTP Status
Error
Root Cause
Action
400
CONTACT_ALREADY_MERGED
Secondary was already merged into another record
Re-fetch secondary; check hs_merged_object_ids for surviving primary ID
400
SAME_OBJECT_MERGE
Both IDs are identical
Remove self-merge pairs from candidate list before executing
400
INVALID_OBJECT_TYPE
One ID belongs to a different CRM object type
Verify via GET /crm/v3/objects/contacts/{id} before merging
404
OBJECT_NOT_FOUND
Contact was deleted between discovery and merge
Re-fetch to confirm existence; skip if deleted
409
MERGE_IN_PROGRESS
A concurrent merge is already running for this contact