One business, many spellings: cross-script identity and deduplication in MapSight

A deep dive into how MapSight decides two listings are the same business: a conservative, versioned text fold that works across Arabic, Persian, Urdu, Devanagari, Thai, Latin and more; blocking keys; and a noisy-OR match score that refuses to pretend it's a probability.

The quality of a market map comes down to one question asked millions of times: are these two listings the same business? Get it wrong one way and a garage appears three times, inflating the market and spamming your sales team. Get it wrong the other way and two different garages merge, and one disappears from your prospect list.

MapSight is built to work anywhere, which makes the question harder. The same shop might be written with or without accents, on an Arabic or a Persian keyboard, with Eastern Arabic or Western digits. This post covers how MapSight handles that. It's a companion to the MapSight architecture post.

What went wrong before

We measured the failures before writing the fix, city by city, in a series of audit sweeps. The kind of thing we found:

  • Accents decided identity. “Auto Pièces” and “AUTO PIECES” scored 0.5 on name similarity, so they became two businesses. And a CRM typed without accents never matched the map.
  • Arabic spelling variants decided identity. An alef written with or without its hamza, or a final ta marbuta written as ha, made the same name look different.
  • Keyboards decided identity. The same name typed on a Persian and an Arabic keyboard uses different code points for “ya” and “kaf”, and scored 0.0.
  • Combining marks split words. Python's \w excludes Unicode marks, so a shadda broke an Arabic word into fragments, and Devanagari names lost the vowel signs between their letters. That made two genuinely different businesses look alike.
  • Digits decided identity. ١٢ and 12 produced different street keys for the same address.

The fold

The fix is one function, fold_identity in text.py: the single form that two writings of one name share. Everything that asks “is this the same name?” uses it: the entity resolver, the customer matcher, the industry classifier, the place index and the geocoder. Some of its patterns are written so that Postgres reads them too, so the SQL side folds identically.

The identity foldA business name is normalised with NFKC and case folding, decomposed with NFKD, stripped of marks that don't distinguish businesses (Latin diacritics, Arabic vowel marks, tatweel, Hebrew points, invisible joiners), has digits from any script converted to ASCII, is recomposed with NFC, and then has script-specific letter variants unified. The result is stamped with a fold version.Name as writtenany script, any keyboardNFKC + casefoldcompatibility forms, caselessNFKD → drop marksLatin accents · Arabic harakat · tatweel · joinersDigits → ASCII١٢ and ۱۲ and १२ all become 12NFC, then unify variantsArabic · Persian · Urdu · Devanagari · Thai · ı → iKept on purposeDevanagari vowel signs · Japanese voicingnothing is transliteratedFolded key · fold-6shared by resolver, matcher, classifier,place index and geocoder · Python and SQL
One conservative fold, versioned, so a stored key is never compared with one folded under different rules. Tap to zoom.

The rules are deliberately conservative:

  • Remove what doesn't tell businesses apart. Latin diacritics, Arabic vowel marks and the tatweel, Hebrew points, zero-width joiners, and the invisible direction marks a right-to-left spreadsheet carries.
  • Unify letters that writers use interchangeably. Arabic alef variants, ta marbuta and alef maqsura; Persian and Urdu forms of ya, kaf and heh; the Devanagari nukta, candrabindu and candra vowels; Thai and Lao SARA AM typed as one character or as its parts; Ethiopic homophone families; Turkish dotless ı; letters such as ł, ø, æ and đ that have no Unicode decomposition.
  • Never transliterate, and keep what's part of a letter. Devanagari vowel signs and Japanese voicing marks change meaning, so they stay.
  • Know the exceptions. Persian's zero-width non-joiner between two letters is a half-space that ends a word, so it becomes a space rather than being deleted. An apostrophe after o or g joins an Uzbek letter, but not before a final s, so “Doug's” stays a possessive.

Word splitting has its own rules. A word character is a letter or digit of any script or a mark attached to one. Scripts written without spaces (Chinese, Japanese, Korean, Thai, Lao, Khmer, Burmese) never join across a boundary, so “トヨタ” is found inside a longer shop name. Arabic and Hebrew prefixes that attach to the next word (“and”, “for”, the article) are recognised, so a trade word fused to one still counts.

The fold has a version

Folded text is stored, for example in the place index's search text. If the fold changes and the index isn't rebuilt, searches silently find nothing. So the fold carries a version (currently fold-6), stamped on everything stored folded. A test pins the fold's output to the version it claims, so changing one without the other fails the build.

From records to entities

The fold makes names comparable. Resolution decides what to do with them:

From records to entitiesRecords are grouped into blocks by shared phone, domain, street key or brand token, so only plausible pairs are compared. Each pair gets a feature vector and a noisy-OR score in which distance attenuates phone, domain, name and brand evidence but not a shared address. A source-pair policy applies thresholds, and pairs the rules accept but that score below the review floor go to a human review queue.Observed placesfrom every source in the marketBlocking keysph:phone · dm:domain · ad:street · tk:brandsame keys for places and imported customersPair featuresname Jaccard over folded tokens · phonedomain · street key · brand · distanceNoisy-OR score1 − Π(1 − wᵢ) · distance attenuatesphone, domain, name, brand, not addressSource-pair policymerge: name ≥ 0.60 within 200 mgroup departments within 120 mMergeone entityreversibleReviewscore < 0.55a human decidesSeparatedifferentbusinessesEvery decision stores its features and scoreand re-scores to exactly what it recorded
Blocking keeps the work near-linear; the score explains itself; anything marginal goes to a person. Tap to zoom.

Blocking

Comparing every pair is O(n²), so records are only compared when they share a blocking key: a phone number, a website domain, a normalised street key, or a brand token from the name. These are the same identity-bearing agreements the scorer reads. Imported customer accounts use the exact same blocking function as places, so there can't be a pair the pipeline compares that the importer never sees.

Features

Each candidate pair gets a feature vector. Name similarity is Jaccard overlap of the folded, identity-bearing tokens, taking the better of a couple of readings on each side. Phones are compared after parsing international trunk prefixes and extensions. Street keys refuse postcodes, unit labels and numbered districts, which would otherwise make neighbours look identical. Legal-form suffixes (“Ltd”, “GmbH”, “sp. z o.o.”, “Công ty TNHH” and many more) are set aside before comparing names.

The score

Evidence is combined as a noisy-OR: each signal independently argues that the two records aren't different businesses.

score = 1 − Π (1 − wᵢ)      over the signals present

same phone       0.85   ┐
same domain      0.80   │ scaled by proximity
name similarity  0.70×s │ 1 / (1 + (d / 250 m)²)
shared brand     0.15   ┘
same address     0.50     not scaled

Noisy-OR suits this problem better than a weighted average. It saturates rather than clipping: phone and domain together beat either alone, and nothing exceeds 1. And missing evidence is neutral rather than negative. A listing with no website isn't penalised for lacking a domain match, which matters because sparse records are the norm.

Distance changes what the evidence means. A shared phone number 3 km apart is usually a chain's central line, not one shop, so phone, domain, name and brand are scaled down with distance. A shared street address is not scaled, because it's already a claim about location, and two records of one business often sit a few hundred metres apart because of geocoding noise. Some examples:

EvidenceDistanceScore
Same phone, name overlap 0.5100 m0.81
Same phone, name overlap 0.53 km0.01
Same phone, same street address3 km0.50

Unknown distance (one side has no coordinates, as with most imported CRM rows) doesn't attenuate anything. Otherwise every imported account would be unmatchable.

A score, not a probability

This is the part we're strictest about. The weights are chosen, not fitted: we don't have enough independently labelled data to fit them honestly. So the number orders candidates and nothing more. Wherever it surfaces, it carries a measures field saying it is “rule-derived evidence strength, not a calibrated probability”, and a test enforces that wording. Every score also carries its working: the list of signals that contributed and how much.

Scores are versioned too (match-score-1). Every stored decision keeps its features and score, and must re-score to exactly what it recorded. If the weights change, re-deriving an old score raises an error instead of quietly showing a reviewer a number they never saw.

Policy, and a human in the loop

Thresholds come from a source-pair policy: by default, a merge needs name similarity of at least 0.60 within 200 m, and departments of one venue group within 120 m. The policy is keyed on where the two records came from, because two different providers agreeing is independent corroboration, and two records from the same provider agreeing is not. No per-provider overrides exist yet, deliberately, because no data justifies one. Any future override has to carry a written rationale or the registry test refuses it.

A pair the rules would accept but that scores below 0.55 isn't merged automatically. It becomes a review case for an operator. Rejected pairs are remembered so the queue drains instead of refilling, and every merge can be reversed.

Honest limits

The fold has been exercised across many scripts and cities, and each fix was mutation-tested, but entity resolution is never finished. The weights are hand-chosen. Precision and recall haven't yet been measured on an independently labelled set; until they are, we report what the system did and why, not an accuracy figure. Expect the fold version to keep climbing as new markets turn up new spellings.

See also: MapSight · MapSight's architecture · How the arcade fits together

See MapSight All posts