The problem
If you sell to local businesses (parts to garages, supplies to clinics, services to restaurants), your market map is usually a spreadsheet someone scraped together. It has duplicates, wrong categories and no idea which of those businesses are already your customers or which ones your branches can actually reach. It takes weeks to build and starts going stale the day it's finished.
MapSight's job is to build that map properly, keep it honest, and tell you when it changes. The first vertical is the automotive aftermarket. This post is about the data engineering underneath, which is where most of the value lives.
A note on AI: language models do two narrow jobs in MapSight. They compile the prompt bar into a typed query plan and draft new industry packs for a human to review. Deduplication, classification, matching and scoring are deterministic, versioned code. Keeping it that way is what makes the results explainable and reproducible.
The pipeline
Any place on Earth, priced before it's built
A market starts as a geography: any administrative boundary (searchable in many languages), a drawn polygon, or a radius. The market stores its own multipolygon and every stage clips to it with PostGIS. London, Ontario is a test fixture, not an assumption. Geographic differences are only allowed to live in source adapters, jurisdiction metadata, pack assets and configuration.
Before anything runs, the estimator prices each source using the same conditions the dispatcher would enforce, so it never quotes optimistically. Approval happens on the server and is checked against a hash of the estimate. The approved amount is held against spending caps before a single paid call is made. A shared acquisition ledger means a probe someone has already paid for can be reused, where the source's terms allow it.
Durable builds and sources as manifests
Builds run as Temporal workflows (MarketBuildWorkflow fanning out to one SourceAcquisitionWorkflow per source), so a failed step retries instead of restarting a long build. Each source is registered by a declarative manifest in source-manifests/: Overture Maps places and divisions, Foursquare Open Source Places, OpenStreetMap through Overpass, OpenAddresses, Statistics Canada, plus Google Places used only for live verification, never bulk-retained. Policy is checked before capture. Rate limiting and circuit breakers sit in Redis, and bulk reads use DuckDB over GeoParquet.
Discovery over a region uses adaptive H3 cells: saturated cells are subdivided, cold ones stop early, and the build reports how complete it thinks it is instead of pretending to be exhaustive.
Observations, not overwrites
Every sighting of a business is stored as an append-only, hashed place_observation, broken into per-field field_claims that record which source said what, and when. The "current" business is a projection over those claims. That's what makes provenance possible on every field, and it's what lets MapSight compare snapshots later.
One business, one record, in any script
Entity resolution starts with blocking keys (phone, a normalised street key, a normalised name) so we only compare plausible pairs. Each pair gets a feature vector, and the evidence is combined as a noisy-OR of independent signals. Proximity weakens phone or domain matches but not a shared address. A score is explicitly not a probability, and a test enforces that. Marginal pairs go to an operator review queue, rejections are remembered so the queue drains, and every merge can be reversed.
The part we're proudest of is text.py: one conservative text fold shared by the resolver, the customer matcher, the classifier, the place index and the geocoder. It removes Latin diacritics and Arabic vowel marks, unifies Arabic, Persian and Urdu letter variants, reads digits in any script, and deliberately leaves Devanagari vowel signs and Japanese voicing alone because removing them changes meaning. On top of that, the features strip legal forms for dozens of jurisdictions and parse international phone formats. We tested it by walking real cities on several continents, and each fix was mutation-tested.
We go much deeper into the fold and the match score in our cross-script identity deep dive.
Industry packs: verticals as data
What counts as a "mechanic" versus a "tyre shop" versus noise lives in an industry pack: YAML for classification rules, brand dictionaries, exclusions and scoring weights, interpreted by a small rule engine. Each classification records the pack version and the rule that fired. To keep vocabulary out of the code, a test-only second pack (cafés) fails the build if automotive terms leak into Python.
Immutable versions
A market only publishes after a structural quality gate. Publishing writes an immutable market_dataset_version with the exact source releases pinned. Overture's schema is pinned per release, and unknown releases fail closed. We measured that the release metadata we'd hoped to rely on was empty, so the system pins explicitly. Every prospect ranking records the version it ranked, so an old result reopens exactly as it was built.
Your customers, your reach, your prospects
Customer lists come in as CSV or Excel. Column mapping is suggested from headers in several languages, and accounts and branches are matched to the market with the same resolver, with uncertain matches sent to a review only your organisation can see. Customer data is protected with forced row-level security per tenant.
Reach is measured by road, not as the crow flies. A self-hosted Valhalla engine built from OpenStreetMap computes drive times from each branch, with a cache keyed to the routing graph version and a cheaper H3 cell screen for large markets. Each result is one of within, outside, unreachable or unknown, and unknown never sorts as reachable.
Prospect scoring runs over a pinned version, excludes existing customers, and breaks every score down into raw value, normalised value, weight and contribution. If a feature can't be measured for a business, it's left out and named rather than filled with a made-up neutral value. The market's caveats travel into every exported CSV row. Exports are also protected against spreadsheet formula injection.
Market memory
When MapSight re-acquires a market, differences between snapshots become candidates, not events:
Corroboration counts distinct sources. A business that "disappears" while its source is having an outage triggers nothing. Changes in the dataset (a source renumbering its ids, for example) are kept separate from changes in the market (a garage actually closing).
Stack
| Layer | What we use |
|---|---|
| Backend | Python 3.12, FastAPI, Pydantic v2, SQLAlchemy 2 with GeoAlchemy2, Alembic (additive-only migrations) |
| Data | Postgres with PostGIS, pgvector and pg_trgm; DuckDB and PyArrow over GeoParquet; H3 and Shapely; S3-compatible object storage |
| Orchestration | Temporal workflows with a transactional outbox; Redis for rate limits and circuit breakers |
| Routing | Self-hosted Valhalla on OpenStreetMap extracts, behind a provider-neutral interface |
| Web | Next.js 16, React 19, MapLibre GL, Turf, Zustand, Tailwind 4, API types generated from OpenAPI |
| Quality | pytest with a per-test database (the full suite runs in about a minute), mypy, ruff, import-linter layer contracts, Playwright with accessibility checks |
What's next
MapSight is pre-release and working with design partners. What's still ahead: accuracy measured against independently labelled data (today's checks are against our own development fixtures), street-level geocoding for customer addresses, vector tiles for very large markets, CRM sync and Slack or webhook alerts, and more verticals beyond automotive. See how MapSight fits with the other products.