Proprietary data
Passport
- Rock
- Assets
- Depth
- 2.5 · Drill rig
- Time to dig
- 3–10 years
- Capital
- ◐ · medium
- Solo
- ✓ solo-reachable
- AI
- → eroding
- Rent
- ~ partly
Sample
- Share of apps
- 11.5%
- No-rate
- 57%
- Median price
- $29
Figures from the canivibecodeit sample. No-rate is the share of apps carrying this tag that cannot be vibe-coded — a proxy for structural strength. Only the first thirteen mechanics were measured.
Essence
Data that cannot be reassembled at a sane price: indexes, crawls, live feeds, archives, maps, unique datasets. The formula of the moat is capital × time: even with the money, a rival needs years to accumulate the history.
How it is built
Ahrefs and Semrush — crawling as capital construction
Ahrefs has crawled the web continuously with its own bot for more than ten years and keeps a historical index of links. To repeat it you would have to build the infrastructure of a small search engine and then wait years for the history to accrue — today's backlinks can be collected, ten years of their movement cannot. The moat is double: hardware and software, plus time that cannot be reproduced.
Waze — the users generate the data
Waze never bought traffic data; drivers create it by the simple act of driving with the app open. That is a data flywheel: more users → more accurate traffic → a more useful product → more users. The moat reinforces itself at no collection cost, while a challenger first has to find users it cannot attract without the data.
AllTrails and Strava — a user-generated archive
Years of user tracks, reviews, photographs and heatmaps. Each unit is cheap, the aggregate cannot be reproduced: you cannot pay for fifteen years of other people's hikes. User-generated content plus time is the cheapest way for a small team to build a data moat.
How it is bypassed
OpenStreetMap against commercial maps — open data covers 80% of the cases
Google and TomTom spent billions on cartography, but for a great many products OSM is enough: the community built a free alternative, and an entire industry grew on top of it (Mapbox among others). A data moat only protects the cases that need the genuinely unique part of the data; everything else is commoditised over time by open equivalents — Common Crawl does the same to proprietary crawls.
Waze against TomTom and Navteq — change the collection technology
Navteq and TomTom built their moat with fleets of scanning cars, hundreds of millions in capital expenditure. GPS in every smartphone zeroed that investment out: Waze collected roads and traffic for free, through its users' hands. The lesson: a data moat is tied to the technology used to collect it, and a new technology — today, LLM synthesis from public sources pressing on directories like ZoomInfo — can make an expensive archive unnecessary.
Buy the raw material wholesale
Wholesalers always grow around big data: DataForSEO and its kind sell crawl and SEO data as an API. The attacker does not need to rebuild the Ahrefs index — it buys the raw material and competes on interface, workflow and price. The moat does not disappear, but it shifts up the stack: only the part of the data nobody sells is still protected.
Verdict
Build it with a flywheel — data as a by-product of use — plus time. Bypass it with open data, a new collection technology, or by renting the raw material from wholesalers.