Data Valuation

What Is a Data Moat, and Does Proprietary Data Still Create One?

14 min read · 2026-08-19
9xmore daily search queries Google gets than all rivals combined, per its 2025 antitrust remedies ruling
93%of US outpatient prescription activity captured through IQVIA's pharmacy and payer data pipeline
300Testimated tokens left in the entire public web's usable text stock, the ceiling on what scraping can offer any rival
2019the year a16z first published its case against the data-moat thesis, years before the LLM boom made the debate mainstream
Key takeaways
  • 1A data moat is not the same as a large dataset: it's a competitive advantage a rival is structurally unable to replicate, whether through scale, contractual access, or law, not one they could close by writing a bigger check.
  • 2The classic data-flywheel thesis has real, credentialed pushback: Andreessen Horowitz's 2019 essay argued there is generally no inherent network effect from merely having more data, and Sequoia Capital's own 2023 essay reversed its earlier bet that a startup's usage data alone builds sustainable defensibility.
  • 3Some data moats are real and have been tested in court: Judge Amit Mehta's 2025 remedies ruling in US v. Google found the company's roughly 9x daily query-volume advantage over rivals fed its ranking systems in a way competitors could not close, forcing a data-sharing remedy.
  • 4Four conditions still make data defensible in the foundation-model era: continuously updating interaction or outcome data, regulatory or contractual access barriers, data that's genuinely expensive or illegal for a rival to recreate, and live-product feedback loops. Generic scraped text, anything buyable from the same marketplace as a rival, and frozen snapshots do not qualify.
  • 5Most startup data, especially a wind-down's, fails the moat test the moment the product stops running. The honest move at that point is to treat the data as a monetizable asset to license or sell now, not a defensibility story worth continuing to defend.
The short version

"Data moat" gets used as a synonym for "we have a lot of data," and that usage is mostly wrong. A real moat is a competitive advantage a well-funded rival is structurally unable to close, not one they'd merely need a checkbook to close. Venture firms that once championed the data-flywheel thesis have walked much of it back: Andreessen Horowitz argued in 2019 that most data advantages are just scale effects, and Sequoia Capital's own 2023 retrospective admitted its portfolio companies' usage data does not create an insurmountable moat now that foundation models generalize across domains and can absorb synthetic data for breadth. But the skeptics don't have the last word either: a 2025 federal antitrust ruling found Google's scale advantage in search-interaction data real enough that a court ordered it shared with rivals. The honest answer is that some data still defends a business and most doesn't, and the difference comes down to four testable conditions, not the size of the pile.

On this page ▾

What is a data moat?

"Economic moat," the term Warren Buffett popularized for a business advantage competitors can't easily cross, gets narrowed to "data moat" whenever the advantage is supposed to come from a company's accumulated data. The narrowing causes most of the confusion. A large dataset is an asset. A moat is a specific, stronger claim: that a competitor with comparable funding, talent, and time still cannot close the gap. Confusing the two is why so many pitch decks list "proprietary data" as a moat slide and so many investors have grown allergic to the phrase.

The sharpest formal definition of the underlying mechanism comes from network-effects research: a data network effect exists when a product's value rises as more data accumulates, and when ordinary usage of the product is what generates that data. [2]

When a product's value increases with more data, and when additional usage of that product yields data, then you have a Data Network Effect.
NFX: The Network Effects Manual

That definition sounds like it should apply to almost any data-driven product, and that's exactly the problem investors have identified since the term entered wide use. Most products satisfy the letter of the definition without ever getting a durable advantage from it, because the value each additional data point adds shrinks fast, and because a well-capitalized rival can often buy, license, or simulate a close enough substitute. Whether your data clears that higher bar is the actual question this article answers, category by category, rather than treating "we have data" as a self-evident yes.

The original thesis: usage begets data begets a better product

Before large language models generalized across domains, the standard pitch for a data moat ran as a closed loop: more usage produces more data, more data trains a better model or ranking system, a better product attracts more usage, and the loop compounds while competitors start from zero. Search ranking, recommendation systems, and fraud detection all plausibly worked this way for a decade, and venture firms underwrote the loop as a real source of defensibility.

Sequoia Capital said as much explicitly when generative AI took off. In 2023 the firm published its thesis that the strongest new AI application companies would win on exactly this mechanism, later summarizing its own prediction plainly: it had bet that the best generative AI companies could build a sustainable competitive advantage through a data flywheel. [2] That prediction was not a fringe idea. It was the working assumption behind a large share of application-layer AI investment for several years, on the theory that whichever product accumulated the most real usage data would pull permanently ahead of anyone who launched later with a comparable base model.

Why the loop looked so convincing
In a world where base models were roughly interchangeable and expensive to fine-tune, a founder's own usage data looked like the one input a copycat genuinely couldn't buy off the shelf. The loop wasn't a bad theory in 2021. It ran into a different world by 2023.

Why the data-moat thesis is now being actively questioned

Two forces broke the clean version of the loop. First, foundation models got general enough to perform competently across many domains without needing a company's own vertical data to bootstrap quality, narrowing the gap a proprietary dataset used to buy. Second, synthetic data became good enough to substitute for real data on breadth and volume, even where it can't substitute for outcome ground truth, so a challenger no longer has to collect years of real usage before reaching a workable baseline.

Andreessen Horowitz made the skeptical case earliest and most bluntly, years before the LLM wave made the argument fashionable. Its 2019 essay on the topic put the core claim in one line: there generally isn't an inherent network effect that comes from merely having more data. [2] The firm's broader point was about economics, not semantics: acquiring more data usually gets more expensive per unit as a dataset grows, while the marginal value each new data point adds to model quality usually shrinks, the opposite of the compounding curve the flywheel story assumes.

There generally isn't an inherent network effect that comes from merely having more data.
Andreessen Horowitz: The Empty Promise of Data Moats

Sequoia's 2023 reversal is the more striking data point precisely because the firm was reversing its own prior bet, in public, rather than critiquing someone else's. Its retrospective concluded that the data application companies generate does not create an insurmountable moat, and that the next generations of foundation models may well obliterate whatever data moats startups had built, redirecting its own view of defensibility toward workflow depth and user networks instead. [2]

The data that application companies generate does not create an insurmountable moat, and the next generations of foundation models may very well obliterate any data moats that startups generate.
Sequoia Capital: Generative AI's Act Two

Even companies built on genuinely enormous, expensive-to-collect proprietary data are hedging the same way. Waymo has run its self-driving fleet for well over a decade and disclosed in early 2026 that the Waymo Driver has logged nearly 200 million fully autonomous real-world miles, a dataset no new entrant could recreate quickly at any price. [2] Yet Waymo's own explanation of its newest simulation system says a system trained from scratch on only its own on-road data "learns from limited experience," and describes building its new world model on top of a general-purpose video foundation model instead of its fleet logs alone. [4] If the company with the deepest real-world driving dataset on earth doesn't trust that dataset alone to carry its simulation quality, the pure flywheel story is clearly incomplete, even where the underlying data is real and hard to copy.

What still makes data genuinely defensible

None of this means proprietary data has stopped mattering. It means the bar for calling something a moat is narrower than "we collected a lot of it." Four conditions still hold up, and each has a real, verifiable example behind it rather than a hypothetical.

1. Interaction or outcome data that keeps updating

Data that captures what actually happened, a click, a resolution, a correction, and keeps arriving as long as the product runs, is far harder to substitute than static content. The clearest tested example is Google's search click-and-query data. Judge Amit Mehta's September 2025 remedies ruling in the DOJ's antitrust case against Google found that Google receives roughly nine times more daily search queries than all of its rivals combined, an advantage the court concluded fed the company's ranking systems in a way rivals structurally could not close on their own. [4] The court's own description of that interaction data was blunt.

User-interaction data is the raw material.
Truth on the Market: Comparing the EU DMA to the Search-Query Data-Sharing Remedy in US v Google

The remedy itself is the tell: rather than telling Google to compete harder, the court ordered it to share the underlying click-and-query datasets feeding its ranking systems with qualified rivals. [2] Regulators don't order data-sharing remedies for advantages a checkbook could erase. They order them for the ones it can't.

2. Regulatory or contractual access barriers

Some data moats aren't about volume at all. They're about being one of a small number of parties with the contracts, compliance infrastructure, and regulatory standing to receive the data legally in the first place. IQVIA's prescription-claims business is a clean, disclosed example: the company reports capturing 93% of US retail outpatient prescription activity, built from direct data-sharing arrangements with pharmacies, payers, software vendors, and transactional clearinghouses. [4] No rival can scrape that pipeline into existence. It requires the specific commercial contracts, healthcare compliance infrastructure, and years of trust with data furnishers that IQVIA already has, which is exactly the kind of access barrier a16z's essay identified as one of the few conditions where data does create real defensibility. [6]

3. Data that's genuinely expensive or illegal for a rival to recreate

Some datasets are defensible because reproducing them costs more money, more time, or more legal clearance than almost any competitor will spend. Waymo's roughly 200 million real-world autonomous miles fall squarely in this category: a decade-plus of fleet operations, sensor hardware, safety drivers, and regulatory approvals across multiple cities, none of which a challenger can shortcut by writing a bigger check this quarter. [2] As the previous section noted, Waymo doesn't treat that dataset as sufficient on its own, but it treats it as necessary, still anchoring its simulation and safety validation in real-world miles that took over ten years to accumulate and combining them with general-purpose pretraining rather than replacing them with it. [4] Hard-to-recreate and sufficient-by-itself are different claims, and only the first one is generally true.

4. Feedback loops from a live product

The purest version of the classic flywheel still exists where a live product keeps generating fresh signal about what users accept, reject, or correct. GitHub disclosed in March 2026 that Copilot's interaction data, including data generated by active use of private repositories, feeds back into model training by default unless a developer opts out. [2] That loop only exists because GitHub already hosts the code and the developer relationship; a rival coding assistant can license public code, but it cannot buy the ongoing stream of accept/reject signal from millions of developers actively using a specific tool inside a specific platform. That stream, not the static training corpus underneath it, is the part that keeps compounding.

The common thread
In all four cases, the advantage survives because a rival can't pay to skip the years, the contracts, or the live usage base. If a competitor could replicate your position by writing a check this quarter, what you have is an asset, not a moat.

What no longer counts as a moat

Three categories get pitched as data moats constantly and rarely survive scrutiny.

  • Generic scraped or licensable text. If a well-funded rival can acquire an equivalent corpus by scraping the same public web or licensing from the same aggregators you did, it was never exclusive to you. Public text is also a shrinking resource in absolute terms: Epoch AI estimates the entire usable stock of human-generated public text at roughly 300 trillion tokens, with exhaustion plausible between 2026 and 2032 on current scaling trends, meaning every serious lab is chasing the same finite pool rather than any one company holding a private supply. [4]
  • Anything a rival can buy from the same source. Data sold non-exclusively on an open marketplace, including Dayda's, stops being a moat for the buyer the moment it's purchased, precisely because other buyers can acquire the same asset. That's not a flaw in the marketplace model; it's the honest description of what a non-exclusive license is. A moat requires the rival to be structurally unable to get the data, not merely unwilling to pay for it yet.
  • Stale snapshots with no new inflow. A dataset frozen at the moment a company stopped operating can't feed any of the four defensible categories above: it doesn't update, it isn't protected by an active contract, and there's no live product generating new feedback. It may still be a valuable one-time asset, but calling it a moat overstates what a buyer is actually getting.

Moat types compared: what holds up and what doesn't

Moat typeStill defensible in the foundation-model era?Real example
Generic scraped or public web textNo. Any funded lab can scrape or license the same corpus.The shared, shrinking pool of public text tracked by Epoch AI [2]
One-time proprietary dataset, no new inflowNo. Value is real but fixed, and it decays the moment it stops updating.A wound-down startup's frozen support-transcript archive
Anything buyable on the same open marketplaceNo, for the buyer. Non-exclusive access by definition isn't exclusive.Non-exclusive datasets licensed to multiple buyers
Continuously updating interaction/outcome dataYes, while usage keeps flowing and the loop stays proprietary.Google's search click-and-query logs [2]
Regulatory or contractual access barrierYes, as long as the contracts and compliance infrastructure hold.IQVIA's ~93% coverage of US outpatient prescriptions [2]
Data expensive or slow to physically recreatePartially. It's a real edge, but foundation-scale pretraining can substitute for part of it.Waymo's ~200 million real-world autonomous miles [2]
Live-product feedback loopYes, if the product retains a durable, active usage base.GitHub Copilot's opt-out interaction-data training pipeline [2]

A 7-question checklist: is your data a moat, or an asset to sell once?

Run your own dataset through these seven questions honestly. A founder who scores mostly toward "moat" has a genuine reason to keep building on the data rather than monetizing it away. A founder who scores mostly toward "asset" is holding something real, valuable, and worth selling now, just not a defensibility story.

  • 1Is new data still arriving from a live process, or is this a fixed historical snapshot? A flywheel needs an active loop, not a photograph of one.
  • 2Would a well-funded rival need years and comparable access, not just money, to reproduce your pipeline? If a check alone would close the gap, it isn't structural.
  • 3Is a meaningful share of the value in outcomes or feedback, rather than raw content? Resolution labels, accept/reject signals, and correction data are harder to substitute than the underlying text.
  • 4Is the data locked behind a regulatory, contractual, or physical barrier a rival can't casually clear? Contracts, compliance certifications, and hardware fleets qualify. A public API does not.
  • 5If you licensed or sold it today, would competitors lose access to something they otherwise couldn't get, or just gain something they could have bought anyway?
  • 6Could a frontier foundation model already get close to your result without your data, through general pretraining or synthetic substitutes? If yes, your data's defensive value is thinner than it looks.
  • 7If your company shut down tomorrow, does the data keep compounding on its own, or does its value stop growing the moment the product does?
The tell that ends most arguments
Question 7 is the fastest gut check. A true moat produces value independent of your continued operation, because it's structural. An asset stops appreciating the moment the underlying company stops running it, which is precisely the position most winding-down startups are in.

A worked example: two founders, two very different answers

Consider two founders who both built customer-support software and both accumulated large transcript datasets, to see how differently the checklist scores them.

Founder A still runs a live SaaS product with 40,000 active daily conversations. Every conversation generates a resolution label, a CSAT score, and an agent-override signal that feeds back into the company's own ranking model within hours. A competitor could buy a comparable-sized transcript corpus on the open market, but they could not buy the live feedback loop, because it only exists inside Founder A's running product. Founder A scores "moat" on questions 1, 3, and 7: the data keeps arriving, a meaningful share of its value is outcome-based, and the value stops compounding for a rival the instant they stop using Founder A's own product. The correct move for Founder A is to keep the data inside the product and, at most, license narrow, non-competing slices of it, not sell it wholesale.

Founder B is winding the same kind of business down. The transcripts are 18 months old, resolution labels stopped updating the day the support team was let go, and there is no live product left to generate new signal or absorb a competitor's use of the old data. Founder B fails questions 1, 6, and 7 outright: nothing new is arriving, a synthetic or newer real dataset could substitute for most of the training value, and shutting the company down did not cost the dataset any of its value because that value was never structural in the first place. Founder B's data is a real, sellable asset, potentially a valuable one, but it stopped being a moat the day the product did. The economically rational move is to license or sell it now, while it's still recent enough to be worth something, rather than let it decay further while defending a defensibility claim nobody in the room actually believes.

Checklist questionFounder A (live product)Founder B (winding down)
New data still arriving?Yes, continuouslyNo, frozen
Outcome/feedback share of value?High (resolution + CSAT labels)Low, static content only
Value survives if the company shuts down?No, tied to the running productAlready effectively zero either way
Correct moveKeep it in-product; license narrow slices onlySell or license it now, before it decays further

The honest verdict on data moats

Both camps in this debate are right about different data. The skeptics are right that most startup data was never a moat, just a scale effect that looked impressive until a rival matched it or a foundation model made it less necessary. [2][4] The believers are right that some data genuinely is defensible, real enough that a federal court ordered one of the world's largest companies to share it with rivals. [6] The difference isn't the size of the dataset. It's whether the data keeps updating from a live process, sits behind a genuine access barrier, costs real years and capital to recreate, or feeds an active feedback loop a competitor cannot buy their way into.

If you're building a company that's still running, that distinction should shape what you protect and what you license out. If you're winding one down, it should shape something more immediate: the product that generated your data no longer exists, which means whatever moat properties that data had are already gone. What's left is an asset, not a defensibility story, and assets are worth the most before they decay any further.

Find out what your data is actually worth.

If your product has stopped running, your data has stopped being a moat, but it hasn't stopped being valuable. Get a free, no-obligation assessment from the team that brokers these deals every day.

List your data on Dayda