AI Data Buying

Data Broker vs. Data Marketplace: What's the Difference?

13 min read · 2026-09-02
6 statesnow require data broker registration or will by 2027 (CA, VT, OR, TX, CT, NJ)
$6,000+annual California data broker registration fee, plus processing costs
70%+of a large sample of AI training datasets found to lack clear licensing or provenance info
2026–2032window when public human training text may run out, driving demand for proprietary data
$278Bestimated size of the global data broker market in 2024
Key takeaways
  • 1Data broker is a defined legal category. California (Civil Code §1798.99.80(c)) and Vermont (9 V.S.A. §2430) both define it as a business that sells or licenses a consumer's personal information to third parties without a direct relationship with that consumer. An AI training data marketplace brokering consented, first-party, or provenance-documented company data generally does not meet that definition.
  • 2Six states now regulate data brokers directly: California, Vermont, Oregon, and Texas require annual registration today, with Connecticut (2027) and New Jersey (registration effective March 2027, sensitive-data-sale ban already in force) recently added. {cite:tonkon}{cite:nj} California charges roughly $6,000 a year to register and now requires brokers to process consumer deletion requests through its new DROP platform within 45 days. {cite:ca-cppa} Marketplaces that broker consented business data carry none of these obligations because they aren't reselling third-party consumer profiles.
  • 3The FTC's own definition of a data broker turns on consumer invisibility: a business that collects and resells personal information about people it never interacts with, who usually don't know it holds their data. {cite:ftc2014} An AI training data marketplace inverts that by design, gating access behind NDAs and documenting exactly where each dataset came from before a buyer ever samples it.
  • 4The business models are structurally different, not just differently worded: a broker's revenue scales with resale volume across many buyers of the same anonymous profile, while a marketplace's revenue scales with curated, often exclusive matches between one seller and the specific buyers whose model needs that data.
  • 5The distinction has teeth. In December 2024 the FTC settled with data brokers Mobilewalla and Gravy Analytics for selling sensitive location data without verifying consumer consent, including the agency's first-ever ban on collecting data from real-time ad-bidding exchanges for purposes beyond the auction itself. {cite:ftc-mobilewalla} Any company deciding where to list data, or where to source it, should treat 'would this survive an FTC consent investigation' as the real test.
The short version

A data broker and an AI training data marketplace can look similar from a distance: both connect a party that holds data to a party that wants it, for a fee. The resemblance stops at the business model. A data broker is a legally defined category in states like California and Vermont, built around collecting and reselling personal information about consumers the broker never directly dealt with, usually without those consumers' knowledge. An AI training data marketplace brokers access to consent-documented, provenance-tracked, often first-party datasets under NDA-gated sampling and negotiated licenses, matching one seller to the specific buyers who need that data rather than reselling the same profile to whoever pays. The regulatory treatment reflects the difference: brokers face state registration laws built explicitly to police anonymous resale of personal profiles, while marketplaces are governed by ordinary contract, IP, and privacy law applied to a specific, disclosed transaction. Neither category is inherently virtuous. A marketplace that skips provenance diligence can still functionally behave like a broker, and the honest answer to 'isn't a marketplace just a nicer word for a broker' is: it depends entirely on what gets verified before a dataset changes hands, not on which word is on the homepage.

On this page ▾

What exactly is a data broker, legally speaking?

"Data broker" isn't a vague industry label. In the states that regulate the term, it's a precise legal category. California's Civil Code §1798.99.80(c) defines a data broker as a business that "knowingly collects and sells to third parties the personal information of a consumer with whom the business does not have a direct relationship." [4] Vermont's statute uses nearly identical language: a business that "knowingly collects and sells or licenses to third parties the brokered personal information of a consumer with whom the business does not have a direct relationship." [8] The load-bearing phrase in both is no direct relationship. The broker never signed you up, never served you, and never got your attention before adding you to its file.

The FTC's landmark 2014 study of the industry described the same pattern from the consumer's side: data brokers "collect data from numerous sources, largely without consumers' knowledge," and because they never interact with the people whose data they hold, most consumers don't know the broker exists, let alone what it has compiled about them. [4] That's the defining trait, not the fact that money changes hands for data. Plenty of legitimate businesses sell data. A broker specifically sells data about people it never met.

Data brokers collect data from commercial, government, and other publicly available sources... Because data broker companies generally never interact with consumers, consumers are often unaware of their existence, much less the variety of practices in which they engage.
Federal Trade Commission: Data Brokers: A Call for Transparency and Accountability
  • Consumer data aggregators that compile purchase histories, demographics, and inferred interests from hundreds of sources for marketing lists.
  • Ad-tech and location-data firms that resell device identifiers, app usage, and precise location signals harvested from ad auctions and SDKs embedded in unrelated apps.
  • People-search and background-check sites that assemble public-record and scraped-web profiles and sell lookups on named individuals.
  • Risk and identity-verification brokers that resell compiled consumer files to detect fraud or verify identity for a fee, per transaction.
The one fact that defines the category
Every state definition and the FTC's own framing converge on the same test: does the business have a direct relationship with the person whose data it's selling? If not, it's a data broker, regardless of how the company describes itself.

What is an AI training data marketplace, and how is it structured differently?

An AI training data marketplace brokers something categorically different: datasets, not individual consumer profiles, typically originating from a company's own operations (support transcripts, product telemetry, domain-specific documents, annotated examples) or from data subjects who consented specifically to training-data use. A seller lists a described, sampled dataset. A buyer, usually an AI lab or an enterprise building a domain model, reviews documentation, signs an NDA, samples a held-out slice, and negotiates a license that specifies exactly what it may do with the data and for how long.

The transaction the marketplace brokers is a licensed grant to a described, provenance-documented asset, closer in structure to an M&A data-room process than to a marketing-list purchase. Nobody buys "a person's profile" on a training data marketplace. They buy the right to train on a specific, disclosed corpus whose lawful basis for existing and for being transferred has already been checked.

The one-line mental model
A data broker sells what it knows about people who never agreed to be known. An AI training data marketplace sells access to what a company or its consenting users actually produced, under terms both sides can point to in writing.

The regulatory distinction: why data broker laws don't (usually) apply to marketplaces

Six states now police the data broker category by name, and the list is growing. California, Vermont, Oregon, and Texas all require annual registration from any business that meets their data broker definition. [2] Connecticut's registration requirement takes effect in 2027, and New Jersey enacted one of the most aggressive versions yet in June 2026, a law that bans the sale of sensitive data (health, biometric, precise geolocation, children's data) immediately and adds a registration requirement, with fees scaling up to $1.5 million depending on volume, effective March 27, 2027. [6]

California's version is the one most companies actually encounter. Under the Delete Act, a business meeting the data broker definition must register every year by January 31, pay a fee north of $6,000, and, as of 2026, check the state's new Delete Request and Opt-Out Platform (DROP) at least every 45 days and honor consumer deletion requests submitted through it. [6] Vermont's fee is smaller in dollar terms but paired with real teeth: $900 a year plus a $20,000 surety bond, tightened further by 2026 amendments that narrow what counts as a "direct relationship" and take effect January 1, 2027. [12]

JurisdictionWho must registerCost / obligationStatus
CaliforniaBusinesses meeting Civ. Code §1798.99.80(c)$6,000+/yr; DROP deletion compliance every 45 daysIn force; DROP live since Jan 2026 [2]
VermontBusinesses meeting 9 V.S.A. §2430$900/yr + $20,000 surety bondIn force; amended definition effective 2027 [2]
Oregon / TexasState-specific data broker definitionsRegistration + opt-out and security disclosuresIn force [2]
New JerseyData brokers and "data collectors" selling to themFees up to $1.5M/yr by volume; sensitive-data sale banSale ban in force now; registration March 2027 [2]

None of these laws mention AI training data marketplaces because none of them were written with training data in mind. They target the specific harm of anonymous, no-relationship resale of consumer profiles. A marketplace that brokers a company's own first-party data, under a license naming a specific buyer and a specific permitted use, generally doesn't meet the statutory data broker definition in the first place, because the seller either has a direct relationship with the people in the data (its own customers or users) or the data isn't personal information about identifiable consumers at all. That's a structural difference in what's being sold, not a loophole.

The test that actually matters
Don't ask "am I a marketplace or a broker" as a matter of branding. Ask the statutory question directly: does the entity selling this data have a direct relationship with the people described in it, or is it reselling profiles of strangers it aggregated from elsewhere? The answer determines which regime applies, not the name on the deal.

The regulatory gap exists because the underlying transactions are built differently. A data broker's supply chain is designed for scale and anonymity: scraped public records, purchased mailing lists, ad-auction bid-stream data, app SDKs collecting on behalf of dozens of unrelated apps. Consent, if it exists at all, is usually buried in a third party's terms of service the consumer never read, several hops removed from the broker actually reselling the data. That's precisely the gap the FTC targeted in its December 2024 action against Mobilewalla, alleging the company failed to meaningfully verify that its own data suppliers had obtained real consumer consent before Mobilewalla resold the data downstream. [4]

An AI training data marketplace's supply chain is designed for traceability. A serious marketplace requires a seller to document where each record came from, under what consent or lawful basis it was collected, and whether any third party holds rights in it, before the listing ever goes live. That matters because the industry's baseline is bad: an academic audit of 1,800+ widely used AI training datasets found that over 70% lacked clear licensing information and roughly half had it recorded incorrectly. [6] A marketplace's entire value proposition to a buyer is closing that gap, not replicating it.

We find widespread miscategorization of licenses on popular dataset hosting sites, with license omission rates above 70% and error rates above 50%, constituting a crisis in misattribution and informed use of some of the most popular datasets in AI.
Longpre et al.: The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI
45 days
the window California data brokers now have to process a consumer's DROP deletion request
California Privacy Protection Agency

Business model: volume resale vs. curated matching

The economics diverge as sharply as the legal treatment. A data broker's revenue model rewards breadth: license the same underlying profile of a person or household to as many marketers, risk teams, and analytics buyers as possible, because the marginal cost of reselling a record is close to zero and the seller doesn't need any one buyer's use case to be a good fit. That volume logic is why the traditional data broker industry is estimated at roughly $278 billion globally in 2024, an industry built on reselling the same data many times over. [6] An AI training data marketplace's revenue model rewards fit: a dataset is worth the most to the specific buyer whose model gap it fills, and pricing depends heavily on whether the buyer gets exclusive or field-of-use rights, not on how many buyers can be signed up for the same record.

DimensionTraditional data brokerAI training data marketplace
Typical data typeAggregated consumer profiles: demographics, location, purchase and browsing historyFirst-party company data or explicitly consented corpora: support transcripts, domain documents, annotated examples
Consent modelIndirect or absent; consumer usually never interacted with the brokerDocumented at collection; provenance file required before listing
Regulatory treatmentState data broker registration laws (CA, VT, OR, TX, CT, NJ); FTC scrutiny of consumer-data resale [2][4]General contract, IP, and privacy law applied to a specific, negotiated license
Business modelVolume resale of the same profile to many buyers at low marginal costCurated, often exclusive matching between one seller and a small set of qualified buyers
Buyer relationshipAnonymous, transactional; buyer typically never interacts with the seller directlyNDA-gated sampling, negotiated Data Purchase Agreement, direct seller-buyer relationship
Exclusivity is a marketplace concept, not a broker one
A broker's business model breaks if it can't resell the same record repeatedly. A marketplace's highest-value deals often go the other way: a buyer pays a premium precisely to be the only one with access, which only makes sense when the seller isn't also trying to resell the same data to a dozen other parties.

A worked example: the same underlying data, two very different paths

Picture a mid-sized fitness app with five years of workout logs, device sensor data, and in-app chat transcripts between users and human coaches. The company is winding down its consumer product and deciding what to do with that data. Two paths are on the table.

Path A: sell to a data broker or ad-tech aggregator. The company licenses its raw location and device-usage logs to a data aggregator that resells them, alongside dozens of other apps' feeds, to advertising and analytics buyers who never see or care which app the data came from. Because the fitness app has a direct relationship with its own users, the app itself likely isn't a data broker under California or Vermont's definitions, but the downstream aggregator reselling that data to strangers almost certainly is, and it inherits registration and DROP-compliance obligations. [6] If the app's own consent language didn't clearly disclose resale for this kind of secondary use, the deal also creates real exposure of the kind the FTC pursued against Mobilewalla for inadequately verified consent. [8]

Path B: license the coach-chat transcripts through an AI training data marketplace. The same company instead lists its human-coach chat transcripts, cleaned, de-identified, with documented user consent for AI-training use, on a marketplace. A health and fitness AI startup signs an NDA, samples a held-out batch to confirm quality, and licenses the corpus under a field-of-use exclusive: the buyer gets sole rights to train coaching models on it for 18 months, while the seller retains rights to license the raw workout-log data (non-personal, aggregated) separately for analytics use. No data broker registration is triggered, because the transaction is a documented, consented, direct license, not anonymous resale to an unrelated third party.

Same company, same underlying five years of data. One path routes through a regulatory category built to police anonymous resale and invites FTC and state-registration exposure if consent wasn't airtight. The other routes through a marketplace structure built around documentation and stays outside that category entirely, while very plausibly clearing a higher price because the buyer gets exclusivity and a defensible provenance file.

Isn't a marketplace just a nicer word for a broker?

This is the fair objection, and it deserves a direct answer rather than a dismissal. Both are intermediaries. Both take a cut for connecting a data holder to a data buyer. And plenty of company history in this space earns the skepticism: some businesses have relabeled themselves from "data broker" to "data marketplace" or "data partner" over the past decade without changing a single underlying practice, precisely because the new label carries less regulatory and reputational baggage.

The honest distinction isn't the label. It's whether three things are actually true of a given deal: the seller has a direct relationship with the people described in the data, or the data isn't personal information about identifiable people at all; consent or a documented lawful basis exists for the specific use being licensed; and the buyer knows exactly what they're getting and from whom, rather than an anonymous feed blended with dozens of other sources. A platform that calls itself a marketplace but skips provenance checks, accepts scraped or resold consumer profiles, and lets buyers train without disclosure is functionally a broker regardless of its homepage copy, and would likely meet the statutory data broker definition on the merits. [4] Conversely, a company that calls itself a broker but only ever sells data it directly collected from its own consenting customers, under named, disclosed licenses, is closer to a marketplace in substance. The category that matters is the one a regulator or a court would apply to the actual transaction, not the one on the pitch deck.

The genuine overlap
Where the line blurs most is B2B data with a thin consent trail: a company's own operational data that includes third-party personal information collected incidentally (a support ticket that names a customer's spouse, a location log embedded in a photo). That data needs the same provenance rigor as consumer-facing data, and a marketplace that waves it through without checking is taking on broker-shaped risk whatever it calls itself.

Why the distinction matters in practice, for sellers and for buyers

For a company deciding whether to list its data, the category determines both risk and price. Selling into the broker model means the buyer, or a downstream aggregator, may need to register in up to six states, may face fee obligations that scale with volume (New Jersey's law reaches $1.5 million a year at the high end), and puts consent gaps directly in the FTC's line of sight, since the agency has now twice in one enforcement cycle penalized data brokers for inadequate consent verification. [6][8] Selling through a documented marketplace transaction instead avoids most of that exposure and, in practice, commands a better price: buyers pay more for a dataset they can defend in their own due diligence, not less.

For an AI lab or enterprise deciding where to source training data, the category is a proxy for how much diligence work you'll do yourself. Data sourced through broker-style aggregation arrives with the provenance problems the Data Provenance Initiative documented at scale, license omission above 70% across widely used datasets, and it exposes the buyer to the same consent risk the seller took on. [4] Data sourced through a marketplace with NDA-gated sampling and a provenance file shifts that verification earlier in the process, before money changes hands, which is exactly the difference between a dataset that survives a labeling audit six months into a training run and one that doesn't. As public, freely licensed text data gets scarcer, Epoch AI's own projection puts the exhaustion window at 2026 to 2032, that verification step only grows more valuable, because more of what gets trained on will necessarily come from exactly this kind of intermediated, proprietary source. [6]

  • 1If you're selling: confirm whether your data, or your buyer's downstream use of it, meets a state's data broker definition before signing anything. A direct relationship with the people in the data and a documented lawful basis for the specific licensed use are what keep you out of that category.
  • 2If you're selling: favor a documented, NDA-gated marketplace transaction over an anonymous resale deal. It's usually both lower-risk and higher-priced, because exclusivity and provenance are worth paying for.
  • 3If you're buying: treat a seller's inability to produce a provenance and consent file as disqualifying, not negotiable, regardless of whether the seller calls itself a broker or a marketplace.
  • 4If you're buying: remember that inheriting a broker's consent gap is your legal exposure too. The buyer, not just the seller, answers for a model trained on data that was never lawfully collected in the first place.

Two different businesses that happen to sit in the same sentence

Data broker and AI training data marketplace both describe an intermediary between someone who holds data and someone who wants it, and that surface similarity is why the two get confused. Underneath, they're regulated by different statutes, built on different consent models, and paid on different economics. A data broker's entire business depends on being able to resell a stranger's profile to as many buyers as will pay. A training data marketplace's entire business depends on being able to prove, to one buyer at a time, exactly where a dataset came from and who agreed to what.

That difference isn't marketing. It shows up as a $6,000 California registration fee and a 45-day deletion window for one category, and a documented, negotiated Data Purchase Agreement for the other. [2] It shows up in FTC enforcement priorities, twice in the same enforcement cycle, aimed squarely at brokers with thin consent verification. [4] And it shows up in price: buyers pay more, not less, for data whose provenance they can actually defend. If you're deciding whether to list your company's data or where to source training data for your own model, the question worth asking isn't which word sounds better. It's which structure your actual transaction fits.

List data the marketplace way, not the broker way.

Dayda gates every listing behind provenance review, NDA-based sampling, and a negotiated license, so sellers and buyers both know exactly what they're getting into before anything changes hands.

List your data on Dayda