AI Data Buying

Where to Buy AI Training Data: Marketplaces, Vendors, and Licensing

14 min read · 2026-06-26
983,296datasets hosted on the Hugging Face Hub
363,694datasets on the U.S. government's data.gov
$14.3BMeta's investment for a 49% stake in labeling vendor Scale AI
$250Mreported OpenAI-News Corp direct content licensing deal
3-5×typical exclusive vs. non-exclusive marketplace pricing premium
Key takeaways
  • 1There are six real channels for buying AI training data, and each fits a different buyer: managed marketplaces, labeling/annotation vendors, direct licensing with a data holder, public/open datasets, synthetic data vendors, and crowdsourcing platforms.
  • 2Labeling vendors carry vendor-concentration risk alongside price risk: after Meta invested $14.3 billion for a 49% stake in Scale AI in June 2025, OpenAI and Google reportedly dropped Scale AI, and Meta's own researchers preferred rival vendors Surge and Mercor on quality grounds.
  • 3Public and open datasets are effectively free (983,296 on the Hugging Face Hub, 363,694 on the U.S. government's data.gov as of 2026) but structurally unsuited to competitive advantage, since any competitor can pull the same data for free.
  • 4Direct licensing produces the largest headline numbers ($250 million reported for OpenAI-News Corp, $203 million disclosed across Reddit's licensing book), but that scale is negotiated between platforms and a handful of well-capitalized labs, not sold at retail to most buyers.
  • 5The do-it-yourself crowdsourcing channel is contracting: Amazon Mechanical Turk, long the default for self-run microtask labeling, is no longer accepting new customers as of 2026, pushing smaller buyers toward managed vendors, marketplaces, or newer platforms like Prolific.
The short version

Buyers looking for AI training data in 2026 have six real channels to choose from, not one generic marketplace. Managed marketplaces like Dayda broker vetted, NDA-gated proprietary datasets with provenance and legal review already done. Labeling and annotation vendors such as Scale AI, Surge AI, Mercor, and Toloka don't sell existing datasets; they run a workforce that produces or labels data to a buyer's spec. Direct licensing means negotiating straight with a publisher, platform, or company for its content, the channel behind headline deals like OpenAI-News Corp and Google-Reddit, but one mostly reserved for frontier labs with real leverage. Public and open datasets on the Hugging Face Hub, data.gov, and academic archives are free and instant but offer zero competitive edge. Synthetic data vendors like Mostly AI generate privacy-safe, model-produced data as a cheap complement to real data, never a full substitute. And crowdsourcing platforms let a buyer run their own labeling operation directly, a channel that shrank when Mechanical Turk stopped accepting new customers, leaving Prolific and Toloka as the more current options. The right channel depends on whether a buyer already holds raw data, how fast they need it, how much provenance certainty they require, and whether exclusivity actually matters for their use case.

On this page ▾

The six channels where AI training data actually gets bought

Ask a buyer where to get AI training data and most answers collapse into "a marketplace" or "a vendor," as if those were the only options. In practice, buyers move through six distinct channels, and they are not interchangeable. Each has its own cost structure, timeline, provenance guarantees, and exclusivity options, and picking the wrong one can burn through months as fast as it burns through budget.

This guide is a shopping map, not a process guide. It assumes you already know what data you need (if you don't, see the companion buyer's process guide for how to define that first) and answers the next question: given that spec, which venue do you actually go to? [2]

ChannelTypical costSpeedQuality/provenance controlExclusivityBest-fit buyer
Managed marketplace (e.g. Dayda)$20K-$400K+ per dataset, scaled by volume and exclusivityWeeks: NDA and sample workflow already builtHigh: provenance vetted before a listing goes liveSold as a pricing tier, typically 3-5× non-exclusiveNeeds proprietary, domain-specific data fast, without building legal/BD capacity
Labeling/annotation vendor$10-15/hr commodity to $50-1,000+/hr expertWeeks to months, scales with volume and task complexityDepends on the vendor's own QA; you're vetting a process, not a finished assetDe facto exclusive: it's commissioned to your specAlready holds raw data needing labels, or has a well-defined RLHF/eval task
Direct licensing$60M/yr to $250M+ over multiple years for top-tier deals; smaller bespoke deals exist but rarely at retailMonths of legal negotiationStrong if the counterparty is reputable, but no built-in vetting layerFully negotiable, native to the dealFrontier labs or buyers with real leverage: money, distribution, or a product integration
Public/open datasetsFreeImmediate, self-serveHighly variable: uncurated, license terms differ dataset to datasetNone. Open to every competitor by definitionPretraining, benchmarking, prototyping. Not competitive differentiation
Synthetic data vendorNear-zero compute cost plus a QA review layerDays to weeks once a pipeline existsNo real-world provenance chain; must augment, not replace, real dataPrivate by default if self-generated; check vendor reuse termsFilling edge cases or producing privacy-safe stand-ins for regulated data
Crowdsourcing platform (build-your-own)Pay workers directly, generally below managed-vendor markupFast to start, slower to reach production-grade qualityYou own quality control, consent, and payment tracking entirelyFully exclusive: you own the output outrightHas in-house data-ops expertise and wants full control over the process
For cost detail, see the dedicated pricing guide
The ranges above are directional. For real, cited dollar figures per channel, hourly labeling rates, disclosed deal sizes, marketplace price bands by data type, and a fully worked budget, see the companion cost guide. This piece stays focused on which venue fits which buyer, not the last dollar of pricing.

Managed, vetted marketplaces

A managed marketplace, the category Dayda operates in, brokers proprietary datasets that companies already hold: support transcripts, clinical notes, marketplace transaction logs, legal filings. The marketplace's job is to do the vetting a buyer would otherwise have to do alone: verify provenance and consent before a listing goes live, structure an NDA-gated sample so a buyer can evaluate before committing, and match licensing terms (exclusive vs. non-exclusive, perpetual vs. time-limited) to the buyer's actual use case.

The advantage over a raw cold-outreach deal is that the legal and quality work is front-loaded. A dataset that clears a managed marketplace's listing bar has already answered the questions a buyer's counsel would otherwise spend weeks asking: where did this come from, under what consent, and is the chain of custody documented. That compresses a process that can take months of direct negotiation into a matter of weeks.

What a managed marketplace can't do
A marketplace only lists what sellers bring to it. If your domain is thin on supply (a narrow regulatory niche, a small vertical), a marketplace may not have inventory yet, and you'll need to combine it with direct outreach or a labeling vendor for the gap.

Data labeling and annotation vendors

Labeling and annotation vendors are a fundamentally different channel from a marketplace, even though buyers often lump them together. A vendor like Scale AI, Surge AI, Mercor, or Toloka doesn't sell you an existing dataset. It runs a workforce, crowdsourced, managed, or expert, that produces or labels data to your specification on demand. [10] You bring the raw material or the task definition; the vendor supplies the labor and the pipeline.

This channel carries a risk that's easy to underweight: vendor concentration and neutrality. In June 2025, Meta invested roughly $14.3 billion for a 49% stake in Scale AI at a $29 billion valuation, and Scale's CEO and co-founder Alexandr Wang left to join Meta directly. [2] Within months, reporting found that OpenAI and Google had reportedly stopped working with Scale AI following the investment, and that Meta's own internal AI researchers preferred working with Scale's competitors, Surge and Mercor, over Scale itself, on data-quality grounds. [4]

TBD Labs is working with third-party data-labeling vendors other than Scale AI to train its upcoming AI models... Those third-party vendors include Mercor and Surge, two of Scale AI's largest competitors.
TechCrunch: Cracks are forming in Meta's partnership with Scale AI

The lesson here is about ownership and client concentration, not a verdict on Scale AI's work: a labeling vendor's investor relationships and client roster are live commercial risks a buyer should ask about before signing, the same way a buyer would ask a marketplace about provenance. A vendor backed by a buyer's competitor, or quietly losing its biggest customers to rivals, is material to a purchasing decision.

Vet the vendor relationship as closely as the price sheet
Before committing volume to a labeling vendor, ask who their largest investors and customers are, and whether any create a conflict with your competitive position. The Scale-Meta-OpenAI-Google episode shows this isn't hypothetical: it happened to some of the best-funded buyers in the industry within a few months of a major investment.

Direct licensing with a data holder

Direct licensing means going straight to whoever holds the content, a publisher, a platform, a company, and negotiating a deal with no intermediary. It's the channel behind the largest disclosed dollar figures in AI data. The Wall Street Journal reported OpenAI's multi-year content licensing deal with News Corp at roughly $250 million, covering archives from the Wall Street Journal, the New York Post, and other News Corp titles. [6] Reddit disclosed in regulatory filings an aggregate data-licensing contract value of $203.0 million as of January 2024, with a separate deal reported at roughly $60 million a year. [12]

In January 2024, we entered into certain data licensing arrangements with an aggregate contract value of $203.0 million and terms ranging from two to three years.
TechCrunch: Reddit Says It's Made $203M So Far Licensing Its Data

Those numbers are the ceiling of the channel, not the floor, but the ceiling is instructive. Deals at that scale get negotiated between platforms with genuinely unique content and a small number of buyers who can offer either a large check or real reciprocal value, referral traffic, product placement, or a strategic partnership. There's no public price list for licensing a major publisher's archive at retail terms, and cold outreach from a buyer without comparable leverage rarely produces a deal on a useful timeline.

Direct licensing still makes sense below that scale when a buyer has an existing relationship or a specific, hard-to-substitute need, a niche trade publication's archive, a regional platform's user data, a single company's operational logs. In those cases the negotiation is smaller but the mechanics are the same: it's slow, it's lawyer-heavy, and the buyer is doing all the vetting work a marketplace would otherwise front-load.

Public and open datasets

The public-data channel is the biggest by raw count and the cheapest by definition: free. The Hugging Face Hub hosts 983,296 datasets as of 2026, spanning translation, speech, vision, and text tasks, searchable and filterable by license, language, and task. [6] On the government side, the U.S.'s data.gov portal lists 363,694 datasets, spanning agencies, geography, health, and economics. [12] Academic corpora, benchmark sets, and repositories like Common Crawl round out the category.

The catch is structural, not incidental. Anything hosted on an open hub is, by definition, available to every competitor running the same search. Licenses on the Hugging Face Hub vary dataset to dataset, some are permissive, others carry non-commercial or attribution restrictions that quietly disqualify a dataset from a commercial fine-tune, and buyers routinely skip reading the license until legal review flags it late. [2] Public data is the right channel for pretraining, benchmarking, and early prototyping. It is structurally the wrong channel for building a defensible, differentiated model, because a competitor can replicate your data acquisition with the same free download.

Public data answers a different question
"Where do I get data cheaply" and "where do I get data that gives me an edge" are different questions with different answers. Public and open datasets answer the first one well. For the full taxonomy of where training data originates, public web, proprietary, user-generated, licensed, synthetic, and brokered, see the companion supply-chain map.

Synthetic data generation as a buy alternative

When real data can't be sourced fast enough, isn't sensitive enough to license, or doesn't exist at the volume needed, synthetic data vendors are a genuine buy channel, not a workaround. MOSTLY AI is a representative example: it generates privacy-safe synthetic datasets designed to statistically mimic real data without exposing the underlying records, aimed at organizations that need to train or test models on sensitive data (financial records, healthcare data) without the legal exposure of moving the real thing. [4]

The economics of this channel are unusual: compute cost is nearly free (see the cost guide for the full math), but the buyer still pays in QA review, someone has to check that the synthetic output actually reflects real-world patterns rather than drifting into the generating model's own biases and blind spots. Synthetic vendors solve for volume and privacy, not for the ground-truth human signal a model needs to actually learn a new domain, which is why this channel works best as a supplement layered on top of real data rather than a replacement for it. The full decision framework for when synthetic is appropriate versus when it actively hurts a model lives in the companion synthetic-vs-real comparison.

Build-your-own via crowdsourcing platforms

The last channel is running the operation yourself: posting tasks directly to a crowdsourced workforce rather than paying a managed vendor's markup. This channel has visibly contracted. Amazon Mechanical Turk, the platform that effectively defined do-it-yourself microtask labeling for over a decade, now displays a banner on its own site stating it is no longer accepting new customers. [4] For a buyer without an existing MTurk account, that option is closed as of 2026.

Two platforms have absorbed the resulting demand, though they aren't drop-in replacements for MTurk's original model. Toloka now positions itself explicitly as an AI training data company rather than a generic crowd marketplace, combining human annotators with tooling aimed squarely at LLM, coding-assistant, and agent training data. [4] Prolific is confirmed active and onboarding new customers as of 2026, positioned around collecting high-quality data from a verified participant pool, a fit for preference tuning, safety evaluation, and research-style data collection rather than high-volume commodity labeling. [8]

Running a crowdsourcing platform directly only pays off when a buyer already has the data-ops muscle to build task instructions, quality-control sampling, and payment logistics in-house. Skip that step and a crowdsourcing platform quietly becomes more expensive than a managed vendor, because the buyer is now paying, in engineering and management time, for the QA layer a vendor would otherwise bundle into its price.

Choosing a channel: a worked example

Consider an early-stage legal-tech startup building a contract-clause classification model. The team has no internal corpus of contracts at meaningful scale, a hard deadline six weeks out tied to a pilot customer, and a seed-stage budget that rules out a seven-figure commitment. They evaluate two channels seriously.

  • Direct licensing with a law firm or legal publisher. This is the channel that would, in principle, produce the highest-quality, most exclusive corpus: a firm's actual redlined contracts or a publisher's annotated clause library. In practice, the startup has no existing relationship, no meaningful check to write, and no reciprocal value to offer a counterparty this size. Legal review alone on a first-time deal with an unfamiliar counterparty routinely runs past the startup's entire timeline, before negotiation even starts.
  • A managed marketplace listing of proprietary contract data. A vetted, NDA-gated dataset of clause-level contract text, already reviewed for provenance and consent, is available on a marketplace within the startup's price range and timeline. It isn't the startup's own future customers' contracts, so it won't be perfectly calibrated to their exact use case, but it clears the acceptance bar for training a first working model, and a non-exclusive license leaves room to license the same or a similar corpus to a competitor.

The startup buys the marketplace dataset non-exclusively, ships the pilot on time, and revisits direct licensing later, once it has revenue, a real relationship in the industry, and a case for exclusivity that justifies the multi-month negotiation. The general pattern holds beyond this example: direct licensing wins on ultimate quality and exclusivity but demands leverage and time few early buyers have; a managed marketplace trades a small amount of specificity for speed, cost, and a legal package the buyer didn't have to build alone.

Match the channel to the constraint, not the other way around

None of these six channels is universally correct. Each one solves for a different constraint: a managed marketplace solves for speed and provenance certainty, a labeling vendor solves for turning raw data you already hold into a labeled asset, direct licensing solves for exclusivity and quality at a price only a well-capitalized buyer can pay, public data solves for cost, synthetic vendors solve for privacy and volume, and crowdsourcing solves for control at the price of doing the QA work yourself.

The practical move is to name the binding constraint first, budget, timeline, exclusivity, or existing raw data, and let that dictate the channel, rather than defaulting to whichever channel is most familiar. The Scale-Meta episode is a useful reminder that even the most well-capitalized channel carries real, current risk: bigger doesn't automatically mean safer or higher quality. [2]

See the vetted, NDA-gated channel firsthand.

Dayda brokers proprietary datasets with provenance pre-checked, samples structured, and licensing terms matched to your use case, one credible channel among the six above. Tell us what you need to build.

See how buying works on Dayda