- 1There are six real channels for buying AI training data, and each fits a different buyer: managed marketplaces, labeling/annotation vendors, direct licensing with a data holder, public/open datasets, synthetic data vendors, and crowdsourcing platforms.
- 2Labeling vendors carry vendor-concentration risk alongside price risk: after Meta invested $14.3 billion for a 49% stake in Scale AI in June 2025, OpenAI and Google reportedly dropped Scale AI, and Meta's own researchers preferred rival vendors Surge and Mercor on quality grounds.
- 3Public and open datasets are effectively free (983,296 on the Hugging Face Hub, 363,694 on the U.S. government's data.gov as of 2026) but structurally unsuited to competitive advantage, since any competitor can pull the same data for free.
- 4Direct licensing produces the largest headline numbers ($250 million reported for OpenAI-News Corp, $203 million disclosed across Reddit's licensing book), but that scale is negotiated between platforms and a handful of well-capitalized labs, not sold at retail to most buyers.
- 5The do-it-yourself crowdsourcing channel is contracting: Amazon Mechanical Turk, long the default for self-run microtask labeling, is no longer accepting new customers as of 2026, pushing smaller buyers toward managed vendors, marketplaces, or newer platforms like Prolific.
Buyers looking for AI training data in 2026 have six real channels to choose from, not one generic marketplace. Managed marketplaces like Dayda broker vetted, NDA-gated proprietary datasets with provenance and legal review already done. Labeling and annotation vendors such as Scale AI, Surge AI, Mercor, and Toloka don't sell existing datasets; they run a workforce that produces or labels data to a buyer's spec. Direct licensing means negotiating straight with a publisher, platform, or company for its content, the channel behind headline deals like OpenAI-News Corp and Google-Reddit, but one mostly reserved for frontier labs with real leverage. Public and open datasets on the Hugging Face Hub, data.gov, and academic archives are free and instant but offer zero competitive edge. Synthetic data vendors like Mostly AI generate privacy-safe, model-produced data as a cheap complement to real data, never a full substitute. And crowdsourcing platforms let a buyer run their own labeling operation directly, a channel that shrank when Mechanical Turk stopped accepting new customers, leaving Prolific and Toloka as the more current options. The right channel depends on whether a buyer already holds raw data, how fast they need it, how much provenance certainty they require, and whether exclusivity actually matters for their use case.
On this page ▾
- The six channels where AI training data actually gets bought
- Managed, vetted marketplaces
- Data labeling and annotation vendors
- Direct licensing with a data holder
- Public and open datasets
- Synthetic data generation as a buy alternative
- Build-your-own via crowdsourcing platforms
- Choosing a channel: a worked example
- Match the channel to the constraint, not the other way around
The six channels where AI training data actually gets bought
Ask a buyer where to get AI training data and most answers collapse into "a marketplace" or "a vendor," as if those were the only options. In practice, buyers move through six distinct channels, and they are not interchangeable. Each has its own cost structure, timeline, provenance guarantees, and exclusivity options, and picking the wrong one can burn through months as fast as it burns through budget.
This guide is a shopping map, not a process guide. It assumes you already know what data you need (if you don't, see the companion buyer's process guide for how to define that first) and answers the next question: given that spec, which venue do you actually go to? [2]
| Channel | Typical cost | Speed | Quality/provenance control | Exclusivity | Best-fit buyer |
|---|---|---|---|---|---|
| Managed marketplace (e.g. Dayda) | $20K-$400K+ per dataset, scaled by volume and exclusivity | Weeks: NDA and sample workflow already built | High: provenance vetted before a listing goes live | Sold as a pricing tier, typically 3-5× non-exclusive | Needs proprietary, domain-specific data fast, without building legal/BD capacity |
| Labeling/annotation vendor | $10-15/hr commodity to $50-1,000+/hr expert | Weeks to months, scales with volume and task complexity | Depends on the vendor's own QA; you're vetting a process, not a finished asset | De facto exclusive: it's commissioned to your spec | Already holds raw data needing labels, or has a well-defined RLHF/eval task |
| Direct licensing | $60M/yr to $250M+ over multiple years for top-tier deals; smaller bespoke deals exist but rarely at retail | Months of legal negotiation | Strong if the counterparty is reputable, but no built-in vetting layer | Fully negotiable, native to the deal | Frontier labs or buyers with real leverage: money, distribution, or a product integration |
| Public/open datasets | Free | Immediate, self-serve | Highly variable: uncurated, license terms differ dataset to dataset | None. Open to every competitor by definition | Pretraining, benchmarking, prototyping. Not competitive differentiation |
| Synthetic data vendor | Near-zero compute cost plus a QA review layer | Days to weeks once a pipeline exists | No real-world provenance chain; must augment, not replace, real data | Private by default if self-generated; check vendor reuse terms | Filling edge cases or producing privacy-safe stand-ins for regulated data |
| Crowdsourcing platform (build-your-own) | Pay workers directly, generally below managed-vendor markup | Fast to start, slower to reach production-grade quality | You own quality control, consent, and payment tracking entirely | Fully exclusive: you own the output outright | Has in-house data-ops expertise and wants full control over the process |
Managed, vetted marketplaces
A managed marketplace, the category Dayda operates in, brokers proprietary datasets that companies already hold: support transcripts, clinical notes, marketplace transaction logs, legal filings. The marketplace's job is to do the vetting a buyer would otherwise have to do alone: verify provenance and consent before a listing goes live, structure an NDA-gated sample so a buyer can evaluate before committing, and match licensing terms (exclusive vs. non-exclusive, perpetual vs. time-limited) to the buyer's actual use case.
The advantage over a raw cold-outreach deal is that the legal and quality work is front-loaded. A dataset that clears a managed marketplace's listing bar has already answered the questions a buyer's counsel would otherwise spend weeks asking: where did this come from, under what consent, and is the chain of custody documented. That compresses a process that can take months of direct negotiation into a matter of weeks.
Data labeling and annotation vendors
Labeling and annotation vendors are a fundamentally different channel from a marketplace, even though buyers often lump them together. A vendor like Scale AI, Surge AI, Mercor, or Toloka doesn't sell you an existing dataset. It runs a workforce, crowdsourced, managed, or expert, that produces or labels data to your specification on demand. [10] You bring the raw material or the task definition; the vendor supplies the labor and the pipeline.
This channel carries a risk that's easy to underweight: vendor concentration and neutrality. In June 2025, Meta invested roughly $14.3 billion for a 49% stake in Scale AI at a $29 billion valuation, and Scale's CEO and co-founder Alexandr Wang left to join Meta directly. [2] Within months, reporting found that OpenAI and Google had reportedly stopped working with Scale AI following the investment, and that Meta's own internal AI researchers preferred working with Scale's competitors, Surge and Mercor, over Scale itself, on data-quality grounds. [4]
“TBD Labs is working with third-party data-labeling vendors other than Scale AI to train its upcoming AI models... Those third-party vendors include Mercor and Surge, two of Scale AI's largest competitors.”
The lesson here is about ownership and client concentration, not a verdict on Scale AI's work: a labeling vendor's investor relationships and client roster are live commercial risks a buyer should ask about before signing, the same way a buyer would ask a marketplace about provenance. A vendor backed by a buyer's competitor, or quietly losing its biggest customers to rivals, is material to a purchasing decision.
Direct licensing with a data holder
Direct licensing means going straight to whoever holds the content, a publisher, a platform, a company, and negotiating a deal with no intermediary. It's the channel behind the largest disclosed dollar figures in AI data. The Wall Street Journal reported OpenAI's multi-year content licensing deal with News Corp at roughly $250 million, covering archives from the Wall Street Journal, the New York Post, and other News Corp titles. [6] Reddit disclosed in regulatory filings an aggregate data-licensing contract value of $203.0 million as of January 2024, with a separate deal reported at roughly $60 million a year. [12]
“In January 2024, we entered into certain data licensing arrangements with an aggregate contract value of $203.0 million and terms ranging from two to three years.”
Those numbers are the ceiling of the channel, not the floor, but the ceiling is instructive. Deals at that scale get negotiated between platforms with genuinely unique content and a small number of buyers who can offer either a large check or real reciprocal value, referral traffic, product placement, or a strategic partnership. There's no public price list for licensing a major publisher's archive at retail terms, and cold outreach from a buyer without comparable leverage rarely produces a deal on a useful timeline.
Direct licensing still makes sense below that scale when a buyer has an existing relationship or a specific, hard-to-substitute need, a niche trade publication's archive, a regional platform's user data, a single company's operational logs. In those cases the negotiation is smaller but the mechanics are the same: it's slow, it's lawyer-heavy, and the buyer is doing all the vetting work a marketplace would otherwise front-load.
Public and open datasets
The public-data channel is the biggest by raw count and the cheapest by definition: free. The Hugging Face Hub hosts 983,296 datasets as of 2026, spanning translation, speech, vision, and text tasks, searchable and filterable by license, language, and task. [6] On the government side, the U.S.'s data.gov portal lists 363,694 datasets, spanning agencies, geography, health, and economics. [12] Academic corpora, benchmark sets, and repositories like Common Crawl round out the category.
The catch is structural, not incidental. Anything hosted on an open hub is, by definition, available to every competitor running the same search. Licenses on the Hugging Face Hub vary dataset to dataset, some are permissive, others carry non-commercial or attribution restrictions that quietly disqualify a dataset from a commercial fine-tune, and buyers routinely skip reading the license until legal review flags it late. [2] Public data is the right channel for pretraining, benchmarking, and early prototyping. It is structurally the wrong channel for building a defensible, differentiated model, because a competitor can replicate your data acquisition with the same free download.
Synthetic data generation as a buy alternative
When real data can't be sourced fast enough, isn't sensitive enough to license, or doesn't exist at the volume needed, synthetic data vendors are a genuine buy channel, not a workaround. MOSTLY AI is a representative example: it generates privacy-safe synthetic datasets designed to statistically mimic real data without exposing the underlying records, aimed at organizations that need to train or test models on sensitive data (financial records, healthcare data) without the legal exposure of moving the real thing. [4]
The economics of this channel are unusual: compute cost is nearly free (see the cost guide for the full math), but the buyer still pays in QA review, someone has to check that the synthetic output actually reflects real-world patterns rather than drifting into the generating model's own biases and blind spots. Synthetic vendors solve for volume and privacy, not for the ground-truth human signal a model needs to actually learn a new domain, which is why this channel works best as a supplement layered on top of real data rather than a replacement for it. The full decision framework for when synthetic is appropriate versus when it actively hurts a model lives in the companion synthetic-vs-real comparison.
Build-your-own via crowdsourcing platforms
The last channel is running the operation yourself: posting tasks directly to a crowdsourced workforce rather than paying a managed vendor's markup. This channel has visibly contracted. Amazon Mechanical Turk, the platform that effectively defined do-it-yourself microtask labeling for over a decade, now displays a banner on its own site stating it is no longer accepting new customers. [4] For a buyer without an existing MTurk account, that option is closed as of 2026.
Two platforms have absorbed the resulting demand, though they aren't drop-in replacements for MTurk's original model. Toloka now positions itself explicitly as an AI training data company rather than a generic crowd marketplace, combining human annotators with tooling aimed squarely at LLM, coding-assistant, and agent training data. [4] Prolific is confirmed active and onboarding new customers as of 2026, positioned around collecting high-quality data from a verified participant pool, a fit for preference tuning, safety evaluation, and research-style data collection rather than high-volume commodity labeling. [8]
Running a crowdsourcing platform directly only pays off when a buyer already has the data-ops muscle to build task instructions, quality-control sampling, and payment logistics in-house. Skip that step and a crowdsourcing platform quietly becomes more expensive than a managed vendor, because the buyer is now paying, in engineering and management time, for the QA layer a vendor would otherwise bundle into its price.
Choosing a channel: a worked example
Consider an early-stage legal-tech startup building a contract-clause classification model. The team has no internal corpus of contracts at meaningful scale, a hard deadline six weeks out tied to a pilot customer, and a seed-stage budget that rules out a seven-figure commitment. They evaluate two channels seriously.
- Direct licensing with a law firm or legal publisher. This is the channel that would, in principle, produce the highest-quality, most exclusive corpus: a firm's actual redlined contracts or a publisher's annotated clause library. In practice, the startup has no existing relationship, no meaningful check to write, and no reciprocal value to offer a counterparty this size. Legal review alone on a first-time deal with an unfamiliar counterparty routinely runs past the startup's entire timeline, before negotiation even starts.
- A managed marketplace listing of proprietary contract data. A vetted, NDA-gated dataset of clause-level contract text, already reviewed for provenance and consent, is available on a marketplace within the startup's price range and timeline. It isn't the startup's own future customers' contracts, so it won't be perfectly calibrated to their exact use case, but it clears the acceptance bar for training a first working model, and a non-exclusive license leaves room to license the same or a similar corpus to a competitor.
The startup buys the marketplace dataset non-exclusively, ships the pilot on time, and revisits direct licensing later, once it has revenue, a real relationship in the industry, and a case for exclusivity that justifies the multi-month negotiation. The general pattern holds beyond this example: direct licensing wins on ultimate quality and exclusivity but demands leverage and time few early buyers have; a managed marketplace trades a small amount of specificity for speed, cost, and a legal package the buyer didn't have to build alone.
Match the channel to the constraint, not the other way around
None of these six channels is universally correct. Each one solves for a different constraint: a managed marketplace solves for speed and provenance certainty, a labeling vendor solves for turning raw data you already hold into a labeled asset, direct licensing solves for exclusivity and quality at a price only a well-capitalized buyer can pay, public data solves for cost, synthetic vendors solve for privacy and volume, and crowdsourcing solves for control at the price of doing the QA work yourself.
The practical move is to name the binding constraint first, budget, timeline, exclusivity, or existing raw data, and let that dictate the channel, rather than defaulting to whichever channel is most familiar. The Scale-Meta episode is a useful reminder that even the most well-capitalized channel carries real, current risk: bigger doesn't automatically mean safer or higher quality. [2]
See the vetted, NDA-gated channel firsthand.
Dayda brokers proprietary datasets with provenance pre-checked, samples structured, and licensing terms matched to your use case, one credible channel among the six above. Tell us what you need to build.
See how buying works on Dayda→