AI Data Buying

How Much Does AI Training Data Cost? A Buyer's Pricing Guide

15 min read · 2026-08-11
$1.32-$12.50/hrspread between offshore labeler pay and vendor bill rate on a documented OpenAI data-labeling contract
$105/hr avgaverage pay for expert contractors doing domain-specific RLHF labeling, up to $1,000/hr for niche specialists
$250Mreported value of the News Corp-OpenAI content licensing deal over 5 years
$18.63B → $57.63Bglobal data labeling and annotation market size, 2024 vs. projected 2030
<$0.01compute cost per AI-generated synthetic sample, vs. $5-20 per human-labeled preference pair
Key takeaways
  • 1Human annotation costs span three orders of magnitude: commodity offshore labeling bills around $10-15 an hour, with the workers themselves paid $1.32-$2 an hour on a well-documented OpenAI vendor contract, while expert contractors doing RLHF-style labeling average about $105 an hour and top out near $1,000 an hour for niche domain specialists.
  • 2Publisher and platform licensing deals for AI training content run from roughly $60 million a year for a single major platform's data to $250 million over five years for a top-tier news archive, a scale mostly reserved for frontier labs, not individual buyers.
  • 3Buying an already-collected, proprietary dataset through a marketplace is usually the cheapest path to domain-specific volume: a 500K-1M conversation support dataset benchmarks around $20K-$80K non-exclusive and $60K-$250K exclusive, well below the cost of commissioning the same collection and labeling from scratch.
  • 4Synthetic data generation is nearly free at the compute layer, well under $1,000 in API cost to generate 500,000 synthetic conversations at current LLM pricing, but the real cost hides in the human quality review needed to keep it from drifting away from real user behavior.
  • 5In-house collection is rarely the cheapest option once engineering time is priced honestly: two data engineers building a collection-and-labeling pipeline for four months can run past $120K in loaded salary before a single item is labeled, on top of a multi-month timeline a marketplace purchase would skip entirely.
The short version

The honest answer to how much AI training data costs depends entirely on which of five channels a buyer uses, and the range spans four orders of magnitude for the same underlying need. Human annotation prices data by the hour or the label: commodity labeling bills around $10-15 an hour, while expert RLHF-style review runs $50 to over $1,000 an hour. Publisher and platform licensing deals price data by the archive, with real deals from about $60 million a year to $250 million over five years, a scale that mostly serves frontier labs. Marketplace and broker purchases of proprietary, already-collected datasets price by volume, annotation depth, and exclusivity, typically landing in the tens to low hundreds of thousands of dollars for a domain-specific corpus. Synthetic data generation prices by the token and is nearly free at the compute layer, though quality review adds real cost back in. In-house collection prices by engineering headcount and calendar time, and is routinely the most expensive channel once loaded salary and lead time are counted honestly. This guide gives real, cited numbers for every channel, compares them head to head, and works through a full budget for one concrete scenario: sourcing 500,000 labeled support conversations for fine-tuning.

On this page ▾

Why there's no single answer to this question

Ask five vendors what AI training data costs and you'll get five different numbers, because they're each quoting a different channel, not a different price for the same thing. A buyer can pay a labeling company by the hour, license content from a publisher by the archive, buy an already-collected proprietary dataset from a broker, generate synthetic examples for pennies of compute, or build a collection pipeline in-house and pay in engineering time. These are genuinely different products with genuinely different economics, and conflating them is why so much advice on this topic is useless.

This guide prices out all five channels with real, dated numbers: human annotation and labeling, licensing deals with publishers and platforms, marketplace and broker purchases of proprietary data, synthetic data generation, and in-house collection. It closes with a full worked budget for one concrete scenario, so the numbers aren't abstract.

This is the cost guide, not the process guide
This piece answers what you'll pay. For the step-by-step process of sourcing, vetting, sampling, and negotiating a purchase, see the companion buyer's guide. The two are meant to be read together.

Human annotation and labeling costs

Human labeling is priced almost entirely on two axes: how much judgment the task requires, and where in the world the labeler sits. Those two variables alone create a cost spread of roughly 1,000x between the cheapest and most expensive labeling a buyer can commission.

Commodity labeling: classification, tagging, simple annotation

At the low end sits commodity labeling: sentiment tags, entity spans, image bounding boxes, simple classification. This work is typically outsourced through business-process outsourcing (BPO) firms and labeling platforms such as Appen and Sama, often to lower-cost labor markets. TIME's 2023 investigation into an OpenAI-Sama content-moderation and labeling contract is one of the few instances where the full economics of a real deal became public: Sama's Kenya-based workers took home between $1.32 and $2 per hour after tax, while Sama itself billed OpenAI $12.50 per hour for the work, a 6-9x markup between what the labeler earned and what the buyer paid. [6]

[Agents] take home a total of at least $1.32 per hour after tax, rising to as high as $1.44 per hour if they exceeded all their targets.
TIME: Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic

That $12.50/hour vendor bill rate is a useful anchor for what a buyer actually pays for commodity labeling through a BPO-model vendor, separate from what the vendor pays its workers. Rates vary by region, task complexity, and vendor margin, but a $10-15/hour billed range for straightforward classification and tagging work is a reasonable planning baseline as of 2026.

Enterprise self-serve and managed labeling platforms

Scale AI is the best-known managed labeling platform, and it illustrates how enterprise labeling pricing actually works: there is no public per-label rate card. Scale's self-serve tier lets a buyer bring their own labelers and pay a small per-unit fee after a free allotment, but its core enterprise business is entirely custom-quoted. Procurement-data aggregator Vendr reports an average Scale AI contract value of $93,158 across its buyer base, with deals ranging up to $400,000. [4] That figure describes mid-market commercial buyers, not the largest labs: CNBC has reported that Google alone spent roughly $150 million on Scale AI's human-labeling services in 2024, with plans to spend closer to $200 million the following year, before moving some of that volume to competitors following Meta's investment in Scale. [6]

The takeaway on labeling vendors
A buyer's actual bill depends entirely on scale and negotiating leverage. A startup running a few thousand labeling units a month pays close to card-rate self-serve pricing; a hyperscaler running a frontier training run negotiates into eight or nine figures a year. There is no single "market rate," only a curve.

Expert and RLHF-style labeling costs far more

The moment a labeling task requires judgment rather than classification, the cost curve changes entirely. RLHF (reinforcement learning from human feedback) and other preference-based labeling require a human to compare two or more model outputs and rank them, often against nuanced criteria like helpfulness, factual accuracy, or tone. Nathan Lambert, a researcher who has written extensively on post-training economics, estimates human preference-data labeling at roughly $5 to $20 per comparison, against under $0.01 per sample for AI-generated feedback used as a cheaper substitute. [4]

Domain-expert labeling pushes the hourly rate even higher. Startups like Mercor, Surge AI, and Alignerr now recruit doctors, lawyers, coders, and PhDs specifically to generate and label training data in their area of expertise. Reporting on this market found Mercor's average contractor pay at $105 an hour, rising to $350 an hour for specialists such as psychiatrists; Surge AI advertising rates up to $1,000 an hour for venture capital partners and startup executives; and Alignerr paying up to $200 an hour for chemistry PhDs. [14] Mercor alone reportedly pays out more than $4 million a day across roughly 30,000 contractors. [16]

Labeling typeWho does itTypical rateSource
Commodity classification/taggingOffshore BPO vendors (billed rate)$10-15/hourTIME
General crowdsourced labelingManaged platforms, self-serve tierSmall per-unit fee after free tierVendor pricing pages
RLHF preference comparisonTrained annotators via labeling platforms$5-20 per comparisonInterconnects
Domain-expert RLHF/evalDoctors, lawyers, engineers via expert networks$50-350/hour (up to $1,000 for niche roles)Yahoo Finance / Moneywise

The overall labeling industry has grown to match this demand. Grand View Research estimates the global data labeling solutions and services market at $18.63 billion in 2024, projecting it to reach $57.63 billion by 2030, a 20.3% compound annual growth rate. [6] That growth is a demand signal a buyer should read literally: the labor supply for high-quality labeling is not infinite, and prices for skilled annotation have been rising, not falling, even as commodity labeling gets cheaper through automation and lower-cost geographies.

Licensing deals with publishers and platforms

The second channel is direct licensing: paying a content owner, a publisher, a social platform, an image library, for the right to train on their material. This is the channel that produces the largest disclosed dollar figures in AI data, because it's the channel frontier labs use to acquire content at web scale with a clean legal chain.

News Corp's multi-year licensing deal with OpenAI, announced in May 2024, is reported by the Wall Street Journal to be worth more than $250 million over five years, covering archives from the Wall Street Journal, the New York Post, the Times of London, and other News Corp titles. [6] On the platform side, Reddit disclosed in its 2024 IPO prospectus that it had entered data licensing arrangements with an aggregate contract value of $203.0 million, and separate reporting put a single deal, widely understood to be with Google, at roughly $60 million a year. [14]

In January 2024, we entered into certain data licensing arrangements with an aggregate contract value of $203.0 million and terms ranging from two to three years.
TechCrunch: Reddit Says It's Made $203M So Far Licensing Its Data

Read together, these disclosed deals span roughly $60 million a year for a single platform's ongoing feed up to $250 million for a multi-year archive of premium, brand-name journalism. That's the real order of magnitude for this channel, and it's a scale built for the handful of AI labs training frontier foundation models, not for a startup or mid-market enterprise sourcing data for a fine-tuning project.

This channel usually isn't available to you
Publisher and platform licensing deals are negotiated directly between the content owner and a small number of well-capitalized buyers. There is no open marketplace for licensing the New York Times archive or Reddit's comment corpus at retail terms. If your use case needs that class of content, expect a bespoke, lawyer-heavy negotiation, not a price list.

Marketplace and broker purchases of proprietary data

The third channel is where most buyers who don't have publisher-scale budgets or in-house labeling operations actually land: buying a proprietary, already-collected dataset through a broker or marketplace. This is data produced inside a company, support transcripts, clinical notes, legal matter files, that was never public and typically comes pre-annotated to some degree, sold under a negotiated license rather than a per-hour labor contract.

This channel prices very differently from labeling and licensing, because the buyer is paying for a finished asset, not for the labor to produce one. The benchmark ranges below reflect the pricing logic observed across marketplace transactions for these data types, framed here from the buyer's side of the same transaction sellers see priced in a valuation guide.

Data typeTypical volumeNon-exclusive priceExclusive price
Customer support transcripts500K-1M conversations$20K-$80K$60K-$250K
Domain text corpus (legal/finance)1M-5M documents$50K-$200K$150K-$600K
RLHF / preference data100K-1M pairs$30K-$150K$90K-$450K
Clinical / de-identified notes1M-4M notes$100K-$400K$300K-$1.2M
Behavioral & clickstream10M+ events$15K-$60K$45K-$180K
Code review / dev tooling1M-10M comments$40K-$130K$120K-$400K

The gap between the two right-hand columns is licensing scope, not data quality: exclusive rights, where the buyer is the only licensee, typically command a 3-5x premium over a comparable non-exclusive license, because the buyer is paying to remove competitors from access to the same training signal, not for a materially different dataset.

Why this channel is usually the best value
For a given volume and domain, buying a finished, already-annotated proprietary dataset almost always costs less than paying to collect and label the equivalent volume from zero, because the seller has already absorbed the collection cost and often part of the annotation cost. The worked example later in this guide shows the gap concretely.

Synthetic data generation cost

The fourth channel doesn't involve buying data from anyone: generating it with a large language model. This is the cheapest channel by a wide margin at the compute layer, and current API pricing makes the math easy to run in full.

Using Anthropic's published Claude API pricing as a reference point, a smaller model like Claude Haiku 4.5 costs $0.50 per million input tokens and $2.50 per million output tokens on the batch tier. [4] Generating a synthetic support conversation of roughly 150 tokens of prompt context and 450 tokens of output, across 500,000 conversations, works out to about 75 million input tokens ($37.50) and 225 million output tokens ($562.50): roughly $600 in raw compute for half a million synthetic conversations.

That number is the reason synthetic data gets so much attention in cost conversations, and it's also the reason the number is misleading on its own. Raw generation cost is not the same as usable-dataset cost. Someone still has to write the seed prompts and personas, review a meaningful sample for realism and safety, de-duplicate near-identical outputs, and check for the repetitive, homogenized patterns that model-generated text tends to produce at scale. That review layer, done by a human at even modest hourly rates, typically adds several thousand to a few tens of thousands of dollars back on top of a sub-$1,000 compute bill for a dataset this size.

Synthetic augments, it doesn't replace
Treat synthetic data as a volume and edge-case multiplier layered on top of a real, human-sourced seed set, not as a standalone substitute for it. The cost advantage is real, but it applies to generation, not to the quality assurance that makes synthetic output trustworthy.

In-house data collection cost

The fifth channel is building the pipeline yourself: instrumenting your product to capture the raw signal, building the tooling to route it to labelers or reviewers, and running the ongoing operations to keep it flowing. Buyers frequently underprice this channel because they only count the labeling cost and forget the engineering and management time required to get to labelable data in the first place.

The U.S. Bureau of Labor Statistics reports a median annual wage of $133,080 for software developers as of May 2024. [4] Loaded for benefits, taxes, and overhead at a typical 1.4-1.5x multiplier, that puts a fully-loaded data engineer at roughly $185K-$200K a year, or about $15,500-$16,700 a month. Two data engineers spending four months building a collection pipeline, event schema, storage, and a labeling queue integration, run $124K-$134K in engineering time alone, before a single record is labeled and before counting the product manager, data scientist, or annotation-ops time layered on top.

The other cost of this channel is calendar time, and it's often the more binding constraint. Building the pipeline is only the first step; a buyer then has to wait for the pipeline to actually generate the target volume from real usage, which can take months if the underlying activity doesn't already occur at scale. A marketplace purchase or a labeling vendor contract can close in weeks. An in-house build routinely takes a quarter or more before it produces its first usable batch.

Comparing all five channels

Putting the five channels side by side makes the tradeoffs concrete. None of them is universally cheapest; each wins under different conditions.

ChannelCost driverTypical cost rangeWhen it makes sense
Human annotation (commodity)Per-hour vendor billing$10-15/hour billed; low-cents-to-low-dollars per labelYou already hold raw data and need it classified or tagged at volume
Human annotation (expert/RLHF)Per-hour or per-comparison expert rate$5-20/comparison; $50-1,000+/hourAlignment, safety-critical, or highly technical judgment tasks
Publisher/platform licensingNegotiated content-rights fee~$60M/year to $250M+ over multiple yearsFrontier-scale training runs needing broad, freshly licensed content
Marketplace/broker purchaseVolume, annotation depth, exclusivity$20K-$400K+ per datasetDomain-specific, already-collected proprietary data, needed fast
Synthetic generationLLM API tokens plus QA reviewHundreds to low thousands of $ in compute, plus review costAugmenting real data and covering edge cases cheaply
In-house collectionEngineering and ops headcount$120K-$300K+ in loaded labor, over monthsYou already generate the raw signal and need durable pipeline ownership

Worked example: budgeting 500,000 labeled support conversations

Take a concrete scenario: a startup needs 500,000 labeled support conversations to fine-tune a customer-support model, labeled for intent, resolution status, and quality. Here's what each channel would actually cost to hit that target.

  • 1Commodity BPO labeling. Assumes the startup already holds 500,000 raw conversations and just needs them labeled. At a vendor-billed rate of $10-15/hour and a throughput of roughly 40 conversations labeled per hour for a straightforward schema, cost per conversation lands around $0.25-$0.38. Total: roughly $100K-$180K.
  • 2Expert/RLHF-style labeling. If the task requires nuanced quality judgment rather than simple tagging, expert rates of $50-100/hour and slower throughput (12-15 conversations/hour for careful review) push cost per conversation to $3-8. Total: roughly $1.5M-$4M, which is why buyers scope expert review to a curated subset rather than the full volume.
  • 3Publisher/platform licensing. Not applicable here. There is no open market for licensing someone else's private customer support logs the way there is for news archives or social platform content; this channel simply doesn't fit this data type.
  • 4Marketplace/broker purchase. A comparable, already-collected 500K-1M conversation dataset benchmarks at $20K-$80K non-exclusive and $60K-$250K exclusive. Total: roughly $20K-$250K, a fraction of the cost of commissioning the same volume fresh.
  • 5Synthetic generation. Roughly $600 in raw LLM API compute to generate 500,000 synthetic conversations, plus a QA and review layer. Total: roughly $5K-$50K, depending on how rigorously the output is validated against real conversation patterns.
  • 6In-house collection. Two data engineers building the pipeline for about four months run $120K-$135K in loaded engineering time, plus the labeling cost from item 1 or 2 above once conversations are flowing, plus months of calendar time to actually accumulate the volume. Total: roughly $220K-$300K+, before counting lead time.
ChannelAssumptionsEstimated costTime to usable data
Commodity BPO labelingRaw conversations already exist; $10-15/hr; ~40/hr throughput$100K-$180K4-8 weeks
Expert/RLHF-style labeling$50-100/hr; ~12-15/hr throughput for nuanced review$1.5M-$4M+3-4 months
Publisher/platform licensingNo market fit for this data typeNot applicableNot applicable
Marketplace/broker purchase500K-1M conversation dataset, non-exclusive to exclusive$20K-$250K2-6 weeks
Synthetic generationLLM API compute plus human QA review$5K-$50K1-3 weeks
In-house build2 data engineers, ~4 months, plus labeling on top$220K-$300K+4-6+ months
The lesson from this example
For the same target dataset, the honest cost spread runs from a few thousand dollars (synthetic, quality caveats aside) to several million (full expert labeling of the whole volume). Most real buyers land on a blend: a marketplace purchase or commodity labeling pass for the bulk of the volume, synthetic generation to cover edge cases cheaply, and expert review reserved for a curated, high-stakes subset.

Building your own budget

The channels above aren't mutually exclusive, and the buyers who get the best outcome rarely pick just one. A realistic budget usually blends a marketplace purchase for the core proprietary volume, synthetic generation to pad edge cases at near-zero marginal cost, and a smaller expert-labeling pass reserved for the highest-stakes examples where judgment quality matters most. In-house collection earns its cost only when a buyer already owns the raw signal at scale and needs long-term pipeline ownership, not a one-time dataset.

Before committing a budget to any single channel, price the alternatives against the same target volume, the way the worked example above does. The gap between the cheapest and most expensive route to the same dataset is routinely an order of magnitude or more, and that gap is exactly where a careful buyer earns back the time spent building the comparison.

Compare the marketplace price before you commission anything.

Dayda brokers vetted, NDA-gated proprietary datasets across the data types in this guide. Get a real quote for the volume and domain you need before you commit to a labeling contract or an in-house build.

See how buying works on Dayda