- 1Labeling price spans nearly two orders of magnitude by workforce alone: commodity outsourced labeling bills around $10-15 an hour, while paid domain experts doing nuanced annotation average $105 an hour and reach $300-$1,000 an hour for specialist judgment. {cite:time}{cite:mercor}
- 2Method changes the price of the identical task more than task type does: Rev's own published rates show AI-only transcription at $0.25 per audio-minute against $1.99 per minute for human transcription of the same audio, an 8x gap for the same unit of work. {cite:rev}
- 3Annotation depth is a direct cost multiplier, not a quality add-on: measuring inter-annotator agreement requires two or more independent labelers on the same item, so a double-checked dataset costs close to double the single-pass rate before any QA overhead is added. {cite:labelstudio}
- 4Crowdsourcing platforms take a real, disclosed cut on top of whatever a buyer pays workers: Amazon Mechanical Turk charges 20% (40% on HITs with 10+ assignments) {cite:aws-mturk}, and Prolific charges 42.8% for corporate accounts {cite:prolific}, and Mechanical Turk itself stopped accepting new customers on July 30, 2026, pushing new crowdsourcing buyers toward its remaining competitors. {cite:techcrunch-mturk}
- 5Scale is the biggest lever on unit price: Scale AI's average enterprise labeling contract runs $93,158 and ranges up to $400,000, pricing that only pencils out at sustained volume, not for a one-off batch a startup needs labeled once. {cite:vendr}
Data labeling costs whatever a specific combination of task complexity, workforce, annotation depth, turnaround, and volume adds up to, and that combination routinely swings the price for the same physical unit of work by 10x or more. Commodity classification and bounding boxes labeled by an outsourced BPO vendor bill around $10-15 an hour, translating to a few cents per image or text row at realistic throughput. Domain-expert labeling, the kind clinical, legal, or technical data actually needs, runs $105 to over $1,000 an hour depending on the specialty. Audio labeling prices by the minute, and the gap between an AI-only pass and a human transcript for the identical file is a documented 8x. RLHF-style preference labeling prices by the comparison, at $5-20 for a general rater and far more once a paid domain expert is doing the judging. Method matters as much as task: an in-house team is cheapest only at sustained volume, an outsourced vendor trades speed for a contract minimum, a crowdsourcing platform is the fastest and least predictable on quality, and an AI-assisted, human-in-the-loop pipeline is usually the cheapest option once a real correction layer is built around it. This guide gives real, cited numbers for all four methods across every major data type, a comparison table, and a fully worked budget comparing two methods on the same job.
On this page ▾
- Why data labeling doesn't have one price
- The five variables that set your price per label
- Real price ranges per unit, by data type
- How the labeling method changes the price of the identical task
- Worked example: budgeting to transcribe and label 1,000 hours of calls two ways
- Comparing the four methods by cost, speed, quality control, and fit
- How to build a realistic labeling budget
- Price the unit economics, not the sticker rate
Why data labeling doesn't have one price
Ask a labeling vendor for a quote and the honest first reply is another question: what task, whose workforce, how many independent passes, and how fast do you need it back. Data labeling is priced by the unit of work, an image, a bounding box, a text row, an audio-minute, a conversation, and the same physical unit can cost ten times more or less depending on who labels it and how carefully. That's the gap this guide fills: real, dated numbers for what the labeling service itself costs, broken out by data type and by the four ways teams actually staff it.
The market this guide prices out is not small. Grand View Research estimates the global data labeling solution and services market at $18.63 billion in 2024, projecting it to reach $57.63 billion by 2030, a 20.3% compound annual growth rate, with outsourced vendors already handling 84.6% of that spend rather than in-house teams. [8] Every dollar in that market is being priced against the same five variables, which is where this guide starts.
The five variables that set your price per label
No vendor rate card captures all of this in one number, but every real quote is a function of the same five inputs. Understand these and a quote stops looking arbitrary.
- Task complexity. A single classification tag takes seconds. A pixel-level segmentation mask, a full transcript, or a read-and-rank preference judgment takes much longer for the same underlying item, and price follows time, not item count.
- Required domain expertise. Generalist labelers can tag sentiment or draw a box around an obvious object. They can't reliably flag a drug interaction in a clinical note or spot a misapplied contract clause, so specialized tasks route to a much smaller, much more expensive labor pool. [4]
- Annotation depth and multi-pass QA. Serious labeling operations don't trust a single annotator's judgment. Measuring inter-annotator agreement, the standard quality check, requires two or more independent people to label the same item before agreement can even be scored, so a double-checked dataset costs close to double the single-pass rate before any dedicated QA reviewer time is added on top. [4]
- Turnaround time. Faster delivery means either more parallel labor or a faster (often less accurate) method, and both cost more. Rev's own pricing makes this concrete: the AI-only tier delivers in under 5 minutes at $0.25/minute, while the human tier is quoted at 12 hours or less at $1.99/minute for the identical audio file. [4]
- Volume and scale. Enterprise labeling contracts are custom-quoted against sustained volume, not a published per-item card rate. Vendr's aggregated buyer data puts the average Scale AI contract at $93,158, ranging up to $400,000, pricing built around ongoing, high-volume relationships rather than a single batch. [8]
Real price ranges per unit, by data type
The table below converts two verified hourly and per-minute rates into per-unit prices: a $12.50/hour commodity BPO billed rate, the same figure a 2023 TIME investigation documented for an OpenAI-Sama annotation contract [4], and real domain-expert hourly rates reported for Mercor, Surge AI, and Alignerr. [6] Where a task type has no direct market throughput data, the per-unit figure is explicitly modeled from the cited hourly rate using a stated, reasonable planning assumption, the same technique used elsewhere on this site to turn a real wage into a real unit cost. Audio and RLHF preference prices are direct, published figures, not modeled.
| Data type | Unit | Commodity/BPO price (modeled) | Domain-expert price (modeled) | Basis |
|---|---|---|---|---|
| Image classification (single tag) | per image | $0.03-$0.05 | $0.25-$0.45 | $12.50/hr at ~300-400/hr vs. $105/hr at ~230-280/hr [2][4] |
| Bounding box | per object | $0.21-$0.31 | $1.50-$2.50 | $12.50/hr at ~40-60 objects/hr vs. expert rate at similar pace [2] |
| Semantic segmentation mask | per mask | $0.50-$0.90 | $4-$7 | $12.50/hr at ~14-25 masks/hr [2] |
| Text row / short document classification | per row | $0.03-$0.06 | $0.35-$0.60 | $12.50/hr at ~200-350 rows/hr vs. $105/hr at similar pace [2] |
| Named-entity / span tagging | per document | $0.15-$0.25 | $1-$2 | $12.50/hr at ~50-80 documents/hr [2] |
| Domain-expert document review (legal, clinical, technical) | per document | not applicable | $50-$120 | $200-$350/hr at ~3-4 documents/hr [2] |
| Audio transcription | per minute | $0.25 (AI) / $1.99 (human) | custom-quoted for specialist terminology | Published Rev rates [2] |
| Conversation / preference pair (RLHF) | per comparison | $5-20 | $21-$167 | Interconnects estimate [2]; upper bound modeled from $105-$1,000/hr expert rates [4] |
The multiplier between columns matters more than any single number in this table: expert labeling runs roughly 7-10x the commodity rate for the same task type, and the gap between a $5 RLHF comparison and a $167 expert one comes almost entirely from who's doing the judging, not from what the task technically requires on paper.
How the labeling method changes the price of the identical task
There are four real ways to get labeling done, and each prices the same task differently before task complexity even enters the picture.
In-house team
An in-house labeler is a fixed cost regardless of volume. Data USA's BLS-derived figures put the average annual wage for data entry keyers, a reasonable proxy for a commodity in-house labeling role, at $38,531 in 2024, or roughly $18.53/hour unloaded. [6] Loaded for benefits and overhead at a typical 1.3-1.5x multiplier, that's closer to $24-$28/hour in real cost, before recruiting, management, and tooling. This only beats an outsourced vendor's per-hour bill rate once volume is high and sustained enough to keep a permanent team busy; on a one-time batch, the fixed cost of hiring and idle time between projects usually erases the hourly-rate advantage entirely.
Outsourced BPO vendor
A BPO vendor bills a blended rate, commonly the $10-15/hour range documented in the TIME investigation [2], and layers in a trained team, existing QA tooling, and management overhead the buyer doesn't have to build. Enterprise contracts with managed platforms like Scale AI custom-quote against volume and average $93,158, ranging up to $400,000. [8] This is the channel that dominates the market by spend, at 84.6% of the total. [10] It wins on predictable throughput and existing infrastructure, and it usually carries a contract minimum that makes it a poor fit for a small, one-off job.
Crowdsourcing platforms
Crowdsourcing platforms let a buyer post small tasks to a distributed, on-demand worker pool and pay per completed unit, with no team to hire and no contract minimum. The platform itself takes a real, disclosed cut on top of whatever the buyer pays workers. Amazon Mechanical Turk charges a 20% fee on worker rewards, rising to 40% total on tasks run with 10 or more assignments (a common setup for consensus labeling), plus a $0.01 minimum fee per assignment and a 5% surcharge for restricting work to its most experienced "Masters" workers. [6] Prolific, positioned more for research-grade tasks, recommends paying participants at least $12/hour and charges a 42.8% service fee for corporate accounts (33.3% for academic and non-profit accounts) on top of that. [12]
“Existing customers can continue to use the service as normal.”
AI-assisted, model-in-the-loop labeling
The fourth method uses a model to produce a first-pass label and routes only low-confidence or sampled items to a human for correction. Rev's own two-tier pricing is a clean, real illustration of the gap this creates for a single task type: $0.25/minute for AI-only transcription against $1.99/minute for a human transcript of the same audio, an 8x difference for an identical unit of work. [8] The mechanics of how model-in-the-loop pipelines route confidence and where human review still can't be automated away are covered in full in our guide to human-in-the-loop AI training; the short version relevant to cost is that the savings are real but only as good as the correction layer built around the model's output.
Worked example: budgeting to transcribe and label 1,000 hours of calls two ways
Take a concrete job: a company needs 1,000 hours (60,000 minutes) of support call recordings transcribed and speaker-labeled to train a voice AI assistant. Here's what that costs under two of the four methods, using Rev's published rates directly rather than a modeled estimate.
- 1Method 1: Outsourced human transcription vendor. At Rev's published human rate of $1.99/minute, transcribing the full batch costs 60,000 × $1.99 = $119,400. [6] Per-file turnaround is quoted at 12 hours or less, but a batch this size still depends on how many files the vendor can run in parallel; throughput at scale is bounded by the vendor's available transcriber pool, not the sticker price.
- 2Method 2: AI-assisted pipeline (ASR pass plus human correction). First, run every file through Rev's own AI tier at $0.25/minute: 60,000 × $0.25 = $15,000, delivered per file in minutes rather than hours. [6] Second, route the AI output to a human corrector rather than a fresh transcriber. Assuming a QA reviewer can correct an already-transcribed file at roughly 3x real-time speed, a reasonable planning assumption since correcting is faster than transcribing from silence, the correction pass needs about 60,000 ÷ 3 = 20,000 labor-minutes, or 333 labor-hours. At the $12.50/hour commodity BPO rate [12], that's 333 × $12.50 ≈ $4,167.
- 3Total the two methods. Method 1 totals $119,400. Method 2 totals $15,000 + $4,167 ≈ $19,167, roughly 6.2x cheaper for the same 1,000 hours of audio, with the AI pass itself landing in minutes per file instead of half a day.
| Method | Cost basis | Total for 60,000 minutes | Turnaround profile |
|---|---|---|---|
| Outsourced human transcription | $1.99/min, published rate [2] | $119,400 | 12 hrs/file, bounded by vendor parallel capacity at volume |
| AI-assisted (ASR + human correction) | $0.25/min AI pass + $12.50/hr correction pass at ~3x real-time [2][4] | ≈ $19,167 | Minutes per file for the AI pass; correction pass scales with labor-hours |
Comparing the four methods by cost, speed, quality control, and fit
| Method | Cost profile | Speed | Quality control | Best-fit scenario |
|---|---|---|---|---|
| In-house team | Fixed cost regardless of volume; ~$24-28/hr loaded for a commodity role [2], cheapest only at sustained high volume | Slowest to start: weeks to hire and ramp before first labels | Tightest feedback loop and deepest product context, but limited by your own hiring budget | Ongoing, sensitive, or competitively core labeling that needs deep institutional context |
| Outsourced BPO vendor | $10-15/hr commodity billed rate [2]; enterprise contracts average $93,158, up to $400,000 [4] | 2-4 weeks from signed contract to first delivered labels | Vendor's own QA program and SLA-backed accuracy targets | Large, sustained volume of moderate-complexity tasks |
| Crowdsourcing platform | Lowest nominal per-task price, plus a 20-42.8% platform fee on top of worker pay [2][4] | Fastest to start: days, once a task spec and gold-standard set exist | Weakest by default; buyer must design gold sets and consensus checks | Simple, high-volume, non-sensitive tasks, or a short one-off project |
| AI-assisted / model-in-the-loop | Lowest overall unit cost once a correction pipeline exists; Rev's own tiers show an 8x gap for one task type [2] | Fastest raw throughput: the AI pass runs in minutes | Entirely dependent on the human correction layer's design and sampling rate | Well-defined, high-volume tasks where a real QA correction step is built in, not skipped |
How to build a realistic labeling budget
The five drivers and four methods above turn into a budget with a short, repeatable process.
- 1Define the unit of work precisely. "10,000 images" is not a scope. "10,000 images, 5 object classes, roughly 5 objects per image, single-pass bounding boxes" is. Cost tracks the second description, not the first.
- 2Score how much domain judgment the task needs. A generalist rate applies to classification and simple boxes. A specialist rate, and a much smaller labor pool, applies the moment a wrong label costs real money or safety. [4]
- 3Decide your QA depth up front. Single-pass labeling is the cheapest and least trustworthy option. Double-pass with inter-annotator agreement scoring roughly doubles labor cost but is close to a floor for anything used to train a production model. [4]
- 4Set a real deadline and price it honestly. A same-day AI-assisted pass and a 12-hour human vendor turnaround are not interchangeable at the same price, as Rev's own two tiers show directly. [4] Don't assume rush pricing is a rounding error.
- 5Model your volume against all four methods before picking one. A one-time batch under a few hundred thousand units rarely justifies an in-house hire or a BPO contract minimum; sustained volume in the millions often does. Run the comparison table above against your actual numbers rather than defaulting to whichever method sounds most sophisticated.
- 6Check whether buying already-labeled data beats labeling it at all. If a comparable, already-annotated dataset exists in your domain, the acquisition cost is often lower than commissioning fresh labeling from any of the four methods above, since the seller has already absorbed the labeling cost.
Price the unit economics, not the sticker rate
Every number in this guide reduces to the same lesson: a labeling quote is only meaningful once you know the task, the workforce, the QA depth, the deadline, and the volume behind it. A $12.50/hour BPO rate and a $105/hour expert rate price genuinely different labor, and neither is wrong for the job it fits. [2][4] The buyers who get burned are the ones who compare a per-hour number across methods without checking what that hour actually buys.
Before committing a budget to any single method, run the numbers the way the worked example above does: price the same target volume under at least two methods, using real published rates rather than a vendor's opening quote. The gap between the cheapest and most expensive route to the identical labeled dataset is routinely 5-10x, and that gap is exactly where a careful buyer earns back the hour spent building the comparison.
Skip the labeling budget entirely.
Dayda brokers vetted, already-labeled proprietary datasets across text, image, audio, and conversational data. If a comparable dataset already exists in your domain, buying it can be cheaper and faster than commissioning labeling from scratch.
See how buying works on Dayda→