AI Data Buying

Best AI Training Data and Labeling Companies in 2026: Compared

17 min read · 2026-09-30
$14.3BMeta's investment for a 49% stake in Scale AI, June 2025
$20Bvaluation Mercor was in talks to raise at by July 2026, after crossing $2B in annualized revenue
$310MEXL's agreed acquisition price for iMerit, announced June 2026
$72MBezos Expeditions-led funding round for Toloka, 2025
$82.8Mannual Google contract Appen lost in January 2024, about 26% of its FY2023 revenue
Key takeaways
  • 1The eight vendors compared here split into three real categories: RLHF/expert-data specialists (Scale AI, Surge AI, Mercor), enterprise and regulated-industry annotation shops (iMerit, Appen, Sama), and self-serve tooling platforms (Labelbox, Toloka). Picking the wrong category costs more than picking the wrong company inside the right one.
  • 2Meta's $14.3 billion investment for a 49% stake in Scale AI in June 2025 reshuffled the market: OpenAI and Google reportedly stopped working with Scale, and Meta's own TBD Labs researchers said they preferred rivals Surge and Mercor on data-quality grounds. {cite:scalecracks}
  • 3Mercor's revenue more than doubled in four months, from $1 billion in annualized revenue in February 2026 to $2 billion by June 2026, while it negotiated a fundraise that would value it at roughly $20 billion, about double the $10 billion valuation it held in September 2025. {cite:mercorraise}
  • 4Legacy vendors are consolidating and pivoting under real pressure: Appen lost an $82.8 million Google contract worth roughly 26% of its FY2023 revenue in January 2024 {cite:appendecline}, Sama exited content moderation entirely after Kenya-based lawsuits {cite:samameta}, and EXL agreed to acquire iMerit for up to $310 million in June 2026. {cite:exlimerit}
  • 5Pricing scales non-linearly with task complexity inside a single vendor's own rate card: Labelbox's usage meter charges 20 Labelbox Units per data row for live multimodal or LLM content versus 1 unit per 60 rows for basic catalog storage, a roughly 1,200x spread. {cite:labelboxbilling}
The short version

There is no single best AI data labeling company in 2026, only a best fit for a specific data problem, and the honest starting point is knowing which of three categories a buyer actually needs. For RLHF, preference ranking, and reward-model data at frontier scale, Surge AI and Mercor have taken visible share from Scale AI since Meta's 2025 investment: Mercor's revenue crossed $2 billion annualized by mid-2026, and Surge has stayed profitable and bootstrapped since 2020 while landing OpenAI, Google, and Anthropic as clients. Scale AI itself remains a credible choice for buyers who want one vendor spanning both data work and a deployed enterprise or government AI application, though its 49%-owned relationship with Meta is a real conflict to weigh if a buyer competes directly with Meta. For computer vision, multimodal, and regulated-industry annotation, iMerit is the strongest specialist on compliance and credentialed expertise, Appen still offers the broadest multilingual crowdsourced reach despite losing a major Google contract, and Sama fits buyers who weight labor practices heavily, having rebuilt its business around direct-hire annotation after its content-moderation history. For teams that want to run labeling with their own workforce and just need the software, Labelbox's usage-based pricing and Toloka's contributor network built for agentic and reinforcement-learning data are the two credible self-serve options. None of these eight companies replace a managed marketplace: all eight get commissioned to produce or label data a buyer doesn't yet have in usable form, while a marketplace like Dayda sells access to datasets that already exist and are already vetted, a different purchase for a different stage of the same underlying problem.

On this page ▾

The vendor landscape in 2026: how the market got reshuffled

Most "top data labeling companies" lists in circulation were written before Meta's investment in Scale AI reordered the field. This piece audits eight vendors buyers actually shortlist today, Scale AI, Surge AI, Mercor, Labelbox, Appen, Sama, iMerit, and Toloka, with every funding, ownership, and positioning claim verified this year. They split into three genuinely different businesses: RLHF and expert-data specialists, enterprise and regulated-industry annotation shops, and self-serve tooling platforms. Treating them as interchangeable, or picking whichever name is most familiar, is the single most common mistake buyers make.

One event set off most of what follows. In June 2025, Meta invested roughly $14.3 billion for a 49% non-voting stake in Scale AI at a $29 billion valuation, and Scale's co-founder and CEO Alexandr Wang departed to lead Meta's own superintelligence lab. [2] Within months, OpenAI and Google reportedly stopped working with Scale AI following the investment, and Meta's own TBD Labs researchers said they preferred working with Scale's rivals, Surge and Mercor, on data-quality grounds rather than Scale itself. [4] A vendor investment that was supposed to signal strength triggered a customer exodus instead.

The cascade is still playing out. Surge AI, bootstrapped and profitable since its 2020 founding, launched its first-ever external fundraise in mid-2025, targeting up to $1 billion at a valuation north of $15 billion. [2] Mercor grew even faster: its annualized revenue crossed $2 billion in June 2026, doubling in four months, while the company negotiated a fundraise that would value it at roughly $20 billion, about double the $10 billion mark it hit just nine months earlier. [4] Two companies that were secondary names before 2025 are now arguably the two most-discussed vendors in the category.

Legacy players moved in the opposite direction. Appen lost an $82.8 million annual Google contract, roughly 26% of its FY2023 revenue, when Google terminated the relationship without warning in January 2024, sending its shares down more than 40%. [2] Sama shut down its content-moderation business entirely, the unit that once made it Meta's primary moderation partner in Kenya, after lawsuits alleging psychological trauma among moderators, and refocused fully on computer vision annotation. [4] iMerit built a durable enterprise niche in regulated-industry annotation and, in June 2026, agreed to be acquired by EXL for up to $310 million. [6] Even the vocabulary is shifting: Turing's CEO Jonathan Siddharth told a podcast audience in December 2025 that "the era of data-labeling companies is over," arguing that agentic and reinforcement-learning systems need domain-expert workflows, not basic tagging, a claim Turing's own repositioning away from labeling toward what it calls "research accelerators" backs up. [8]

The era of data-labeling companies is over. It's now the era of research accelerators.
Business Insider (via AOL): 'The era of data-labeling companies is over,' says the CEO of a $2.2 billion AI training firm

Comparing the eight leading vendors at a glance

The table below is the fast reference: what each vendor actually specializes in, how its workforce is structured, roughly what kind of commitment it expects, and who it fits best. Treat "typical minimum engagement" as directional. None of these companies publish a universal rate card for enterprise work, and every one of them will quote differently depending on task complexity, volume, and how much of the QA burden the buyer wants to own.

VendorSpecialtyWorkforce modelTypical engagementBest-fit buyer
Scale AIRLHF/labeling plus deployed enterprise and government AI applicationsManaged expert workforce, contributors often hold advanced degreesEnterprise SOW, no published minimum, historically large accountsBuyers wanting one vendor for both data and a built AI system, comfortable with Meta's 49% stake
Surge AIHigh-end RLHF: preference ranking, red-teaming, reward-model dataVetted, highly trained contractors, not open crowdsourcingEnterprise-only, pay-as-you-grow, no public minimumFrontier labs needing top-tier human judgment at the hardest tier of RLHF
MercorExpert-sourced RLHF and evals from credentialed professionalsMarketplace of vetted domain experts (PhDs, lawyers, bankers, coders), paid per taskEnterprise engagement, contractors billed hourly ($85+/hr average)Labs needing deep subject-matter expertise for evals or RLHF at high speed and scale
LabelboxSelf-serve annotation tooling, plus optional managed human-data add-on (Alignerr)Software platform; buyer's own team labels, or opts into a managed workforceFree tier (500 LBU/mo); paid self-serve from $0.10/LBU; no contract minimumTeams that want to own their annotation workflow and pay only for usage
AppenCommodity crowdsourced labeling, search-quality rating, multilingual annotationGlobal crowdsourced contributor network across 500+ localesEnterprise SOW, historically large multi-year contractsBuyers needing broad-language commodity labeling at volume, accepting vendor-transition risk
SamaComputer vision, image/video/3D point-cloud annotation and model evaluationDirect-hire, full-time associate workforce, B Corp impact-sourcing modelEnterprise SOWBuyers prioritizing CV annotation with a direct-employment labor model
iMeritExpert annotation for healthcare, autonomous vehicles, geospatial, financeiMerit Scholars: ~25,000 domain experts (physicians, scientists, engineers) across 60+ countriesEnterprise SOW; SOC2/ISO 27001/HIPAA/GDPR/TISAX-certifiedRegulated-industry buyers needing credentialed experts and an audit-ready compliance file
TolokaRLHF, agent-trajectory and RL-environment data, coding data, multilingual crowd and expert tasks200,000+ annotators and professionals, 40+ languages, ISO 27001/SOC2/GDPR/HIPAA-compliantEnterprise SOW, Nebius-backedBuyers needing agent or RL-environment data, or multilingual coverage at expert quality
None of these eight sell existing datasets
Every vendor on this list is commissioned to produce or label data to a buyer's spec. If the actual need is a proprietary dataset a buyer doesn't already hold in raw form, that's a different channel entirely: a managed marketplace, or direct licensing. See the tradeoffs section below for where the line actually falls.

RLHF and human-preference data specialists: Scale AI, Surge AI, and Mercor

Scale AI remains the largest and most recognized name in the category, but its own business mix is changing fast. Interim CEO Jason Droege inherited a company that ran roughly 70% on data labeling and 30% on building AI applications for enterprises, and has been reversing that ratio, expecting the applications business to overtake labeling revenue within about 18 months. Scale's 2025 revenue came in just under $1 billion, up from $870 million in 2024, and the company has landed new enterprise accounts including Mayo Clinic, BP, Allianz, and Ernst & Young, alongside government work through the Pentagon's Chief Digital and Artificial Intelligence Office. [4] Scale is still a defensible choice for a buyer who wants one vendor spanning both raw data work and a deployed AI system with government-grade security clearance. The complication is structural: Meta owns 49% of Scale AI, and a buyer building anything that competes with Meta's own AI roadmap has a real conflict to weigh, not a hypothetical one.

Surge AI took the opposite path to scale: no outside funding at all until mid-2025, built entirely on profitability since its 2020 founding by former Google and Meta engineer Edwin Chen. By the time it launched its first fundraise, targeting up to $1 billion at a valuation north of $15 billion, Surge's own annual revenue already exceeded $1 billion. [4] Surge markets itself specifically around premium, high-end labeling for reinforcement learning, using highly trained contractors rather than an open crowd, and its client list, OpenAI, Google, and Anthropic, reads like the short list of labs with the most demanding data-quality bars in the industry. [6] Meta's own researchers put a fine point on why: reporting found that TBD Labs, Meta's internal AI team, was working with third-party vendors other than Scale AI specifically because researchers judged Scale's data as lower quality.

TBD Labs is working with third-party data-labeling vendors other than Scale AI to train its upcoming AI models... Those third-party vendors include Mercor and Surge, two of Scale AI's largest competitors.
TechCrunch: Cracks are forming in Meta's partnership with Scale AI

Mercor is the fastest-moving name in the category by a wide margin. Founded in January 2023 by three co-founders who met on their high school debate team, Mercor's annualized revenue run rate crossed $2 billion in June 2026, doubling in just four months, and the company was in talks in July 2026 to raise $500 million at a $20 billion valuation, roughly double where it stood the previous September. [4] Mercor's model is closer to an expert staffing marketplace than a traditional labeling shop: it recruits PhDs, lawyers, bankers, and engineers to produce specialized RLHF and evaluation data for frontier labs, paying contributing experts $85 or more per hour on average. [6] For buyers whose bottleneck is genuine subject-matter depth rather than volume, Mercor's positioning is built specifically around that gap, and its growth curve suggests the market agrees.

Computer vision and regulated-industry annotation: iMerit, Appen, and Sama

iMerit has spent over a decade building the deepest compliance posture in the category. Its iMerit Scholars program embeds roughly 25,000 domain experts, physicians, scientists, engineers, and linguists, across more than 60 countries directly into annotation, evaluation, and RLHF workflows for frontier labs and regulated enterprises, and the company holds SOC2, ISO 27001, GDPR, HIPAA, and TISAX certifications. [4] That combination made it an acquisition target: in June 2026, EXL (Nasdaq: EXLS) agreed to acquire iMerit for up to $310 million, split between $170 million upfront and $140 million in earnouts over two years, explicitly to fold iMerit's foundation-model training expertise, its Ango annotation platform and Scholars expert network, into EXL's enterprise AI practice for healthcare, insurance, and banking clients. [6] Buyers already using iMerit, or evaluating it, should ask directly how the EXL integration affects existing pricing, data-handling terms, and account continuity through the Q3 2026 close.

$310M
EXL's agreed acquisition price for iMerit, structured as $170M upfront plus $140M in milestone earnouts
EXL (ExlService Holdings), June 2026

Appen was, for over a decade, the default choice for large-scale, multilingual crowdsourced labeling, running Google's search-quality-rating program among other major contracts. That default status made its 2024 setback more visible than most: Google terminated its roughly $82.8 million annual contract with Appen without warning in January 2024, a deal worth about 26% of Appen's FY2023 revenue, and the stock fell more than 40% on the news. [4] Appen remains an operating public company and has pivoted toward frontier-model alignment and agentic-AI data products, but its recent history is itself the clearest illustration of the vendor-concentration risk buyers should weigh with any single-vendor relationship: a company can lose a quarter of its revenue on one counterparty's decision it had no say in.

Sama built its early business on content moderation, most visibly as Meta's primary moderation partner in Kenya. That relationship ended after lawsuits alleging psychological trauma, poverty wages, and labor-rights violations among Kenyan moderators, and Sama shut its content-moderation division entirely, shifting its business fully to computer vision, image and video annotation, 3D point-cloud labeling, and model evaluation. [4] Sama operates as a certified B Corporation with a direct-hire, full-time associate workforce rather than a gig-labor model, a structure it markets as fair-wage impact sourcing. Given the company's own history, buyers who care about labor practices should treat that positioning as a real evaluation criterion worth verifying, not a marketing line to take at face value.

Self-serve annotation tooling: Labelbox and Toloka

Labelbox sells software, not labor. Teams that already have annotators, in-house, contracted, or borrowed from one of the vendors above, use Labelbox to manage labeling ontologies, QA workflows, and model-assisted pre-labeling. Pricing runs on a usage meter called the Labelbox Unit: a free tier includes 500 LBU per month, paid self-serve costs $0.10 per LBU, and consumption rates vary sharply by workload, from 1 LBU per 60 rows for basic Catalog storage up to 20 LBU per data row for live multimodal or LLM content. [4] Enterprise tiers add SSO and HIPAA compliance. For buyers who want managed human labor rather than pure tooling, Labelbox layers in Alignerr, a sales-quoted human-data service, which blurs the tooling-versus-workforce line the rest of this comparison otherwise assumes.

Toloka sits between Labelbox's pure-tooling model and the expert-marketplace approach of Surge and Mercor: it runs the workforce itself, but positions its output specifically around agent and reasoning-model training rather than classic RLHF alone. The company operates a network of more than 200,000 annotators and professionals spanning 40-plus languages, producing trajectory demonstrations, RL-environment simulations, code-generation workflows, and red-teaming data, and holds ISO 27001, SOC2, GDPR, and HIPAA compliance. [4] Toloka operates as a unit of Nasdaq-listed Nebius Group, and in 2025 took $72 million from Bezos Expeditions to expand its U.S. presence, with Shopify CTO Mikhail Parakhin joining as executive chairman. [6] For buyers building agentic products specifically, rather than classic chat-model alignment, Toloka's positioning is the most directly targeted of the eight.

BPO vendor, self-serve tooling, or managed marketplace: the real tradeoffs

All eight companies above solve one problem: turning a task specification into labeled or evaluated output, using either a vendor-managed workforce or a buyer-managed one on rented software. That is a genuinely different problem from the one a managed marketplace like Dayda solves, and buyers who blur the two waste time evaluating the wrong category of vendor.

  • Traditional BPO-style vendor (Appen, Sama, and the SOW side of Scale and iMerit): the buyer defines the task, the vendor recruits and manages a workforce against it, and pricing runs per labeled unit or fixed statement of work. The strength is turnkey scale: a buyer doesn't have to build recruiting, training, or QA infrastructure. The weakness is that quality is only as good as a vendor's own QA design, which the buyer usually can't fully inspect, and switching vendors mid-project means retraining a new workforce on the same spec from scratch.
  • Self-serve tooling platform (Labelbox, and Toloka's tooling layer): the buyer keeps the workforce, internal or contracted, and pays for software that manages the pipeline. The strength is transparency and flexibility: Labelbox's LBU pricing is public and usage-based, with no lock-in to a fixed contract. The weakness is that the buyer inherits the QA-design burden a managed vendor would otherwise absorb, which is exactly why Labelbox itself sells a managed add-on for teams that decide they don't want to own that work after all.
  • Managed marketplace (Dayda's category): brokers datasets that already exist, already labeled, already provenance-checked, rather than commissioning new labeling work against a spec. This solves a prior-stage problem: a labeling vendor turns raw data a buyer already holds into a usable asset; a marketplace supplies data the buyer doesn't hold at all, with the vetting and licensing already done.

The honest framing: if the problem is "we have raw data and need it labeled," none of the eight vendors above get replaced by a marketplace, that is a labeling problem, and the right question is which of the eight fits the task, workforce model, and compliance needs. If the problem is "we don't have the raw data to begin with," commissioning any of these eight vendors to generate something close to it from scratch is usually slower and pricier than sourcing an existing, vetted dataset through a marketplace. Most serious buyers need both channels at different points in a model's lifecycle, not one instead of the other.

How to actually run a vendor selection process: a 5-step evaluation

Picking a vendor off the comparison table above is a starting shortlist, not a decision. A buyer choosing between, say, Surge AI, Mercor, and Toloka for an RLHF program should run all three through the same structured process before committing budget to any one of them.

  • 1Run an identical pilot batch across every shortlisted vendor. Send the same 200-500 unit sample, with the same instructions, to two or three vendors simultaneously. A vendor's sales conversation reveals almost nothing about actual output quality; a real pilot does.
  • 2Measure inter-annotator agreement, don't eyeball it. Score the pilot output with a real agreement metric (Cohen's kappa for categorical labels, exact-match rate for structured tasks) rather than a spot-check of a handful of examples. Low agreement on a pilot almost always gets worse, not better, at scale.
  • 3Get real cost-per-unit at production scale, not pilot pricing. Pilot batches are frequently priced as loss leaders to win the account. Ask each vendor for its actual rate at the buyer's real projected volume before comparing costs, and multiply out the full engagement, not just a per-unit number in isolation.
  • 4Run a data security and compliance review before signing anything. Verify claimed certifications directly rather than taking a sales deck's word for them, confirm where data is actually processed and stored geographically, and ask for the full subcontractor chain: several vendors in this category route work through additional layers a buyer never sees named on the contract.
  • 5Negotiate exit and data-portability terms up front, not after a problem appears. Confirm in writing what happens to in-progress work, already-delivered labeled output, and any proprietary QA logic or ontology if the engagement ends early. A vendor that won't commit to portability terms before signing is telling a buyer something about how a dispute would go later.
Run the pilot with your worst-case data, not your cleanest sample
Vendors optimize pilot performance for whatever sample they're handed. Include the messiest, most ambiguous examples from the real dataset in the pilot batch, not just the clean cases, since that's where quality differences between vendors actually show up at scale.

What a skeptical buyer should ask before signing

Vendor lock-in. Both workforce vendors and tooling platforms create switching costs, just different kinds. A managed workforce trained on a buyer's specific task spec represents sunk ramp-up time that a new vendor has to rebuild from zero. A tooling platform like Labelbox creates lock-in through its ontology structures and pipeline configuration, not the labor itself. Either way, the fix is the same: negotiate data and configuration portability terms before signing, not after deciding to leave.

Offshore workforce quality and compliance risk. This is not a hypothetical risk to wave off with a certifications page. Sama's own history, lawsuits from Kenyan content moderators alleging psychological trauma and labor-rights violations, and the eventual shutdown of that entire business line, is a real, reported example of what goes wrong when workforce conditions aren't scrutinized. [4] Certifications like iMerit's SOC2 and HIPAA coverage answer a data-security question, not a labor-practices question. Buyers who care about both need to ask about them separately: where the workforce is located, whether it's direct-hire or subcontracted, and what independent verification exists beyond the vendor's own marketing.

Non-linear pricing as task complexity rises. Buyers who anchor on a vendor's advertised entry-level rate consistently underbudget. Labelbox's own public rate card shows the pattern cleanly: basic catalog storage costs 1 LBU per 60 data rows, while live multimodal or LLM content costs 20 LBU per single row, a roughly 1,200x difference in metered cost for a more complex data type on the exact same platform. [4] Expert-network pricing follows the same curve in dollar terms: Mercor's contributing experts average $85 or more per hour for specialized professional judgment, an order of magnitude above commodity crowd-labeling rates. [6] Any budget built on an entry-level quote for a task that will actually require domain expertise or multimodal handling is going to be wrong, often by a full order of magnitude.

Which vendor should you actually use?

Match the vendor to the category first, and the specific company second. For RLHF and human-preference data at frontier scale, Surge AI and Mercor are the two names that have most visibly earned trust from major labs since the Scale-Meta shakeup, and Scale AI itself is still viable for buyers who also want a deployed enterprise application built alongside the data work. For computer vision and regulated-industry annotation, iMerit's compliance depth suits healthcare, autonomous vehicle, and financial buyers directly, Appen still covers the widest multilingual footprint despite its own recent volatility, and Sama fits buyers for whom labor practices are a real evaluation criterion, not an afterthought. For teams that want to run labeling themselves and just need software, Labelbox's transparent usage pricing and Toloka's agent-focused contributor network are the two credible self-serve paths.

The market itself keeps moving. A vendor that looked dominant in early 2025 lost its two most important customers within months, and a company that didn't exist before 2023 is now negotiating a $20 billion valuation. [2][4] Whatever a buyer decides today, the right process is the one in this piece: pilot before committing, measure agreement rather than trusting a sales deck, price at real scale, review security and compliance directly, and lock in exit terms before signing, not after something goes wrong. And if the actual gap is a dataset that doesn't exist yet inside the company at all, rather than raw data waiting to be labeled, that's a different purchase than any of the eight vendors above, and worth recognizing as one before a labeling contract gets signed by default.

Need a dataset, not a labeling contract?

Dayda brokers proprietary, provenance-checked datasets that already exist, NDA-gated sampling and licensing terms included, for the buyers whose real gap is data they don't hold yet, not a workforce to label data they already have.

See how buying works on Dayda