- 1HR and recruiting AI is trained on five data types: resumes/CVs, job descriptions, structured ATS/HRIS fields, interview transcripts and scores, and hire/promotion/attrition outcome labels. The last category is the scarcest and most valuable, because it's the only data that shows what actually happened to a candidate after the decision.
- 2This is the highest-legal-risk vertical in AI training data. NYC Local Law 144 requires an independent bias audit and public posting of results for any automated employment decision tool used on NYC-based or NYC-remote jobs, backed by real penalties. {cite:ll144}
- 3The EU AI Act classifies recruitment and employment-management systems as high-risk under Annex III, but a political agreement on the Digital Omnibus has postponed the compliance deadline for those obligations from August 2, 2026 to December 2, 2027, pending formal adoption. {cite:omnibus}
- 4Federal guidance is unsettled, not repealed: the EEOC pulled its 2023 Title VII/AI adverse-impact guidance from its website in January 2025, but the underlying four-fifths adverse-impact standard remains fully enforceable law. {cite:eeocremoved}{cite:fourfifths}
- 5Historical hiring outcome data is a double-edged sword. Amazon's own resume-screening tool, built on a decade of male-dominated hiring outcomes and scrapped in 2017, penalized resumes containing the word "women's," showing how unaudited outcome data can automate the discrimination it was trained on. {cite:amazon}
HR and recruiting AI, resume parsers, candidate matchers, interview-scoring models, applicant rankers, and internal HR chatbots, are trained on five kinds of data: resumes, job descriptions, structured applicant-tracking fields, interview transcripts and scores, and outcome labels recording who was hired, promoted, or fired. Of those, outcome labels are the genuinely scarce asset, because resumes and job postings are already commoditized across the open web while a company's actual hiring, promotion, and attrition history exists only inside its own systems. That scarcity collides with the field's defining problem: historical hiring data encodes historical discrimination, and training a model on it without auditing for bias risks automating that discrimination at scale, the exact failure that killed Amazon's internal resume-screening tool. The regulatory response is real and active, not theoretical: New York City requires an independent bias audit before an employer can use an automated employment decision tool, Illinois requires notice and consent before AI scores a video interview, the EU AI Act classifies recruitment AI as high-risk, and Title VII's adverse-impact standard applies to algorithmic hiring tools whether or not the EEOC currently publishes guidance saying so. Anyone buying, selling, or licensing HR training data has to treat bias and compliance evaluation as a precondition, not a formality, and this guide walks through exactly how to do that.
On this page ▾
- What data actually trains HR and recruiting AI
- Why outcome labels are the genuinely scarce asset
- The regulatory landscape: NYC, the EEOC, states, and the EU
- Why historical hiring data is a double-edged sword
- A worked example: evaluating an HR dataset before buying it
- Data type, use case, and risk: the full picture
- What separates licensable HR data from a liability
- The bottom line
What data actually trains HR and recruiting AI
Resume parsing, candidate matching, AI interview scoring, applicant ranking, performance-review analysis, and internal HR chatbots are five different products, but they draw on the same underlying pool of data. Understanding what that pool actually contains matters more than the product category, because the same resume dataset that trains a parser can also train a ranking model, and the compliance exposure travels with the data, not the use case label a vendor puts on it.
- Resumes and CVs. Free-text and semi-structured documents: work history, education, skills, certifications. This is the input side of parsing and matching models, and it's the least scarce of the five, since resumes circulate across job boards, LinkedIn, and career sites.
- Job descriptions and postings. The other half of the matching problem: required skills, seniority, compensation bands, and the language a company uses to describe a role. Also widely public, though internal, unpublished JD drafts and leveling rubrics are proprietary.
- Structured ATS and HRIS data. The applicant-tracking and human-resources-information-system fields: application dates, stage transitions, recruiter notes, rejection reasons, source channel, demographic self-identification where collected. This is where a candidate's raw resume becomes a trackable record.
- Interview transcripts and scores. Text or video from screening and panel interviews, plus the scores or rubric ratings assigned to them, whether by a human interviewer or an AI scoring tool. This is the training signal for interview-scoring products specifically.
- Hire, promotion, and attrition outcome labels. The record of what actually happened after a decision: who was hired, at what level, who was promoted and when, who left voluntarily or was terminated, and often why. This is the label, not the feature, and it's the input that turns a matching or ranking model from a keyword filter into something that claims to predict success.
Why outcome labels are the genuinely scarce asset
Resumes and job postings are commoditized. Millions sit on public job boards, professional networks, and career pages, and they've been scraped, licensed, and re-licensed for years. A buyer looking for raw resume text or job-description corpora has many options and correspondingly little reason to pay a premium for any single source.
Outcome data doesn't work that way. Whether a specific candidate was hired, at what level, how they were rated in their first two performance cycles, whether they were promoted or left within eighteen months, exists nowhere except inside the hiring company's own ATS and HRIS. No public crawl reconstructs it, because it was never published anywhere, and it can't be inferred reliably from a resume or a LinkedIn profile after the fact. That is what makes it the scarce, structurally valuable half of HR training data: it is the only data that lets a model learn what separates a candidate a company judged successful from one it didn't, across its actual hiring history rather than a proxy for it.
Demand for that signal is rising alongside broader AI adoption inside HR functions. Recent SHRM survey data puts the share of organizations using AI to support recruiting at just over half, up sharply from the year before, with resume screening and candidate search among the most common applications. [2] Every one of those deployments needs training and evaluation data, and the deployments that go beyond keyword matching into ranking or prediction need outcome labels specifically, which is exactly the layer few companies are willing or able to share.
The regulatory landscape: NYC, the EEOC, states, and the EU
No other vertical in AI training data sits under this much active, overlapping regulation. Four regimes matter most right now, and each is evolving, so treat the specifics below as current as of mid-2026, not permanent.
NYC Local Law 144: the bias-audit standard
New York City's Local Law 144 prohibits an employer or employment agency from using an automated employment decision tool (AEDT) on a candidate or employee for a job located in NYC, or a remote job associated with an NYC office, unless the tool has undergone an independent bias audit within the prior year, a summary of the results is publicly posted, and candidates receive at least 10 business days' notice before the tool is used. [10] The audit has to calculate impact ratios across race/ethnicity and sex categories, using a methodology comparable to the EEOC's four-fifths rule, and it has to be performed by an auditor with no affiliation to the employer or the tool's vendor. Violations carry penalties starting at $500 per violation and rising for repeated or ongoing non-compliance. [12]
For a training-data buyer or seller, LL144 matters even if you never deploy a tool in New York, because it's the closest thing the U.S. has to a legally mandated bias-testing methodology for hiring AI. A dataset that has already been through, or can support, an impact-ratio analysis by protected category is materially more valuable and materially safer to build on than one that can't.
The EEOC, Title VII, and the ADA: guidance pulled, law unchanged
In May 2023 the EEOC published technical guidance applying Title VII's adverse-impact framework directly to algorithmic hiring tools, telling employers to test selection rates across demographic groups using the same four-fifths rule long used for human-run selection procedures: a selection rate for any race, sex, or ethnic group below four-fifths (80%) of the rate for the highest-scoring group is generally treated as evidence of adverse impact. [6] In January 2025 the EEOC removed that guidance, along with its broader AI and Algorithmic Fairness Initiative page, from eeoc.gov. [8]
That removal changed enforcement posture, not the law. Title VII's prohibition on disparate impact, and the Uniform Guidelines on Employee Selection Procedures that codify the four-fifths rule, remain fully in force at 29 CFR Part 1607, guidance page or not. [2] What's genuinely unsettled is how aggressively the current EEOC will pursue algorithmic-discrimination cases; reporting on the guidance removal notes that at least four states have moved to write their own AI-employment rules specifically because the federal posture became less predictable. [4] The ADA runs alongside Title VII here too: interview-scoring tools that infer traits from speech patterns, facial movement, or response timing raise separate disability-discrimination exposure if they screen out candidates whose disabilities affect those signals, regardless of intent.
“A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact.”
Illinois and the state patchwork
Illinois's Artificial Intelligence Video Interview Act requires an employer to notify each applicant, before the interview, that AI may analyze their video interview, explain in general terms how the AI works and what characteristics it evaluates, and obtain the applicant's consent before proceeding. [4] Video may only be shared with people whose expertise is needed to evaluate it, and an employer must delete an applicant's interview, and instruct all recipients to do the same, within 30 days of a deletion request. [8] Employers who rely solely on AI analysis to determine which applicants get an in-person interview also have separate demographic-reporting obligations under the statute. [10]
Illinois isn't unique. Coverage of the EEOC guidance removal points to a wider pattern: several states, not just Illinois, have enacted or advanced their own AI-employment statutes as federal guidance receded. [2] For anyone sourcing interview-transcript or video data specifically, the practical implication is the same one that runs through this entire vertical: the notice-and-consent trail has to be documented per jurisdiction, because a transcript collected without Illinois-compliant notice is a liability no amount of annotation quality offsets.
The EU AI Act: recruitment AI as “high-risk”
Under the EU AI Act, Annex III(4) classifies AI systems intended for the recruitment or selection of natural persons, including targeted job-ad placement, filtering applications, and evaluating candidates, as high-risk, alongside systems used to make decisions on promotion, termination, task allocation, and performance monitoring within an employment relationship. [6] High-risk classification triggers a substantial compliance package: risk management, data governance, technical documentation, record-keeping, human oversight, and accuracy and robustness requirements. [8]
The timeline here is a live, moving target, which is exactly the kind of thing to verify rather than assume. Annex III high-risk obligations were originally set to apply from August 2, 2026. A political agreement reached in May 2026 on the EU's Digital Omnibus package postpones that date for stand-alone Annex III systems, employment AI included, to December 2, 2027, though the deferral is not final until formally adopted and published in the Official Journal, expected before the original August 2026 date. [6] Separate transparency obligations under Article 50 are unaffected and still take effect on the original schedule. [8] Anyone building compliance timelines around this vertical should treat the December 2027 date as the current expectation, not a locked-in fact, until the Official Journal publication lands.
Why historical hiring data is a double-edged sword
The most valuable HR training asset, real outcome data showing who a company hired, promoted, and let go, is also the asset most likely to encode the exact discrimination the regulations above exist to catch. The clearest documented case is Amazon's internal resume-screening project. Between 2014 and 2017 the company built a tool that scored applicants for technical roles from one to five stars, trained on ten years of resumes the company had actually received. [2] Because the tech industry's applicant pool skewed heavily male over that decade, the model learned to treat maleness itself as a signal of fit: it downgraded resumes that included the word “women's,” as in “women's chess club captain,” and penalized graduates of two all-women's colleges. [4] Amazon's engineers edited the model to neutralize those specific terms, then lost confidence it wasn't finding other, subtler proxies for the same pattern, and scrapped the project. [6]
“The system taught itself that male candidates were preferable. It penalized resumes that included the word “women's,” as in “women's chess club captain.” And it downgraded graduates of two all-women's colleges.”
The mechanism matters more than the anecdote. A model trained to predict “who gets hired” from historical outcomes isn't learning merit in the abstract; it's learning the pattern of a specific company's past decisions, including whatever bias, structural or individual, shaped those decisions. Removing the protected characteristic itself, name, gender, race, rarely fixes this, because the model finds correlated proxies instead: school names, gaps in employment history that correlate with caregiving, zip codes, even word choice and phrasing. That's exactly what the four-fifths rule and NYC's bias-audit requirement exist to catch after the fact. [2][4] It's also why a dataset's age and source distribution matter as much as its size: outcome labels drawn from a company's hiring during a period of known demographic skew carry that skew forward into anything trained on them, whether or not anyone intended it.
A worked example: evaluating an HR dataset before buying it
Take a concrete case. A staffing-marketplace startup is winding down. It holds 550,000 candidate profiles with resumes and structured ATS fields, interview scores from about 40,000 completed video interviews, and hire, promotion, and termination outcome labels for 210,000 candidates placed over three years. An AI lab building a candidate-ranking product wants to license it. Here is the review a serious compliance and data-science team runs before making an offer, not after.
- 1Map protected-class exposure and proxies. Does the dataset carry explicit demographic fields (self-identified race, sex, disability status, veteran status)? If not, can proxies (school, zip code, name, employment gaps) plausibly stand in for them? Either way, the buyer needs to know what a downstream model could learn to discriminate on, even unintentionally.
- 2Check the notice-and-consent trail on the interview data. Were the 40,000 video interviews collected with AIVIA-compliant notice and consent, given the startup operated candidates in multiple states? [4] No documented consent for AI analysis of video interviews is a hard stop for that slice of the data, regardless of price.
- 3Run a four-fifths adverse-impact test on the outcome labels. Using whatever demographic fields or reliable proxies exist, calculate selection, promotion, and termination rates by group and compare against the 80% threshold. [4] A dataset that already fails this test on its own historical outcomes will train a model that reproduces the failure unless it's specifically corrected for.
- 4Verify the license actually covers AI training re-use. Candidate consent language written for “internal matching purposes” does not automatically extend to sale for third-party model training; the seller needs to show the license or terms of service that permit this specific transfer.
- 5Weigh the collection period and industry mix. Outcome labels from a period or sector with a known demographic skew (as with Amazon's decade of male-dominated tech applicants) carry that skew forward into anything trained on them. [4] Note the years and roles covered before assuming the data is representative.
- 6Decide, and document the decision. Buy as-is only if the impact-ratio test passes and consent is clean; buy with a bias-mitigation and re-weighting plan if the labels are usable but skewed; or reject the outcome-label slice specifically while still taking the resume and ATS data, if that's the only piece with unresolved risk.
| Review item | What passes | What fails (walk away or discount hard) |
|---|---|---|
| Protected-class / proxy mapping | Fields identified and documented, proxies flagged for testing | No visibility into demographic composition at all |
| Interview consent trail | Documented AIVIA-style notice and consent per jurisdiction | No record of notice or consent for AI-scored interviews |
| Four-fifths adverse-impact test | Selection/promotion rates within or near the 80% threshold, or a documented mitigation plan | Rates well below 80% for a protected group with no mitigation plan |
| License scope | Terms explicitly permit third-party training re-use | Consent limited to internal use only |
| Collection period / industry mix | Recent, demographically representative, or skew is documented and correctable | Old data from a known skewed period presented as representative |
Data type, use case, and risk: the full picture
Put together, the five data types map onto specific products and specific regulatory exposure. Use this as a quick reference, not a substitute for the jurisdiction-by-jurisdiction review above.
| Data type | Primary use case | Key risk / requirement |
|---|---|---|
| Resumes / CVs | Resume parsing, candidate matching | Largely public and low-scarcity; still personal data under GDPR/CCPA if identifiable |
| Job descriptions | Job-ad generation, role/skill matching | Lowest legal risk of the five; watch for historically gendered or exclusionary language patterns |
| Structured ATS/HRIS data | Applicant ranking and scoring | Triggers NYC LL144's AEDT bias-audit requirement when used to substantially assist a hiring decision [2] |
| Interview transcripts / video & scores | AI interview scoring | Illinois AIVIA notice, explanation, and consent requirements; ADA exposure from trait-inference scoring [2] |
| Hire / promotion / attrition outcome labels | Ranking and prediction model training | Encodes historical discrimination if unaudited (the Amazon pattern); Title VII disparate-impact exposure [2][4] |
| Performance-review text | Promotion/output prediction, internal-mobility models | Rater bias baked into the “ground truth” label; falls under the EU AI Act's work-related-decision category [2] |
| Internal HR chatbot transcripts | Employee-assistant and HR-support models | Confidential personnel data; needs the same consent-of-use disclosure as any other employee record |
What separates licensable HR data from a liability
Every regime covered above converges on the same practical requirements. Whether you're a startup preparing HR data for sale or a lab evaluating a purchase, this is the checklist that actually matters:
- Documented notice and consent, jurisdiction by jurisdiction, for any interview or video data, matched to where candidates actually applied. [4]
- A completed or reproducible adverse-impact analysis on outcome labels, using the four-fifths standard or an equivalent impact-ratio methodology. [4]
- An independent bias audit, or the underlying data needed to run one, if the tool trained on this data will touch NYC-based or NYC-remote candidates. [4]
- A license that explicitly covers third-party AI training re-use, not just internal operational use of the original ATS or HRIS system.
- Transparency about the collection period and demographic composition of outcome labels, so a buyer can correct for known skew instead of inheriting it silently. [4]
- A defensible position under the EU AI Act if any candidates or the deploying entity touch the EU, given recruitment AI's high-risk classification and the still-moving compliance timeline. [4][6]
The bottom line
HR and recruiting AI runs on five data types, and the one that actually differentiates a serious model, hire and promotion outcome labels, is also the one most likely to carry forward a company's past hiring bias if nobody checks. That combination of scarcity and risk is what makes this the vertical where compliance work has to happen before a listing goes live, not after a deal closes. New York City, Illinois, and the EU already have binding or soon-to-bind rules that require exactly this kind of check, and Title VII's adverse-impact standard applies regardless of what guidance the EEOC currently publishes on its website. [2][4][6][8]
For sellers, that means the audit and consent work is what turns a shut-down company's ATS export into an asset a reputable buyer can actually acquire. For buyers, it means the review in this guide isn't optional diligence, it's the difference between licensing a training signal and licensing a discrimination lawsuit with a data schema attached.
Get your HR dataset compliance-ready before you list it.
Dayda reviews provenance, consent, and bias-audit readiness before any HR or recruiting dataset goes to buyers. We'll tell you what's defensible, what needs fixing, and what it's worth once it clears.
List your data on Dayda→