Vertical AI

What Training Data Does Government and Public-Sector AI Need?

16 min read · 2026-09-28
1,110AI use cases reported across 11 federal agencies in 2024, nearly double 2023's 571
~40,000algorithm-only fraud determinations issued by Michigan's MiDAS system, 2013–2015, with an error rate later estimated near 85%
365 dayswindow OMB gives agencies to complete minimum risk-management practices for a high-impact AI system before deployment
3AI vendors (OpenAI, Google, Perplexity) that had cleared FedRAMP 20x authorization for government use by early 2026
44.85% vs 23.45%false-positive recidivism-risk rate for Black vs. white defendants found in the COMPAS criminal-justice algorithm
Key takeaways
  • 1Government AI training data sits behind a compliance stack no commercial buyer faces: FedRAMP authorization for the cloud service itself, OMB's April 2025 procurement rules (M-25-21 and M-25-22) for how it's bought, and, for many use cases, EU AI Act Annex III high-risk classification if the deployment touches EU citizens.
  • 2The EU AI Act designates entire categories of public-sector AI as high-risk by default under Annex III, including eligibility determination for public benefits, law enforcement profiling, migration and asylum processing, and administration-of-justice tools, triggering Article 10's data-quality and bias-examination duties.
  • 3Automated fraud and eligibility systems have a documented failure record: Michigan's MiDAS unemployment system issued roughly 40,000 algorithm-only fraud determinations from 2013 to 2015 with an error rate later estimated near 85%, leading to a $20 million settlement approved in January 2024.
  • 4FedRAMP's 2025 AI Prioritization Initiative fast-tracked authorization for a narrow set of enterprise conversational AI tools; three vendors (OpenAI, Google, and Perplexity) cleared FedRAMP 20x authorization by early 2026, but that pathway does not cover most custom or fine-tuned government AI systems.
  • 5A compliant public-sector training data pipeline needs artifacts a commercial ML project never produces: a Privacy Act System of Records Notice before any personal data collection, a documented AI impact assessment within OMB's 365-day window for high-impact systems, and a bias/representativeness audit trail that survives FOIA and GAO review.
The short version

Government and public-sector AI needs the same categories of training data as commercial AI, structured records for eligibility rules, transaction histories for fraud patterns, case narratives for chatbots, but almost none of it can be sourced, procured, or deployed the way commercial teams are used to. The cloud infrastructure has to clear FedRAMP authorization before an agency can touch it. The procurement itself runs through OMB's April 2025 rules (M-25-21 and M-25-22), which require a documented AI impact assessment and ongoing performance monitoring for anything classified as "high-impact." If the system's output reaches an EU resident, entire categories, benefits eligibility, law enforcement, migration, and judicial support, are presumptively high-risk under Annex III of the EU AI Act, triggering Article 10's data-quality and bias-testing duties. And the underlying data itself is constrained by the Privacy Act of 1974's restrictions on interagency sharing and by Section 508 accessibility requirements that commercial vendors rarely design for from day one. The result is that government AI training data pipelines look less like a data-licensing deal and more like a compliance program with a dataset attached, and the agencies and vendors that treat it that way are the ones whose systems survive the audit, the FOIA request, and the inevitable court challenge.

On this page ▾

Why government AI training data is a different problem than commercial AI training data

A benefits-eligibility model and a customer-churn model can be built from structurally similar data: structured records, historical outcomes, demographic fields. What separates them is not the shape of the data but everything wrapped around it. A commercial buyer licenses a dataset, checks provenance and consent, and ships a model. A government agency (or a vendor selling to one) has to clear the same provenance bar, then add a cloud-authorization requirement, a procurement framework with its own risk-management mandates, a due-process standard that commercial products never face, and, if the deployment touches EU residents, a statutory presumption that the use case is high-risk before a single training run happens.

That stack is why public-sector AI adoption keeps growing while public-sector AI training data supply stays thin. The Government Accountability Office found that AI use cases reported across 11 federal agencies nearly doubled from 571 in 2023 to 1,110 in 2024, with generative AI use cases growing ninefold, from 32 to 282, over the same period. [2] Demand is real and accelerating. Supply is constrained by exactly the compliance requirements this article walks through, and by a documented history of failures that makes agencies, courts, and the public rightly cautious about what data trains a system that can deny someone a benefit, flag them as high-risk, or trigger an investigation.

The real use cases and what data they actually need

"Government AI" spans a wider range of data needs than any single commercial vertical, because agencies run everything from citizen-facing chatbots to classified intelligence analysis. The common thread is that almost every use case below touches personally identifiable information, a protected legal process, or both.

  • Benefits and eligibility determination. Unemployment insurance, SNAP, Medicaid, and housing assistance systems need historical application data, income and household-composition records, and documented outcome labels (approved, denied, appealed, overturned) to train eligibility or triage models. This is the single most litigated category of government AI, covered in the worked example below.
  • Fraud detection in public programs. Distinct from eligibility determination: these models look for anomalous claim patterns, duplicate identities, and transaction sequences across unemployment insurance, tax refunds, and procurement payments. They need labeled historical fraud cases, but mislabeled or algorithm-only labels are exactly what produced the MiDAS failure described below.
  • Predictive policing and criminal justice risk assessment. Recidivism scoring, pretrial risk tools, and location-based patrol allocation systems train on arrest records, criminal histories, and case outcomes. This category carries the most documented bias controversy of any government AI use case, discussed in detail below.
  • Public health surveillance. Outbreak detection, syndromic surveillance, and resource-allocation models need de-identified clinical and lab data, often pooled across jurisdictions, which runs directly into the interagency-sharing restrictions the Privacy Act imposes.
  • Defense and intelligence applications. Signal classification, satellite imagery analysis, and threat-pattern detection typically use classified or controlled-unclassified data that never enters a commercial marketplace at all; sourcing runs through cleared contractors under separate authorities, not the FedRAMP/OMB path this article focuses on.
  • Citizen-facing chatbots and case management. Benefits-navigation assistants and case-status tools need historical support transcripts, FAQ resolution data, and case-note text, but any public-facing interface is squarely inside Section 508's accessibility scope. [4]
  • Procurement and contracting analysis. Tools that flag anomalous bids, score vendor risk, or summarize proposals train on historical solicitation and award data, most of which is already public under federal acquisition regulations, making this one of the easier categories to source cleanly.
  • Infrastructure and utility monitoring. Predictive maintenance for water systems, grids, and transit uses sensor and inspection data that is less personally identifiable but often subject to critical-infrastructure security restrictions that limit who can even see the training set.
The pattern across every use case
Every category above needs a labeled outcome, whether that's a benefits decision, a fraud finding, or a recidivism result, and in government contexts that label was very often produced by a prior human process with its own error rate, appeal history, and demographic pattern. Training on those labels without auditing them is how a biased or error-prone human process becomes a biased or error-prone automated one, just faster and at greater scale.

FedRAMP and the OMB procurement rules: the gate before the gate

Before an agency can train, fine-tune, or even run inference against any cloud-hosted AI service, the underlying cloud offering generally needs FedRAMP authorization, the government-wide risk-assessment program that certifies a cloud service's security controls at the Low, Moderate, or High impact level. This is true whether the agency is buying an off-the-shelf model or hosting its own fine-tuned system on a cloud platform. In August 2025, FedRAMP and the General Services Administration launched an AI Prioritization Initiative to fast-track authorization for a narrow class of enterprise conversational AI tools meeting specific criteria: enterprise features like single sign-on and role-based access control, guarantees against unauthorized use of agency data for model training, demonstrated demand from at least five CFO Act agencies, availability through a GSA Multiple Award Schedule, and the ability to reach FedRAMP 20x authorization within two months. [2]

That initiative closed to new entrants in April 2026, and by early 2026 it had authorized exactly three vendors: OpenAI's ChatGPT Enterprise and API Platform, Google's Gemini for Government, and Perplexity Enterprise Pro for Government. [2] That is a meaningful signal for anyone selling into this market: the fast-track path was narrow, closed, and covered general-purpose conversational tools, not custom fraud-detection models, benefits-eligibility systems, or fine-tuned domain models built on agency-specific training data. Most of the systems described in the use-case list above still have to go through standard FedRAMP authorization or an agency-specific authorization-to-operate process, which routinely takes months rather than weeks.

FedRAMP governs the infrastructure. OMB's procurement memos govern how the AI itself gets bought and managed once the infrastructure clears. On April 3, 2025, OMB issued two memoranda that replaced the prior administration's framework: M-25-21, "Accelerating Federal Use of AI through Innovation, Governance, and Public Trust," and M-25-22, "Driving Efficient Acquisition of Artificial Intelligence in Government." Both rescinded their predecessors, M-24-10 and M-24-18, respectively. [2][4]

M-25-21 requires agencies to identify which of their AI systems are "high-impact," a category that captures most of the use cases in this article: anything whose output serves as a principal basis for a decision or action that could have a legal, material, binding, or otherwise significant effect on rights, safety, or access to critical resources or services. For each high-impact system, agencies must complete seven minimum risk-management practices within 365 days of identifying the system as high-impact, unless a waiver applies: pre-deployment testing, a completed AI impact assessment, ongoing monitoring for performance and adverse impacts, adequate human training on the system, human oversight and intervention capability, consistent remedies or appeal channels for affected individuals, and consultation with end users and the public. [2]

M-25-22 governs how agencies actually contract for AI, applying to solicitations issued after September 30, 2025, and to contract options exercised after October 1, 2025. It directs agencies to build cross-functional acquisition teams, require pre-award performance testing for high-impact AI, and protect government data from being repurposed by the vendor: contractors cannot use non-public agency inputs or outputs to further train AI models offered commercially or publicly, without the agency's explicit consent, which functionally requires vendors to isolate government training data from every other customer's data. [2]

LayerWhat it governsWho clears itKey instrument
Cloud infrastructureSecurity controls of the hosting platformCloud/AI vendor, before any agency can use itFedRAMP authorization (Low/Moderate/High or 20x)
ProcurementContract terms, data rights, performance testingAgency contracting officers and the vendorOMB M-25-22
Deployment riskImpact assessment, monitoring, human oversight, appealsAgency Chief AI Officer and governance boardOMB M-25-21
Underlying training dataProvenance, consent, interagency sharing limitsData supplier and agency privacy officePrivacy Act of 1974, agency-specific statutes

The EU AI Act treats most public-sector AI as high-risk by default

Any vendor or agency whose system touches an EU resident, including a U.S. state agency running a citizen-services chatbot accessible to EU nationals abroad, needs to check Annex III of Regulation (EU) 2024/1689, the EU AI Act. Annex III enumerates the categories of AI systems that Article 6(2) classifies as high-risk, and public-sector use cases dominate the list. [2]

  • Category 5, access to essential services. AI systems used by public authorities to evaluate eligibility for essential public assistance benefits and services, including healthcare, are explicitly high-risk, which places benefits-eligibility and fraud-adjacent triage systems squarely in scope. [4]
  • Category 6, law enforcement. Systems used to assess an individual's risk of offending or reoffending, evaluate evidence reliability, or profile people during criminal investigation and prosecution are high-risk, covering the predictive-policing and recidivism-scoring category discussed next. [4]
  • Category 7, migration, asylum, and border control. Risk assessment of people entering a territory, and processing of asylum, visa, and residence applications by or on behalf of public authorities, are high-risk. [4]
  • Category 8, administration of justice. Systems assisting judicial authorities in researching, interpreting, and applying facts and law to a case are high-risk. [4]

A high-risk classification under Annex III activates Article 10's data-governance duties: training, validation, and testing data has to be relevant, sufficiently representative, and examined for the kind of bias that could affect fundamental rights, with documented mitigation. Dayda's companion article on the EU AI Act's training-data rules covers Article 10's exact requirements and the Digital Omnibus timeline in full; the point for public-sector AI specifically is that the classification is not discretionary. If a system falls into one of these Annex III categories and touches the EU market, high-risk obligations apply regardless of whether the agency or vendor thinks of itself as building a "high-risk AI product."

U.S.-only deployment doesn't guarantee you're outside scope
A state agency's benefits chatbot built for domestic residents can still reach an EU national living or traveling abroad who accesses the same public-facing system. Whether that single fact triggers Annex III depends on the specific placement-on-the-market and deployment facts, but any agency or vendor selling government AI internationally should get a scope determination before assuming U.S. deployment means EU rules don't apply.

Why public-sector training data is unusually hard to source

Four structural factors make government training data scarcer and more expensive to prepare than a comparable commercial dataset, independent of the FedRAMP and OMB compliance layers already covered.

  • 1The Privacy Act restricts exactly the kind of data-pooling that improves model performance. The Privacy Act of 1974, 5 U.S.C. § 552a, generally prohibits an agency from disclosing a record about an individual without that individual's consent, subject to a defined list of exceptions including a published "routine use." [4] That default-closed posture is the opposite of how commercial ML teams pool data across products and business units; a fraud-detection model that would benefit from cross-agency transaction patterns often cannot legally see them without a new or amended System of Records Notice.
  • 2The accountability and explainability bar sits higher than any commercial product faces. A denied loan application gets a customer-service escalation. A denied unemployment benefit, under M-25-21's high-impact framework, requires a documented appeal channel and human-oversight capability by design, not as an afterthought. [4] That requirement reaches back into the training data: a model has to be auditable enough that an agency can explain, defend, and if necessary reverse a specific decision, which pushes public-sector buyers toward heavily annotated, outcome-labeled data over black-box embeddings.
  • 3Data-sharing restrictions between agencies are a feature, not a gap to route around. Health, benefits, tax, and law-enforcement data typically sit in separate statutory silos (HIPAA, the Privacy Act, tax confidentiality under 26 U.S.C. § 6103, and Criminal Justice Information Services policies, respectively) each with its own consent and access rules. A vendor building a cross-program fraud model has to clear each silo independently; there is no single blanket authorization.
  • 4Section 508 accessibility requirements shape what "good" training data looks like for citizen-facing systems. Any public-facing chatbot or case-management interface has to meet WCAG 2.0 Level A and AA accessibility standards under 36 CFR Part 1194. [4] That means training data for citizen-facing systems needs coverage of accessibility-relevant interaction patterns, screen-reader users, alternative input methods, plain-language variants, that commercial support-chatbot datasets rarely bother to capture.

Documented failures worth studying before building the next system

Two cases illustrate what happens when public-sector AI ships without the data-quality and human-oversight rigor the current framework now requires. Both are well-documented, both ended in litigation, and both are frequently cited in policy debates precisely because the facts are not in dispute.

Michigan's MiDAS unemployment fraud system

Michigan's Integrated Data Automated System, deployed to detect unemployment insurance fraud, adjudicated fraud determinations without meaningful human review from October 2013 through September 2015. It issued roughly 40,000 algorithm-only fraud findings during that window, and reviews later found error rates near 85% among cases examined, false positives that triggered wage garnishment, seizure of tax refunds, and quadrupled penalties for people who had not committed fraud. [2] The state faced a class action, Bauserman v. Michigan Unemployment Insurance Agency, alleging the automated determinations violated due-process rights by imposing severe financial penalties without adequate notice or a meaningful opportunity to contest the finding before enforcement. Michigan's Court of Claims gave final approval to a $20 million settlement on January 30, 2024, roughly a decade after the false determinations began. [4]

The training-data lesson is specific: MiDAS was not undone by a bad model architecture. It was undone by treating an automated match (a discrepancy between reported and verified income data) as a final fraud determination rather than a lead requiring human adjudication, exactly the kind of labeling and human-in-the-loop failure that M-25-21's mandatory human-oversight and appeal-channel requirements now target directly. [2]

COMPAS and the recidivism-scoring controversy

In May 2016, ProPublica published an analysis of COMPAS, a proprietary recidivism-risk-scoring tool used across the criminal justice system, examining scores assigned to roughly 11,757 defendants in Broward County, Florida, against actual two-year reoffense outcomes. [2] The analysis found that Black defendants who did not go on to reoffend were misclassified as high-risk at nearly twice the rate of white defendants: a 44.85% false-positive rate for Black defendants versus 23.45% for white defendants. [4] The pattern reversed for false negatives: white defendants who did reoffend were more often mistakenly scored as low-risk.

Black defendants who did not recidivate over a two-year period were nearly twice as likely to be misclassified as higher risk compared to their white counterparts.
ProPublica: How We Analyzed the COMPAS Recidivism Algorithm

Northpointe (COMPAS's developer, later renamed Equivant) and a number of academic researchers disputed ProPublica's framing, arguing the tool was calibrated fairly by a different, and mathematically incompatible, definition of fairness: similar predicted-risk scores corresponded to similar actual reoffense rates across race, even though error rates diverged. Both sides are right about their own metric; the disagreement is a real, still-unresolved tension in the fairness literature, not a case where one side has been proven wrong. That tension is exactly why Annex III treats criminal-justice risk-assessment tools as presumptively high-risk and why Article 10's bias-examination duty applies to them specifically: representativeness and fairness in this domain cannot be certified by a single metric, and any procurement or data-sourcing process has to document which fairness definition it is optimizing for and why. [2]

A worked example: building a compliant state unemployment-fraud model

Take a state workforce agency building a fraud-detection model to flag suspicious unemployment insurance claims, explicitly designed to avoid a MiDAS-style failure. The stage-by-stage pipeline looks nothing like a commercial fraud-model build, because of the documentation each stage has to produce, not because the modeling techniques differ.

  • 1Privacy Act clearance before data touches the project. The agency's privacy office confirms an existing System of Records Notice covers the intended use, or publishes an amended notice, before any claimant data is compiled for model development. [4] A commercial fraud-model build has no equivalent gate; it starts with a data inventory, not a federal-register filing.
  • 2High-impact classification and OMB scoping. The agency's Chief AI Officer determines whether the system is "high-impact" under M-25-21, which it almost certainly is, since a fraud flag can trigger benefit suspension and financial penalties. That triggers the 365-day clock to complete all seven minimum risk-management practices. [4]
  • 3Historical label audit, not just label collection. Before training on past fraud determinations as ground truth, the team audits a sample of those determinations for the MiDAS failure mode: how many were confirmed by human review versus closed automatically. Determinations that were never human-reviewed get down-weighted or excluded rather than treated as clean positive labels.
  • 4Representativeness and bias testing across claimant demographics. The team documents demographic coverage (age, geography, language, industry sector) and tests for disparate false-positive rates before deployment, the Article 10 standard applied voluntarily even for a domestic-only system, because it is the same test a court or GAO auditor would run after the fact. [4]
  • 5Procurement and vendor data-isolation terms. If any part of the pipeline uses a third-party vendor or cloud AI service, the contract has to include M-25-22's data-protection terms: the vendor cannot use the agency's claimant data to train models it offers to other customers, and the underlying cloud service needs FedRAMP authorization at the appropriate impact level before it can touch claimant PII. [4][6]
  • 6Human-review and appeal design, built before launch. Following M-25-21's mandatory practices, every model flag routes to a human adjudicator before any penalty attaches, and claimants get a documented channel to contest a determination, the exact process gap that produced the MiDAS settlement. [4]
  • 7Ongoing monitoring and a retention/FOIA plan. The agency documents monitoring cadence for model drift and disparate impact, and separately determines what portions of the training data, model documentation, and impact assessment are subject to public records requests, since a fraud model's methodology (though typically not the underlying claimant PII) is often FOIA-discoverable.
The audit trail is the deliverable
In a commercial fraud-model build, the audit trail is a nice-to-have for internal governance. In a government build, the System of Records Notice, the impact assessment, the label audit, the bias test results, and the FedRAMP authorization letter are the actual work product a court, a legislative oversight committee, or a GAO reviewer will ask to see, often years after deployment. Budget for producing and retaining that documentation as its own line item, not as a byproduct of the modeling work.

Commercial AI data sourcing vs. government AI data sourcing

The comparison below isolates what actually changes between a commercial and a government training-data pipeline for a structurally similar model, holding data type roughly constant.

DimensionCommercial AI data sourcingGovernment AI data sourcing
Cloud/infrastructure gateVendor's standard security review (SOC 2, ISO 27001)FedRAMP authorization at Low/Moderate/High or 20x before any agency use [2]
Procurement pathStandard commercial contract, negotiated bilaterallyOMB M-25-22 acquisition rules: cross-functional review, pre-award testing for high-impact systems [2]
Risk classificationInternal model-risk framework, if any, set by the company itselfMandatory "high-impact" designation under M-25-21 for rights- or safety-affecting systems, with a 365-day compliance clock [2]
Cross-source data poolingGenerally permitted subject to contract terms and privacy policyRestricted by the Privacy Act's no-disclosure default and routine-use exception; often needs a new or amended SORN [2]
Accessibility requirementBest practice, rarely a legal mandateMandatory WCAG 2.0 AA compliance under Section 508 for any public-facing system [2]
EU high-risk exposureCase-by-case under Article 6 and Annex I/IIIPresumptively high-risk by default for benefits, law enforcement, migration, and justice use cases under Annex III [2]
Public disclosure exposureTrade-secret protected; rarely disclosedModel documentation and impact assessments are frequently FOIA-discoverable
Typical time from data-ready to deploymentWeeks to a few monthsOften 12–24+ months once FedRAMP authorization, impact assessment, and procurement cycles stack
Audit requirement after deploymentInternal, discretionaryStatutory: GAO review, agency Inspector General audits, and potential court discovery in due-process litigation

The rule of thumb for anyone building or selling into this market

Government and public-sector AI does not need exotic data. It needs the same eligibility records, transaction histories, case notes, and outcome labels any comparable commercial system would use, prepared to a standard that survives an audit rather than a demo. The compliance stack, FedRAMP for infrastructure, OMB's M-25-21 and M-25-22 for procurement and risk management, the Privacy Act for how personal data moves between systems, Section 508 for citizen-facing interfaces, and the EU AI Act's Annex III for anything with EU reach, is not paperwork bolted onto a good model. It is the direct institutional response to real, well-documented failures: MiDAS's 40,000 unreviewed fraud findings and COMPAS's racially disparate error rates are not hypothetical cautionary tales. They are the reason the current rules exist in their current form.

For a vendor or agency sourcing training data for this vertical, the practical takeaway is to treat documentation as inseparable from the dataset. A clean, well-labeled dataset without a Privacy Act clearance, a bias audit, and a plan for the FOIA request that will eventually come is not a compliant public-sector asset. It is a liability waiting for the first denied claimant to ask how the decision was made.

Sourcing training data for a government or public-sector AI system?

Get a free, no-obligation provenance and compliance review from the team that brokers AI training data deals every day. We'll flag what an agency's procurement and privacy office will ask for before they ask.

See how buying works on Dayda