Vertical AI

What Training Data Powers Legal AI?

14 min read · 2026-09-10
17%–33%hallucination rate of leading AI legal research tools in a preregistered evaluation
$5,000sanction in Mata v. Avianca, the first major AI fabricated-citation case
$15,500fine against lead counsel in an Oregon federal AI-citation sanctions order
6 rulesof the California Rules of Professional Conduct proposed for AI amendment in March 2026
Key takeaways
  • 1Legal AI splits into four applications with four different data bottlenecks: contract review needs clause-level risk annotations, legal research needs citation-grounded authority, e-discovery needs coded review decisions, and litigation prediction needs docket outcomes. No single corpus serves all four.
  • 2Purpose-built legal research tools from LexisNexis and Thomson Reuters still produced hallucinated or misgrounded answers between 17% and 33% of the time in Stanford RegLab's preregistered evaluation, later published in the Journal of Empirical Legal Studies, so retrieval architecture alone does not fix grounding. The underlying corpus does.
  • 3Attorney-client privilege and ABA Model Rule 1.6 make legal data harder to license than healthcare or finance data, because the person who must consent is the client, not the firm holding the file, and ABA Formal Opinion 512 (July 29, 2024) states that boilerplate waivers do not constitute informed consent.
  • 4Courts have escalated sanctions for AI-fabricated citations from the $5,000 imposed in Mata v. Avianca in June 2023 to a $15,500 fine plus dismissal with prejudice in an Oregon federal case in December 2025 and $30,000 against two attorneys at the Sixth Circuit.
  • 5In a worked evaluation of a 1.2 million contract corpus, only about 21,500 documents carried attorney-authored clause-level risk annotations. That 1.8% slice, not the headline volume, is what determines the dataset's price.
The short version

Legal AI runs on four distinct data types, and the useful ones are almost never public. Contract review and redlining models need executed agreements with clause-level risk and fallback-position annotations written by lawyers. Legal research models need citation-grounded authority plus the briefs and memos that show how authority gets applied. E-discovery models need coded review decisions and privilege calls, not just raw document dumps. Litigation prediction needs dockets tied to outcomes. What makes this vertical harder than insurance or healthcare is not volume: it is that the two most valuable assets, attorney-annotated contract work and privilege-reviewed discovery, sit behind attorney-client privilege and the work-product doctrine, so the party who can authorize a training license is the client rather than the firm or vendor holding the data. ABA Formal Opinion 512 confirms that consent to feed client information into a self-learning AI tool must be specific and informed rather than boilerplate. The cost of getting the grounding wrong is now measurable in the sanctions record, which runs from the $5,000 penalty in Mata v. Avianca in 2023 to $15,500 plus dismissal with prejudice in an Oregon federal case in December 2025. This guide covers each data type, the professional-responsibility gate attached to it, current state bar and UPL developments as of 2026, a worked evaluation of a real-shaped contract and discovery corpus, and the value drivers that set price.

On this page ▾

A general-purpose model writes fluent legal English. It will name a doctrine, structure an argument, and produce a clause that reads like something a lawyer wrote. What it cannot do without grounding is tell you whether the authority it cited exists, whether an indemnity carve-out is aggressive for this counterparty in this industry, or whether a document in a production set is privileged. Those judgments live in work product, not in web text, and a model that has only seen web text will produce a confident answer with no reliable relationship to the underlying law.

The clearest evidence that this is a data problem rather than a model problem comes from tools built specifically to solve it. Stanford's RegLab ran the first preregistered empirical evaluation of commercial AI legal research products and found that LexisNexis's Lexis+ AI and Thomson Reuters's Westlaw AI-Assisted Research produced hallucinated or misgrounded answers between 17% and 33% of the time, despite both vendors marketing retrieval-augmented generation as the fix for hallucination. [4] Retrieval on top of a thin or badly structured corpus still fails. What separates a legal model that can be trusted from one that cannot is the depth and fidelity of the legal work product behind it.

Legal AI breaks into four applications, and each is bottlenecked by a different proprietary artifact:

  • Contract review and redlining. Extract obligations, flag risky clauses, propose fallback language, and mark a draft against a negotiation playbook. Needs executed agreements with clause-level typing and attorney risk positions.
  • Legal research and drafting. Find controlling authority, summarize a line of cases, and draft a brief with citations that resolve. Needs case law plus the briefs and memos that show how practitioners actually deploy authority.
  • E-discovery and document review. Classify responsiveness, cluster issues, and predict privilege across millions of documents. Needs coded review decisions, not raw ESI.
  • Litigation prediction and analytics. Estimate outcome, timeline, and damages exposure by court, judge, claim type, and counterparty. Needs dockets linked to dispositions.
The common thread
In every one of these, the training signal is a human lawyer's judgment recorded against a specific document: this clause is unacceptable, this case controls, this email is privileged, this motion was granted. The documents may be plentiful. The recorded judgment attached to them is not.

The data types that actually power legal AI

Legal text is abundant. Legal text with a usable label is rare. Ranked roughly by how much they move model performance relative to how hard they are to obtain, these are the assets buyers ask for.

  • Contracts with clause-level annotations. Executed agreements where each clause is segmented and typed (indemnification, limitation of liability, assignment, change of control, IP ownership, termination for convenience) and tagged with a risk rating, a market-standard comparison, or an approve/escalate decision. This is the single most valuable asset in the vertical.
  • Negotiation and redline histories. The full version chain from first draft through executed final, showing which positions were asked for, which were conceded, and in what order. This is the only data that teaches a model what a realistic counter looks like rather than what an ideal clause looks like.
  • Case law, briefs, and motions. Opinions are largely public. The briefs, memoranda, and motions that apply them, plus the court's disposition of each, are what turn a citation index into reasoning data.
  • Discovery document sets with review coding. Produced ESI paired with the responsiveness calls, issue codes, and confidence levels assigned by human reviewers across a matter.
  • Privilege logs and privilege-review decisions. Documents withheld or redacted, with the asserted basis (attorney-client, work product, common interest) recorded per document. This is the supervised signal privilege-detection models need and the hardest legal data in existence to source.
  • Attorney work product. Research memos, deal issue lists, deposition outlines, and internal analyses. Highest reasoning density per token of anything in the vertical, and the most legally encumbered.
  • Court dockets and outcome data. Filing-level event streams tied to dispositions, damages, timelines, judge, venue, and counsel, which is what litigation-prediction models are actually fit on.

Redline histories deserve a specific note, because buyers routinely underrate them. A corpus of executed contracts teaches a model the distribution of final language. Only the version chain teaches it the distribution of positions: that a mutual cap at 12 months of fees was the opening ask, that the counterparty countered at 24 months with a carve-out for data breach, and that the deal closed at 18 months. A redlining product that cannot model the counter is a template generator, not a negotiation assistant.

Why privilege makes legal data harder to license than any other vertical's

Every vertical claims its data is scarce. Legal data is scarce for a structurally different reason, and the difference decides whether a deal is possible at all. In healthcare, HIPAA sets a de-identification standard, and data that clears it can move. In insurance, PII scrubbing plus documented consent clears most of the path. In legal, the two most valuable assets are governed by protections that survive de-identification and that the holder of the data cannot waive on its own.

Attorney-client privilege belongs to the client, not the firm

The privilege attaches to confidential communications between lawyer and client made for the purpose of legal advice. The consequential detail for data licensing is who owns it: the client does. A law firm sitting on 40 years of matter files cannot license those files for AI training on its own authority, because the firm is a custodian rather than the holder of the right being waived. Getting a valid license means going back to every client whose matter appears in the corpus, per matter, which is why firm-side legal data almost never reaches the open market intact.

Waiver risk compounds this. Under Federal Rule of Evidence 502, an inadvertent disclosure escapes waiver only if the holder took reasonable steps to prevent it and promptly took reasonable steps to rectify the error, and an intentional disclosure can extend waiver to undisclosed communications on the same subject matter where fairness requires. [2] A firm that ships privileged material into a training pipeline is not just breaching a duty. It is creating an argument that the client's privilege over an entire subject matter has been waived in ongoing or future litigation, a harm that lands on the client rather than on the firm that sold the data.

The work-product doctrine covers exactly the annotations buyers want

Federal Rule of Civil Procedure 26(b)(3) shields documents prepared in anticipation of litigation or for trial, and when a court does order production of such materials, it "must protect against disclosure of the mental impressions, conclusions, opinions, or legal theories of a party's attorney or other representative concerning the litigation." [2] That language describes the training label. A privilege log entry, an issue code, a risk rating on an indemnity clause, and a redline comment explaining why a position is unacceptable are all attorney mental impressions recorded against a document. The annotation layer that makes legal data valuable for AI is precisely the layer the doctrine protects most strongly.

Rule 1.6 makes consent specific, per client, and non-boilerplate

ABA Model Rule 1.6 bars a lawyer from revealing information relating to the representation of a client absent informed consent, and it covers all such information regardless of source, not only privileged communications. ABA Formal Opinion 512, issued July 29, 2024, applied this directly to generative AI and reached two conclusions that govern legal data licensing. First, because many self-learning AI tools are designed such that their output could lead directly or indirectly to disclosure of client-representation information, a client's informed consent is required before that information is input into such a tool. Second, and decisively for anyone drafting a data agreement, boilerplate waivers do not constitute informed consent. [4]

California is moving to write this into binding rules rather than guidance. Following an August 22, 2025 directive from the California Supreme Court, the State Bar's Standing Committee on Professional Responsibility and Conduct approved proposed amendments to six rules on March 13, 2026, including a change to Rule 1.6 that would define "reveal" to include exposing confidential information to AI systems where that exposure creates a material risk of misuse. [4] If adopted, feeding client material into a training corpus becomes a revelation under the rule itself, not an analogy to one.

What this means in practice
The cleanest legal training data does not come from law firms. It comes from parties who hold legal documents outside a representation: contract lifecycle management vendors holding executed agreements under customer terms, corporate legal departments licensing their own contracts (the company is the client and can consent for itself), document-review vendors holding produced, non-privileged ESI under a protective order that permits residual use, and public court records. Any seller pitching firm-side matter files should expect the first question to be which client authorized this, and to have an answer per matter.

Unauthorized practice of law: what a legal dataset can and should be used to build

Unauthorized practice of law rules restrict who may give legal advice, and they matter to data buyers because they constrain the product a dataset can lawfully support. A model fine-tuned on attorney-annotated contracts is unambiguously fine as a tool used by a lawyer. The same model shipped as a consumer-facing product that tells a small business owner whether to sign is where the exposure starts.

The state of the rules as of 2026 is genuinely unsettled, and buyers should treat anyone claiming otherwise with suspicion. The Thomson Reuters Institute's survey of the area identifies the core problem plainly: there is no uniform definition of the practice of law across US jurisdictions, which leaves both courts and providers without a workable test. It maps three live regulatory paths, explicit enablement (amending UPL statutes to permit purpose-built legal AI under conditions), regulatory sandboxes (Utah's Standing Order No. 15 being the leading example, evaluated against whether a tool beats the nothing most people currently have), and narrowing UPL to human conduct so that tools fall outside it entirely. It also names the unresolved "personhood gap": whether an AI system or its corporate provider bears responsibility for a UPL violation. [2]

Separately, the lawyer-side rules are tightening fast. ABA Formal Opinion 512 addresses supervision of AI output under Model Rules 5.1 and 5.3, treating AI work product as something a lawyer must review rather than adopt. [2] California's proposed amendments would add an express duty to independently review and verify technology-generated output under Rule 1.1, and would clarify under Rule 3.3 that lawyers must verify the accuracy and existence of every cited authority, including AI-generated ones. [4]

For a dataset seller, three practical consequences follow. Restrict the license by application rather than by industry, since a corpus licensed for "legal AI" is under-specified and a corpus licensed for "attorney-supervised contract review tooling" is not. Expect buyers building consumer-facing products to require heavier warranties, because their UPL exposure is real and they will try to push it upstream. And keep the annotation provenance auditable, because a buyer whose model is challenged will need to show the labels came from licensed attorneys working in a defined jurisdiction.

What ungrounded legal AI costs: the sanctions record

The strongest argument for real, verifiable legal training data over synthetic approximation is not theoretical. It is a growing line of sanctions orders, and the amounts are rising.

Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023) is the origin point. Roberto Mata sued Avianca over an in-flight injury; his counsel filed an opposition brief citing six federal decisions that opposing counsel could not locate because ChatGPT had invented them, including the fictitious "Varghese v. China Southern Airlines" attributed to the Eleventh Circuit. Judge P. Kevin Castel issued a $5,000 sanction on June 22, 2023 against attorneys Peter LoDuca and Steven A. Schwartz and their firm, Levidow, Levidow & Oberman P.C., and ordered corrective letters sent to the real judges whose names had been attached to the fabricated opinions. [6]

Courts in 2025 and 2026 stopped treating this as a novel error. In Couvrette v. Wisnovsky, 2025 WL 4109655 (D. Or. Dec. 12, 2025), the court found 15 fabricated citations and 8 fabricated quotations across three briefs. It struck the briefs without leave to refile, dismissed the plaintiffs' claims with prejudice, fined lead counsel $15,500, and awarded the defendants their fees for both the summary judgment and sanctions briefing. [6] The ABA Journal has documented the same escalation elsewhere: a $30,000 total sanction at the Sixth Circuit against two attorneys for more than two dozen fabricated citations in a case the court called almost entirely frivolous, and a $7,500 sanction in the Southern District of Ohio accompanied by a contempt finding and a referral to the state's disciplinary counsel. [12]

The mechanics of detection are worth understanding, because they explain why partial grounding is worse than none. In Green Building Initiative, Inc. v. Peacock, No. 3:24-cv-00298-SI (D. Or. Oct. 27, 2025), Judge Michael H. Simon examined two authorities cited in a fee brief. One was invented outright. The other was what the court called "almost real": there is a real case named Page v. Parsons discussing the Oregon anti-SLAPP statute, but it is an Oregon Court of Appeals decision at 249 Or. App. 445 (2012), not the federal District of Oregon decision at 249 F. Supp. 3d 998 that the brief cited, and it does not address the proposition it was cited for. [4]

Stell, however, is totally fake. There is no case at 2022 WL 1696093.
U.S. District Court for the District of Oregon (via CourtListener RECAP): Green Building Initiative, Inc. v. Peacock, No. 3:24-cv-00298-SI, Opinion and Order and Order to Show Cause (D. Or. Oct. 27, 2025)

That "almost real" failure mode is the one synthetic legal data cannot fix and can actively worsen. A model trained on generated contract language that approximates market clauses will produce plausible clauses. A model trained on generated case law will produce citations that have the right reporter format, a real-sounding party name, a plausible pincite, and no existence. Verifiability is the governing property of legal training data, and it can only come from real documents whose citations resolve to real authority.

Data type, AI use case, and the professional-responsibility issue that governs it

Data typePrimary AI use caseKey regulatory / professional-responsibility consideration
Contracts with clause-level risk annotationsContract review, obligation extraction, risk scoringRule 1.6 confidentiality if drafted inside a representation; for CLM vendors, whether customer terms actually grant model-training rights [2]
Negotiation and redline historiesRedlining, fallback-position and playbook modelingWork product: redline comments embody counsel's mental impressions and conclusions [2]
Case law, briefs, and motionsLegal research, brief drafting, citation groundingFilings are largely public, but sealed exhibits and court terms of use limit bulk reuse; Rule 3.3 candor obligations sit downstream [2]
Produced discovery documents (ESI)Responsiveness classification, technology-assisted reviewProtective orders typically bar use outside the matter; Rule 502(d) orders define the waiver posture [2]
Privilege logs and privilege-review decisionsPrivilege detection and second-pass review modelsThe privilege belongs to the client, not the firm; log entries can themselves reveal work product [2][4]
Attorney work product (memos, issue lists)Legal reasoning fine-tuning, preference-pair dataRule 1.6 survives the engagement; informed consent must be specific rather than boilerplate [2]
Court dockets with outcome dataLitigation prediction, judge and venue analyticsDocket terms of use and sealed-record handling; UPL exposure if outcome predictions are sold as advice to non-lawyers [2]

Worked example: evaluating a contract and discovery corpus

Consider a contract lifecycle management vendor winding down. It holds 1.2 million executed commercial agreements uploaded by 3,400 customer accounts over six years, and it separately acquired a document-review services arm that retains 8.4 million documents across 60 completed litigation matters. The headline numbers look enormous. Here is how the evaluation actually runs.

  • 1Start with the licensing gate, not the data. For the contract side, read the customer terms of service and data processing agreements in force when each account uploaded. If only 62% of accounts were on terms granting the vendor rights to use customer content for product improvement and model development, the licensable population is about 744,000 contracts, not 1.2 million, before a single quality question gets asked. Accounts on the older terms are unlicensable without individual outreach.
  • 2Segment by annotation depth. Of the licensable contracts, how many have machine-readable text with clean party, date, and governing-law metadata? Assume most do. How many were run through clause segmentation and typing? Assume 410,000. How many carry an attorney-authored risk rating or approve/escalate decision at the clause level? That number is usually one to three percent of the corpus. At 21,500, it is the subset a contract-AI buyer is actually paying for.
  • 3Isolate the redline chains separately. Contracts where the full version history survived, first draft through execution, support negotiation modeling that the executed-only population cannot. If 96,000 agreements retain a complete chain, that is a distinct product line, priced on its own, not a footnote on the contract count.
  • 4Treat the discovery corpus as guilty until proven clean. Of 8.4 million documents, the only portion that can move is material that was produced (not withheld), is covered by a protective order that permits residual or de-identified use, and belongs to a client who has consented. Documents on privilege logs and their associated basis codes are the most valuable slice for privilege modeling and typically the one that cannot be licensed at all, because the log is itself a record of counsel's mental impressions. [4]
  • 5Value the review coding, not the documents. Where the protective order and client consent do allow use, the asset is the second-pass responsiveness and issue coding attached to each document, since raw ESI without coding is undifferentiated corporate email. If 2.1 million documents carry full review coding but only 640,000 sit inside matters with clean consent and a permissive protective order, 640,000 is the number.
  • 6Check jurisdiction and recency. Contracts executed after 2023 carry AI, data-use, and model-training clauses that older agreements do not, which makes them disproportionately valuable to buyers building tools that need to reason about current market terms. Multi-jurisdiction corpora also need a governing-law breakdown, since a model trained overwhelmingly on New York and Delaware agreements will underperform on English or Singapore law deals.
SubsetVolumePrimary buyer use case
Contracts under terms permitting model-training use~744,000Baseline corpus for clause extraction and search
Contracts with clause segmentation and typing410,000Obligation extraction, clause classification
Contracts with attorney clause-level risk annotations21,500Risk scoring, review triage, playbook automation
Contracts with complete redline version chains96,000Negotiation modeling, fallback-position generation
Discovery documents with review coding and clean consent640,000Responsiveness classification, TAR model training
Privilege log entries with basis codesNot licensableWould support privilege detection; blocked by work-product protection [2]
The lesson
A 9.6 million document headline collapses to roughly 21,500 documents of genuinely differentiated, attorney-annotated training signal plus two mid-sized supporting corpora. That is a correctly scoped dataset, and scoping it honestly before a buyer does is what gets a deal closed rather than repriced in diligence.

How legal training data gets priced

Legal data prices against the same seven drivers Dayda applies across every vertical: volume, quality and annotation depth, domain scarcity, metadata richness, recency, licensing terms, and legal cleanliness. Three of them do most of the work here, and one behaves differently in legal than anywhere else.

Annotation depth dominates. The gap between a segmented contract and an attorney-annotated contract is the gap between a search index and a training set. Clause typing is increasingly automatable, which erodes its premium. A licensed attorney's judgment that a specific limitation-of-liability construction is off-market for this deal size is not automatable, which is exactly why it holds value. The worked example above shows the practical shape of this: the annotated 1.8% carries more of the price than the other 98% combined.

Legal cleanliness is a gate rather than a multiplier, and the gate is higher here. In most verticals, a buyer with a compliance function wants documented consent and clean provenance before purchase. In legal, a defect does not just create regulatory exposure for the buyer, it can destroy the seller's clients' privilege in unrelated litigation and expose the seller's lawyers to discipline under Rule 1.6. [4] Serious buyers will not proceed without per-source documentation of who held the data, under what terms, which party consented, and whether any of it originated inside a representation. Sellers who arrive with that documentation assembled do not get a bonus. They get to have the conversation at all.

Recency compounds faster than in other verticals. Contract language turns over as market terms shift, and agreements executed since 2023 contain AI use restrictions, data-training carve-outs, and model-output warranties that simply did not exist before. A case law corpus that stops in 2022 will not ground a research assistant on the questions practitioners are asking now. Volume from a decade ago is worth materially less per document than volume from last year.

Licensing structure moves price the way it does everywhere else. A non-exclusive license lets a CLM vendor or legal-tech company monetize the same corpus across multiple buyers, while an exclusive license removes a distinctive annotated corpus from the market and commands a substantial premium over the non-exclusive baseline. Legal adds one wrinkle: exclusivity is more defensible here than in most verticals, because the scarcity is real. A buyer who locks up the only sizable attorney-annotated redline corpus on the market has acquired something a competitor cannot readily replicate by spending more on scraping.

The annotation layer is the asset, and it is the layer the rules protect

Legal AI has a supply problem shaped by its own professional-responsibility rules. The documents that would make legal models reliable exist in enormous quantity inside firms, corporate legal departments, CLM platforms, and review vendors. The layer that makes them trainable, an attorney's recorded judgment about risk, relevance, position, and privilege, is the same layer that Rule 1.6 and the work-product doctrine protect most tightly. That tension is not going to resolve in favor of looser rules; California's proposed amendments point the other direction. [2]

What does resolve it, deal by deal, is sourcing from parties who can actually consent and documenting that authority before a buyer asks. A company licensing its own contract portfolio consents for itself. A CLM vendor whose customer terms grant model-development rights can license what those terms cover. A review vendor operating under a protective order that permits residual use, with the client's sign-off, can license the coded output. Everything else needs per-matter client consent that is specific rather than boilerplate. [2] The sanctions record makes clear what buyers are paying to avoid: a filing struck from the docket, a claim dismissed with prejudice, and a five-figure fine attached to a lawyer's name. [4][6]

Get your legal data evaluated.

Dayda brokers vetted, NDA-gated contract, docket, and review datasets between legal-tech companies, corporate legal departments, and the AI labs and enterprises that need them, with provenance, consent authority, and privilege exposure checked before any deal.

List your data on Dayda