AI Data Supply Chain

Are We Running Out of AI Training Data?

13 min read · 2026-06-12
300Testimated usable stock of high-quality, human-generated public text (90% CI: 100T–1,000T tokens)
2026–2032Epoch AI's projected window for exhausting that stock at current scaling trends
~4 epochspoint at which repeating the same training data stops adding meaningful value
8 monthshow fast training-dataset sizes have been doubling
Key takeaways
  • 1Epoch AI estimates the usable stock of high-quality, human-generated public text at roughly 300 trillion tokens and projects frontier labs will fully use it sometime between 2026 and 2032. That range already moved once: the group's original 2022 analysis put the shortfall at 2024, before better data filtering and multi-epoch training pushed the estimate out by years.
  • 2The debate is not settled. OpenAI cofounder Ilya Sutskever told a NeurIPS audience in December 2024 that pre-training as the field knows it will end because "we have but one internet," and Elon Musk agreed weeks later that human-generated training data was already exhausted. Epoch's own 90% confidence interval spans a tenfold range, 100 trillion to 1,000 trillion tokens, which is itself evidence that credible researchers still disagree by years on the actual date.
  • 3The real constraint is quality-adjusted, deduplicated tokens, not the raw size of the web. Epoch's own methodology revision found that filtering web data carefully can match curated corpora and that models can train on the same data for several epochs before returns flatten, and a separate 2023 study found repeating data up to roughly 4 epochs barely changes model loss versus using unique tokens.
  • 4The industry is responding on five fronts at once: synthetic data generation, direct licensing deals with publishers and platforms, expansion into image, video, and audio data, proprietary and private data acquisition, and squeezing more value out of existing data through repeated training epochs and reinforcement learning. No single fix carries the load.
  • 5Scarcity is sharpest for public, human-generated text used in frontier-scale pretraining specifically. Other modalities have more raw headroom by Epoch's own account, and as public text keeps depleting, freshly generated, proprietary, and licensable data becomes the scarcer input, which raises its price and negotiating leverage for whoever holds it.
The short version

The honest answer is not yet, but the easy, free version of the data supply is close to gone. Epoch AI, the research group that first modeled this question, estimates the usable stock of high-quality human-generated public text at roughly 300 trillion tokens and projects that frontier labs will fully use it somewhere between 2026 and 2032. That window has already shifted once: the group's original 2022 analysis put the shortfall at 2024, before two methodology updates, better web-data filtering and multi-epoch training, pushed the estimate out by years. OpenAI cofounder Ilya Sutskever told a NeurIPS audience in December 2024 that pre-training as the field knows it will end because the internet, unlike compute, is not growing, and Elon Musk agreed a few weeks later that the industry had already burned through the cumulative sum of human text. Neither claim is universally accepted; Epoch's own 90% confidence interval spans a tenfold range, and its overtraining scenarios alone move the exhaustion date across seven years. What is not in dispute is where the real constraint sits: not the raw size of the web, but the supply of deduplicated, high-quality, non-repeated tokens. That is why the industry's actual response spans five separate levers, synthetic data, publisher and platform licensing, multimodal expansion, proprietary data acquisition, and squeezing more epochs out of what has already been collected, rather than any single fix. For a company sitting on real, freshly generated proprietary data, every one of those levers points the same direction: rising leverage over price and terms, not urgency to give the data away.

On this page ▾

What does "peak data" actually mean?

"Are we running out of AI training data?" has become one of the most-searched and most-argued questions in the field, and most answers to it collapse into a headline instead of a real number. The serious version of the question is narrower and more useful: is the supply of public, human-generated text usable for frontier-scale pretraining running out, on what timeline, and does that even determine whether models keep improving? Each of those three sub-questions has a different, better-supported answer, and conflating them is exactly what makes this debate feel louder and less resolved than it is.

The group that put real numbers behind the question is Epoch AI, a research nonprofit that has modeled the size of the usable public text corpus since 2022 and revised its own estimate since. Its current, most-cited figure: the effective stock of quality- and repetition-adjusted human-generated public text sits at roughly 300 trillion tokens, with a wide 90% confidence interval of 100 trillion to 1,000 trillion. On present scaling trends, the group projects an 80% chance that stock is fully used somewhere between 2026 and 2032. [10]

Why the range matters more than the midpoint
A tenfold uncertainty band on the underlying stock, and a seven-year window on the exhaustion date, is not a rounding error. It means the honest state of this research is genuine uncertainty, not a settled countdown clock. Anyone citing a single confident year is oversimplifying Epoch's own numbers.

The core estimate, and how it already moved once

The 2026-2032 window is not the number Epoch started with. Its original 2022 analysis, published as the same paper later revised, projected that high-quality text data would be fully utilized by 2024. [4] That earlier, tighter estimate is worth sitting with, because it is exactly the kind of confident near-term prediction that gets quoted as fact and then quietly falls out of date as the underlying research improves.

Two findings pushed the timeline out. First, the researchers found that web data, filtered carefully rather than discarded as noisy, can perform as well as manually curated corpora, which expanded their usable high-quality data estimate roughly 5x since web sources dwarf curated ones in volume. Second, they found models can be trained on the same data across several epochs without meaningful quality loss, which multiplies effective data availability by roughly 2-5x on top of that. [8] Combined, those two revisions are the entire reason the projected shortfall moved from 2024 to a 2026-2032 range, a shift of two to eight years driven by better methodology, not by the web suddenly growing faster.

Demand is also part of the equation, and it isn't standing still. Stanford's 2025 AI Index reports that training-dataset sizes have been doubling roughly every 8 months, alongside training compute doubling roughly every 5 months. [4] A fixed stock, even a generously revised one, gets consumed faster when the appetite for it keeps compounding at that pace, which is part of why Epoch's window has a 2032 edge as well as a 2026 one.

Overtraining, running a model on far more tokens per parameter than compute-optimal scaling calls for, compresses the timeline again in the other direction. Epoch's model shows that training compute-optimally leaves the stock intact until roughly 2028, moderate 5x overtraining brings that to around 2027, and heavy 100x overtraining would exhaust it by 2025. [2] That range is not hypothetical: Epoch's own researchers note that Llama 3-70B was trained with roughly 10x overtraining, comfortably inside the band that produces earlier exhaustion. [6]

Training regimeWhat it means in practiceEpoch AI's projected exhaustion
Compute-optimal (no overtraining)Token count sized to match compute, the Chinchilla-style approach~2028
Moderate overtraining (5x)More tokens per parameter than compute-optimal, common for models meant to run cheaply at inference~2027
Heavy overtraining (100x)Extreme token-per-parameter ratios, rare but not unheard of~2025
Documented real example: Llama 3-70BReported at roughly 10x overtrainingFalls within the 2026-2032 band

Is the debate settled, or just loud?

It is loud, and not settled. At NeurIPS in December 2024, OpenAI cofounder Ilya Sutskever told a packed room that the industry had reached what he called peak data, and framed the shift in blunt terms. [2]

Pre-training as we know it will unquestionably end. Data is the fossil fuel of AI: it was created somehow, and now we use it, and we've achieved peak data and there'll be no more.
Techmeme (reporting on The Verge): Ilya Sutskever says "Pre-training as we know it will end" at NeurIPS 2024

Elon Musk agreed publicly a few weeks later, in a livestreamed conversation carried in January 2025. [2]

We've now exhausted basically the cumulative sum of human knowledge … in AI training. That happened basically last year.
TechCrunch: Elon Musk agrees that we've exhausted AI training data

Both statements read as confident and final. Neither one is a peer-reviewed measurement, and both were made by people with a direct interest in the industry's next chapter, Sutskever running a new AI lab premised partly on moving past pretraining scale, Musk running a competing lab that trains its own frontier models. That doesn't make the underlying observation wrong. It does mean the loudest voices in this debate are advocates as much as they are analysts, and the actual research, Epoch's own numbers, carries far more built-in uncertainty than either quote suggests. A 90% confidence interval spanning 100 trillion to 1,000 trillion tokens, and an exhaustion window that moves by seven years depending on training regime, describes a field that is genuinely unsure, not one counting down to a known date. [2]

What "exhausting the stock" does not mean
Even Epoch's own paper does not claim model progress stops once the public text supply is used up. It names synthetic data generation, domain-specific transfer learning, and better data efficiency as plausible paths past the ceiling. [2] Running out of free, generic pretraining fuel is a real constraint on one input. It is not the same claim as AI capability growth stopping.

The real bottleneck: quality, not raw token count

Treating "the internet" as a single fixed number is the most common error in this debate, and it's exactly what Epoch's own revision disproves. The same underlying web produced a 5x larger usable-data estimate once the researchers filtered it more carefully instead of discarding anything that wasn't already curated. [2] The ceiling was never the size of the web. It was how much of it counted as clean, deduplicated, genuinely informative text, and that number moves with better filtering, not with the web growing.

Reuse compounds the same point. A 2023 study training 400 language models up to 9 billion parameters on as many as 900 billion tokens found that repeating a fixed pool of data for multiple passes barely hurts model quality up to a point, then the value of additional compute decays sharply. [2]

Training with up to 4 epochs of repeated data yields negligible changes to loss compared to having unique data.
Muennighoff et al.: Scaling Data-Constrained Language Models

Past roughly four epochs, the same study found the value of throwing more compute at the same fixed dataset eventually collapses to zero. [2] That is a precise, load-bearing number: it tells a lab exactly how much runway a fixed pool of proprietary or licensed data actually has before more passes over it stop paying off, which is a far more useful planning input than any headline about the internet running dry.

Curation can substitute for scale even more directly. Microsoft Research's Phi-1 model, published in 2023, trained a 1.3-billion-parameter model on just 7 billion tokens, most of it filtered "textbook-quality" web content plus a smaller synthetic set, and reached 50.6% pass@1 on the HumanEval coding benchmark, competitive with far larger models trained on orders of magnitude more raw data. [2] A dataset that is 1/1000th the size of a typical pretraining corpus, chosen for quality instead of volume, produced a genuinely capable model. That is the clearest evidence available that the constraint the industry is actually running into is curated signal, not gigabytes.

What this means for anyone holding a dataset
The scarcity that matters to a buyer is scarcity of clean, deduplicated, well-labeled, non-repeated tokens, not row count. A smaller dataset with clear provenance and genuine novelty will outcompete a larger, messier one in a market that has already learned this lesson from its own research.

How the industry is actually responding to scarcity

No lab is betting on a single fix. The response splits across five distinct levers, each addressing a different piece of the constraint, and each with its own real-world example already in production.

StrategyWhat it actually addressesReal, documented exampleWhere it falls short
Synthetic data generationFills gaps where real data is thin, especially for reasoning and math/code training signalDeepSeek-R1's reasoning ability, trained largely through reinforcement learning on self-generated trajectories rather than additional human-labeled data [2]Reliably improves narrow domains more than open-ended, general knowledge; risks compounding a model's own blind spots if overused
Direct licensing dealsConverts previously free or scraped content into legally clean, high-provenance corporaPublisher and platform licensing agreements now routinely worth tens to hundreds of millions of dollars per dealExpensive and slow to negotiate at scale; doesn't expand the underlying supply, just formalizes access to what already exists
Multimodal expansionShifts training budget toward modalities with more raw headroom than textMarketplaces like Versos AI licensing specific, rights-cleared video clips (down to requests like "10 hours of white wedding balloons") directly from creators [2]Curation and rights-clearance, not raw volume, are the actual bottleneck; a video frame carries far less semantic signal than a paragraph of text
Proprietary and private data acquisitionSecures data a rival cannot simply scrape or re-derive from the public webAI labs increasingly building or buying direct relationships with data holders rather than relying on open scrapesAccess is exclusive to the buyer, not the field, so it doesn't relieve scarcity industry-wide, only for whoever secures the deal
More efficient reuse of existing dataSqueezes more usable training signal out of a fixed pool through repeated epochs and careful curationMuennighoff et al.'s finding that ~4 epochs of repetition costs almost nothing versus unique data [2]; Phi-1's curated-over-massive result [4]Has a hard ceiling; returns to added compute decay to zero past the optimal number of epochs

For the deal economics behind the licensing row specifically, including what OpenAI and Google actually paid for publisher and platform content, see our breakdown of AI data licensing deals. None of those agreements exist because scraping stopped working; they exist because scarcity, legal risk, and the value of clean provenance all moved in the same direction at once.

Synthetic data's real role, and where it stops helping

Sutskever's own NeurIPS remarks did not stop at declaring pretraining over. He named synthetic data and agentic systems as the two most likely paths past peak data. [2] DeepSeek-R1 is the clearest production example of that path taken seriously: rather than requiring more human-labeled reasoning demonstrations, its reasoning capability was trained largely through large-scale reinforcement learning, letting the model generate and refine its own reasoning trajectories, with a documented multistage pipeline layering rejection sampling and fine-tuning on top. [4] That is synthetic data in its most productive form: not text manufactured to pad a corpus, but signal generated by the model's own trial and error, verified against a ground truth (a correct answer, a passing test) that keeps the loop honest.

That distinction, between synthetic data that manufactures volume and synthetic data that's verified against real outcomes, is also where synthetic data's risk lives: unverified model-generated text recycled back into training is the mechanism behind model collapse, where each generation of synthetic data drifts further from real-world distributions. For a full framework on when synthetic data is a safe substitute for real data and when it degrades a model instead, see our guide to synthetic vs. real training data, which covers that decision in depth. The short version relevant here: synthetic data is a genuine, working answer to parts of the scarcity problem, specifically reasoning, math, and code, and not yet a general substitute for the human judgment and outcome data that alignment and open-domain quality still depend on.

Is this a text problem, or does it apply to every modality?

Scarcity, as Epoch measures it, is specifically about public, human-generated text used for frontier-scale pretraining. The same paper is explicit that other modalities sit on different ground: images and video are comparatively easy to generate and collect at volume next to text. [4] That single line is doing a lot of work in this debate. It means the 2026-2032 exhaustion window is not a claim about AI running out of material broadly. It's a claim about one specific, historically dominant input running thin faster than the others.

That doesn't make video and image data free for the taking, though. The constraint there is different in kind: not raw volume, but specificity and clean rights. Forbes profiled Versos AI in May 2026 as a licensing marketplace built precisely around that gap, gathering footage, attaching a C2PA provenance manifest, and paying the original creator, because labs training video models don't need more generic footage, they need narrowly specific, rights-cleared clips: the article's own example is a lab requesting "10 hours of white wedding balloons" or a close-up of a hand tying a shoelace. [4] Video and audio aren't scarce the way text is. They're scarce in the sense that matters to a buyer: pre-cleared, precisely tagged, legally licensable at the level of an individual clip rather than a whole archive.

The practical read for a non-text dataset
If your proprietary data is video, audio, sensor logs, or images rather than text, the peak-data timeline above does not directly apply to you. What does apply is the same lesson from the text side: curated, well-labeled, provenance-clean data commands a real premium over raw, undifferentiated volume, in any modality.

What this means if you're sitting on proprietary data

Put the pieces together and the direction of travel is consistent, even where the exact dates aren't. Public text is being consumed faster than quality-filtering and multi-epoch training can keep expanding the usable stock, real dollar-figure licensing deals for publisher and platform content are already commonplace, and the industry itself is spending real engineering effort on five separate workarounds rather than one. Every one of those facts describes the same underlying shift: freshly generated, proprietary, legally clean data is becoming the scarcer and more negotiable input, while generic scraped text keeps commoditizing.

That is not the same claim as "your data is a permanent competitive moat." Whether a given dataset defends a business long-term depends on whether it keeps updating, sits behind a real access barrier, and survives if the product generating it stops running, a separate and much higher bar covered in our analysis of data moats. Scarcity is a market-price argument, not a defensibility argument. A dataset does not need to be a moat to be worth a real, growing number in a licensing negotiation; it only needs to be data a well-funded buyer cannot simply scrape or synthesize instead.

Three questions turn that general claim into something you can act on for your own dataset.

  • 1Is it text, or another modality? Text-specific scarcity, and the licensing activity it has already produced, means text corpora are appreciating on a faster and more visible timeline than image, video, or sensor data right now, per Epoch's own analysis. [4]
  • 2Is it fresh, or a frozen snapshot? A dataset that keeps generating new, current tokens is worth more against a backdrop of accelerating scarcity than one frozen at a point in time; recency compounds the scarcity premium instead of just riding it once.
  • 3Could a lab plausibly synthesize a substitute instead of buying yours? Reasoning, math, and code are increasingly reproducible synthetically, per DeepSeek-R1's own results. [4] Outcome-labeled human judgment, and messy, real-world domain-specific text, are much harder to fake convincingly, which is exactly why they hold their price better.

For the full map of where different kinds of training data sit in today's supply chain, including which sources are growing and which are shrinking, see our complete guide to where AI training data comes from. The scarcity trend described here is the reason that map keeps shifting toward the closed, proprietary categories rather than the open, scraped ones.

The honest bottom line

The field is not running out of data in the way a headline implies. It is running out of one specific, historically dominant input, free, generic public text, on a timeline that credible researchers still disagree about by years, using a methodology that has already revised itself once. What's genuinely happening underneath the debate is narrower and more durable than either side's confident framing: quality, deduplication, and provenance now matter more than raw volume, and the industry's own response, five parallel strategies rather than one silver bullet, is the clearest confirmation of that.

None of that requires picking a side between Sutskever's certainty and Epoch's wide confidence interval to draw a practical conclusion. Public data is getting scarcer and more expensive to use well, real licensing deals already prove buyers will pay for clean alternatives, and freshly generated proprietary data, especially text, especially outcome-labeled, especially recent, is worth more in this market than it was two years ago. If you're holding that kind of data, the trend favors acting from a position of leverage now, not waiting for the debate above to resolve itself.

Find out what your data is worth in this market.

Public data is getting scarcer and more expensive to substitute. If you hold proprietary, human-generated data, get a free, no-obligation assessment from the team that brokers these deals every day.

List your data on Dayda