All articlesTopic · 6 articles

AI Data Supply Chain

14 min read·2026-08-09

Types of AI Training Data Explained (With Examples)

A complete, cited taxonomy of AI training data types across two axes: modality (text, image, audio, video, tabular, multimodal) and function (labeled vs. unlabeled, train/validation/test, synthetic vs. real, structured vs. unstructured), with real examples, formats, labeling methods, cost tiers, and a full comparison table.

Read the guide
13 min read·2026-08-07

How AI Training Works: From Raw Data to Deployed Model

A complete, plain-English map of how AI training actually works: the seven-stage pipeline from data collection through cleaning, labeling, pretraining, fine-tuning, RLHF alignment, evaluation, and deployment, and exactly what role training data plays at each stage.

Read the guide
14 min read·2026-08-03

What Data Do LLMs Like GPT, Claude, and Llama Actually Train On?

A source-by-source answer built from the actual model and system cards: the open corpora that anchored LLM pretraining, what OpenAI, Anthropic, Google DeepMind, and Meta each disclose about their current flagship models, and exactly what every one of them leaves out.

Read the guide
13 min read·2026-07-06

Where Does AI Training Data Actually Come From? A Map of the Data Supply Chain

The definitive map of the AI data supply chain: public web corpora, proprietary and held-out datasets, user-generated platform data, licensed and purchased datasets, synthetic data, and the brokers and marketplaces that connect supply to demand. Cited and original.

Read the guide
13 min read·2026-06-27

Are We Running Out of AI Training Data?

The definitive answer to AI's most-debated supply question: Epoch AI's peak-data research and how its estimate has already shifted once, why quality and deduplication matter more than raw token count, and the five real strategies labs are using to respond.

Read the guide
14 min read·2026-06-25

What Is AI Training Data? A Complete Guide

A plain-English, fully cited guide to what AI training data actually is: the six data modalities models learn from, labeled vs. unlabeled data, and how pretraining, fine-tuning, and evaluation datasets differ and interact.

Read the guide