Data Licensing

Inside AI's Data Licensing Deals: What OpenAI and Google Actually Paid For

14 min read · 2026-08-21
$250Mreported ceiling of OpenAI's 5-year News Corp deal, the largest publicly reported figure
$60M/yrreported value of Reddit's data-licensing deal with Google
~$70M/yrreported estimate for Reddit's separate OpenAI deal, derived from Reddit's disclosed financials
$104MShutterstock's total AI data-licensing revenue in 2023 across all buyers
Key takeaways
  • 1Disclosed figures cluster far apart by content type: Reddit's forum data reportedly clears $60M-$70M a year per buyer, while most news-archive deals with individual publishers are undisclosed and, where estimated by reporters, land in the tens of millions annually; News Corp's deal is the largest publicly reported figure at up to $250M over five years.
  • 2Most deals never disclose a number at all. Of the roughly dozen major AI content-licensing deals with named terms this post tracks, only three publishers (News Corp, Reddit's two buyers, and Shutterstock's disclosed revenue total) have any dollar figure attached to them in reputable reporting; the rest are confirmed to exist but explicitly undisclosed.
  • 3Leverage tracks defensibility, not fame. Reddit and Stack Overflow, platforms whose best content sits behind a login or inside a moderated community and can't be cheaply re-scraped, have negotiated disclosed, recurring, per-year fees; most open-web news publishers have settled for undisclosed lump sums instead.
  • 4The market is shifting from training-only grants toward inference-time licensing. 2025 deals increasingly separate rights to train a model from rights to cite and summarize content in real time (RAG/search), and some publishers, like The Washington Post with OpenAI, have licensed only the latter.
  • 5A new licensing infrastructure layer, the RSL Standard launched in September 2025 by Reddit, Yahoo, and a dozen other publishers, is trying to standardize machine-readable pricing so smaller data holders don't have to negotiate a bespoke contract with every AI lab.
The short version

Every AI lab now claims it licenses training data responsibly, but the actual dollar figures behind those claims are scattered across press releases, SEC filings, and reporter estimates that rarely agree. This piece pulls together the real, verifiable terms of the AI data licensing deals that have closed since 2023, OpenAI's arrangements with the Associated Press, Axel Springer, News Corp, Condé Nast, and other publishers, Google's and OpenAI's separate licensing deals with Reddit, Stack Overflow's OverflowAPI partnership, and Shutterstock's multi-buyer image and video licensing business, and marks every figure as either confirmed, reported-but-unconfirmed, or undisclosed. The pattern that emerges: platforms with community or archive content that can't simply be scraped off the open web command disclosed, recurring fees, while most news publishers have accepted undisclosed lump sums, and the market is quietly moving from training-only grants toward pay-per-citation, inference-time licensing. For a startup or smaller data holder wondering what its own dataset might command, the lesson is not to benchmark against the $250M headline. It's to ask which of these two buckets your data falls into.

On this page ▾

Why a real map of these deals is hard to find

Search for what OpenAI or Google actually paid for training data and you get a wall of headlines with wildly different numbers attached to similarly-described deals: some cite hundreds of millions, some cite tens of millions, most cite nothing at all. That's not sloppy reporting. It's because almost every AI company and content owner has agreed, deliberately, not to disclose the price. The handful of numbers that leak out come from earnings-call disclosures, SEC filings, or reporters citing people familiar with the deal, not from press releases.

This matters for a specific, practical reason: if you're a founder or data holder trying to figure out what your own dataset might be worth, benchmarking against a headline like News Corp's reported $250M deal is close to useless. That figure describes a five-year, multi-title, global newspaper archive negotiated by one of the largest media companies in the world. [2] Your dataset is not that, and neither is almost anyone else's. What's actually useful is understanding the pattern beneath the headlines: which kinds of content get a disclosed, recurring number attached to them, which kinds settle for silence, and why. That's what this piece maps, deal by deal, with every figure traced to where it came from.

How to read the figures below
Three categories appear throughout this piece: confirmed (a company disclosed it directly, e.g. in an SEC filing or earnings call), reported (credible reporters cite sources or documents for a specific number, but neither party confirmed it publicly), and undisclosed (no number exists in any reputable source). None of the figures below were invented or estimated by us.

OpenAI's news publisher deals: one disclosed ceiling, everything else undisclosed

OpenAI has spent the last three years building the widest publisher-licensing portfolio in the industry, and it has done so while disclosing almost no numbers. The first deal, with the Associated Press in July 2023, set the template: a two-year, non-exclusive license to part of AP's text archive dating back to 1985, in exchange for AP getting access to OpenAI's technology and product expertise. AP built in a first-mover safeguard, a clause letting it revise its terms if a later publisher negotiated a better deal, which is itself a tell that nobody, including AP, knew what this content was worth yet. Financial terms were never disclosed. [6]

Five months later, Axel Springer (Politico, Business Insider, Bild, Welt) signed a multi-year, non-exclusive deal covering both free and paywalled content, for training and for real-time summarization with attribution and links inside ChatGPT. Neither company disclosed payment size or frequency; the Wall Street Journal reported the arrangement would generate what it called 'substantial revenue' for Axel Springer, but that's a characterization, not a number. [4]

The one deal in this category with an actual figure attached is News Corp, announced in May 2024: a multi-year global partnership covering the Wall Street Journal, Barron's, MarketWatch, the New York Post, and News Corp's UK and Australian titles. The Wall Street Journal itself, News Corp's own paper, valued the deal at approximately $250 million, reportedly payable partly in cash and partly in OpenAI technology credits over roughly five years. [6] That makes it the largest figure publicly attached to any AI content-licensing deal to date, and it's worth being precise about what kind of number it is: a press estimate of a deal's total ceiling, not a confirmed line-item OpenAI or News Corp put in a filing.

Everything else in OpenAI's publisher portfolio, Condé Nast (The New Yorker, Vogue, Wired, GQ, Vanity Fair, Condé Nast Traveler, Architectural Digest), The Atlantic, Vox Media (Vox, The Verge, New York Magazine), Time, and the Financial Times, followed the same pattern of confirmed existence, confirmed multi-year term, and no public dollar figure. Condé Nast's August 2024 deal grants OpenAI access to display and train on its titles' content through SearchGPT and ChatGPT, with attribution back to the source; financial terms and exact duration were not disclosed by either party. [12] Time's deal grants access to 101 years of archive alongside real-time content, again without a published number.

PublisherAnnouncedScopeExclusivityDisclosed terms
Associated PressJul 2023Text archive back to 1985, 2-year termNon-exclusiveUndisclosed
Axel SpringerDec 2023Politico, Business Insider, Bild, Welt; free + paywalledNon-exclusive, multi-yearUndisclosed (WSJ: 'substantial revenue')
News CorpMay 2024WSJ, NY Post, Barron's, MarketWatch, UK/AU titlesMulti-year, ~5 yr~$250M reported ceiling (WSJ estimate)
Financial TimesApr 2024FT archive and real-time contentNon-exclusiveUndisclosed
The Atlantic / Vox MediaMay 2024Full editorial archives, both companiesNon-exclusiveUndisclosed
Condé NastAug 2024New Yorker, Vogue, Wired, GQ, Vanity Fair, and moreNon-exclusiveUndisclosed
TimeJun 2024101 years of archive plus real-time contentNon-exclusiveUndisclosed
Every 'estimated' figure needs a source, not a vibe
Several blog aggregators repeat a supposed $5M-$10M/year figure for the Financial Times deal. We could not trace that number to any named reporter, filing, or company statement, so it does not appear in the table above. When a figure can't be traced, the honest move is to drop it, not round it into something that sounds plausible.

Reddit's two deals: the clearest disclosed numbers in the market

Reddit is the sharpest contrast to the publisher pattern above, and it's not a coincidence. In February 2024, on the same day it filed for its IPO, Reddit disclosed a content-licensing agreement with Google worth a reported $60 million a year, giving Google continuous access to Reddit's Data API and quarterly bulk transfers of Reddit data for use in AI products including Gemini. [6] Three months later, Reddit struck a separate, non-exclusive deal with OpenAI, reported at roughly $70 million a year, an estimate analysts derived from Reddit's own disclosed revenue figures rather than a number either company stated outright. [12]

Why did Reddit get disclosed, recurring, per-year numbers when News Corp only got a five-year press estimate and everyone else got silence? Reddit's content lives behind rate limits and API terms, not a freely crawlable archive; a huge share of its highest-value discussion, particularly recent threads, is logged-in, community-moderated, and updated continuously in ways a one-time scrape can't replicate. That scarcity, not Reddit's brand recognition, is what gave it the leverage to negotiate a metered-feeling annual fee instead of a lump sum, and to disclose it as part of an IPO prospectus where investors would have demanded the color anyway.

As of mid-2026, the Google-Reddit deal is up for renewal, and reporting indicates Reddit executives are pushing for a structure with usage-based fees rather than a flat annual number, a sign that a platform with real leverage renegotiates upward rather than accepting a renewal at the original price. That's the opposite of what tends to happen with a one-time archive license: once a publisher signs a flat multi-year deal, there's no natural point to revisit the price until the term ends.

Stack Overflow, Shutterstock, and Google's quieter posture

Stack Overflow's May 2024 deal with OpenAI, built on its OverflowAPI, licenses access to more than 59 million coding questions and answers, plus new posts as they're created, for use in training and in ChatGPT's coding responses with community attribution. [6] The arrangement is explicitly non-exclusive: Google had separately licensed the same Stack Overflow corpus a few months earlier. Financial terms for either deal were not disclosed. What Stack Overflow does disclose is the pricing mechanism: OverflowAPI is a subscription product any company can buy into, which is a meaningfully different structure from a bespoke bilateral contract. It's closer to a licensing product than a one-off negotiation.

Shutterstock took a similar 'sell the same asset to everyone' approach, but it's the one company in this market that has disclosed real aggregate revenue. Its six-year OpenAI agreement, expanded in 2023, licenses Shutterstock's image, video, and music libraries plus their metadata for model training, in exchange for priority access to OpenAI's generative tools and continued funding of Shutterstock's Contributor Fund, which pays creators whose work trains the models. [4] Shutterstock has also licensed the same library to Meta, Google, Amazon, and Apple. CEO Paul Hennessy has disclosed that Shutterstock's total AI-licensing business generated $104 million in revenue in 2023 across all of those buyers combined, and that its early anchor agreements with individual Big Tech customers ran on the order of $10 million a year each. [12]

The pattern inside this pair
Stack Overflow and Shutterstock both avoided negotiating a single bespoke mega-deal. Instead they built a repeatable licensing product (an API subscription; a standard content-library agreement) and sold it to multiple AI companies in parallel. That's a structure worth studying if you have one well-organized dataset and more than one plausible buyer.

Google shows up throughout this market as a buyer, in the Reddit deal, in Stack Overflow's OverflowAPI, in Shutterstock's multi-buyer roster, but it has made far fewer standalone, publicly named licensing announcements than OpenAI has. That's partly optics: Google already operates the crawler most of the open web depends on for search traffic, so a splashy new licensing deal invites the obvious question of why Google needs to pay for what it can already index. It's also partly strategy: Google has reportedly approached a larger, more diffuse set of national news organizations for AI-training licensing on a case-by-case basis, without disclosing the individual outlets or terms, rather than announcing marquee bilateral partnerships the way OpenAI has with Axel Springer or News Corp.

The through-line is the same regardless of how quiet Google is about it: even the company with the largest existing crawl of the public web still pays for Reddit's community data and Stack Overflow's structured Q&A, because scale of crawling doesn't substitute for access to content that sits behind a login, a rate limit, or a moderation wall. That's the single clearest signal in this entire market for what actually commands a price: not fame, not sheer volume, but whether the content can be reproduced by anyone with a scraper.

How content type actually prices, once you strip out the headlines

Line up every deal in this piece and a clear hierarchy shows up, not in exact dollar terms, since most aren't disclosed, but in which content types get any disclosed number at all, and which structure that number takes.

Content typeExampleTypical structureDisclosure pattern
Logged-in community / forum textRedditAnnual recurring fee, API access + bulk transferDisclosed, reported by name ($60M-$70M/yr)
Structured Q&A / code knowledge baseStack OverflowSubscription-style API license, non-exclusive, multi-buyerScope disclosed; price undisclosed
Licensed stock imagery/video/audioShutterstockMulti-year library license, non-exclusive, multi-buyerAggregate revenue disclosed; per-deal price mostly undisclosed
Open-web news archive, single publisherAP, Axel Springer, Condé Nast, Atlantic, Vox, FT, TimeMulti-year archive + real-time license, non-exclusiveExistence and scope disclosed; price undisclosed
Open-web news archive, large conglomerateNews CorpMulti-year global partnership, non-exclusivePress-estimated ceiling disclosed (~$250M)

The forum and structured-knowledge rows get real numbers because the content can't be cheaply reconstructed: a scraper can grab a publisher's homepage, but it can't replicate Reddit's continuously updated, moderated, logged-in discussion threads or Stack Overflow's accepted-answer voting signal at the same fidelity. The news-archive rows, by contrast, largely involve content that was already public and crawlable before the deal, which weakens the seller's negotiating position and, not coincidentally, correlates with less willingness on either side to publish a number: a publisher that reveals it accepted a modest fee for content anyone could once scrape for free invites every other publisher to ask why theirs was worth less.

Size of the archive isn't the driver
News Corp's reported $250M figure is often misread as proof that news archives are worth more than forum data. It's more precise to say News Corp negotiated as one of the largest media conglomerates on earth, across dozens of mastheads and multiple countries, in a single deal. Per-publisher, the far more typical outcome for news content has been an undisclosed figure that individual reporting pegs well below Reddit's disclosed per-buyer rate.

The shift from training rights to inference-time licensing

The earliest deals in this market, AP, Axel Springer, News Corp, were framed almost entirely around training: giving OpenAI the right to use a publisher's archive as raw material to build its models. By 2025, that framing had started to split into two separable rights: training rights and inference-time or search/citation rights, meaning the AI product can look up and quote a publisher's live content in real time (retrieval-augmented generation) without that content necessarily being baked into the model's weights. [8]

The Washington Post's 2025 licensing arrangement with OpenAI is reported to be structured this way: it grants search-and-citation rights inside ChatGPT without the same explicit grant of model-training rights that OpenAI's earlier News Corp and Axel Springer deals carried. [2] That's a meaningful shift for anyone thinking about licensing their own data, because it means 'AI licensing deal' is no longer one product. A buyer might want to train on your data permanently, cite your data live without training on it, or both, and each of those is a separate right worth pricing separately rather than bundling by default.

Unbundle your own rights before you negotiate
If you're structuring your own data license, don't assume a buyer wants everything. Ask explicitly whether they want training rights, inference/retrieval rights, or both, and price them as separate line items. The market's own biggest players are already doing this.

RSL: an attempt to standardize pricing instead of negotiating from scratch

Every deal covered so far required a bespoke, lawyered, bilateral negotiation, which is exactly why smaller data holders rarely see anything like a Reddit- or News Corp-sized deal: nobody at OpenAI or Google is going to spend months negotiating a custom contract for a dataset with modest volume. On September 10, 2025, Reddit, Yahoo, Medium, O'Reilly Media, Quora, Ziff Davis, and a dozen other publishers and platforms launched the RSL Standard (Really Simple Licensing) along with a nonprofit RSL Collective, an open, machine-readable protocol, modeled on RSS, that lets any site attach licensing and royalty terms, including free, attribution-only, subscription, pay-per-crawl, and pay-per-inference, directly to its robots.txt file. [6]

RSL matters for this market for a specific reason: it's an attempt by the exact companies that already have disclosed, individually-negotiated licensing deals, Reddit chief among them, to build infrastructure that lets smaller publishers and data holders set standardized terms without needing Reddit's negotiating leverage or legal budget. Whether it achieves real adoption from the AI labs that would need to honor those machine-readable terms is still unresolved as of this writing, but its existence is itself evidence that the current one-off, bespoke, mostly-undisclosed negotiation model is seen as broken even by the platforms that have won the biggest deals under it.

What this means if you're not a household name

If your dataset isn't Reddit-scale or News Corp-scale, the deals above are still useful, just not as a price benchmark. Use them as a diagnostic instead. Ask which side of the defensibility line your data sits on: is it the kind of content a well-funded scraper could reconstruct from the open web in a few weeks, or does it live somewhere a scraper can't reach, behind a login, inside a proprietary workflow, gated by consent, moderated by humans, updated continuously?

  • 1If your data is scrapable in principle, even if nobody has bothered yet, expect a negotiation that looks like the news-publisher rows: modest, likely undisclosed terms, and real pressure to accept whatever a buyer offers because they have a credible walk-away option.
  • 2If your data can't be reconstructed from the public web, proprietary support transcripts, internal workflow data, community content behind a login, expert-annotated records, you're negotiating from the same position Reddit and Stack Overflow did. That's where a recurring, disclosed-to-you-if-not-to-the-public fee and real leverage become possible.
  • 3Separate training rights from inference rights explicitly, the way the market has started to. A buyer who only needs to cite your data live pays for something different than a buyer who wants to bake it into model weights permanently; don't let one deal quietly grant both by default.
  • 4Watch for standardized-pricing infrastructure like RSL as a signal, not necessarily a tool you'll use directly yet. Its existence tells you the market is aware that bespoke negotiation doesn't scale down to smaller holders, which is exactly the gap a managed marketplace and broker relationship is built to close.
  • 5Don't anchor your ask to a headline number from a different content type. A $250M five-year News Corp deal and a $60M/year Reddit deal describe two different kinds of leverage; neither is a ceiling or a floor for a mid-size proprietary dataset in a different domain.

The real lesson isn't the dollar figures, it's the leverage map

Strip away the headlines and this market has a simple structure. A handful of platforms with content that can't be freely re-scraped, Reddit above all, have negotiated real, disclosed, recurring money. A much larger group of open-web publishers have accepted undisclosed terms for content that, whatever its editorial value, was already sitting on the crawlable internet before the AI labs came calling. And the whole market is quietly splitting one bundled 'AI license' into separately priced training and inference rights, while a subset of the same platforms tries to build standardized infrastructure so the next tier of data holders doesn't have to negotiate from zero.

None of that requires guessing at a number nobody has disclosed. It requires knowing where your own data sits on the defensibility spectrum, and structuring your ask around the rights a buyer actually wants, not the rights a template contract happens to bundle together.

Find out what your data is actually worth.

Get a free, no-obligation review from the team that structures these deals for a living. We'll help you figure out where your data sits on the leverage spectrum and what terms are realistic for it.

List your data on Dayda