Data Is Not Oil (a four-part series)

science

On the left an oil pump jack and pipeline; sparks drifting from the pipe's mouth form a glowing ring set with petri dishes, a DNA double helix and network nodes

Part 1 | Data is not oil

In 2006 the British mathematician Clive Humby said something that has been quoted ever since: data is the new oil. Humby designed the loyalty-card scheme for the British supermarket Tesco. The phrase is often followed by a gloss: like crude, data is valuable, but it is of no use until it is refined. Eleven years later, in May 2017, The Economist ran a cover package titled “The world’s most valuable resource is no longer oil, but data”, and the metaphor became common sense in the technology industry.

The metaphor is half right. The right half is that data really is valuable, really does need processing, and the companies that hold it really have drawn antitrust attention just as the oil giants once did. The half that may be wrong matters more: run a data business on the logic of oil — extract, refine, sell by volume — and you pick the wrong business model. That wrong half is what this series is about.

Seven differences

Compare data and oil item by item and the differences pile up.

Dimension Oil Data Implication for the business model
Consumption Burnt and gone; a barrel can be sold to only one buyer Can be copied and reused; the same data can serve many users at once Selling by volume doesn’t pay; the buyer’s cost of copying is almost zero
Homogeneity Standardised, priced by the barrel Highly varied; value depends on what it is combined with and what question it answers No single price; a liquid market is hard to form
Inspection Examined and assayed, the oil is still there To judge what data is worth you have to see it, and once seen its value is gone A verification mechanism before the trade matters more than the quote
After use Burnt into energy and exhaust Once trained into a model, its value settles in the parameters and the data owner cannot get it back A one-off licence sells long-term value cheaply in a single package
Property rights Mineral rights are clear Involves patients, hospitals, platforms and others; privacy and compliance boundaries are blurred Needs technical and contractual arrangements where “the model moves, the data stays”
Scarcity Geological reserves are finite Public high-quality text is running out, but data can be actively produced The most valuable data is increasingly generated on demand
Shelf life Oil is oil however long it sits It goes stale; distributions drift Needs continuous supply, not one-off delivery

The first difference is the most fundamental. In a 2020 paper the economists Charles Jones and Christopher Tonetti looked specifically at the non-rivalry of data: one person’s location history, medical record or driving data can be used by many companies at once without being depleted. They conclude that, even accounting for privacy costs, using data more widely often brings substantial social gains — while firms, worried about being overtaken by competitors, may hoard data, which is inefficient. Oil is the opposite: a barrel given to you cannot be given to anyone else, and possession itself is where the value lies.

The third and fourth differences together are what make data trading so awkward. The buyer wants to see whether the data is any good; the seller fears that once seen, its value is gone. After the deal, the data is trained into a model and the seller can never get it back. These two problems mean data can hardly change hands like oil, cash on delivery. The next part takes them up in detail.

Oil runs out; data can be made

The sixth difference is the most interesting and the most easily overlooked.

As far as public data goes, “new oil” seems ever more apt. The research group Epoch AI estimates that the internet holds about 300 trillion tokens of high-quality human text usable for training, and that on current trends large models will use it up between 2026 and 2032. Some hope synthetic data generated by models will fill the gap, but a 2024 study in Nature found that if each generation of model is trained indiscriminately on content generated by the previous one, the models develop irreversible defects, with the rare information in the original data disappearing first. Other research finds that the degradation can be contained as long as real data keeps being retained and accumulated rather than replaced by synthetic data. Whichever conclusion holds, real data is getting scarcer; public data really is like a well running dry.

In science, though, things are entirely different. Most of the most valuable data in biomedicine is not a stock lying around waiting to be extracted; it is produced by doing experiments. Add a compound to cells, knock out a gene, change the culture conditions, and a new batch of data appears. And the value of that data depends heavily on what question was asked. Ten thousand random experiments may yield less information than a hundred key experiments chosen by a model.

This brings us to the deepest dividing line between data and oil. Oil’s value is fixed while it is still underground; extraction just brings it out. Much of data’s value is created in use: the model poses a question, an experiment generates data, the data corrects the model, and the model poses a better question. Whoever controls that loop controls the value.

Where the metaphor goes wrong

If this series’ argument had to fit in one sentence, it would be: oil’s value is realised once, at extraction and sale; data’s value is realised only through repeated use and continuous feedback.

A data business designed on the logic of oil has a built-in ceiling. The seller gets paid only at the moment of delivery, and none of the value the data later creates in the buyer’s hands comes back to him — so he will always ask a high price at delivery. The buyer, at delivery, does not yet know what the data will bring — so he will always push the price down. Both worries are reasonable, and the result is that many deals that should happen don’t, and many of those that do leave one side feeling short-changed.

The next three parts take up three questions in turn: why data is almost impossible to price at the outset; how the oil industry itself got out of a similar bind; and which existing business models could be borrowed to break today’s deadlock in data deals.

Part 2 | Why data can’t be priced at the outset

At the negotiating table of a data deal, the scene is usually this. The data owner names a high price, on the grounds that the data took years and a great deal of money to build up. The model company can’t pay that much, and can’t say how much improvement the data will actually bring, so it pushes back. After a few rounds, the talks either break down or end in a fixed price neither side is happy with. One side is afraid of overpaying, the other of underselling.

My view is that this deadlock is not a matter of negotiating skill. At the moment of signing, a “correct price” for the data simply does not exist. There are at least four reasons.

Value travels a long road before it is realised

Take biomedicine. A batch of data first trains a model; the model makes predictions; predictions become hypotheses; hypotheses go through experiments that screen out candidate molecules; candidates then go through preclinical work and Phase III trials, and only at the end might a drug reach the market and earn revenue. The road usually takes more than ten years, and only about one in ten projects that enter the clinic is eventually approved.

In other words, the value of data at signing is a lottery ticket whose draw is more than ten years away, whose odds of winning are low and whose prize is uncertain. Worse, even if it wins, it is hard to say how much of the win was due to this data and how much to the model, the lab team, the trial design and luck.

Each side holds half the information

The seller knows the data best: where the samples came from, how good they are, where the flaws lie. The buyer knows best how useful the data is to him: that depends on what data he already has, which piece his model is missing and what problem he is trying to solve. Each holds key information the other can’t see, and each has a motive not to tell the whole truth — the seller will overstate the data’s quality, the buyer will understate its usefulness to him.

There is a classic result in economics about exactly this situation. In 1983 Roger Myerson and Mark Satterthwaite proved that when buyer and seller each privately know their own valuation, and it is not known in advance whose is higher, no trading mechanism can achieve four things at once: both sides willing to take part, both sides having an incentive to tell the truth, no need for an outside subsidy, and every trade that benefits both sides actually going through. That means that, without outside money, however the negotiation rules are designed, some trades that ought to happen will fail.

So “one afraid of overpaying, one afraid of underselling” is not just a matter of mindset; it is structural. As long as the two sides must settle a price once, under asymmetric information, some deals will always fall through.

See it and you needn’t buy; buy it and it can’t be returned

In 1962 Kenneth Arrow identified a fundamental paradox of information goods: to judge what information is worth, the buyer must know what it contains — and once he knows, he no longer needs to pay for it. Data fits this description exactly. The seller can’t hand over all the data for the buyer to try first, and the buyer won’t pay a high price for data he hasn’t seen.

In the age of large models the paradox has another layer. Once data is trained into a model, its value settles in the model’s parameters; the seller cannot get it back, and can hardly prove how much the buyer used. So the seller leans even more towards demanding a high price once, at delivery, because it may be the only time he gets paid.

Value doesn’t simply add up

The same data can be worth orders of magnitude more to one buyer than to another. If the buyer already has similar data, the increment from new data is close to zero; if the buyer happens to be missing exactly this piece, it may be what takes the model from unusable to usable. Data’s value also changes with combination: two batches of data, each unremarkable on its own, may together yield information neither has alone.

That means data can hardly have a single market price. Oil can be listed on an exchange by the barrel because every barrel is the same to every buyer. Data cannot.

What the market shows

All four difficulties are clearly visible in today’s market.

Almost all the large data deals made public so far use fixed fees. OpenAI’s five-year deal with News Corp is reportedly worth more than $250 million; Reddit’s deal with Google is about $60 million a year. A fixed fee is how the buyer caps his risk: however much value the data ends up bringing, this is what I pay. But the arrangement is already loosening: Reddit executives reportedly want to move to usage-based pricing when the Google deal is renewed.

China looks much the same. According to the National Data Administration, the national data market was expected to exceed 160 billion yuan in 2024, of which exchange-based trading (including registered trades) exceeded 30 billion; most trades are still bilateral, off-exchange negotiations. Rules in force since 2024 on recognising data as assets on the balance sheet require qualifying data resources to be booked as intangible assets or inventory, initially measured at historical cost — that is, documented spending on purchase, processing and registration. Valuation methods such as the income and market approaches are used only in asset appraisals and do not reach the balance sheet. That says a lot: what the accounts can recognise is how much the data cost, not what it is worth.

A more vivid example: in 2018 GSK invested $300 million in the genetic-testing company 23andMe in exchange for a collaboration to develop drugs using its user data. After 23andMe went public in 2021, its valuation briefly reached about $6 billion. In 2025 it filed for bankruptcy, and almost all its assets were eventually bought for $305 million — the core of them being the genetic data of more than 15 million users — at less than a twentieth of its peak valuation. The company’s valuation included far more than data, but the gap is enough to show how wildly the price of the same data asset can vary at different times and in different circumstances.

From finding a price to designing a contract

If data has no correct price at signing, then the effort to find that price is futile. The more useful question is: can a deal be structured so that the price is settled step by step as information is revealed?

Concretely, that means paying little while the data’s value cannot yet be seen, and paying more, as agreed, when the model’s performance is validated, when a candidate molecule enters the clinic, when the drug finally reaches the market. The buyer does not have to put up a large sum at the outset, and the seller need not worry about selling long-term value cheaply in one go.

The idea is not new. The oil industry went through a similar transition in the mid-twentieth century, and biomedical licensing long ago made “pay as it progresses” standard practice. The next part starts with oil.

Part 3 | The oil industry’s fifty-fifty moment

Today’s data owners are in a position much like that of the oil-producing countries in the first half of the twentieth century: holding a valuable resource, unable to say what it is worth, forced to accept one-off or fixed payments while watching the resource create enormous value in someone else’s hands. The oil industry took roughly half a century to get out of that situation, and its history is worth the data industry’s close attention.

The concession era

From the early 1900s to the 1940s, most oil in the Middle East and Latin America was produced under concessions. A producing country granted a foreign oil company long-term rights over a large area; the company handled exploration, production and sales, and the country collected royalties, mainly by volume. In 1933 Saudi Arabia granted Standard Oil of California rights over a vast area for sixty years; the company that ran the concession was later renamed the Arabian American Oil Company — Aramco.

At signing, the arrangement was not unreasonable. Exploration was extremely risky — a well might not strike oil — and the producing country had neither capital nor technology, and did not know how much oil lay underground. The oil company bore the risk in exchange for long-term rights, the country got a certain income, and it looked like a fair bargain.

The problem came later. As giant fields were discovered one after another, the oil turned out to be worth far more than anyone imagined at signing, while the producing countries’ income was still calculated on the old fixed terms. In Iran’s case, royalties as a share of the oil export price fell from 33% in 1933 to 9% in 1947. The more valuable the fields became, the smaller the share that went to the owner of the resource.

How similar to today’s data deals: the data owner doesn’t know the data’s value at signing, licenses it at a fixed price, and none of the value the model later creates comes back to him.

Fifty-fifty

The turning point came in 1948, when Venezuela was first to establish a principle: the profits oil companies made in the country would be split equally between company and government. On 30 December 1950 Saudi Arabia and Aramco signed a new agreement that brought the same fifty-fifty split to the Middle East. The design was clever: Aramco kept paying its old royalties, but they were treated as a prepayment credited against a new 50% income tax. With this arrangement, Saudi oil revenue nearly quadrupled.

What followed fell like dominoes.

Year Event
1933 Saudi Arabia grants Standard Oil of California a sixty-year concession
1948 Venezuela establishes the fifty-fifty profit-sharing principle
1950 Saudi Arabia and Aramco sign a fifty-fifty agreement
1951 After long failing to get the Anglo-Iranian Oil Company to accept similar terms, Iran’s parliament passes a law nationalising the oil industry; Kuwait and Iraq also adopt fifty-fifty within a year or two
1960 After the oil companies unilaterally cut posted prices, Venezuela, Saudi Arabia, Iran, Iraq and Kuwait found OPEC in Baghdad
1971 The Tehran and Tripoli agreements, among others, raise producers’ profit share to about 55%
1970s Producing countries gradually take ownership of the oil companies through equity stakes; Saudi Arabia eventually owns Aramco outright

From fixed royalties, to profit-sharing, to equity stakes and finally full ownership, the resource owners went step by step from sellers of raw material to partners sharing profits and taking part in the business.

What made sharing possible

Looking back, several forces drove the transition, and each has a counterpart in the data industry.

First, information was revealed. At signing no one knew how much oil lay underground; as giant fields were found, the value of the resource became plain, and the old fixed payments looked less and less reasonable. When Saudi Arabia asked to renegotiate in 1950, one of its reasons was that the oil in Aramco’s concession far exceeded expectations. The data industry is going through the same process: the success of large models has shown everyone the value of high-quality data, and the price of one-off licences is starting to look too low.

Second, precedent spread. Once Venezuela’s fifty-fifty was in place, it became the reference point for every producing country’s negotiations. After the Saudi deal, internal US State Department documents noted that it made Iran feel that demanding a “Venezuelan-style” split was not asking too much. In data, precedents for usage-based sharing are appearing too. More than 2,200 members of the News/Media Alliance in the US can opt into an agreement with ProRata that pays out half the platform’s revenue according to each publisher’s content’s contribution to AI answers.

Third, the books were opened. A core demand in Iran’s dispute with Anglo-Iranian was that the company open its accounts; without that, there was nothing to base a split on. Sharing data revenue likewise depends on auditable usage records: who used how much, and with what result, must be verifiable.

The fourth is the easiest to overlook: one important reason the split could be agreed was that someone helped carry the cost. The Saudi fifty-fifty was designed as an income tax, which Aramco could credit against the tax it owed in the United States; the company paid little more in practice, and a good part of the cost was borne by the US Treasury. Data deals need a third party like that too: if a financial institution is willing to buy the data owner’s future share in advance, the buyer need not put up a large sum at the outset. The Myerson–Satterthwaite theorem from the last part provides a theoretical footnote here: bringing in outside money is precisely one of the conditions that lets more trades go through.

Where data and oil part ways

This history, though, differs from data in one respect — and it is the respect that matters most.

The international-business scholar Raymond Vernon proposed the idea of the “obsolescing bargain”: once a foreign company has sunk large fixed assets in a host country, it can no longer move them, and bargaining power gradually shifts from the company to the host. That is exactly what happened with oil: fields, pipelines and refineries all stayed on the producing country’s soil, and the country could demand renegotiation at any time, or simply nationalise.

With data it is the other way round. Once data is trained into a model, it has been carried away; the data owner cannot get it back, and cannot apply pressure by “shutting off the well”. Before signing, the data owner’s leverage is at its greatest; after signing and delivery, it falls to almost nothing.

That means data owners cannot expect to do what the producing countries did — sign a bad contract and renegotiate once the value emerges. Every term about revenue sharing, milestones and usage billing has to be settled before the data is delivered. In that sense, what data owners need is to get to “fifty-fifty” in one step, not to walk the fifty years from concession to equity again.

Another difference is how dispersed the resource owners are. Oil was concentrated in a handful of countries, which could band together to form OPEC. Data is spread across thousands of hospitals, research institutions and platforms, and any single owner’s bargaining power is weak. News publishing points to one direction: negotiate collectively through an industry alliance, with members opting into a common revenue-sharing agreement. Medical and biological data may need a similar consortium.

Where we stand

Measured against oil’s history, today’s data deals are still broadly in the concession era: fixed fees, one-off licences, and all the value going to the buyer after delivery. But signs of a fifty-fifty moment have appeared: content platforms sharing by usage, a health-data company in which the health systems that supply the data share in the returns as shareholders, a social platform that wants to switch to usage-based billing at renewal.

The oil industry got from concessions to profit-sharing through revealed information, spreading precedent, open books and a third party sharing the cost. To get out of the fixed-fee era the data industry needs all four — and because leverage falls to zero once data is delivered, they must be built into the contract at the very start of the deal. The next part discusses exactly how.

Part 4 | Nine business models to borrow

The first three parts come down to one sentence: data has no correct price at the moment of the deal, so the way out is not to calculate a price but to design a contract in which the price is settled step by step as information emerges. The good news is that such contracts need not be invented from scratch. Biomedical licensing, M&A, finance, media and the music industry have long dealt with assets whose value is uncertain and slow to realise, and have built up a mature toolkit.

The deadlock in data deals breaks down into three specific problems: the buyer can’t pay that much now; neither side dares to set the price first; and data’s value is produced continuously over a long time, so one-off settlement is bound to be unfair. The nine models to borrow map onto these three problems.

Group 1: the buyer can’t pay now

Upfront payment, milestones and sales royalties. This is the standard structure of biomedical licensing deals, built for assets whose value is slow to realise. In AI drug discovery, for example, Isomorphic’s deal with Lilly has an upfront payment of $45 million, development milestones of up to $1.7 billion, and tiered royalties up to the low double digits on sales. The upfront is less than 3% of the potential total; most of the value is staked on the future. The structure is mainly used today for deals on AI platforms and candidate molecules and is still rare among publicly disclosed data licences, but the logic carries straight over: a small upfront, then staged payments tied to model performance, pipeline progress and eventual sales.

Data for equity. The data owner takes shares rather than cash. In 2018 GSK invested $300 million in 23andMe and the two agreed to split the costs and profits of drug development, so the data owner got both equity and a share at the product level. Truveta went further: 17 health systems, together with Illumina and Regeneron, put $320 million into its preferred stock at a valuation above $1 billion. The health systems that supply the data are themselves shareholders, which ties together the interests of data owners, researchers and drug companies.

A third party buys future returns in advance. In drugs there are funds that specialise in buying royalties; in mining there is “streaming”: a financial institution pays a sum up front in exchange for a percentage of future output or revenue. This can be carried straight over to data. The data owner sells its future milestones and royalty rights to a financial institution and gets cash early; the buyer need not pay a large upfront. As the last part noted, one important reason the Saudi fifty-fifty could be agreed was that the US Treasury in effect bore part of the cost; third-party capital can play a similar role in data deals. Such a “data royalty fund” has no established precedent yet, and I think it is something worth someone’s while to build.

Group 2: neither side dares to price first

Option-style licensing. The buyer pays a small option fee for a period of exclusive evaluation rights, with the price and terms for exercising the option agreed in advance. During the evaluation the data’s value becomes clearer, and the buyer then decides whether to exercise. This is common in biotech licensing; applied to data, it turns “price first” into “evaluate first”.

Measure the gain first, then pay for results. In a privacy-preserving computing environment, the buyer’s model trains without touching the raw data; measure how much the model improves when this data is added, and pay according to the improvement. It is the same logic as advertising’s move from paying per impression to paying per click or per conversion: don’t pay for possible value, pay only for measured results. It has already been validated at scale. In the MELLODDY project, from 2019 to 2022, ten competing pharma companies trained models jointly without sharing data, using more than 2.6 billion confidential experimental data points, and every company’s predictive models improved overall.

Earn-outs and contingent value rights. In M&A, when buyer and seller can’t agree on a valuation, they often use an earn-out or contingent value rights: close at a conservative price, and pay more later if agreed targets are met. Data deals can do likewise: if a project based on the data reaches a given clinical stage, the buyer pays an additional sum.

Group 3: value has to be shared over time

Usage-based revenue sharing. ProRata uses attribution technology to calculate each publisher’s contribution to a given AI answer and pays half the platform’s revenue to content owners on a per-use basis. Truveta’s approach is similar: according to its founder, the key to sustaining the model is paying the health systems and sequencing partners whenever de-identified data is used by others. Collective management of music rights and pooled revenue distribution in streaming are older, more mature predecessors of this kind of model.

Pricing by time window. The UK Biobank exome sequencing consortium is a good example. In 2018, led by Regeneron, the first five pharma companies each put in $10 million to have sequencing of 500,000 participants finished three years early; in return, the funders got an exclusive period of 6 to 12 months, after which the data returned to the database and opened to researchers worldwide. It works like the release window in film: charge a premium for the premiere, then widen distribution. It neatly sidesteps the hard question of what the data is worth by pricing something else instead: getting to use it a while before everyone else.

Data for models. The three-year agreement Tempus signed with AstraZeneca and Pathos in 2025 is non-exclusive: once the model is built, all three get a copy; the latter two bear most of the compute costs, and Tempus also received $200 million in data licensing and model development fees. The data owner gets not only cash but a model that feeds back into its own business, and keeps the right to license the data to others.

Putting the nine together

These models are not mutually exclusive. Combined, they give a layered structure suited to biomedical data deals, each layer corresponding to a point on the value chain where information is revealed.

Layer Trigger Payment Problem solved Borrowed from
Evaluation Model gain measured in a privacy-preserving environment Small evaluation fee See it and you needn’t buy Pay-for-performance advertising
Option Buyer wants an exclusive evaluation period Option fee, locking in later terms Neither side dares to price first Biotech option licensing
Model milestones Model reaches an agreed benchmark improvement Staged payments Buyer’s limited ability to pay upfront Drug licensing milestones
Pipeline milestones Target validation, candidate nomination, IND filing, entry into Phase II Staged payments Value realised slowly Drug licensing milestones, earn-outs
Long-term sharing Product launch, or the model earns commercial revenue Sales royalties or revenue share; multiple data sources split an attribution pool Value produced continuously Royalties, usage-based sharing

Two options can be layered on top: part of the consideration paid in equity, making the data owner a long-term stakeholder; and the data owner selling its pipeline-milestone and long-term-sharing rights to a financial institution to cash out early.

Back to the deadlock at the negotiating table. Under this structure, the buyer pays only evaluation and option fees during the most uncertain stage, with no need to stake a large sum at the outset. The seller gets less up front but keeps the right to share in all the value that follows, and need not fear selling long-term value cheaply in one go. The price is not fixed at signing; it is realised layer by layer, with model performance, clinical progress and sales.

What it takes to work

Four conditions must hold for this to work in practice.

First, every term must be settled before the data is delivered. As the last part showed, once data is trained into a model it can’t be taken back, and the data owner’s bargaining power falls to almost nothing at the moment of delivery; there is essentially no leverage for renegotiating afterwards.

Second, usage must be auditable. Usage-based sharing and milestone payments both depend on reliable records of how the data is used and how projects progress. Just as the producing countries demanded that the oil companies open their books, data owners need a verifiable mechanism for usage and provenance.

Third, attribution has to be pragmatic. Methods that precisely calculate each piece of data’s contribution to a model, such as the game-theoretic Shapley value, are the fairest in theory, but exact computation grows exponentially with the number of data owners and in practice can only be approximated; research also suggests such valuations can be strategically gamed. Early on, simple transparent rules are better — say, weighting by volume and quality grade and distributing from a common revenue pool — refined gradually as methods mature.

Fourth, counterparty risk must be guarded against. Over a ten-year payback period, anything can happen to either side. 23andMe is a case in point: after its exclusive research period with GSK ended in July 2023, its research-services revenue fell; two years later the company filed for bankruptcy and its data changed hands at auction. Long-term payment arrangements need escrow, guarantees or pledged rights to back them; otherwise the promise to “share future value” may come to nothing.

Finally

“Data is the new oil” has been popular for nearly twenty years, and its biggest distortion is to make people think the data business is about extraction and sale. The oil industry itself took half a century to learn that the resource owner’s real way forward is to share profits and take part in the business, not to sell at a higher price. The data industry does not need another fifty years. The tools it needs have long since matured, separately, in biomedical licensing, M&A, finance and the copyright industries; what is missing is to put them together, and to write them into the contract before the data is delivered.

For two parties negotiating a data partnership, perhaps there is a different question to open with: rather than arguing over what this data is worth, agree first on how we will split its value as, bit by bit, that value comes to light.

References

Part 1

  1. Clive Humby, on the 2006 phrase “data is the new oil”; the gloss that it “cannot be used unless refined” is often thought to be a later elaboration.
  2. The Economist. The world’s most valuable resource is no longer oil, but data. May 2017.
  3. Charles I. Jones & Christopher Tonetti. Nonrivalry and the Economics of Data. American Economic Review, 110(9), 2020.
  4. Pablo Villalobos et al. Will we run out of data? Limits of LLM scaling based on human-generated data. ICML 2024 (as reported by AP / VOA; estimates from Epoch AI’s summary of the analysis).
  5. Ilia Shumailov et al. AI models collapse when trained on recursively generated data. Nature, 631, 2024; for a different view, see Matthias Gerstgrasser et al. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413, 2024.

Part 2

  1. Eric Vallabh Minikel et al. Refining the impact of genetic evidence on clinical success. Nature, 629, 2024.
  2. Roger B. Myerson & Mark A. Satterthwaite. Efficient mechanisms for bilateral trading. Journal of Economic Theory, 29(2), 1983.
  3. Kenneth J. Arrow. Economic Welfare and the Allocation of Resources for Invention. In The Rate and Direction of Inventive Activity, Princeton University Press, 1962.
  4. Summary of AI training-data licensing deal prices, Quartz, 2026.
  5. Reddit Weighs Cutting Google AI Access as $60M Deal Expires, as reported by The Wall Street Journal.
  6. National data market exceeds 160 billion yuan in 2024: figures from the National Data Administration at the national data work conference.
  7. Ministry of Finance. Interim Provisions on Accounting Treatment of Enterprise Data Resources, in force from 1 January 2024; on initial measurement at historical cost, see the commentary by Zhong Lun Law Firm.
  8. 23andMe; an account of the 23andMe bankruptcy case.

Part 3

  1. International Petroleum Agreements: Politics, oil prices steer evolution of deal forms, Oil & Gas Journal.
  2. Foreign Relations of the United States, 1951, Vol. V, on the 1950 Saudi–Aramco agreement.
  3. The Anglo-Iranian Oil Company, 1951: Britain vs. Iran, Seven Pillars Institute.
  4. Saudi Arabian Oil: The Obsolescing Bargaining Model, Becker, 2018.
  5. The History of Saudi Arabian Oil, on the tax-credit arrangement behind the fifty-fifty split.
  6. Posted oil price, on the founding of OPEC and the 1971 change in profit shares.
  7. Raymond Vernon. Sovereignty at Bay: The Multinational Spread of U.S. Enterprises. Basic Books, 1971.
  8. News/Media Alliance, ProRata AI Sign Content Licensing Deal, MediaPost.

Part 4

  1. Lilly, Novartis sign AI partnership with Alphabet’s Isomorphic, PharmaLive, 2024.
  2. 23andMe Gets $300 Million Boost From GlaxoSmithKline, Bio-IT World, 2018.
  3. Leading US health systems launch the Truveta Genome Project, 2025; reported by GeekWire.
  4. Wouter Heyndrickx et al. MELLODDY: Cross-pharma Federated Learning at Unprecedented Scale Unlocks Benefits in QSAR without Compromising Proprietary Information. Journal of Chemical Information and Modeling, 2024.
  5. ProRata Invents Generative AI Attribution Technology, Business Wire, 2024.
  6. Regeneron announces major collaboration to exome sequence UK Biobank, UK Biobank, 2018; on the exclusivity period, see GenomeWeb.
  7. Tempus Signs Expanded Strategic Agreements with AstraZeneca and Pathos, 2025; deal details from Benzinga.
  8. Anne Wojcicki will acquire 23andMe for $305M following bankruptcy, MobiHealthNews, 2025.
  9. Amirata Ghorbani & James Zou. Data Shapley: Equitable Valuation of Data for Machine Learning. ICML 2019.