AI Affairs, home

Tuesday 29 September 2026

Corporate

Oxford allowed OpenAI to train models on Bodleian historical texts

The partnership was announced as a digitisation project. By June 2025, the Bodleian had shared 125,000 scanned dissertation images with OpenAI.

Interior of the Divinity School in the Bodleian Library, Oxford
Photo: Diliff, CC BY-SA 3.0, via Wikimedia Commons (cropped)

The University of Oxford allowed OpenAI to use historical Bodleian Library texts to train its models under a March 2025 partnership, with 125,000 scanned dissertation images shared by June that year, the Guardian reported on 26 September 2026.

Key points

  • Oxford announced a digitisation partnership in March 2025 without mentioning model training.
  • Internal documents say digitised Bodleian material entered OpenAI’s training set.
  • Oxford says the project covers out-of-copyright material and that the Bodleian retains rights to the scans.

Oxford’s March announcement and OpenAI’s training set

Oxford presented the partnership in March 2025 as a way to digitise Bodleian texts using OpenAI software and make them more accessible to students and researchers. The announcement did not mention their use in model training, the Guardian reported.

Internal university documents obtained through a freedom of information request say material digitised by OpenAI has been used to “populate the OpenAI training set”. That gives the company a use for the texts beyond the public-access purpose Oxford described when it announced the partnership.

A university spokesperson rejected the suggestion that Oxford had hidden the training use from students or the public. The spokesperson said digitisation was Oxford’s principal aim and that staff had been open about the project’s contribution of training data.

OpenAI said it was “proud” to use present-day AI models to preserve historical knowledge. A spokesperson said that, with more than a billion people using the technology in daily life, the models should reflect a range of cultures, histories and perspectives.

125,000 Bodleian images shared by June

By June 2025, the Bodleian had shared 125,000 images scanned from historical dissertations with OpenAI. The material included PhD theses written at European and American universities in the 19th and 20th centuries. The Bodleian’s wider holdings comprise 23m items.

Texts scanned under the project also include 10,000 broadside ballads from the 16th century, containing lyrics and musical notation. Such ballads were once distributed on Tudor streets. University staff have also discussed digitising 18th-century Irish state papers, letters by the novelist Marie Edgeworth and Dorothy Hodgkin’s penicillin notebooks.

Oxford’s spokesperson described the amount being digitised as “modest in scale” and said it comprised only out-of-copyright texts. The spokesperson said OpenAI’s use was not exclusive and that the Bodleian retained the rights to make the scans available itself.

The Bodleian’s physical collections remain intact under the agreement. OpenAI rival Anthropic has spent tens of millions of dollars buying books, removing their spines for scanning and then pulping them. Anthropic has said it does not buy and destroy rare or antiquarian books.

NextGenAI agreements and Oxford staff concerns

OpenAI has reached similar agreements with Boston Public Library, Caltech, MIT and the University of Michigan through NextGenAI. Oxford is the project’s only UK member. The arrangements give OpenAI access to academic collections at a time when websites scraped for training data increasingly contain AI-generated material.

That shift has drawn developers towards physical collections, including historical books, as sources of training material. The Bodleian arrangement combines that access with Oxford’s stated aim of putting digitised texts within reach of more students and researchers.

Minutes from Oxford meetings record concerns among staff, including members of the Bodleian governance committee, about reputational exposure from the OpenAI partnership and its effect on the university’s environmental commitments, the Guardian reported. The concerns concerned a deal involving energy-intensive technology.

The same meetings included discussion of an “Ask the Bod” chatbot. Oxford’s spokesperson said the library would begin publishing material digitised with OpenAI openly online within months, as it does with material from other digitisation partnerships.

Topics: Copyright, Foundation models