The University of Oxford has granted OpenAI permission to train its artificial intelligence models using historical texts from the Bodleian Library. This collaboration, announced in March 2025, involves OpenAI software digitising texts to make content more widely available for students and researchers.
Internal documents indicate that the Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”. This process allows AI models to recognise patterns in words and learn to construct sentences and perform other cognitive tasks.
While the university stated its primary interest was digitisation, minutes from Oxford University meetings, obtained via a freedom of information request, reveal concerns from staff, including members of the Bodleian governance committee. These concerns relate to the reputational risk of partnering with OpenAI and the environmental impact of using energy-intensive AI technology.
An OpenAI spokesperson stated the company is “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future.” The university spokesperson clarified that the amount of text being digitised is “modest in scale” and covers only out-of-copyright material. The Bodleian retains the rights to the scans and plans to publish them openly online within months.