AI models need more data about biology, and OpenAI is paying to create it

The OpenAI Foundation is funding a $40 million initiative called Data for Public Health to improve AI capabilities in medicine. The project aims to overcome data bottlenecks by digitizing and utilizing clinical trial data and biotech archives.
Why it matters
Access to high-quality, proprietary biological data is currently the primary constraint for developing AI models capable of medical breakthroughs.
The OpenAI Foundation is funding a new effort called Data for Public Health.
Last year, the clinical trial policy analyst Ruxandra Teslo posted an idea for super-charging medical AI systems: use data from failed biotech companies.
By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—information usually kept hidden as valuable trade secrets. She called these documents “ biotech’s lost archive ” and said they could be used to help train AIs that would act as powerful co-pilots in the often opaque drug approval process.
Today, the OpenAI Foundation, the nonprofit parent of OpenAI , said it would fund her idea as part of a new effort it calls Data for Public Health, which aims to help artificial intelligence make big leaps in medicine by funding the creation of “high-quality scientific datasets.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in