AI-based protein structure models could be improved by incorporating data from pharmaceutical companies. Credit: Miyako Nakamura/Getty For drug discovery, protein folding models like AlphaFold have a data problem: There aren’t enough of them in public databases. Some scientists argue that improving the performance of these AI-based tools will require additional data that provides examples of how proteins and drugs interact. Protein structures, locked away by the thousands in the vaults of pharmaceutical companies, offer a promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding markedly improves model performance. The group used OpenFold3, an open source replication of AlphaFold 3, to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparables trained solely on public data and those trained on isolated data sets from individual companies. The study, described in a blog post, has not been peer-reviewed and the model is not publicly available. “If you add all this data, you get a pretty big increase in performance,” says Mohammed AlQuraishi, a computational biologist at Columbia University in New York City, who was part of the effort. AlphaFold is running out of data, so pharmaceutical companies are building their own version. The findings, he says, strengthen the case for generating similar publicly available data sets to boost protein folding. AI. One such project, called OpenBind and backed by up to £8 million ($10.8 million) in UK government funding, launched hundreds of new protein structures last month, and thousands more are in the pipeline. An untapped vein The Protein Data Bank (PDB), an open repository of more than 200,000 experimentally determined protein structures, was the basis of AlphaFold 2’s training data. It allowed the tool to predict protein structures with surprising accuracy, a breakthrough recognized with the 2024 Nobel Prize in Chemistry. Successors to the model, including AlphaFold 3, added the ability to predict how proteins will interact with other molecules, including potential drugs. But the PDB has relatively few examples of experimentally determined structures interacting with drug-like molecules: perhaps only 10,000, says Paul Mortenson, vice president of computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, United Kingdom. That lack of data is a problem for drug discovery efforts. Research has suggested that the accuracy of AlphaFold 3 and other ‘cofolding’ models, which predict the structure of proteins that interact with each other, falls off a cliff when the models are challenged to predict interactions between molecules very different from those on which they were trained1. For these tools to be more useful in drug discovery, the researchers say, they need access to more data, which is why they have turned to pharmaceutical companies’ vaults of molecular structures. revolution These data are generated during drug discovery programs, using techniques such as X-ray crystallography and cryo-electron microscopy. Many of the protein structures have never been deposited in public databases because they are related to proprietary drug development efforts. The total size of these vaults is unknown, but some have estimated that they could contain more data than the PDB contains. “The data missing from the PDB is exactly the data that is present in our internal data,” John Karanicolas, head of computational drug discovery at pharmaceutical company AbbVie in Chicago, Illinois, told Nature last year. Better predictions To test whether their data could be useful for protein folding models, AbbVie, Astex and several other pharmaceutical companies last year formed a collaboration called the AI Structural Biology Network (AISB). It involved “tuning” OpenFold3 (previously trained on PDB data alone) on an additional 20,167 structures that capture proteins bound to potential drugs or ligands. The structures came from five companies and were provided to the model in such a way that the proprietary data remained private. The AISB study found that the additional data improved predictions. When tested on 1,056 protein-ligand structures that deviated from the training data, the AISB model predicted more than half of them with a high level of accuracy. By contrast, the public version of OpenFold3 achieved the same performance on only a third of the structures, and a competing open source model called Boltz-2 achieved about 40%. The team plans to submit a paper describing the work to a peer-reviewed journal. The fact that the AISB model also outperformed co-folding tools that were trained only on each company’s individual data highlights the benefits of pooling information, Karanicolas says.