Key takeaways

  • Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover…
  • The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in…
  • ” Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in…

What happened

Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage. Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years—a phenomenon known as Eroom’s Law. 5 billion, with failure rates upward of 90%. AI has become the pharmaceutical industry’s biggest bet on bringing success rates up and timelines down.

Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. Because models have access to the same data, they all reach similar conclusions, with diminishing returns over time. Additionally, the datasets weren't built with AI in mind, meaning they lack the structure, labeling, and diversity needed to keep models accurate and free of bias. Publication bias reinforces the problem.

“Most publicly available datasets and scientific publications focus exclusively on positive results,” says Belcher. “No one wants to share their failures. This bias is almost like having one hand tied behind your back. ” The data Belcher believes would markedly improve models—the failed experiments, the compounds that don’t bind—remains frustratingly difficult to come by. “We often joke that there should be a journal of negative data,” he says.

” This lack of negative data creates a fundamental problem: Without access to a broad range of data, models can’t be adequately trained to avoid bias. “In all machine learning applications, the model’s performance relies heavily on the quality and scope of the training data,” notes Belcher. Fabrication has also become much easier with AI, compounding concerns around data integrity. Take Western blots, for example.

Why it matters

The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in development. “The main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial,” says Paul Belcher, director of protein research strategy at global life sciences company Cytiva.

” Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in lab systems. One of the most promising early-stage applications of AI in drug discovery is in hit identification. This involves screening libraries of molecular entities against a disease-related target, such as a protein, to find molecules that bind to it.

A successful hit gives researchers a starting point for further testing and refinement, with the aim of eventually developing a viable drug. Belcher has seen a shift from empirical screening to predictive design: Instead of physically screening libraries, drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing anything to research and development (R&D).

This means companies are no longer limited by how much they can physically screen to identify starting points. “AI does away with that,” says Belcher. ” What AI can’t do yet is reliably predict kinetics or developability of new compounds, says Belcher. This means every AI-generated candidate still needs to be validated in the lab.

Traditional screening workflows were built to identify hits at scale, not to profile large numbers of complex candidates in detail. This is placing more pressure on lab teams, who now have to test, characterize, and purify a growing volume of more diverse, AI-generated compounds.

“The current techniques used in hit identification can screen hundreds of thousands, sometimes millions of compounds, using binary or threshold-based techniques producing low-fidelity data—yes-or-no responses,” Belcher explains. “AI can increase the number of hits you get and potentially give you better quality hits as well. ” As AI has accelerated demand for data-rich lab systems, it has also highlighted a fundamental need for better, more complete data.

What to watch

These are part of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. This was back in 2016, before generative AI made fabrication trivial.

“Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences,” says Belcher.