AI Is Accelerating Drug Discovery, but Patients Will Prove Its True Value
AI has cut drug discovery timelines, but its ultimate impact will depend on clinical trial outcomes.
It’s no secret that many drug candidates entering clinical trials today will never reach patients.
The vast majority fail, in part because human biology is far more complex than what can be modeled in preclinical work.
The emergence of new approach methodologies and AI marks a new era in which we no longer need to rely solely on traditional models. In particular, AI tools not only predict outcomes—increasing the chances of success in the later stages of development—but also enable the discovery of novel targets that were once inaccessible, and the development of therapeutics that were previously unattainable.
Having previously explored the role of digital twins in drug development, Dr. Jean-Philippe Vert, co-founder and chief executive officer at Bioptimus. Vert spoke with Technology Networks to discuss AI approaches used in drug discovery, their impact so far, and the barriers limiting their full potential.
This article is the first in a series exploring the role of AI across the drug discovery and development pipeline, covering preclinical study conduct, virtual control groups, and clinical trials.
What are the main AI approaches currently used in drug discovery, and which therapeutic areas have seen the most meaningful impact?
One big thing AI has enabled is the design of new small and large molecules. In terms of the techniques used, it's quite similar to those in large language models (LLMs). For example, a few years ago, AlphaFold (an LLM for proteins) was developed by DeepMind to predict the structure of proteins from their sequence. The model has evolved since then. Now we can give it a sequence, and it can predict protein structure and interactions, and even design new proteins with target structures.
The main tools used for protein design combine LLMs with diffusion networks or generative AI. Together, these are the core technologies used to both understand the language of proteins and generate new, therapeutic ones.
Combining AI approaches for drug discovery
An LLM uses statistical patterns and mathematical probability to generate a response. Outputs are generated step-by-step or word-by-word, and are typically text or code. Diffusion models create an output by starting with a sheet of random mathematical noise, almost like a screen of television static, and then work spatially to predict what to subtract until the image becomes clear.
These tools are strong when you have a target, but there are many diseases for which we have no idea what a good target is. In oncology, we have seen significant progress, in part because we understand the biology. This might not be true across all cancers, but in many cancers, there is a clear target—perhaps a protein in your body that becomes super active—and we know that if we can interfere with it, we can slow or stop cancer. So, I think oncology is probably the one area that benefits most today.
What are the most significant benefits AI has delivered in drug discovery to date?
AI is accelerating drug discovery, from target identification to getting a candidate drug into clinical trials. The most impactful effect of AI so far has been to reduce the time required.
It's not as if we couldn’t do these things before, but companies say it used to take five to seven years—now it's less than two.
The time is compressed in these stages, and spending is reduced, but it's creating a new problem. Too many candidates will hit clinical trials. We already see difficulties in conducting clinical trials: finding enough patients and making the best use of their time. We want to make sure we give patients the best chance. While AI has already revolutionized target discovery to preclinical work, the next phase is really, how can AI de-risk, optimize, and accelerate clinical trials? Once we test the drug on patients, how can we improve? Understanding patient responses is much more complex.
When we send a drug to clinical trials, even with the best minds and biggest budget, we still have a 90% failure rate. This is where AI excels at understanding, simulating, and generating hypotheses. Biology is a complex domain, and when you have a lot of complex data, AI is good at making connections.
Until we have a tsunami of new drugs on the market, I want to remain cautious that what we have today is expectations.
What support do AI-generated hypotheses require to progress into preclinical studies, and are we seeing—or when might we expect—a meaningful reduction in resource use?
We are already seeing a reduction in time, and often time equals budget as well. The goal is to screen more efficiently and to identify candidates with strong potential.
But, while AI is almost autonomous in the computational world, you would never go to clinical trials based on that alone. This is where the “lab-in-the-loop” approach comes in. For an AI to be efficient in the drug discovery process, it must be coupled with something real. A lab can generate proteins and molecules to test them in cells, and so on.
So, in terms of time and money, the computational aspect is faster, but the rest largely stays the same because we still need to confront the AI predictions in the real world.
How can AI models for drug discovery be effectively trained and iterated in a timely way?
The field is trying to find faster ways to iterate, rather than waiting the 10 years it takes for a clinical trial to succeed or fail.
Ultimately, a clinical trial's success is still the only thing that matters. At the end of the day, this is how you get treatment to patients who need it. But it's not fast enough for AI to iterate. So, we have learned that this can be done during the pre-clinical period, by feeding the LLMs data collected in cell cultures, organoids, animal models, and so on.
Then, at each stage of a clinical trial, there is an opportunity for iteration.
This helps improve decision-making earlier in the process and enables more accurate AI predictions. There's lots of research happening there, trying to pick up on early signals that a drug may or may not be working. That's the only way we can make the right decisions earlier.
How can the field address the AI “black box” challenge and reduce data silos to support greater collaboration?
There are some models for which the black box aspect is acceptable. In this instance, I mean that we know the network's architecture, its physical properties, and so on. The output we want is going from a question to a hypothesis. In this case, it is hard to interpret exactly what the AI did, just as with LLMs. When you ask a question, you get an answer. It's not interpretable. But if the answer is good, that's enough. In such cases, I would say there is no issue with the black-box nature of AI.
Black boxes in AI
In the context of AI, a black box can mean two things. Firstly, the non-explainability of AI systems, as black box models mean inputs and outputs are seen but not the processing steps between. Secondly, it can also refer to data silos, as AI systems and their proprietary information are typically not shared.
There are other domains, including diagnostics or disease pathophysiology, where it is more useful to gain an explanation for that output. For example, if we ask, “Will the patient respond to the treatment?”, we care about the explanation. Having the answer is useful, but it would be even more useful if you could say, “The patient does not respond because a specific pathway or gene is activated and therefore…”
Once you have the explanation, you build trust in the system, and you can also make progress. In that domain, contrary to protein design, interpretability is important. This part is a technical problem, prompting AI solution developers to prioritize interpretability.
On the other hand, if the question is also about closed versus open models, this is not so much a technical question as a company decision. This is where you can also see data silos as well. Personally, I think models should be more open.
There would be so many benefits for society, patients, and scientific progress if data were shared.
This is more of an economic problem: How can we encourage individuals who have paid a lot of money to generate data to share data if there is no benefit for them?