Before the Molecule: Why Drug Development Needs Causal Human Biology Earlier
Considering causal biology early on could improve the chances of trial success.
Drug development failure is often discussed as a clinical trial problem. A study did not meet its endpoint. The patient population was too broad. The biomarker was not informative. The effect size was smaller than expected.
Those explanations may be true, but they often describe the moment failure becomes visible, rather than the moment failure begins.
In many programs, the decisive risk is introduced much earlier, when a target is selected without enough evidence that it causally drives disease in humans. From that point forward, the organization may do many things well. The chemistry may improve. The assay package may become more elegant. The clinical operations may be strong. But if the target does not cause the disease, or its progression, a drug development program will always fail.
The field has become remarkably good at generating therapeutic candidates. Medicinal chemistry, biologics engineering, RNA technologies, protein design, high-throughput screening, and AI-enabled discovery have expanded what can be built, and at an unprecedented speed. But better molecules still depend on whether the drug’s target actually causes the disease. The first question should not be whether a target is interesting; it should be whether the target is likely to matter in humans.
Relying on the literature does not work
At 5 Prime, we have undertaken a thorough curation of over 100,000 clinical trials to understand simple questions: Did the trial fail? Why did it fail? What predicts whether a drug development team launches a clinical trial? What predicts whether a drug makes it to market?
What we found was that, by far, the most important determinant of whether a target is pursued in a clinical trial was the number of publications focused on the target. People pursue targets that they know, or that the community is familiar with.
The problem is that the number of publications focused on the target has absolutely no effect on whether the trial was successful.
Picking targets based on the literature is a reliable way to convince your colleagues and investors to support a trial. Unfortunately, it has no effect on trial success. This can be improved upon.
Association is not the same as causation
Modern biology produces an enormous number of associations. A gene may sit near a disease-associated locus. A protein may correlate with disease severity. A metabolite may be elevated in patients. A pathway may appear repeatedly in model systems or published literature.
Each of these observations can be valuable. None is automatically causal.
That distinction matters because drug development is an intervention, not an observation. A therapeutic program is based on the belief that changing a target will change the course of disease. If the evidence only shows that a target travels with disease, rather than helps drive it, the program may be built on a correlate of pathology rather than a lever of pathology.
This is a common problem in translational research. Disease processes generate many downstream signals. Inflammation, tissue injury, metabolic disruption, and immune activation can all produce measurable changes that look biologically compelling. But treating a downstream consequence may not alter the disease mechanism that matters most for patients.
Human genetics can help separate these possibilities. It is not perfect, but it offers a way to ask whether naturally occurring variation in or near a target is associated with human disease outcomes in a way that supports therapeutic modulation.
Genetic evidence improves odds, but interpretation is everything
Multiple studies have shown that drug targets with human genetic support are more likely to succeed in clinical development. This has become an important idea in the industry, and rightly so. But it is sometimes reduced to a slogan: genetic support is good.
The reality is more nuanced. Genetic support can mean many different things. A common variant may implicate a broad genetic locus, but not identify the causal gene. A rare coding variant may clarify gene function, but be too rare to provide statistical power across many outcomes. A protein quantitative trait locus may support a circulating protein as causal for a phenotype, but only if the instrument is well understood and the direction of effect fits the therapeutic hypothesis.
The practical value of genetic evidence depends on its ability to sharpen a decision. It should help determine whether a gene is likely causal, whether inhibition or activation is more plausible, whether the biology relates to disease onset or progression, and whether on-target effects suggest clinical liabilities.
Weakly interpreted genetics can create false confidence. Well-interpreted genetics can prevent programs from advancing on attractive but fragile assumptions.
Mendelian randomization as a translational tool
Mendelian randomization has become especially useful because it can address some of the confounding and reverse causation that limit observational biology. By using genetic variants as proxies for biological perturbation, it can test whether differences in a protein, metabolite, or other measurable trait are likely to influence disease outcomes. This is because genetic variation is randomly assigned at conception, much like treatment assignment in a randomized controlled trial. Such randomization prevents confounding. Since genotypes are always assigned prior to disease onset, they aren’t subject to reverse causation.
This makes it relevant to drug development in a very practical way. If genetic variants that lower a protein are associated with lower disease risk, that may support a therapeutic strategy aimed at inhibiting the target. If variants that raise a biomarker do not increase disease risk, then the biomarker may be a correlate rather than a cause. If the disease appears to influence the biomarker, rather than the biomarker influencing the disease, then the therapeutic hypothesis may need to be reconsidered.
The strength of Mendelian randomization is not that it replaces randomized trials. It does not. Its strength is that it can expose weak causal assumptions before years of work and substantial capital are committed to testing them clinically.
Direction matters as much as selection
Choosing the right target is only the first step. The next issue is how to modulate it.
Many programs fail to make this distinction clearly enough. A target may be involved in disease biology, but that does not mean every form of intervention will be beneficial. In some settings, reduced function may be protective. In others, reduced function may be harmful. The same pathway can behave differently across tissues, disease stages, and patient subgroups.
Human genetic variation can provide useful clues. Loss-of-function variants may resemble partial therapeutic inhibition. Coding variants can reveal the phenotypic consequences of altered gene function. Protein quantitative trait loci can help connect genetically influenced protein levels to disease outcomes. Expression and tissue-enrichment data can add context about where the biology may be most relevant.
These data do not make the development path obvious. They do, however, reduce the chance that teams advance a target without knowing whether they should inhibit it, activate it, or approach it in a more selective way.
Disease progression is often the overlooked biology
Most genetic studies focus on disease susceptibility: who is more likely to develop a condition. That has generated many valuable insights, but it does not always answer the question a drug developer needs to ask.
Many medicines are developed for patients who already have disease. In those cases, the more relevant biology may be progression, not susceptibility. A gene that influences whether someone develops disease may not be the same gene that determines how quickly the disease worsens, which organs are affected, or which patients respond to treatment.
This distinction is particularly important in autoimmune conditions and neurodegeneration. This is because these disease classes often incur a tipping point phenomenon, after which disease progression factors are different from disease susceptibility factors. A target linked to disease susceptibility may be less useful if it does not alter the biology of established disease. Conversely, progression biology may reveal targets that would not stand out in a conventional susceptibility-focused analysis.
Bringing progression biology into target evaluation can change the shape of a program. It can influence endpoint selection, patient enrichment, biomarker strategy, and the level of confidence around clinical translation.
Safety should be considered before the first patient is dosed
Human genetics is often discussed in the context of efficacy, but it can also inform safety. Variants that mimic partial activation or inhibition of a target may reveal phenotypic consequences of long-term target perturbation. Phenome-wide analyses can identify traits or diagnoses associated with genetically influenced target function.
This type of evidence should not be overinterpreted. Lifelong genetic variation is not identical to short-term pharmacologic intervention. Genetic instruments vary in quality. Some associations will reflect biology that is not clinically meaningful.
Even with those caveats, the information can be useful. It can highlight plausible on-target liabilities, guide early monitoring plans, and help teams distinguish expected biology from unexpected toxicity. In some cases, it may reveal a safety concern serious enough to change the indication, dose strategy, or development plan.
The point is not to replace preclinical toxicology or clinical observation. The point is to make safety planning less reactive.
The standard should be actionability
The best use of human genetics and multi-omics is not to generate a larger evidence package. It is to produce evidence that changes action.
A useful analysis should influence whether a program advances, pauses, or stops. It should clarify whether the therapeutic hypothesis is biologically coherent. It should identify which biomarkers are mechanistically connected rather than merely convenient. It should help define the patients most likely to show a signal. It should make the safety plan more specific.
That is a higher standard than statistical significance alone. A result can be statistically interesting and still not be decision-relevant. Drug development needs evidence that can survive contact with portfolio committees, clinical teams, translational scientists, and ultimately, patients.
Moving causal evidence upstream
Causal human biology is most powerful when it is considered early. Once a program has already entered expensive clinical development, the ability to change course becomes more limited. Teams become invested in the target, the asset, the indication, and the clinical plan.
Earlier evaluation creates more freedom. A weak target can be stopped before it consumes years of work. A strong target can be advanced with greater confidence. A biomarker can be selected because it reflects the mechanism, not because it is easy to measure. A trial can be designed around the patients most likely to benefit.
Drug development will always involve uncertainty. Biology is complex, and not every successful medicine will have a simple genetic story. But uncertainty is not the same as guesswork.
The industry has invested heavily in making better molecules. The next step is to be equally rigorous about the human causal evidence behind the targets those molecules are built to modulate. The best medicines begin before the molecule, with a biological hypothesis strong enough to justify everything that follows.