We've updated our Privacy Policy to make it clearer how we use your personal data. We use cookies to provide you with a better experience. You can read our Cookie Policy here.

Advertisement

Why 90% of Cancer Drugs Fail: The Data Gaps in Oncology

A clinician holds a tablet with icons relating to data in healthcare floating above it.
Credit: iStock.
Read time: 8 minutes

Over 90% of cancer drug candidates fail during development. Promising candidates often advance through preclinical studies and early trials, yet fail during clinical development.  


Biological data can help to shape everything from target identification to clinical trial design in cancer drug development. Incomplete or poorly contextualized datasets in early discovery can lead to incorrect assumptions about target expression, patient selection, or expected drug response, derailing seemingly promising candidates later down the line. 


As datasets grow in both scale and complexity, so do the challenges of ensuring they are predictive, translatable, and clinically meaningful. 


Dr. Tammer Farid, general manager, Data at Champions Oncology, navigates the intersection of tumor biology, model systems, and translational data. Technology Networks spoke with Farid to discuss the most common data failures in oncology R&D, the limits of current preclinical models, the promise of multiomics and AI, and the evolving data ecosystem shaping the next decade of cancer drug development. 

Building reliable biological datasets for oncology drug discovery 

What do you think are the most common issues with biological data in early oncology drug development, and how does this cause drug candidates to fail later down the line? 

Early‑stage oncology programs often struggle with two foundational data challenges: generating enough high‑quality information to build robust analytical models and ensuring that the data collected truly reflects patient biology. 


While patient‑derived samples offer the most direct clinical relevance, it can be challenging to obtain enough material to generate multiple types of data. In contrast, model systems such as patient-derived xenografts (PDXs), organoids, and cell lines can be characterized more comprehensively, but their utility hinges on whether the insights they provide are translatable to patients.  

“Incomplete information can lead to false assumptions about target patient populations or likely drug responses.” —Dr. Tammer Farid 

Farid explained that incomplete or poorly contextualized datasets can lead to incorrect assumptions about drug targets. For example, transcriptional data may not accurately predict durable protein expression in tumors, creating a false sense of confidence in a target’s druggability. 


“The leading cause of drug candidates failing later down the line is due to losing the expected efficacy from the model system to the patient,” he explained. “Fundamentally, this is an issue of low predictive power in the patient context.” 


Drivers of data quality in oncology R&D: 

  • Limited patient‑derived material restricts multi‑modal data generation 
  • Deeply characterized model systems still require proof of translatability 
  • Misaligned datasets can distort assumptions about target expression and response 
  • Predictive failure in patients remains the dominant cause of late‑stage attrition 

Where preclinical models fall short in capturing tumor complexity 

How well do current preclinical models capture the complexity and heterogeneity of human tumors, and where do you think the biggest translational gaps still lie? 

Preclinical models vary widely in their ability to reflect human tumor biology. Cell lines remain key for early discovery due to their scalability and cost‑effectiveness, but they lack the heterogeneity and structural complexity of real tumors. 


Advertisement

PDXs and organoids offer a closer approximation of tumor biology, with PDXs preserving patient-specific genomic architecture and histology, and organoids retaining intrinsic tumor genotypes and phenotypes. 


Despite these advantages, no model fully captures the dynamic interplay of the tumor microenvironment (TME), immune interactions, or clonal evolution. Farid emphasized that these gaps are central to why many therapies fail to translate. 


“No model fully recapitulates the complexity and heterogeneity seen in human biology.” — Dr. Tammer Farid 


The most significant limitations today involve modeling the TME, immune–tumor interactions, and the emergence of resistance. These factors profoundly influence therapeutic response but remain difficult to replicate in preclinical models. 


Translational gaps in tumor modeling: 

  • Cell lines lack heterogeneity and structural complexity 
  • PDXs and organoids improve relevance but still miss immune and microenvironmental factors 
  • Tumor‑immune interactions remain poorly modeled 
  • Clonal evolution and resistance mechanisms are difficult to capture over time in model systems 

Advertisement

Using biological insight to improve study design and reduce failure risk 

How can drug developers use a deeper understanding of tumor biology to shape better drug study design and reduce the risk of failure during development? 

Biological data can improve drug development at multiple stages. “Upstream of study design, the right biological data can help identify novel drug targets, prioritize potential targets, and elucidate mechanisms of action,” Farid explained.

  

Insights into cancer biology can also guide the selection of in vitro or in vivo models, helping developers choose systems most likely to generate translatable results. This, in turn, supports more precise patient stratification and trial design.  


For successful drug candidates, understanding their mechanisms can reveal additional indications or combination strategies worth exploring. 


By grounding decisions in mechanistic evidence, developers can improve the translatability of drug candidates from preclinical models to patients and increase the likelihood of clinical success. 


Biology‑driven study design advantages: 

  • Strengthens target discovery and prioritization 
  • Improves model selection for translatable preclinical studies 
  • Enables more precise patient stratification and trial design 
  • Supports indication expansion for successful candidates 

Overcoming AI limitations through better data integration 

From your perspective, what limits the performance or reliability of AI in oncology R&D, and how might this be overcome? 

Advertisement

AI models are advancing rapidly, but their performance is still constrained by the fragmentation of high‑quality biological data. 


“The greatest challenge today is the integration of data from multiple sources,” said Farid. “AI algorithms themselves are rapidly improving our ability to resolve this.”  

 

While algorithms continue to improve, Farid noted that the biggest bottleneck is the lack of integrated, multi‑source datasets that reflect the full complexity of disease.  


“High-quality data tends to be siloed across many organizations, limiting the ultimate potential of disease models.”—Dr. Tammer Farid 

Data silos—spread across institutions, companies, and research groups—limit the ability of AI systems to build comprehensive, predictive models. Farid suggests that “companies should explore partnership models to integrate deeper data and improve the predictive power of the resulting models.”  


Advertisement

“Patients will derive the greatest benefits if holders of different data modalities can effectively layer their resources for higher-powered analyses and greater model sophistication,” he noted. 


By layering diverse data modalities and pooling resources, more powerful, clinically relevant AI models that better predict patient outcomes can be created. 


AI performance is constrained in oncology through: 

  • Fragmented datasets that limit model accuracy 
  • Challenges in integrating multiple data modalities 
  • A lack of partnerships, which are essential to build richer, more predictive datasets 
  • Limitations to algorithms and data accessibility 

The promise and challenges of multiomics in cancer research 

Multiomics datasets hold promise for improving our understanding of tumor biology and accelerating drug discovery. What do you see as the biggest opportunities and barriers in this space? 

Multiomics offers a path to deeper, causal understanding of tumor biology by integrating genomic, transcriptomic, proteomic, and functional data.  


“The biggest opportunities lie in integrating multiple omic data layers for each model system, enabling the pairing of observational and functional data to build true, causal understanding of biological systems,” Farid explained.  


He also highlighted how emerging multiomic data types—such as comparisons between whole-cell and cell surface-specific proteomes—along with greater access to high‑resolution data, are enabling the creation of functional datasets like chemical and genetic Perturb-seq. 

Advertisement


These datasets can then be used to reveal disease mechanisms, drug targets, and drug mechanisms of action.  “If these can be paired with rich multiomic data, their utility increases,” Farid said. 

 

However, the same factors that create opportunity also create barriers. Multiomic integration is only as strong as the underlying data sources, making it essential to characterize systems that accurately reflect patient populations, especially those with unmet needs. 


Multiomics opportunities and obstacles: 

  • Integrated datasets enable causal biological insights 
  • High‑resolution proteomics and functional assays expand discovery potential 
  • Perturb‑seq and related tools strengthen mechanistic understanding 
  • Success depends on high‑quality, clinically relevant model systems 

How oncology data ecosystems are evolving 

What are some of the biggest changes you have seen in oncology data ecosystems, and how do you think they will continue to evolve over the next five to ten years? 

The oncology data landscape has expanded dramatically in both breadth and depth over recent years. Researchers now have access to more samples, more studies, and more data layers per sample than ever before. 

 

Advertisement

Yet, as Farid pointed out, no dataset excels in both granularity and scale simultaneously—large datasets often lack depth, while deeply characterized datasets remain small. 


He anticipates that “the rapid advancements in interpretation power delivered by AI models will create demand for more fully characterized models, with additional focus on patients with the greatest unmet need.” This includes characterizing refractory tumors, rare cancers, and cases where standard‑of‑care treatments fail. 


As data ecosystems mature, the focus is likely to shift toward completeness, clinical relevance, and the ability to model complex patient populations more accurately. 


Future directions in oncology data ecosystems: 

  • Simultaneous expansion of dataset breadth and depth 
  • Growing demand for fully characterized, multi‑layered models 
  • Increased focus on complex patient populations 
  • AI‑driven pressure to improve data completeness and relevance 

 

Biological data is becoming increasingly central to cancer drug discovery, but its value depends on quality, integration, and clinical relevance. While challenges of data scarcity, model translatability, and siloed information persist, opportunities are emerging through multiomics, mechanistic insights, and AI‑enabled data interpretation. As datasets grow and technologies evolve, the field is moving toward more comprehensive, patient‑aligned models that can better predict therapeutic success. 

Key takeaways:

  • Early‑stage failures often stem from incomplete or non‑translatable biological data 
  • Mechanistic biological insight strengthens study design, model selection, and patient stratification 
  • Patient-derived models improve relevance but still miss key tumor‑immune dynamics 

  • AI- and multiomics-based approaches hold promise for cancer drug discovery, but still face challenges 


This content includes text that has been created with the assistance of generative AI and has undergone editorial review before publishing. Technology Networks' AI policy can be found here.


Google News Preferred Source Add Technology Networks as a preferred Google source to see more of our trusted coverage.