Integrating Screening and Omics Data in Drug Discovery
Combining HTS data with genomics, transcriptomics, and proteomics to build richer drug discovery workflows.
High-throughput screening (HTS) and multi-omics technologies have each transformed drug discovery independently; their integration represents the next evolution in how laboratories move from compound libraries to validated therapeutic hypotheses. The drive to integrate screening and omics data reflects a fundamental limitation of either approach in isolation: HTS identifies compounds that alter a biological readout, but rarely explains why, while single-omics datasets capture one molecular layer of a complex system without the experimental context that screening campaigns provide. Together, they offer a more mechanistically grounded path from hit identification to target validation.
Modern drug discovery pipelines increasingly incorporate genomics, transcriptomics, proteomics, metabolomics, and epigenomics alongside functional screening data to generate richer, more translatable insights. As reviewed by Du et al., the shift from single-omics to integrated multi-omics analysis has become a defining trend in drug target identification, driven by the recognition that complex disease phenotypes cannot be adequately explained at any single molecular level.1 For an overview of the principles underpinning HTS and the analytical frameworks used to manage its data output, see High-Throughput Screening: Principles, Applications and Advancements and AI and Machine Learning in High-Throughput Screening.
Genomics and HTS: linking genetic evidence to compound screening
Genome-wide association studies (GWAS) have generated thousands of statistically robust disease–variant associations, but identifying the causal genes underlying GWAS signals remains a central challenge. Multi-omics integration - particularly the combination of GWAS data with expression quantitative trait loci (eQTL) and molecular QTL from other omics layers - has substantially improved the resolution of target identification from genetic data. Lessard et al. applied a combination of variant annotation, activity-by-contact maps, Mendelian randomization, and molQTL colocalization across 4,611 disease GWAS, demonstrating that genes prioritized by integrated multi-omics approaches are enriched for drug targets with successful clinical trial outcomes.2
Once genetic evidence supports a candidate target, HTS campaigns can be designed to screen for modulators of that specific gene product or pathway, with the genomic data informing both compound library selection and assay design. Conversely, active compounds identified in HTS can be mapped back to genetic databases to assess whether their primary targets coincide with disease-associated loci, providing an independent layer of translational confidence. This bidirectional relationship between genomics and HTS is increasingly formalized in systems biology approaches to target prioritization.
Transcriptomics and connectivity mapping in screening workflows
Transcriptomic profiling of drug-treated cells provides a genome-scale readout of how a compound perturbs gene expression - a mechanistic signature that can be compared across compounds, cell types, and disease states. The Connectivity Map (CMap), introduced by Lamb et al. in 2006, established the foundational framework for this approach by constructing a reference library of gene-expression profiles from cells treated with bioactive small molecules and using pattern-matching to identify functional connections between drugs, genes, and diseases.3
The concept was substantially scaled by the LINCS L1000 initiative, which generated over one million perturbation profiles across more than 29,000 chemical and genetic perturbagens in 98 cell lines, enabling large-scale connectivity analysis for drug repurposing and mechanism-of-action elucidation.4 In the context of integrated screening workflows, transcriptomic profiling of confirmed HTS hits allows researchers to interrogate whether active compounds share expression signatures with known drugs, act through anticipated pathways, or produce unexpected transcriptional responses that may indicate off-target activity or toxicity. This positions transcriptomics as a post-HTS triage layer capable of substantially narrowing the hit-to-lead gap.
Statistical frameworks for multi-omics and drug screening data integration
Combining HTS outputs with one or more omics datasets requires analytical frameworks capable of handling data of fundamentally different types, scales, and dimensionalities. Simple correlation between, for example, gene expression and drug sensitivity across cell lines is informative but does not capture the latent structure present when multiple omics layers are considered jointly. Probabilistic data integration methods have been developed to address this challenge.
El Bouhaddani et al. presented a computational workflow combining multi-omics integration - specifically transcriptomics and proteomics - with drug screening data from cell line models of synucleinopathies, using a probabilistic partial least squares approach to identify genes and proteins distinguishing affected and control cells and prioritize potentially druggable pathways.5 The same group's OmicsPLS software package has since been adopted as a practical tool for joint omics data integration in a range of biological contexts.8 This framework illustrates how statistical integration of omics and screening data can generate mechanistic hypotheses from otherwise disconnected datasets, moving beyond hit lists to pathway-level understanding of disease biology and compound activity.
Network-based approaches represent another major class of integration methods. A systematic review by Jiang et al. catalogued four primary types of network-based multi-omics integration methods used in drug discovery - network propagation/diffusion, similarity-based approaches, graph neural networks, and network inference models - each suited to different combinations of data types and discovery objectives.6 These methods can integrate protein–protein interaction networks, gene expression profiles, and compound activity data to identify novel drug targets, predict drug response, and identify repurposing candidates in ways inaccessible to any single-omics or single-assay analysis.
Table 1. Omics layers and their contributions to integrated high-throughput screening workflows.
Omics layer | Information contributed | Role in integrated screening workflows |
Genomics | Disease-associated variants; candidate target genes from GWAS | Prioritizes targets with genetic validation; directs compound library design |
Transcriptomics | Gene expression changes under drug or genetic perturbation | Connectivity mapping; mode-of-action profiling; hit contextualization |
Proteomics | Protein abundance, modifications, and interaction networks | Confirms target engagement; identifies off-target effects; refines SAR |
Metabolomics | Small-molecule metabolites reflecting cellular metabolic state | Reveals downstream phenotypic consequences of hits; supports biomarker discovery |
Epigenomics | Chromatin accessibility, histone marks, DNA methylation patterns | Identifies regulatory context of GWAS loci; informs cell-state-specific target biology |
High-throughput multi-omics platforms and single-cell applications
The development of high-throughput multi-omics platforms has enabled the simultaneous acquisition of multiple molecular data types from the same biological sample or cell population, removing the batch and sample mismatch problems that previously complicated cross-modal integration. Mass cytometry (cytometry by time-of-flight; CyTOF), high-dimensional imaging technologies, and genomic cytometry approaches such as CITE-seq and REAPseq allow simultaneous measurement of protein expression and genomic features at single-cell resolution, generating data directly relevant to mechanistic screening programs.
Zielinski et al. reviewed how high-throughput single-cell multi-omics platforms have enabled unprecedented characterization of cellular phenotypes, immune effector functions, and metabolomic states in the context of clinical trial evaluation and drug discovery, noting that these approaches are constructing whole atlases of cell types and interaction networks relevant to therapeutic development.7 For screening applications, this means that active compounds can be profiled not only for aggregate population effects but for cell-state-specific responses — an important capability when the relevant biology occurs in rare or phenotypically heterogeneous subpopulations.
Challenges in translational bioinformatics and integrated data analysis
Despite the scientific appeal of integrating screening and omics data, significant analytical and practical challenges limit the routine adoption of these approaches in drug discovery settings. Data heterogeneity - differences in scale, noise structure, missingness, and biological context across omics layers - makes joint analysis technically demanding. Batch effects, which arise from experimental variation across time points, instruments, or cell passage number, can introduce systematic noise that confounds biological signal when datasets are merged across studies or platforms.
The dimensionality of integrated multi-omics datasets frequently exceeds the number of biological samples available, creating statistical power limitations that are compounded when rare disease models or patient-derived cell lines are used. Data standardization across institutions and platforms remains a recognized bottleneck: omics measurements from different laboratories may not be directly comparable without extensive normalization, and the absence of universally accepted integration benchmarks makes it difficult to evaluate method performance across contexts.
Computational infrastructure requirements for large-scale integration workflows are substantial, and the specialized bioinformatics expertise needed to deploy, validate, and interpret these methods is not uniformly available across the drug discovery community. Addressing these barriers - through improved data standards, open-source tooling, and collaborative data-sharing frameworks - is a recognized priority for translational bioinformatics as the field matures.
Key considerations for integrating screening and omics data:
- Genomic evidence from GWAS, when combined with eQTL and other molecular QTL data, can substantially enrich HTS target selection and increase the probability of clinical translation.
- Transcriptomic connectivity mapping provides a mechanism-agnostic method for contextualizing HTS hits, linking active compounds to known biology and disease-relevant gene expression programs.
- Statistical integration of multi-omics and drug screening data requires methods capable of handling high dimensionality, data type heterogeneity, and incomplete overlap between samples.
- High-throughput single-cell multi-omics platforms enable cell-state-specific resolution of compound activity, revealing biology inaccessible to bulk population assays.
- Batch effects, data standardization gaps, and computational infrastructure demands remain practical barriers to routine integration and should be addressed in workflow design from the outset.
Integrating screening and omics: implications for drug discovery workflows
The integration of HTS data with genomics, transcriptomics, proteomics, and other omics layers represents a genuine expansion of what drug discovery campaigns can achieve with the same experimental investment. Screening hits that would previously have been triaged on activity alone can now be filtered and prioritized using molecular context - their consistency with genetic evidence, their mechanistic alignment with disease-relevant pathways, and their capacity to engage targets in a therapeutically meaningful way.
Realizing this potential requires investment in both analytical infrastructure and data governance: the methods for integration exist and are improving rapidly, but their deployment in routine discovery workflows depends on standardized data formats, interoperable databases, and multidisciplinary teams capable of bridging experimental screening and computational biology. As these foundations mature, the combination of functional screening and multi-omics profiling is positioned to deliver mechanistically richer hit lists, better-validated targets, and more predictive translational models than either approach can provide in isolation.
This content includes text that has been created with the assistance of generative AI and has undergone editorial review before publishing. Technology Networks' AI policy can be found here.
1. Du P, Fan R, Zhang N, Wu C, Zhang Y. Advances in integrated multi-omics analysis for drug-target identification. Biomolecules. 2024;14(6):692. doi: 10.3390/biom14060692
2. Lessard S, Chao M, Reis K, et al. Leveraging large-scale multi-omics evidences to identify therapeutic targets from genome-wide association studies. BMC Genomics. 2024;25(1):1111. doi: 10.1186/s12864-024-10971-2
3. Lamb J, Crawford ED, Peck D, et al. The connectivity map: using gene-expression signatures to connect small molecules, genes, and disease. Science. 2006;313(5795):1929–1935. doi: 10.1126/science.1132939
4. Subramanian A, Narayan R, Corsello SM, et al. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell. 2017;171(6):1437–1452.e17. doi: 10.1016/j.cell.2017.10.049
5. El Bouhaddani S, Höllerhage M, Uh HW, et al. Statistical integration of multi-omics and drug screening data from cell lines. PLoS Comput Biol. 2024;20(1):e1011809. doi: 10.1371/journal.pcbi.1011809
6. Jiang W, Ye W, Tan X, et al. Network-based multi-omics integrative analysis methods in drug discovery: a systematic review. BioData Min. 2025;18(1):27. doi: 10.1186/s13040-025-00442-z
7. Zielinski JM, Luke JJ, Guglietta S, Krieg C. High throughput multi-omics approaches for clinical trial evaluation and drug discovery. Front Immunol. 2021;12:590742. doi: 10.3389/fimmu.2021.590742
8. El Bouhaddani S, Uh HW, Jongbloed G, et al. Integrating omics datasets with the OmicsPLS package. BMC Bioinformatics. 2018;19(1):371. doi: 10.1186/s12859-018-2371-3