Inside the Dark Proteome: Biology’s Hidden Protein Universe
Researchers are uncovering the “dark proteome”—proteins that remain structurally and functionally hidden.
Despite decades of genomics, proteomics, and structural biology research, a substantial portion of the protein universe remains stubbornly unexplored.
This elusive realm is known as the dark proteome, and much like dark matter in physics, its presence is inferred from what we cannot fully explain.
For researchers working in this space, the dark proteome is not just a technical limitation; it is a signal that fundamental aspects of protein biology may still be missing from our understanding.
To understand what remains hidden, researchers are beginning to combine large-scale proteogenomics with advances in computational and structural biology.
Among them are Dr. Jinchuan Xing, a professor at Rutgers University, whose work focuses on uncovering previously unrecognized proteins through integrative proteomic approaches, and Dr. Nelson Perdigão, a former collaborator at the Institute for Systems and Robotics, Instituto Superior Técnico, University of Lisbon, whose work has helped map the extent of the dark proteome. Together, their research spans both the discovery of new protein forms and the broader challenge of understanding why so much of the proteome remains structurally uncharacterized.
“Despite the large amount of proteomics information the field has, our understanding of the proteome is far from complete,” said Xing.
Technology Networks spoke with Xing and Perdigão to understand what the dark proteome is, why it has proven so difficult to characterize, and what its discovery could mean for the future of biology and medicine.
What is the dark proteome?
The dark proteome refers to proteins or protein regions that lack identifiable structural homologs, have no experimentally resolved structures, or cannot be reliably predicted using current computational approaches. These are sequences that, when compared against known databases, yield little to no insight into how they fold—or whether they fold at all.
“The landmark 2015 paper that coined the term ‘dark proteome’ revealed that a very large fraction of proteins lacked known structures and could not be modeled by homology,” said Perdigão.
This definition overlaps with the concept of intrinsically disordered proteins (IDPs). While IDPs are characterized by their lack of stable three-dimensional structure, not all dark proteins are disordered. Some may fold only under specific cellular conditions, interact with other molecules to adopt structure, or represent entirely new structural classes not yet captured in existing databases.
The dark proteome vs IDPs
While often used interchangeably, the dark proteome and IDPs are not the same.
- IDPs are proteins or regions that lack a stable three‑dimensional structure under physiological conditions, yet remain functional—often through flexibility and dynamic interactions.
- The dark proteome is a broader concept, referring to proteins or regions that lack structural or functional characterization, whether experimentally or computationally.
The result is a heterogeneous category that defies simple classification—some dark regions may eventually be resolved as structured proteins, while others may challenge the very notion that stable structure is a prerequisite for function.
What makes the dark proteome interesting is its abundance. It spans across all domains of life, including humans, microbes, and viruses. Estimates suggest that a substantial portion of proteins, up to 30–50%, across organisms fall into this category.
In higher eukaryotes, the proportion of proteins lacking structural characterization is highest, with some studies estimating the figure to sit over 50%, and even in well-studied systems, a significant fraction of proteins remains “dark,” with ~30% of proteins in the model organism Caenorhabditis elegans lacking a functional annotation in UniProt.
Some researchers believe the dark proteome may also represent a reservoir of evolutionary innovation, where new protein functions can emerge more rapidly than in well-conserved structural families.
Why is the dark proteome difficult to study?
If the tools of modern biology are so powerful, why does the dark proteome persist?
The answer lies in a combination of computational, experimental, and biological challenges.
On the computational side, many dark proteins do not resemble any known proteins. They may lack sequence similarity to characterized proteins, belong to small or rapidly evolving families, or contain unusual amino acid compositions. Even powerful AI-driven tools are limited by the data they are trained on—and dark proteins, by definition, sit outside those boundaries.
Experimental approaches add another layer of complexity. Proteins in the dark proteome are often difficult to express in laboratory systems, unstable outside their native environments, or incompatible with traditional structural techniques such as X-ray crystallography. Even newer methods, such as cryo-electron microscopy (cryo-EM), are not universally applicable.
But the challenge is not purely technical. In many cases, it is biological.
“In the past, because of the sample availability and technology limitations, we could only identify the proteome at a bulk/tissue level, with a limited coverage of diseases. In addition, small peptides (those that contain <50 amino acids) are often ignored. Many proteins, especially those that have a dynamic presence (e.g., expressed at a certain cell type/time/disease), are often overlooked,” explained Xing.
This bias means that entire categories of proteins—especially low-abundance ones—may not appear in standard analyses.
Perdigão emphasized that recent advances have not fully resolved the issue: “Even with the advances brought by cryo-EM, deep learning, AlphaFold2, and AlphaFold3, there are still many structural and conceptual reasons why the dark proteome remains large. Intrinsically disordered proteins, experimental bias in structural biology, functions driven by dynamics rather than static structures, small evolutionary families, explosive growth of sequence data (even with partial mitigation from the AlphaFold Protein Structure Database in 2022 and 2024), continue to pose significant challenges relevant to efforts to illuminate the dark proteome.”
“Structural biology was built around the assumption that determining a protein's structure provides key insight into its function. Computational methods have mostly addressed the technical barriers to structure prediction, yet a substantial fraction of the proteome remains dark,” Perdigão added.
“This persistence could indicate something deeper: where a large portion of proteins that are intrinsically disordered or context-dependent, whose function manifests not from a stable three-dimensional form, but from conformational ensembles that existing frameworks can only partially or incompletely capture,” he said.
In other words, some proteins may not have a single structure at all—and this may be central to how they function.
“This raises an important and pressing question for the field: how do we develop faster yet precise approaches to illuminate and characterize the function of the molecules that persist in the dark?” said Perdigão.
Why the dark proteome matters for biology and medicine
Many dark or partially characterized proteins play roles in cellular processes, including signaling pathways, immune responses, and host–pathogen interactions.
Several are also implicated in human disease, including cancer and neurodegenerative conditions, or are thought to be involved in shaping microbial communities in soil, aquatic, and extreme environments.
“In medicine, IDPs have already been implicated in cancer, neurodegenerative diseases such as Alzheimer's and Parkinson's, and cardiovascular conditions, yet remain largely inaccessible as therapeutic targets. A deeper characterization of these molecules may therefore uncover novel disease mechanisms, expand the druggable proteome, and improve the interpretation of genomic data, ultimately contributing to a more complete understanding of protein function in health and disease,” said Perdigão.
“An accurate understanding of the proteomic diversity will allow us to better understand how biological processes are controlled in different cells under different physiological/disease conditions,” added Xing. “For cancer research, identifying proteomes that are specific to tumor cells can help determine the tumor/individual-specific drug targets and treatment plans.”
While tools such as AlphaFold have successfully predicted structures for millions of proteins, they often assign low confidence to regions that fall within the dark proteome. Rather than marking the end of the problem, these tools have clarified its boundaries.
Perdigão sees this as a sign that the field must evolve: “The scale of the dark proteome suggests that our current approaches to studying proteins capture only part of the landscape, and this limitation is now less technological than conceptual.”
“Improved mapping of the dark proteome might shift the focus from well-structured proteins towards dynamic and context-dependent molecules,” Perdigão explained.
Toward a complete proteome
The dark proteome’s persistence suggests that the challenge is not solely one of better tools, but of deeper understanding. Many dark proteins appear to operate outside classical frameworks, relying on flexibility, context, and dynamics rather than fixed structures.
“The greatest challenge may lie in the fact that many of these dark proteins operate through dynamic, context-dependent conformations, thereby demanding both the development of experimental and computational approaches that are robust, precise, and efficient, and new ways of thinking about protein function,” said Perdigão.
At the same time, progress is underway.
“With the advancement in technology and the domain knowledge, the field is getting much better at collecting diverse proteomic information,” said Xing.
Taken together, these perspectives highlight both the scale of the challenge and the momentum behind efforts to address it. These developments point toward a future in which the dark proteome becomes less opaque—but perhaps no less intriguing.