We've updated our Privacy Policy to make it clearer how we use your personal data. We use cookies to provide you with a better experience. You can read our Cookie Policy here.

Advertisement

Inside a National Compound Library Facility

Vials with gold and silver tops placed into numbered slots within a rack, in a compound library facility.
Credit: Julia Koblitz / Unsplash.
Read time: 8 minutes

As demand for new medicines rises, compound library facilities remain essential for enabling reliable, collaborative drug discovery.


While scientists work to develop novel therapeutics for rare diseases and overlooked indications, they also endeavor to overcome the limitations of existing drugs. High-throughput screening (HTS) underpins much of this work, allowing teams to screen hundreds of thousands to millions of compounds each day.


And the infrastructure behind HTS is compound libraries: meticulously organized, large-scale collections of molecules.


Few scientists have better insight into how these systems operate at scale than Dr. Matthew Hall. As scientific director at the National Center for Advancing Translational Sciences (NCATS), part of the National Institutes of Health (NIH), Hall spends his days coordinating expert groups behind the discoveries that later become real-world therapies.


In this interview with Technology Networks, Hall brought the day-to-day of compound library facilities to life.

Collaboration is an essential component in compound library facilities

As scientific director at NCATS, what does a typical day look like for you?

Traditional “open access” laboratories are physical facilities that offer external researchers access to research space and equipment. NCATS operates slightly differently, providing a government-funded hub where external scientists can partner with in-house experts to conduct research, rather than using facilities themselves.


Hall described how the facility supports underserved scientific areas: “We partner with academics, biotech start-ups, and foundations, all of whom are usually working on understanding rare diseases that need treatments. We de-risk and enable the discovery and development of therapeutics, primarily for diseases that the private sector wouldn’t tackle.”


“That’s what the NIH is helpful for. To make sure that all Americans, and ultimately everyone in the world, can have treatments.” — Dr. Matthew Hall.


Hall explained that researchers who approach the facility bring disease expertise, while NCATS provides the infrastructure and capabilities—forming an integrated model for leveraging centralized screening resources. “An academic researcher, or a rare disease foundation, is not going to have all the infrastructure or expertise to start a drug discovery and development program,” he added.


The set-up and maintenance of these programs require coordination of several disciplines:


  • Acquisition management: Identify, evaluate, and procure compounds from commercial and collaborative sources.
  • Assay development: Design and validate assays to reliably measure compound activity in screening campaigns.
  • HTS and automation: Execute large-scale, automated screening of compounds to identify potential hits.
  • Compound management: Handle the management of compound collections, including storage, tracking, preparation, and distribution across workflows.
  • Analytical chemistry: Confirm compound identity and purity, ensuring that screening results are accurate and reproducible.
  • Medicinal chemistry: Optimize hits by designing and synthesizing analogs to improve potency, selectivity, and drug-like properties.
  • Regulatory and toxicology experts: Evaluate safety, compliance, and translational requirements to prepare compounds for preclinical and clinical progression.


“We have a team-based environment with project management. We take on projects and develop them to the point where it's either successful, or we terminate the project because it’s not working for one reason or another,” explained Hall.

Collaboration is key:

  • The compound library facility at NIH operates as a partnership engine.
  • Multidisciplinary work is essential in drug discovery projects.
  • Centralized expertise enables individual labs to carry out research in neglected areas, such as rare and infectious diseases.

The drug discovery workflow: from acquisition to quality control

How do expert groups come together in drug discovery efforts?

A national compound library has a handful of collection types, each serving distinct purposes. Hall outlined key categories:


  • Drug repurposing collection: Includes drugs that are approved in Europe, Australia, the United Kingdom, and the United States. Every time a drug is approved, it is added to this collection, with the potential for repurposing.
  • Annotated collection: A larger collection that includes molecules with known mechanisms of action. Sometimes this library works in reverse, revealing unexpected insights into disease during campaigns. These libraries are often a starting point for screening.
  • Diversity-oriented collection: Large-scale collection of molecules with unknown mechanisms used to explore entirely new chemical space.
  • Retired/miscellaneous collections: Includes all molecules ever made or investigated that were not hits within the primary project. Researchers may return to these at a later date.


Hall described diversity-oriented collections as where “the real action” happens in HTS. “We’d be screening between 200,000 and 300,000 molecules,” he noted.


Maintaining each collection requires coordination. For example, the acquisitions team sources compounds based on guidance from the chemoinformatics team, who are responsible for library design and filling gaps.

Advertisement


“Then, you need a compound management facility,” added Hall. The scale of the compound management at NCATS is similar to what you might expect at a large pharmaceutical company.


“They store almost a million samples. They’re a hub for gathering and distributing all the molecules,” he continued, later describing how this team manages samples, creates drug dilutions, plates assays, and so on.


Once you’ve designed a library, acquired compounds, and can maintain them, quality control is critical. Hall put it simply: “If you buy a molecule and it's not very pure, then you are going to discover activity that you can never reproduce.”


He explained that the analytical chemistry team typically validates purity post-HTS, as this is more cost- and time-effective than validating thousands of molecules, when only a handful may be of interest.


And just because NCATS is not a traditional open-door facility, it doesn’t mean physical collaboration doesn’t happen: “We ship compounds and plates all over America, for collaborators to test and demonstrate reproducibility.” This adds another layer of validation, this time external, to the process.

Collaboration is key:

  • A compound library facility integrates acquisition, storage, quality control, and distribution.
  • Types of libraries vary, but high-level workflow overviews are generally consistent.

Lifecycle decisions: adding, retiring, and reprioritizing compounds

How are decisions made to add, retire, or reprioritize compounds within a library?

Compound libraries are dynamic systems that evolve as science advances. A library that seemed interesting 10 years ago may have proven unsuccessful in screening attempts.


“It’s almost like fashion. I can almost look at a molecule in a library and say: ‘That’s from an old library’ or ‘That’s not a very cool molecule.’ It’s like me wearing clothes in the 90s vs now.” — Dr. Matthew Hall.


Advertisement

Lifecycle management is the process of tracking, maintaining, and controlling a collection from acquisition to retirement. The latter is influenced by several factors:


  • Exhaustion: If a library has been repeatedly screened to no avail, it may be retired.
  • Lack of novelty: This might not mean that the library is of no use, but if the space has already been explored extensively, it can deter interest.
  • Intellectual property (IP) considerations: The more a library is used, the more work around the library is published, and the less ability there is to patent molecules. This can make compounds less intriguing from an IP perspective.


Hall explained how NCATS works with pharmaceutical companies to ease IP concerns: “We might work together on a disease that wouldn’t be commercially interesting to them, but that they feel motivated to partner on from a greater-good point of view.”


“They provide blinded libraries. Only at the end, when we find and verify hits, will they reveal the structure of those molecules. The remaining ones are still a secret,” he revealed. “It’s like not knowing what’s in the Coca-Cola® formulation—they preserve the IP value of their library, while helping provide world-class chemical matter to help patients with rare of infectious diseases.”


Pharmaceutical companies sometimes donate libraries to academic discovery groups when they are no longer commercially viable. At NCATS, the library can also move toward retirement. As is the case with other exhausted libraries, which are either put into long-term storage to be “dusted off” again one day or donated to academic institutions.


In some cases, retirement of an entire library may not be required, as untoward results may be due to a handful of “bad actors.” Hall explained how false-positive compounds arise and how they are handled: “You might have a library where a group of molecules from one manufacturer is continually ‘active’, so you resynthesize them. Sometimes they’ve got an impurity from the manufacturing process in there that makes things look active.”


“We flag those molecules and no longer see data from them,” he continued, highlighting why the chemistry team resynthesizes hits to validate them post-screen.

Compound library lifecycle considerations:

  • Factors to consider include scientific value, reliability, and IP.
  • Retired libraries are donated or archived, not discarded.

Ensuring reproducibility in screening campaigns

What day-to-day processes are most critical for ensuring reliability and reproducibility?

Hall outlined the key steps that ensure reproducibility:

Advertisement


  1. Biologists validate assay robustness before screening and confirm feasibility with the automation team.
  2. Following HTS, large datasets chemoinformaticians process and analyze data.
  3. Promising hits are identified and re-tested for confirmation.
  4. For confirmed hits, chemists re-synthesize the molecule to confirm there is a pure hit.
  5. Counterassays are also used to “weed out bad actors.”


“You want to be very confident before you ask people to put time and energy into [a molecule]. You want to do best by the patients as well.” — Dr. Matthew Hall.


“By the time we're interested in a molecule, it's been tested multiple times, then the chemists make brand new material, and then we'll test that,” he added.

Achieving reproducible results is embedded throughout the workflow:

  • Reproducibility relies on iterative testing and validation.
  • Confirmatory synthesis is essential before committing resources to drug development.

Finding balance in shared compound library facilities

How do you balance access, governance, and scientific rigor in a shared facility?

One of the most complex aspects of running shared chemical repositories is balancing openness with the need to protect IP.


Hall is known as an advocate for sharing in science, which is embedded in NCATS’ own projects. “All data we generate eventually sees the light of day. We even make data available before we publish a paper,” he noted.

Advertisement


Hall highlighted that during the COVID-19 pandemic, NCATS shared screening data in real time: “There were a lot of people doing the same thing. Why not just make all the data available, so people can check rather than waste resources?”


However, openness has limits—particularly during early-stage discovery when collaborating with academic institutions or pharmaceutical companies.


“When you're creating a brand new molecule, you can’t post it on the internet, you can't show people the structure, you can’t give it out. If you do that at the beginning, and it's not patented, there is no asset, and there’s nothing to develop… Patients won’t get a medicine,” he explained.


Naturally, this creates tension between the benefits of data sharing and the risks to patent protection. Despite this, Hall noted that NCATS aims to publish work as quickly as possible and encourages its collaborators to do the same once a suitable point is reached in a project.

Potential for patients vs patent value:

  • Open science accelerates discovery, which is particularly important during global crises.
  • IP protection is essential for a drug to be deemed viable in the private pharmaceutical industry.

The future of compound library facility operations

How is AI influencing everyday operations in compound library facilities?

AI and machine learning are already embedded in modern compound library facility workflows. Hall explained that AI facilitates molecular predictions at the discovery stage and streamlines data analysis for chemoinformaticians later down the line.


“We’re predicting different structures that should be active that are not in our physical screening libraries, and then we either make or buy those molecules, then test them.” — Dr. Matthew Hall.


However, he emphasized that AI cannot replace experimental screening: “Having a physical compound library is more important now than ever as the data generated drives AI.”


But the lack of high-quality experimental datasets is currently impeding the full potential of predictive approaches. This results from processes before the era of in silico work.


Advertisement

“When scientists performed a drug screen, they were not trying to create the perfect data set. They were just trying to find the best molecules. The curation and the cleanup of those data sets never happened, so scientists are returning to those now to clean them up and maximize their value,” added Hall.


This challenge is exacerbated by the aforementioned lack of data sharing. Both data silos and the disregard of “failure” datasets deprive models of potentially critical information, limiting their ability to generate fully informed outputs.

AI is embedded, but not yet reaching its potential:

  • AI complements, rather than replaces, physical screening.
  • Data quality is the key limiting factor in AI-driven discovery.
  • Clean-up and integration of historical datasets is required to utilize AI tools to their full potential.

Compound library facilities: the heart of translational science

Hall closed by making one thing clear: “Small molecule libraries are just as important now with AI as they ever have been. They generate the datasets needed to feed the beast.”


It is evident that national compound library facilities like NCATS are far more than repositories. They are dynamic, collaborative engines driving discoveries that will later serve patients.


As the field evolves, so do the libraries, with Hall hinting at the potential for “new library flavors” as modalities emerge and grow.

Key takeaways:

  • National compound libraries underpin collaborative drug discovery, especially in underserved indications.
  • Reproducibility depends on robust workflows and coordinated efforts.
  • Future progress will rely on integrating high-quality data with AI tools and achieving increased openness. 


This content includes text that has been created with the assistance of generative AI and has undergone editorial review before publishing. Technology Networks' AI policy can be found here.

Google News Preferred Source Add Technology Networks as a preferred Google source to see more of our trusted coverage.