The Road to the Connected Laboratory
David Manning explores how connected, AI-ready labs depend on harmonized data and traceability.
Laboratories are generating more data than ever before, yet much of that information remains fragmented across disconnected instruments, software platforms, and legacy systems. While many organizations have invested heavily in digitization, simply connecting systems is not enough to unlock the full value of scientific data or prepare laboratories for AI-driven workflows.
Technology Networks spoke with David Manning, global lead of digital transformation at Thermo Fisher Scientific, about what it takes to build a connected lab that gives scientists faster access to trusted data while preserving existing workflows and regulatory compliance.
In this interview, Manning discusses the challenges of overcoming data fragmentation, why metadata harmonization is essential for AI readiness, and how regulatory expectations are shaping digital transformation strategies.
From your perspective, what are the biggest sources of data fragmentation in labs today?
Often when customers come to us, they’re struggling with similar challenges. Most labs deploy a variety of systems, sometimes thousands of instruments, and hundreds of applications that are each built for a specific purpose. Each of these systems generates data in different ways, leaving labs with results that lack uniformity and context. The data, too, can be largely unstructured, living in a variety of formats, such as PDFs, presentation decks, or spreadsheets.
On top of that, the software driving these instruments is often instrument-specific, rather than system-agnostic, leaving the lab or the scientist to convert output into something workable. That’s a manual, error-prone process that takes an extraordinary amount of time at scale.
Beneath it all lies a landscape of legacy instruments that can be difficult to replace if they are core to lab operations. Take pharmaceutical labs, for example. Once drugs have come through those systems and been approved by regulators, those products are validated on them. Replacing those systems would disrupt a validation trail that requires significant rework to reconstruct.
That complexity is compounded by how these organizations are structured. From drug discovery through development to manufacturing and product release, each function tends to operate in its own silo—even within a single company. When pharma companies merge with or acquire other companies, there’s a multiplier effect. Acquired companies bring their own systems, workflows, and data formats. There’s often no practical path to creating a common model across all of it.
What this all adds up to is metadata inconsistency. That’s the piece that makes true integration so difficult. The inability to contextualize data across all sources is at the core of the problem.
The instinct most organizations follow is to treat this as a connectivity problem, using time and resources to link systems together.
However, connectivity is only half of the battle. The other half, and arguably the harder half, is contextualizing the data and establishing governance around it. The goal should be to create a unified scientific data platform that allows data to flow more seamlessly, which means thinking about how these workflows will impact scientists, not just the IT teams.
As I mentioned, labs are often bound by the fact that they can’t replace existing systems without significant downstream impact, and there’s a need to connect and integrate while keeping validated systems of record intact.
The diversity of technologies and integration of legacy systems into new ones add to the urgency of building an open, connected scientific ecosystem so that data is accessible across the organization. There’s enormous pressure to adhere to validation boundaries, maintain the audit trail, and avoid standard operating procedure rewrites that could invite regulatory scrutiny. The whole process has to happen without disrupting what regulators have already accepted.
Lastly, if the metadata model isn’t harmonized as part of this work, labs haven’t moved the needle. The real goal is moving from unstructured data to structured, fully searchable, fully traceable data, and the metadata model is what makes that possible.
There are several maturity models that try to map this journey, but I tend to prefer the BioPhorum Digital Plant Maturity Model (DPMM®). It helps organizations assess where they are today and define a roadmap toward more connected, data-driven operations across manufacturing, laboratories, quality, and IT.
Based on our experience working with thousands of labs, system integration is where they tend to get stuck. Many organizations have done the work of connecting systems, but the metadata is still inconsistent.
The assumption was that connectivity would solve the problem, but as we’ve dug in with their teams, we’ve discovered that the data harmonization journey is where the real work is needed. This is also where projects most often stall.
I’ve found that the scale of the journey is routinely underestimated. Thinking about integrating 10 or 20 instruments may sound manageable, but in most of our customers’ lab environments, the instrument population is in the thousands, and all with different vendors and software versions. That’s a fundamentally different challenge, and it’s one that causes organizations to get stuck or fail.
However, once an organization gets through the integration work, a different challenge can emerge: who actually owns the scientific data strategy? Is it IT, the business, or the scientists? Everyone has legitimate claims, but that tension can grind progress to a halt. Labs must work through governance in order to succeed.
AI readiness comes back to traceability. The entire data model needs to support the use of AI, and it’s also what regulators will increasingly require. Metadata is the other piece that’s foundational, and it can make or break an AI program. If the metadata isn’t clean and harmonized before you start building AI applications on top of the data, you are building on an unstable base.
The dimension of AI readiness that often gets overlooked, though, is the human one. Many scientists today approach AI with a fair amount of skepticism. They trust their own data, their own methods, and their own judgment. For a lab to be truly AI-ready, scientists must be willing to trust and use the technology, and that begins with ensuring traceability and keeping a human in the loop whenever the technology is used.
The foundation has to create authentic, verifiable provenance. From there, there’s an opportunity for data quality monitoring directly into the platform so the system can surface missing records, flag duplicates, or identify files that aren’t stored properly. The platform can continuously self-cleanse and improve the quality of the data that AI works with. That ongoing quality loop is part of what builds and maintains scientists' confidence over time.
The evolving regulatory environment is shaping how connectivity decisions are made. They’re no longer purely technical questions when there are compliance implications running through every design and system choice.
As the auditor’s perspective shifts from individual systems and point-in-time states and integration becomes more common, regulators will increasingly want end-to-end traceability: where did this data originate, who touched it, and what happened to it along the way. As labs evolve, traceability becomes a requirement.
The practical consequence is that validated systems of record can’t simply be displaced in the name of modernization. Any integration effort has to be architected around that, and those systems must be connected in ways that maintain audit integrity.
A truly connected lab environment is
one in which data flows into a common structure that’s searchable, traceable,
and consistently contextualized. At the infrastructure level, labs are leveraging
a unified scientific data platform where the underlying metadata model is
harmonized across sources. Harmonization is what allows information from
different instruments, different sites, and different vendor systems to be
treated as a coherent whole.
What changes for the scientist in that environment is time allocation and how that impacts downstream outcomes and performance. Right now, scientists in many labs spend as much as half their time looking through various systems for the information they need just to do their work. That is a staggering amount of wasted capacity.
The goal of a connected lab environment is
to give scientists
back their time so they can focus on what they do best: science. Ultimately,
having the right insights available when they need it helps accelerate decision-making,
improve reproducibility, and speed up the time to result.
The “people side” of this journey is consistently underestimated. It gets assumed rather than planned for, and that’s where a lot of well-designed implementations fall short. Technology changes the workflow, but it doesn’t automatically change how people relate to their work or each other.
A connected environment, properly designed, changes that. The scientist’s cognitive energy goes toward designing experiments, interpreting results, and advancing their research. When you put the change management conversation in that frame, it becomes a compelling case that speaks to scientists directly.
The excitement around AI fits into the same frame. AI amplifies the value of connected, harmonized data, and it extends what scientists can do without extending how many hours they work. But that framing has to come upfront. The organizations that get change management right are the ones that start with the scientist experience, rather than leading with the technical architecture and hoping adoption follows.
This introduction to this interview includes text that has been created with the
assistance of generative AI and has undergone editorial review before
publishing. Technology Networks' AI policy can be found here.