We've updated our Privacy Policy to make it clearer how we use your personal data. We use cookies to provide you with a better experience. You can read our Cookie Policy here.

Advertisement

Data Sharing in Life Sciences: How To Share Research Data Without Sharing What You Shouldn't

AI-generated researcher uploading structured datasets to a secure cloud repository in a life science lab.
Credit: AI-generated image created using Google Gemini (2026).
Read time: 7 minutes

Research data sharing in life sciences has moved from a voluntary courtesy to a funder condition, a journal requirement, and often a term of continued grant funding. Life scientists now have to weigh privacy protections, intellectual property, competitive timing, and technical interoperability standards at once, all while making genuinely reusable data available to reviewers, funders, and the wider research community.

Key takeaways

  • Major funders, including the National Institutes of Health (NIH), UK Research and Innovation (UKRI), and the European Commission, now require data management and sharing plans as a condition of funding, not as an optional best practice.
  • Journal data availability statements, driven largely by International Committee of Medical Journal Editors (ICMJE) requirements, ask authors to specify what data can be shared, when, and under what access conditions.
  • The General Data Protection Regulation (GDPR) requires researchers working with human-derived data to build deidentification and consent scope into any sharing plan from the outset, not as a retrofit.
  • A justified, documented embargo can protect a lab's priority claim on a discovery while still satisfying funder and journal data sharing obligations.
  • Persistent identifiers, standardized formats, controlled vocabularies, and documented provenance are the technical requirements that turn a shared dataset into one other researchers can actually reuse.

Why research data sharing is now a life science requirement

Research data sharing in life sciences has shifted from a discretionary act of goodwill into a formal condition attached to funding, publication, and institutional compliance. Major funding bodies and journals now treat a dataset's availability as part of the scientific record itself, not an optional addendum released only if time allows after a paper is accepted.


The shift reflects a broader push toward reproducibility and reuse across biomedical research. UKRI research councils operate common data sharing principles that treat publicly funded research data as a public good to be made openly available with as few restrictions as possible, a framing now echoed across most major national and international funders.


For researchers, this means data management and sharing plans, availability statements, and repository deposit requirements now function as forcing mechanisms that push labs toward more careful documentation, whether or not the work is described in those terms. The practical challenge is doing this without exposing information that should stay protected, from unpublished intellectual property to identifiable participant data.

NIH, UKRI, and Horizon Europe mandates for research data sharing

Funder mandates for research data sharing now span the major bodies that support life science research, each with a slightly different mechanism but a shared underlying expectation. The NIH requires a data management and sharing plan for essentially all NIH-funded research that generates scientific data, and that plan becomes a formal condition of the award once a grant is issued.


In the European Union, Horizon Europe imposes a comparable but distinct requirement. Beneficiaries must produce a data management plan within months of a grant's start, guided by the FAIR (Findable, Accessible, Interoperable, and Reusable) principles and the European Commission's open science expectations, which follow the principle of making data as open as possible and as closed as necessary.


None of these frameworks require unconditional openness. Each explicitly allows justified restrictions, whether for legal, ethical, or competitive reasons, provided the limitation is documented in the plan rather than left unexplained. That flexibility is what allows a lab to comply with a mandate while still protecting a genuine, time-limited need for confidentiality.


Table 1: Core research data sharing requirements across major life science funders and journal bodies.

Body

Mechanism

Core requirement

NIH

Data management and sharing plan

Plan submitted at application; data shared at publication or award end, whichever comes first

UKRI

Common data sharing principles

Publicly funded data made open with as few restrictions as possible

Horizon Europe

Mandatory data management plan

FAIR-aligned plan due within months of grant start; justified restrictions allowed

ICMJE member journals

Data sharing statement

Statement filed at manuscript submission and in trial registration

Journal data availability statements and research data sharing

Journal data availability statements add a second, independent layer of pressure beyond funder mandates, since a dataset submitted for peer review has to be genuinely usable by someone outside the original lab. The ICMJE data sharing requirement has applied to manuscripts reporting clinical trial results submitted to member journals since mid-2018, and it extended to trial registrations for studies enrolling participants from 2019 onward.


The statement itself does not mandate that data be shared unconditionally. It requires transparency: authors must state whether deidentified individual participant data will be available, what will be shared, and under what access conditions, so editors and readers can evaluate the claim rather than take it on faith.


For researchers outside clinical trials, many journals now apply a similar expectation informally, asking for a data availability statement as a standard manuscript component regardless of study design. Treating this requirement as a planning step early in a project, rather than a late-stage scramble before submission, avoids the common failure mode of promising access to data that was never structured well enough to share.

GDPR requirements for research data sharing with human subjects

GDPR adds a further, non-negotiable constraint for any life science data involving human subjects. The regulation's core text requires that any sharing plan account for lawful basis, consent scope, and appropriate safeguards from the outset, rather than as something addressed only once a repository asks for it.


Anonymization is the most direct route to lifting data outside GDPR's scope entirely, but achieving it is harder than it sounds. The UK's Information Commissioner's Office has published anonymization guidance that frames anonymization as sufficiently remote risk of re-identification rather than a theoretical impossibility, and it introduces a motivated intruder test to evaluate whether a determined party could still identify individuals from a dataset that looks anonymous on its face.


For most human-derived life science data, full anonymization is difficult to guarantee, which pushes many labs toward pseudonymization combined with controlled, tiered access instead. Consent forms drafted with future data sharing in mind, specifying the scope of reuse a participant has agreed to, are far easier to work with later than consent language written without that possibility in view.

Embargo strategies and technical standards for research data sharing

A justified embargo lets a lab delay public release of a dataset for a defined period, most often to align with a related publication or a patent filing, without violating funder or journal data sharing terms. Every major framework reviewed here permits this kind of restriction, provided the rationale and the timeline are documented in the data management plan rather than applied informally after the fact.


Advertisement

Meeting the technical side of a sharing obligation is a separate challenge from meeting the policy requirements. The Global Alliance for Genomics and Health has developed international standards for genomic data sharing that illustrate what shareable, interoperable life science data actually requires in practice, beyond simply depositing files in a repository.

AI-generated flowchart of the research data sharing process from capture to repository deposit.

Figure 1: A five-stage flowchart of the research data sharing process, from identifying requirements through repository deposit and eventual public access. Credit: AI-generated image created using Google Gemini (2026).


Those requirements translate into a consistent set of practical needs regardless of data type:

  • A persistent identifier, such as a digital object identifier, that makes a dataset citable and locatable independent of where it is hosted.
  • A standardized file format and, where relevant, a controlled vocabulary that lets the data be combined with other datasets without manual reconciliation.
  • Documented provenance, including instrument, protocol, and sample information, that lets another researcher trust and trace the result.
  • Clear, machine-readable metadata describing any access restrictions, so automated tools and human reviewers alike can tell what is available and under what terms.


Building a compliant data sharing plan is easiest as a sequence rather than a single late-stage task:

  1. Identify which funder, journal, and regulatory requirements apply to the specific project before data collection begins.
  2. Decide what will be shared, what will be restricted, and for how long, documenting the justification for any restriction.
  3. Select a repository that supports persistent identifiers and the metadata standards relevant to the data type.
  4. Address consent scope and deidentification requirements for any human-derived data before the first sample is collected.
  5. Revisit the plan at submission and at the point of data deposit, since requirements and project scope can shift between application and publication.

Making research data sharing part of every project plan

Research data sharing in life sciences is no longer a separate compliance exercise bolted onto a finished project; it is a condition that shapes funding, publication, and institutional standing from the outset. Funder mandates, journal statements, GDPR, and embargo provisions all point toward the same underlying practice: deciding early what will be shared, documenting any restriction, and building the technical groundwork for reuse before a dataset is ever requested.


Labs that treat this planning as routine, rather than a hurdle to clear only when a funder or editor asks, spend less time reconciling gaps later and produce data that other researchers, and increasingly the AI and data science tools now used across life science research, can actually put to use. That discipline connects directly to the broader practice of structured research data management and to implementing FAIR principles in day-to-day lab work, of which sharing is the final, most visible step.


This content includes text that has been created with the assistance of generative AI and has undergone editorial review before publishing. Technology Networks' AI policy can be found here.

Google News Preferred Source Add Technology Networks as a preferred Google source to see more of our trusted coverage.