Library preparation is critical in next-generation sequencing workflows, where bias, contamination, and inefficiencies directly impact coverage, sensitivity, and reproducibility.
Laboratories must balance throughput demands with the need for accurate quantitation, consistent fragment sizing, and minimized artefacts. Poorly optimized workflows can lead to uneven genome coverage, reduced sequencing yield, and compromised detection of low-frequency variants.
This guide outlines practical strategies to strengthen library quality, reduce bias, and improve overall sequencing performance.
Download this guide to discover:
- How to minimize bias and maintain uniform genome coverage across complex samples
- Strategies to reduce artefacts, contamination risks, and library loss during preparation
- Approaches to improve quantitation, multiplexing balance, and sequencing efficiency
How To Prepare High-Quality
Samples for Next-Generation
Sequencing
Michael Quail, PhD
Library preparation (otherwise known as sample prep) is arguably the most critical part of the next generation
sequencing (NGS) workflow, since errors and artefacts will irreversibly affect sequencing quality
and outcome.
Good libraries will give rise to even genome coverage with little bias or artefact and will also give both optimal
sequencing yields and even sample-to-sample representation when multiplexing. Libraries should
ideally be as complex as possible to minimize duplicates and, in some applications, enable detection of
low frequency variants.
1. Maintain evenness of coverage and minimize bias
The main steps that can introduce bias into NGS libraries are shearing and library amplification. For
shearing, acoustic fragmentation is widely considered the gold standard because it avoids many of the
drawbacks of conventional sonication-based methods. Traditional sonication can generate localized heat
and oxidative stress, which may contribute to DNA damage, irreversibly denaturing AT-rich regions.
Enzymatic fragmentation can work almost as well as acoustic shearing and has the advantage that it
doesn’t require any expensive capital equipment. Since it can be performed in the same labware as the
rest of the library prep, DNA losses can be minimized. However, some enzymatic shearing formulations
can be sensitive to contaminants such as EDTA (check manufacturers’ product sheet), give variable fragment
sizes with varying DNA amount and GC content, and in some cases introduce bias due to non-random
cleavage.
Tip: Use a PCR-free protocol if possible.
NGS libraries contain a mixture of fragments of differing length, repeat, and GC content. Hence, PCR
amplification can introduce bias into NGS libraries. The sequencing data with the most even genome
coverage is therefore obtained from PCR-free library preparation approaches.
How To Guide
HOW TO PREPARE HIGH-QUALITY SAMPLES FOR NEXT-GENERATION SEQUENCING 2
In PCR-free library prep, full-length adapters containing all of the platform adapter sequences are ligated
to DNA fragments. Note this process is not 100% efficient so the only way to accurately quantify adapter-
ligated sequenceable fragments is by qPCR. Fragments that do not have adapters ligated on both ends
will not be amplified and clustered, so they will not generate sequence. However, due to random sampling
of library fragments, differences in mappability, and platform-specific biases, even PCR-free libraries will
not give rise to perfectly even representation of all positions across the genome.
Tip: If PCR is required, limit cycles and use an enzyme that gives minimal bias.
Most “standard” Taq polymerases will preferentially amplify smaller, more GC-neutral fragments, such
that with multiple PCR cycles, libraries will contain shorter fragments that are depleted in the more ATand/
or GC-rich regions of the genome. These biases typically worsen with lower inputs and increasing
numbers of PCR cycles.
Since the introduction of NGS, high-fidelity PCR formulations optimized for amplification of complex fragment
mixtures have been developed. The best-performing enzymes have been shown to provide coverage
uniformity comparable to PCR-free approaches while introducing minimal size bias. Some high-fidelity
formulations have also demonstrated effective amplification of longer fragments (up to ~20 kb) for
long-read sequencing applications.
2. Watch out for bubble products
Overamplification due to excessive PCR cycle numbers can also give rise to bubble product artefacts.
After too many cycles, primers are exhausted, and the fragments can no longer amplify; they anneal via
their common adapter ends, but the central regions are unlikely to anneal when the library contains a
highly complex genome’s worth of fragments.
Bubble products are often observed as libraries that give a bimodal size distribution when analyzed on an
electrophoretic platform. The individual adapter-ligated fragments in these molecules will still sequence,
but the central single-stranded region will not bind double-strand fluorescent dyes (such as SYBR green)
effectively, resulting in abnormal migration during electrophoresis and inaccurate quantification using
fluorescent DNA dyes. In these circumstances, one can either repeat with fewer PCR cycles or recondition
by adding extra primers and running a few extra cycles.
3. Reduce adapter dimers
While adapters used in NGS library prep are typically T-tailed to minimize adapter to adapter ligation,
adapter dimers can form, nonetheless. This is particularly prevalent when low amounts of input material
are used. Primer dimers can also accumulate during PCR. Being smaller, these adapter/primer dimers are
clustered and sequenced preferentially and can result in considerable loss of mapped sequencing yield.
Surplus adapters and primers, as well as dimers, are normally effectively removed by solid-phase reversible
immobilization (SPRI) bead-based cleanup at the post-ligation and post-PCR stages, respectively.
For PCR-based workflows, using a short “stubby” adapter, a 0.8x or 0.9x SPRI bead cleanup at these stages
is normally sufficient. Being larger, full-length PCR-free adapters and dimers thereof are more difficult
to remove. Here, a 0.7x SPRI cleanup and sometimes a secondary cleanup may be necessary.
Adapter and PCR dimers are seen as sharp spikes that are smaller than the main library fragments on
QC electropherograms. If observed, it is advisable to do an extra SPRI cleanup if the library is of sufficient
concentration (e.g., >10 nM) to tolerate an extra cleanup.
4. Achieve good sample-to-sample representation when
multiplexing
Most sequencers generate enough sequence per lane or flow cell to give acceptable coverage of multiple
genomes/samples. It is therefore common practice to add a sample-specific sequence, or barcode,
to samples during library prep so that they can be mixed and sequenced together, thus spreading the
sequencing run cost over multiple samples. This normally works, but in practice, achieving a balanced
representation of those mixed multiplexed samples can be challenging.
Errors in quantitation, pooling, as well as fragment size differences between the libraries will result in
variable representation. This has an economic cost if extra sequencing is required in order to get sufficient
sequence from the least represented sample. It has been observed that small fragments get preferentially
amplified and sequenced, such that libraries with smaller fragment size ranges in a pool will end
up with more sequence reads than those with larger fragments. It is therefore good practice to ensure
that libraries to be pooled are as matched as possible.
Tip: Standardize libraries to be multiplexed in a pool.
To achieve even sample representation in a multiplexed pool, the input amount should be standardized.
Where possible, samples should be from the same genome type and prepared using the same DNA
purification method. Acoustic shearing has been shown to be more reproducible in giving the same size
range, as opposed to enzymatic methods, where size range can vary depending on sample context.
Reproducibility can be increased by automation, and pipette servicing will ensure accurate pipetting. Particular
care should be taken to ensure that equal SPRI bead volumes are used for all samples. SPRI bead
formulations are often viscous, so it is advisable to pipette and dispense slowly and to check that drops of
excess SPRI beads are not clinging to the outside surfaces of the tips.
5. Increase library complexity
High complexity can be achieved by using higher levels of input DNA and by minimizing losses. Exact
input amounts will vary by kit, but ideally, for a human genome should be >200 ng for a PCR-based workflow
and >500 ng for a PCR-free workflow.
Tip: Choose the most efficient labware and library reagents.
Low-retention tips, plates, and tubes should be used, especially for low-input samples, since DNA can
bind to plastic surfaces. NGS library kits do not add adapters to both ends of input DNA fragments with
100% efficiency. It has been demonstrated that this efficiency can vary between kits, so sensitive applications
should seek higher efficiency kits. Many NGS library protocols have been written to minimize total
workflow time and often use a 15-minute ligation step. We have found that increasing this to 30 or 60
minutes can increase adapter ligated fragment yield by 50%.
Tip: Optimize SPRI bead recovery.
The main losses during NGS library prep occur during SPRI bead cleanup. Higher yields are obtained with
higher ratios, but lower ratios are often required to ensure removal of smaller fragments. Use the highest
bead:sample ratio that works for your application. Recovery is maximized by thorough mixing (vortexing
or >20x pipette mixes) of beads and sample, and by not overdrying the bead pellet after the ethanol
washes (low recovery will be obtained from cracked and dried bead pellets). Elution from SPRI beads is
complete within 2 minutes for short <1000 base fragments, but extended elution times (~10 minutes) with
elution buffer heated to 37 °C may be required for large fragments and libraries for long-read platforms.
DNA losses can be minimized further by using the “with-bead” approach, where the DNA is eluted but not
pipetted away from the SPRI beads. Reactions are then carried out in the presence of the beads and bead
binding for the next cleanup initiated by the addition of PEG+NaCl solution.
Conclusion
For the highest-quality sequencing data, it is vital to prepare the best quality NGS library. A variety of
highly efficient formulations are available for library preparation, and by following the recommendations
outlined in this guide, you can minimize bias, maximize yield, standardize your workflow, and achieve
optimal results.