We've updated our Privacy Policy to make it clearer how we use your personal data. We use cookies to provide you with a better experience. You can read our Cookie Policy here.

Advertisement

AI Model Tracks DNA Signals Behind RNA Splicing

Diagram of a DNA sequence analysis workflow using SpliceSelectNet to predict RNA splicing outcomes from genomic sequences.
A hierarchical Transformer model analyzes DNA sequences spanning up to 100,000 base pairs, linking distant regulatory elements with splice-site activity while highlighting biologically important regions through attention maps. Credit: Professor Kenta Nakai / The Institute of Medical Science, The University of Tokyo, Japan.
Read time: 2 minutes

Scientists have developed a hierarchical Transformer-based artificial intelligence model that predicts RNA splice sites across DNA sequences up to 100,000 base pairs long while preserving single-nucleotide resolution. The system captures long-range regulatory interactions that conventional approaches often miss and visualizes the sequence regions influencing its predictions. The approach achieved state-of-the-art performance on multiple benchmark datasets and may support future studies in genetic diseases, cancer, precision medicine, genomic medicine, and therapeutic development.

Accurate RNA splicing is essential for gene expression and human health, yet predicting how DNA sequence variations affect splicing remains a major challenge. Although recent artificial intelligence (AI) models have improved splice site prediction, many struggle to capture regulatory signals located thousands of DNA bases away from the sites they influence. This limitation restricts our ability to understand disease-causing mutations and the complex mechanisms governing RNA processing, particularly in disorders ranging from genetic diseases to cancer.

To address these challenges, Professor Kenta Nakai from the Human Genome Center, the Institute of Medical Science and Ms. Yuna Miyachi, a Ph.D. student, from the Department of Computer Science, Graduate School of Information Science and Technology, both at The University of Tokyo, Japan, developed SpliceSelectNet (SSNet), a hierarchical Transformer-based deep learning framework for splice site prediction. Their study, published in Nucleic Acids Research on June 22, 2026, introduces a computational approach capable of analyzing DNA sequences spanning up to 100,000 base pairs while maintaining single-nucleotide resolution. By combining local and global attention mechanisms, SSNet efficiently captures both nearby and distant regulatory signals that contribute to RNA splicing.

Many existing computational tools struggle to model long-range genomic interactions because the computational cost increases rapidly with sequence length. To overcome this limitation, SSNet divides long DNA sequences into smaller blocks, analyzes local patterns within each block, and then integrates information across the entire sequence through a hierarchical attention process. This design allows the model to preserve dense attention while remaining computationally efficient. In addition, the researchers enabled visualization of attention scores, allowing them to identify which DNA regions the model considered important during prediction.

The model was trained and evaluated using several large genomic datasets and benchmarked against leading splice prediction systems. Across multiple validation datasets, SSNet achieved state-of-the-art performance for splice site prediction and aberrant splicing detection. The researchers also showed that the model could capture the effects of distant regulatory sequences beyond the effective range of conventional convolutional neural network approaches. In simulations using the DMD gene and evaluations of pathogenic variants from ClinVar, SSNet maintained sensitivity to regulatory signals located many thousands of base pairs from the affected splice site.

"The key achievement of this work is that we successfully modeled ultra-long-range genomic interactions while preserving high computational efficiency and single-nucleotide resolution," says Prof. Nakai. "We also demonstrated that the regions highlighted by the model closely correspond to biologically meaningful regulatory elements, helping to bridge predictive accuracy and biological interpretability."

The study suggests that hierarchical Transformer architectures could become valuable tools beyond splice site prediction. The same framework may support future research into promoter-enhancer interactions, three-dimensional genome organization, and broader DNA language models. The researchers also expect opportunities for collaboration with researchers in clinical and genomic medicine, where the technology could help screen variants in non-coding regions that currently have uncertain significance. In pharmaceutical research, the approach could assist in designing oligonucleotide therapeutics that target abnormal splicing.

"Many existing AI models for DNA analysis were adapted from natural language processing, but DNA has fundamentally different properties," explains Ms. Miyachi. "By redesigning the architecture to account for long-range genomic interactions and strict sequence resolution, we aimed to create a system better suited to biological reality."


Reference: Miyachi Y, Nakai K. SpliceSelectNet: a hierarchical Transformer-based deep learning model for splice site prediction. Nucleic Acids Res. 2026;54(12):gkag625. doi: 10.1093/nar/gkag625

This article has been republished from the following materials. Note: material may have been edited for length and content. For further information, please contact the cited source. Our press release publishing policy can be accessed here.

Google News Preferred Source Add Technology Networks as a preferred Google source to see more of our trusted coverage.