Junyi Li 0004

dblp:28/6612-4 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0001-8045-5264ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 MZSGO: multimodal zero-shot protein function annotation via evolutionary signals and textual semantics
abstract
MOTIVATION: Although deep learning has significantly advanced the field of protein function prediction, current approaches are limited by their reliance on a narrow set of modalities. Specifically, they primarily rely on sequence patterns and treat protein domain data and functional labels merely as categorical tags. Consequently, they fail to capitalize on the semantic richness embedded within their textual definitions. These constraints hinder their ability to generalize to novel labels. To tackle this issue, we present MZSGO, a multimodal zero-shot framework that fuses evolutionary signals from protein language models with semantic features derived from large language models (LLMs). By employing an adaptive gated fusion mechanism, MZSGO effectively aligns sequence-based and text-based modalities to enable robust predictions for unseen labels. RESULTS: By unifying protein representations and functional annotations, we bridge the semantic gap that limits current approaches. Results indicate that while our model remains competitive on supervised benchmarks, it demonstrates a marked advantage over existing methods in zero-shot tasks. It specifically excels at recognizing previously unseen long-tail and novel Gene Ontology (GO) terms. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/toxic-byte/MZSGO.
Boyue Cui, Yujuan Li, Shiqu Chen, Jiaming Wei, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.7
2026 GR2ST: spatial transcriptomics prediction based on graph-enhanced multimodal contrastive learning
abstract
MOTIVATION: Spatial transcriptomics techniques capture gene expression data and spatial coordinates, while simultaneously correlating them with tissue section images. This advantage makes Spatial transcriptomics data highly valuable for research, such as investigating disease mechanisms and cancer prognosis. However, the extended time and high cost of spatial transcriptomic sequencing currently limit further advancements in this field. The development of numerous deep learning methods aimed at predicting spatial transcriptomics from histology images has advanced significantly. However, these approaches often lack the ability to effectively integrate histology images with spatial transcriptomic data. Here, we propose GR2ST, a deep learning model that learns the underlying connections between image features and gene expression to predict spatial transcriptomics. RESULTS: GR2ST leverages a large pre-trained pathology model to extract high-level histological features. We designed a dual-branch graph architecture, consisting of a dynamic threshold-based functional graph and a radius-constrained spatial graph, to capture complex spot interactions within heterogeneous tissues. The model aligns histology images with gene expression representations through a multimodal contrastive learning framework. It achieves adaptive gene expression generation via a Cell-Type Guided Multi-Branch Regression Head supervised by a context-aware weighting network, which is further integrated with cross-sample retrieval to construct an ensemble prediction. The performance of the model is evaluated on three cancer-related spatial transcriptomics datasets, including cutaneous squamous cell carcinoma and two human breast cancer cohorts, to demonstrate its effectiveness and robustness. AVAILABILITY: https://github.com/zjl1109294570/GR2ST.
Jingli Zhou, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.6
2025 EMST: An Interpretable Multi-Modal Model for Spatial Transcriptomic Data Analysis
abstract
Spatial transcriptomic data provides gene expression, spatial positions, and histopathological images for each spot. However, due to technical difficulties in obtaining high-quality spatial transcriptomic data, it is crucial to enhance the quality of raw gene expression data through computational methods. Existing methods can not effectively fuse gene features and morphological features, and deep learning-based methods lack interpretability. Based on this, we propose EMST, an interpretable multi-modal model for spatial transcriptomic data analysis. We first extract morphological features of spots using a pre-trained large biological model, and then integrate morphological and gene features through a co-attention-based multi-modal fusion module. The fused features undergo an interpretable contrastive learning model to complete feature learning and enhance original gene expression data. Comparative experimental results on multiple datasets show that EMST can effectively enhance gene expression. We also conducted an interpretability analysis to make the training process of the model clearer.
Weiwei Yuan, Xuan Wang 0002, Junyi Li 0004
BIBM4
2025 E-MSNGO: Explainable Multi-species Protein Function Prediction Model Based on Aggregated Networks
Shiqu Chen, Junyi Li 0004
ICIC (26)4
2025 MKDTI: Predicting Drug-Target Interactions via Multiple Kernel Fusion on Graph Attention Network
Yuhuan Zhou, Yaqiu Wang, Yulin Wu 0001, Qian Chen 0028, Weiwei Yuan, Xuan Wang 0002, Junyi Li 0004
ICIC (25)7
2025 MSNGO: multi-species protein function annotation based on 3D protein structure and network propagation
abstract
MOTIVATION: In recent years, protein function prediction has broken through the bottleneck of sequence features, significantly improving prediction accuracy using high-precision protein structures predicted by AlphaFold2. While single-species protein function prediction methods have achieved remarkable success, multi-species approaches still face challenges such as difficulties in multi-source data integration and insufficient knowledge transfer between distantly-related species. How to integrate large-scale data and provide effective cross-species label propagation for species with sparse protein annotations remains a critical and unresolved challenge. To address this problem, we propose the MSNGO (Multi-species protein Structures and Network to predict GO terms) model, which integrates structural features and network propagation methods. Our validation shows that using structural features can significantly improve the accuracy of multi-species protein function prediction. RESULTS: We employ graph representation learning techniques to extract amino acid representations from protein structure contact maps and train a structural model using a graph convolution pooling module to derive protein-level structural features. After incorporating the sequence features from ESM-2, we apply a network propagation algorithm to aggregate information and update node representations within a heterogeneous network. The results demonstrate that MSNGO outperforms previous multi-species protein function prediction methods that rely on sequence features and protein-protein networks. AVAILABILITY AND IMPLEMENTATION: https://github.com/blingbell/MSNGO.
Boyue Cui, Shiqu Chen, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.6
2024 MCDHGN: heterogeneous network-based cancer driver gene prediction and interpretability analysis
abstract
MOTIVATION: Accurately predicting the driver genes of cancer is of great significance for carcinogenesis progress research and cancer treatment. In recent years, more and more deep-learning-based methods have been used for predicting cancer driver genes. However, deep-learning algorithms often have black box properties and cannot interpret the output results. Here, we propose a novel cancer driver gene mining method based on heterogeneous network meta-paths (MCDHGN), which uses meta-path aggregation to enhance the interpretability of predictions. RESULTS: MCDHGN constructs a heterogeneous network by using several types of multi-omics data that are biologically linked to genes. And the differential probabilities of SNV, DNA methylation, and gene expression data between cancerous tissues and normal tissues are extracted as initial features of genes. Nine meta-paths are manually selected, and the representation vectors obtained by aggregating information within and across meta-path nodes are used as new features for subsequent classification and prediction tasks. By comparing with eight homogeneous and heterogeneous network models on two pan-cancer datasets, MCDHGN has better performance on AUC and AUPR values. Additionally, MCDHGN provides interpretability of predicted cancer driver genes through the varying weights of biologically meaningful meta-paths. AVAILABILITY AND IMPLEMENTATION: https://github.com/1160300611/MCDHGN.
Lexiang Wang, Jingli Zhou, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.5
2024 circ2DGNN: circRNA-Disease Association Prediction via Transformer-Based Graph Neural Network
abstract
Investigating the associations between circRNA and diseases is vital for comprehending the underlying mechanisms of diseases and formulating effective therapies. Computational prediction methods often rely solely on known circRNA-disease data, indirectly incorporating other biomolecules' effects by computing circRNA and disease similarities based on these molecules. However, this approach is limited, as other biomolecules also play significant roles in circRNA-disease interactions. To address this, we construct a comprehensive heterogeneous network incorporating data on human circRNAs, diseases, and other biomolecule interactions to develop a novel computational model, circ2DGNN, which is built upon a heterogeneous graph neural network. circ2DGNN directly takes heterogeneous networks as inputs and obtains the embedded representation of each node for downstream link prediction through graph representation learning. circ2DGNN employs a Transformer-like architecture, which can compute heterogeneous attention score for each edge, and perform message propagation and aggregation, using a residual connection to enhance the representation vector. It uniquely applies the same parameter matrix only to identical meta-relationships, reflecting diverse parameter spaces for different relationship types. After fine-tuning hyperparameters via five-fold cross-validation, evaluation conducted on a test dataset shows circ2DGNN outperforms existing state-of-the-art(SOTA) methods.
Keliang Cen, Zheming Xing, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 NIEE: Modeling Edge Embeddings for Drug-Disease Association Prediction via Neighborhood Interactions
Yu Jiang 0002, Jingli Zhou, Yulin Wu 0001, Xuan Wang 0002, Junyi Li 0004
ICIC (3)6
2023 HNetGO: protein function prediction via heterogeneous network transformer
abstract
Protein function annotation is one of the most important research topics for revealing the essence of life at molecular level in the post-genome era. Current research shows that integrating multisource data can effectively improve the performance of protein function prediction models. However, the heavy reliance on complex feature engineering and model integration methods limits the development of existing methods. Besides, models based on deep learning only use labeled data in a certain dataset to extract sequence features, thus ignoring a large amount of existing unlabeled sequence data. Here, we propose an end-to-end protein function annotation model named HNetGO, which innovatively uses heterogeneous network to integrate protein sequence similarity and protein-protein interaction network information and combines the pretraining model to extract the semantic features of the protein sequence. In addition, we design an attention-based graph neural network model, which can effectively extract node-level features from heterogeneous networks and predict protein function by measuring the similarity between protein nodes and gene ontology term nodes. Comparative experiments on the human dataset show that HNetGO achieves state-of-the-art performance on cellular component and molecular function branches.
Xiaoshuai Zhang, Huannan Guo, Xuan Wang 0002, Kaitao Wu, Shizheng Qiu, Bo Liu 0023, Yadong Wang 0001, Yang Hu 0008, Junyi Li 0004
Briefings Bioinform.10
2023 Struct2GO: protein function prediction based on graph pooling algorithm and AlphaFold2 structure information
abstract
MOTIVATION: In recent years, there has been a breakthrough in protein structure prediction, and the AlphaFold2 model of the DeepMind team has improved the accuracy of protein structure prediction to the atomic level. Currently, deep learning-based protein function prediction models usually extract features from protein sequences and combine them with protein-protein interaction networks to achieve good results. However, for newly sequenced proteins that are not in the protein-protein interaction network, such models cannot make effective predictions. To address this, this article proposes the Struct2GO model, which combines protein structure and sequence data to enhance the precision of protein function prediction and the generality of the model. RESULTS: We obtain amino acid residue embeddings in protein structure through graph representation learning, utilize the graph pooling algorithm based on a self-attention mechanism to obtain the whole graph structure features, and fuse them with sequence features obtained from the protein language model. The results demonstrate that compared with the traditional protein sequence-based function prediction model, the Struct2GO model achieves better results. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available at https://github.com/lyjps/Struct2GO.
Peishun Jiao, Xuan Wang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
Bioinform.6
2023 SLGNN: synthetic lethality prediction in human cancers based on factor-aware knowledge graph neural network
abstract
MOTIVATION: Synthetic lethality (SL) is a form of genetic interaction that can selectively kill cancer cells without damaging normal cells. Exploiting this mechanism is gaining popularity in the field of targeted cancer therapy and anticancer drug development. Due to the limitations of identifying SL interactions from laboratory experiments, an increasing number of research groups are devising computational prediction methods to guide the discovery of potential SL pairs. Although existing methods have attempted to capture the underlying mechanisms of SL interactions, methods that have a deeper understanding of and attempt to explain SL mechanisms still need to be developed. RESULTS: In this work, we propose a novel SL prediction method, SLGNN. This method is based on the following assumption: SL interactions are caused by different molecular events or biological processes, which we define as SL-related factors that lead to SL interactions. SLGNN, apart from identifying SL interaction pairs, also models the preferences of genes for different SL-related factors, making the results more interpretable for biologists and clinicians. SLGNN consists of three steps: first, we model the combinations of relationships in the gene-related knowledge graph as the SL-related factors. Next, we derive initial embeddings of genes through an explicit message aggregation process of the knowledge graph. Finally, we derive the final gene embeddings through an SL graph, constructed using known SL gene pairs, utilizing factor-based message aggregation. At this stage, a supervised end-to-end training model is used for SL interaction prediction. Based on experimental results, the proposed SLGNN model outperforms all current state-of-the-art SL prediction methods and provides better interpretability. AVAILABILITY AND IMPLEMENTATION: SLGNN is freely available at https://github.com/zy972014452/SLGNN.
Yuhuan Zhou, Yang Liu 0039, Xuan Wang 0002, Junyi Li 0004
Bioinform.5
2023 ScCCL: Single-Cell Data Clustering Based on Self-Supervised Contrastive Learning
abstract
The growing maturity of single-cell RNA-sequencing (scRNA-seq) technology allows us to explore the heterogeneity of tissues, organisms, and complex diseases at cellular level. In single-cell data analysis, clustering calculation is very important. However, the high dimensionality of scRNA-seq data, the ever-increasing number of cells, and the unavoidable technical noise bring great challenges to clustering calculations. Motivated by the good performance of contrastive learning in multiple domains, we propose ScCCL, a novel self-supervised contrastive learning method for clustering of scRNA-seq data. ScCCL first randomly masks the gene expression of each cell twice and adds a small amount of Gaussian noise, and then uses the momentum encoder structure to extract features from the enhanced data. Contrastive learning is then applied in the instance-level contrastive learning module and the cluster-level contrastive learning module, respectively. After training, a representation model that can efficiently extract high-order embeddings of single cells is obtained. We selected two evaluation metrics, ARI and NMI, to conduct experiments on multiple public datasets. The results show that ScCCL improves the clustering effect compared with the benchmark algorithms. Notably, since ScCCL does not depend on a specific type of data, it can also be helpful in clustering analysis of single-cell multi-omics data.
Linlin Du, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 PSPGO: Cross-Species Heterogeneous Network Propagation for Protein Function Prediction
abstract
How to use computational methods to effectively predict the function of proteins remains a challenge. Most prediction methods based on single species or single data source have some limitations: the former need to train different models for different species, the latter only to infer protein function from a single perspective, such as the method only using Protein-Protein Interaction (PPI) network just considers the protein environment but ignore the intrinsic characteristics of protein sequences. We found that in some network-based multi-species methods the networks of each species are isolated, which means there is no communication between networks of different species. To solve these problems, we propose a cross-species heterogeneous network propagation method based on graph attention mechanism, PSPGO, which can propagate feature and label information on sequence similarity (SS) network and PPI network for predicting gene ontology terms. Our model is evaluated on a large multi-species dataset split based on time and is compared with several state-of-the-art methods. The results show that our method has good performance. We also explore the predictive performance of PSPGO for a single species. The results illustrate that PSPGO also performs well in prediction for single species.
Kaitao Wu, Lexiang Wang, Bo Liu 0023, Yang Liu 0039, Yadong Wang 0001, Junyi Li 0004
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 Prot2GO: Predicting GO Annotations From Protein Sequences and Interactions
abstract
Protein is the main material basis of living organisms and plays crucial role in life activities. Understanding the function of protein is of great significance for new drug discovery, disease treatment and vaccine development. In recent years, with the widespread application of deep learning in bioinformatics, researchers have proposed many deep learning models to predict protein functions. However, the existing deep learning methods usually only consider protein sequences, and thus cannot effectively integrate multi-source data to annotate protein functions. In this article, we propose the Prot2GO model, which can integrate protein sequence and PPI network data to predict protein functions. We utilize an improved biased random walk algorithm to extract the features of PPI network. For sequence data, we use a convolutional neural network to obtain the local features of the sequence and a recurrent neural network to capture the long-range associations between amino acid residues in protein sequence. Moreover, Prot2GO adopts the attention mechanism to identify protein motifs and structural domains. Experiments show that Prot2GO model achieves the state-of-the-art performance on multiple metrics.
Xiaoshuai Zhang, Hucheng Liu, Xiaofeng Zhang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
IEEE ACM Trans. Comput. Biol. Bioinform.7
2022 NSAP: A Neighborhood Subgraph Aggregation Method for Drug-Disease Association Prediction
Qiqi Jiao, Yu Jiang 0002, Yang Zhang 0057, Yadong Wang 0001, Junyi Li 0004
ICIC (2)5
2022 Deepgmd: A Graph-Neural-Network-Based Method to Detect Gene Regulator Module
abstract
Regulatory module mining methods divide genes into multiple gene subgroups and explore potential biological mechanisms from omics data. By transforming gene expression profile data into gene co-expression network, we transform the task of gene module detection into the problem of finding community structure in the graph, and introduce the latest network representation learning method-graph neural network to optimize this problem. In order to systematically evaluate whether the algorithm allows overlap to affect such problems, we make two variants of the output of the algorithm, Deepgmd_cluster and Deepgmd. The difference between them is whether overlap is allowed. By comparing the known modules and the modules generated by the algorithm, we can evaluate the quality of the algorithm. We use this method to compare our algorithm with some current mainstream methods. The results show that our method has greater advantages. In the end, we analyze some typical modules from the modules found by the algorithm for visualization, and use the GO database and KEGG database to perform enrichment analysis and pathway analysis on these modules.
Yulin Wu 0001, Jiangsheng Pi, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
IEEE ACM Trans. Comput. Biol. Bioinform.7
2021 ScSSC: Semi-supervised Single Cell Clustering Based on 2D Embedding
Naile Shi, Yulin Wu 0001, Linlin Du, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
ICIC (3)6
2021 Factor graph-aggregated heterogeneous network embedding for disease-gene association prediction
abstract
BACKGROUND: Exploring the relationship between disease and gene is of great significance for understanding the pathogenesis of disease and developing corresponding therapeutic measures. The prediction of disease-gene association by computational methods accelerates the process. RESULTS: Many existing methods cannot fully utilize the multi-dimensional biological entity relationship to predict disease-gene association due to multi-source heterogeneous data. This paper proposes FactorHNE, a factor graph-aggregated heterogeneous network embedding method for disease-gene association prediction, which captures a variety of semantic relationships between the heterogeneous nodes by factorization. It produces different semantic factor graphs and effectively aggregates a variety of semantic relationships, by using end-to-end multi-perspectives loss function to optimize model. Then it produces good nodes embedding to prediction disease-gene association. CONCLUSIONS: Experimental verification and analysis show FactorHNE has better performance and scalability than the existing models. It also has good interpretability and can be extended to large-scale biomedical network data analysis.
Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
BMC Bioinform.5
2021 IIMLP: integrated information-entropy-based method for LncRNA prediction
abstract
BACKGROUND: The prediction of long non-coding RNA (lncRNA) has attracted great attention from researchers, as more and more evidence indicate that various complex human diseases are closely related to lncRNAs. In the era of bio-med big data, in addition to the prediction of lncRNAs by biological experimental methods, many computational methods based on machine learning have been proposed to make better use of the sequence resources of lncRNAs. RESULTS: We developed the lncRNA prediction method by integrating information-entropy-based features and machine learning algorithms. We calculate generalized topological entropy and generate 6 novel features for lncRNA sequences. By employing these 6 features and other features such as open reading frame, we apply supporting vector machine, XGBoost and random forest algorithms to distinguish human lncRNAs. We compare our method with the one which has more K-mer features and results show that our method has higher area under the curve up to 99.7905%. CONCLUSIONS: We develop an accurate and efficient method which has novel information entropy features to analyze and classify lncRNAs. Our method is also extendable for research on the other functional elements in DNA sequences.
Junyi Li 0004, Huinian Li, Qingzhe Xu, Xiaozhu Jing, Wei Jiang 0023, Qing Liao 0001, Bo Liu 0023, Yadong Wang 0001
BMC Bioinform.1
2021 deGSM: Memory Scalable Construction Of Large Scale de Bruijn Graph
abstract
The de Bruijn graph, a fundamental data structure to represent and organize genome sequence, plays important roles in various kinds of sequence analysis tasks. With the rapid development of HTS data and ever-increasing number of assembled genomes, there is a high demand to construct the very large de Bruijn graph for sequences up to Tera-base-pair level. Current approaches may have unaffordable memory footprints to handle such a large de Bruijn graph. We propose a lightweight parallel de Bruijn graph construction approach: de Bruijn Graph Constructor in Scalable Memory (deGSM). The main idea of deGSM is to efficiently construct the Burrows-Wheeler Transformation (BWT) of the unipaths of the de Bruijn graph in constant RAM space and transform the BWT into the original unitigs. The experimental results demonstrate that, just with a commonly available machine, deGSM is able to handle very large genome sequence(s), e.g., the contigs (305 Gbp) and scaffolds (1.1 Tbp) recorded in GenBank database and Picea abies HTS dataset (9.7 Tbp). Moreover, deGSM also has faster or comparable construction speed compared with state-of-the-art approaches. With its high scalability and efficiency, deGSM has enormous potential in many large scale genomics studies. The deGSM is publicly available at: https://github.com/hitbc/deGSM.
Hongzhe Guo, Yilei Fu, Yan Gao 0031, Junyi Li 0004, Yadong Wang 0001, Bo Liu 0023
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 GONET: A Deep Network to Annotate Proteins via Recurrent Convolution Networks
abstract
Finding out the functions of protein in life activities precisely is nontrivial, which is the core of current proteomics research. Gene Ontology standardizes the function of protein into a series of GO terms, each of which belongs to exactly one of the three subontologies: Biological Process (BP), Cellular Component (CC), and Molecular Function (MF). The prediction of protein function can be considered as a multi-label classification problem. Traditional methods often spend a lot of costs to extract handcrafted features and plenty of domain knowledge is needed when solving these tasks, while using deep learning technology can overcome these shortcomings. Here, we propose a deep model GONET based on recurrent convolutional neural networks, which annotates protein in an end-to-end manner. Our model combines protein sequences and protein-protein interaction (PPI) network data, and utilizes representation learning to learn distributed representation of proteins to overcome the sparse nature and semantic independence problem. Moreover, we adopt a quite deep CNNRNN-Attention model, which is able to effectively extract high-order features of protein sequences. We have carried out experiments on several datasets, which achieve the state-of-the-art in some metrics compared with the existing competitive methods.
Junyi Li 0004, Xiaoshuai Zhang, Bo Liu 0023, Yadong Wang 0001
BIBM1
2020 HGAlinker: Drug-Disease Association Prediction Based on Attention Mechanism of Heterogeneous Graph
Xiaozhu Jing, Wei Jiang 0023, Zhongqing Zhang, Yadong Wang 0001, Junyi Li 0004
ICIC (2)5
2019 Prediction of Human LncRNAs Based on Integrated Information Entropy Features
Junyi Li 0004, Huinian Li, Qingzhe Xu, Xiaozhu Jing, Wei Jiang 0023, Bo Liu 0023, Yadong Wang 0001
ICIC (2)1
2019 rMETL: sensitive mobile element insertion detection with long read realignment
abstract
SUMMARY: Mobile element insertion (MEI) is a major category of structure variations (SVs). The rapid development of long read sequencing technologies provides the opportunity to detect MEIs sensitively. However, the signals of MEI implied by noisy long reads are highly complex due to the repetitiveness of mobile elements as well as the high sequencing error rates. Herein, we propose the Realignment-based Mobile Element insertion detection Tool for Long read (rMETL). Benchmarking results of simulated and real datasets demonstrate that rMETL enables to handle the complex signals to discover MEIs sensitively. It is suited to produce high-quality MEI callsets in many genomics studies. AVAILABILITY AND IMPLEMENTATION: rMETL is available from https://github.com/hitbc/rMETL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Jiang 0021, Bo Liu 0023, Junyi Li 0004, Yadong Wang 0001
Bioinform.3
2019 Integrated entropy-based approach for analyzing exons and introns in DNA sequences
abstract
BACKGROUND: Numerous essential algorithms and methods, including entropy-based quantitative methods, have been developed to analyze complex DNA sequences since the last decade. Exons and introns are the most notable components of DNA and their identification and prediction are always the focus of state-of-the-art research. RESULTS: In this study, we designed an integrated entropy-based analysis approach, which involves modified topological entropy calculation, genomic signal processing (GSP) method and singular value decomposition (SVD), to investigate exons and introns in DNA sequences. We optimized and implemented the topological entropy and the generalized topological entropy to calculate the complexity of DNA sequences, highlighting the characteristics of repetition sequences. By comparing digitalizing entropy values of exons and introns, we observed that they are significantly different. After we converted DNA data to numerical topological entropy value, we applied SVD method to effectively investigate exon and intron regions on a single gene sequence. Additionally, several genes across five species are used for exon predictions. CONCLUSIONS: Our approach not only helps to explore the complexity of DNA sequence and its functional elements, but also provides an entropy-based GSP method to analyze exon and intron regions. Our work is feasible across different species and extendable to analyze other components in both coding and noncoding region of DNA sequences.
Junyi Li 0004, Huinian Li, Qingzhe Xu, Renjie Tan, Bo Liu 0023, Yadong Wang 0001
BMC Bioinform.1
2018 DeepDNA: a hybrid convolutional and recurrent neural network for compressing human mitochondrial genomes
Yan-Shuo Chu, Yongtian Wang, Mingrui Sun, Junyi Li 0004, Tianyi Zang, Yadong Wang 0001
BIBM7
2017 Differential regulation analysis reveals dysfunctional regulatory mechanism involving transcription factors and microRNAs in gastric carcinogenesis
Quanxue Li, Junyi Li 0004, Wentao Dai
Artif. Intell. Medicine2
2017 rMFilter: acceleration of long read-based structure variation calling by chimeric read filtering
abstract
MOTIVATION: Long read sequencing technologies provide new opportunities to investigate genome structural variations (SVs) more accurately. However, the state-of-the-art SV calling pipelines are computational intensive and the applications of long reads are restricted. RESULTS: We propose a local region match-based filter (rMFilter) to efficiently nail down chimeric noisy long reads based on short token matches within local genomic regions. rMFilter is able to substantially accelerate long read-based SV calling pipelines without loss of effectiveness. It can be easily integrated into current long read-based pipelines to facilitate SV studies. AVAILABILITY AND IMPLEMENTATION: The C ++ source code of rMFilter is available at https://github.com/hitbc/rMFilter . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bo Liu 0023, Tao Jiang 0021, Siu-Ming Yiu, Junyi Li 0004, Yadong Wang 0001
Bioinform.4