Hao Lin 0001

dblp:89/3472-1 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0001-6265-2862ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021
YearPublicationVenuePosition
2026 ac4C modification sites prediction in human mRNA: a complete review
abstract
ac4C alteration in RNA is a conserved epigenetic mark that is critical for post-transcriptional control, mRNA stability, translational efficiency, and human immune function regulation. In the meantime, the conventional experimental procedures for predicting ac4C alteration sites are costly, time-consuming, and difficult. The precise recognition of ac4C modification sites in human mRNA has been greatly aided by computational prediction techniques using sequence data, machine learning (ML), deep learning (DL), and large language models (LLMs). The application of ML, DL, and LLM-based techniques for the identification of ac4C modification sites in human mRNA has been evaluated and contrasted in this review. Distinctively, we have also addressed the shortcomings of the existing methods and tools, as well as potential future developments. We anticipate that this study will provide sufficient information and awareness for ac4C modification research.
Hasan Zulfiqar, Ramala Masood Ahmad, Hao Lin 0001, Xiao-Long Yu
Briefings Bioinform.3
2026 HybridGNN: a graph neural network approach for human miRNA-disease association prediction
abstract
MOTIVATION: MicroRNAs (miRNAs) are small non-coding RNAs, typically 18-24 nucleotides in length, that play a pivotal role in RNA silencing and the post-transcriptional regulation of gene expression by targeting messenger RNAs (mRNAs). Dysregulation of these miRNAs has consistently been implicated in the onset and progression of a variety of complex human diseases. RESULTS: In this study, we propose a novel HybridGNN model that integrates a Graph Convolutional Network (GCN), a Graph Attention Network (GAT), and Matrix Decomposition with Matrix Factorization (MDMF) to predict potential miRNA-disease associations (MDAs). We incorporate five types of similarity in which three are derived from miRNAs and two are derived from diseases, to comprehensively explore and optimize multi-source feature information. The complementary interactions among these modules also help to mitigate the oversmoothing problem. The model utilizes neighboring nodes in a heterogeneous network to generate node embeddings via a message-passing mechanism. To improve computational efficiency, we employ a mini-batch gradient descent approach that partitions the graph into smaller sub-graphs, thereby enhancing the model's accuracy, speed, and scalability. As a result of these advanced techniques, HybridGNN achieved an area under the receiver operating characteristic curve (AUC-ROC) of 0.9715 using a dot-product classifier, outperforming several existing methods and underscoring its potential as a robust and accurate tool for predicting MDAs. AVAILABILITY: Code and data are freely available at https://github.com/mbasharatahmad/HybridGNN-miRNA-disease/.
Basharat Ahmad, Muhammad Hammad Musaddiq, Aboma Temesgen Sebu, Bakanina Kissanga Grace-Mercure, Huma Fida, Hao Lin 0001, Ye-Chen Qi
Bioinform.6
2026 DPS-Tool: an online service platform for disease perturbation scoring
Changchun Wu, Xueqin Xie, Ziru Huang, Hao Lin 0001, Jian Huang 0004
Frontiers Comput. Sci.4
2025 Navigating the 3D genome at single-cell resolution: techniques, computation, and mechanistic landscapes
abstract
The 3D organization of the genome is critical for gene expression regulation, cellular identity, and disease progression. Traditional methods that analyze bulk genomic data often obscure cell-to-cell heterogeneity, limiting the resolution of intrinsic variability within complex biological systems. To overcome this, single-cell 3D genomics has emerged, revealing chromatin architecture at the individual cell level. Advanced experimental approaches enable genome-wide chromatin contact mapping, while computational frameworks reconstruct dynamic chromatin topologies from high-dimensional data. Building on these breakthroughs, recent advances in single-cell 3D genomics have led to transformative progress in epigenetics, linking 3D genome architecture with gene regulation, cellular identity, and disease phenotypes. This review focuses on the breakthroughs in single-cell 3D genomics, demonstrating how integrated experimental, computational, and mechanistic approaches decode chromatin architecture. These insights have deepened the understanding of genome function at the single-cell level and lay the foundation for future advances in precision medicine and topology-guided therapeutic strategies.
Feitong Hong, Kaiyuan Han, Yuduo Hao, Xueqin Xie, Qiuming Chen, Yijie Wei, Xinwei Luo, Sijia Xie, Benjamin Lebeau, Crystal Ling, Hao Lv 0007, Hao Lin 0001, Fu-Ying Dao
Briefings Bioinform.15
2024 Muli-Task Dual-view Model for Drug Combination Analysis: In Vivo and In Vitro Perspectives
abstract
Drug-drug interactions (DDIs) and drug compatibility interactions (DCIs) are both critical components of drug combination (DC), essential for ensuring safe and effective therapeutic strategies. DDIs typically occur in vivo, while DCIs occur in vitro. However, most existing drug combination prediction models focus solely on either DDIs or DCIs, lacking a comprehensive approach to predict DC in vivo and in vitro perspectives, which can lead to incomplete understanding and management of DC risks. To address this problem, we propose a multi-task dual-view model for drug combination (MDDC) that simultaneously predicts DDIs and DCIs by using pre-trained molecular sequence and graph representations. MDDC utilized the pretrained models ChemBERTa-2 and MolCLR to capture diverse view molecular features and fine-tune them in a multi-task learning framework. This approach demonstrated notable performance improvements, with AUROCs of 0.914 for DDI task, and 0.826 for DCI tasks on independent testing dataset. MDDC offers a more comprehensive understanding of potential drug interactions in medication guidance, thereby reducing adverse effect risks and enhancing therapeutic safety.
Zhao-Yue Zhang 0002, Xueqin Xie, Kejun Deng, Hao Lin 0001
BIBM6
2024 Attention is all you need: utilizing attention in AI-enabled drug discovery
abstract
Recently, attention mechanism and derived models have gained significant traction in drug development due to their outstanding performance and interpretability in handling complex data structures. This review offers an in-depth exploration of the principles underlying attention-based models and their advantages in drug discovery. We further elaborate on their applications in various aspects of drug development, from molecular screening and target binding to property prediction and molecule generation. Finally, we discuss the current challenges faced in the application of attention mechanisms and Artificial Intelligence technologies, including data quality, model interpretability and computational resource constraints, along with future directions for research. Given the accelerating pace of technological advancement, we believe that attention-based models will have an increasingly prominent role in future drug discovery. We anticipate that these models will usher in revolutionary breakthroughs in the pharmaceutical domain, significantly accelerating the pace of drug development.
Yang Zhang 0125, Caiqi Liu, Mujiexin Liu, Hao Lin 0001, Cheng-Bing Huang, Lin Ning 0002
Briefings Bioinform.5
2024 Detecting key genes relative expression orderings as biomarkers for machine learning-based intelligent screening and analysis of type 2 diabetes mellitus
Xueqin Xie, Changchun Wu, Cai-Yi Ma, Dong Gao 0002, Jian Huang 0004, Kejun Deng, Dan Yan, Hao Lin 0001
Expert Syst. Appl.9
2023 A comprehensive review of bioinformatics tools for chromatin loop calling
abstract
Precisely calling chromatin loops has profound implications for further analysis of gene regulation and disease mechanisms. Technological advances in chromatin conformation capture (3C) assays make it possible to identify chromatin loops in the genome. However, a variety of experimental protocols have resulted in different levels of biases, which require distinct methods to call true loops from the background. Although many bioinformatics tools have been developed to address this problem, there is still a lack of special introduction to loop-calling algorithms. This review provides an overview of the loop-calling tools for various 3C-based techniques. We first discuss the background biases produced by different experimental techniques and the denoising algorithms. Then, the completeness and priority of each tool are categorized and summarized according to the data source of application. The summary of these works can help researchers select the most appropriate method to call loops and further perform downstream analysis. In addition, this survey is also useful for bioinformatics scientists aiming to develop new loop-calling algorithms.
Kaiyuan Han, Huimin Sun, Dong Gao 0002, Qilemuge Xi, Lirong Zhang, Hao Lin 0001
Briefings Bioinform.8
2022 Detection of transcription factors binding to methylated DNA by deep recurrent neural network
abstract
Transcription factors (TFs) are proteins specifically involved in gene expression regulation. It is generally accepted in epigenetics that methylated nucleotides could prevent the TFs from binding to DNA fragments. However, recent studies have confirmed that some TFs have capability to interact with methylated DNA fragments to further regulate gene expression. Although biochemical experiments could recognize TFs binding to methylated DNA sequences, these wet experimental methods are time-consuming and expensive. Machine learning methods provide a good choice for quickly identifying these TFs without experimental materials. Thus, this study aims to design a robust predictor to detect methylated DNA-bound TFs. We firstly proposed using tripeptide word vector feature to formulate protein samples. Subsequently, based on recurrent neural network with long short-term memory, a two-step computational model was designed. The first step predictor was utilized to discriminate transcription factors from non-transcription factors. Once proteins were predicted as TFs, the second step predictor was employed to judge whether the TFs can bind to methylated DNA. Through the independent dataset test, the accuracies of the first step and the second step are 86.63% and 73.59%, respectively. In addition, the statistical analysis of the distribution of tripeptides in training samples showed that the position and number of some tripeptides in the sequence could affect the binding of TFs to methylated DNA. Finally, on the basis of our model, a free web server was established based on the proposed model, which can be available at https://bioinfor.nefu.edu.cn/TFPM/.
Hao Lin 0001, Guohua Wang 0001
Briefings Bioinform.4
2022 iRice-MS: An integrated XGBoost model for detecting multitype post-translational modification sites in rice
abstract
Post-translational modification (PTM) refers to the covalent and enzymatic modification of proteins after protein biosynthesis, which orchestrates a variety of biological processes. Detecting PTM sites in proteome scale is one of the key steps to in-depth understanding their regulation mechanisms. In this study, we presented an integrated method based on eXtreme Gradient Boosting (XGBoost), called iRice-MS, to identify 2-hydroxyisobutyrylation, crotonylation, malonylation, ubiquitination, succinylation and acetylation in rice. For each PTM-specific model, we adopted eight feature encoding schemes, including sequence-based features, physicochemical property-based features and spatial mapping information-based features. The optimal feature set was identified from each encoding, and their respective models were established. Extensive experimental results show that iRice-MS always display excellent performance on 5-fold cross-validation and independent dataset test. In addition, our novel approach provides the superiority to other existing tools in terms of AUC value. Based on the proposed model, a web server named iRice-MS was established and is freely accessible at http://lin-group.cn/server/iRice-MS.
Hao Lv 0007, Yang Zhang 0125, Jia-Shu Wang, Shi-Shi Yuan, Fu-Ying Dao, Zheng-Xing Guan, Hao Lin 0001, Ke-Jun Deng
Briefings Bioinform.8
2022 PSnoD: identifying potential snoRNA-disease associations based on bounded nuclear norm regularization
abstract
Many studies have proved that small nucleolar RNAs (snoRNAs) play critical roles in the development of various human complex diseases. Discovering the associations between snoRNAs and diseases is an important step toward understanding the pathogenesis and characteristics of diseases. However, uncovering associations via traditional experimental approaches is costly and time-consuming. This study proposed a bounded nuclear norm regularization-based method, called PSnoD, to predict snoRNA-disease associations. Benchmark experiments showed that compared with the state-of-the-art methods, PSnoD achieved a superior performance in the 5-fold stratified shuffle split. PSnoD produced a robust performance with an area under receiver-operating characteristic of 0.90 and an area under precision-recall of 0.55, highlighting the effectiveness of our proposed method. In addition, the computational efficiency of PSnoD was also demonstrated by comparison with other matrix completion techniques. More importantly, the case study further elucidated the ability of PSnoD to screen potential snoRNA-disease associations. The code of PSnoD has been uploaded to https://github.com/linDing-groups/PSnoD. Based on PSnoD, we established a web server that is freely accessed via http://psnod.lin-group.cn/.
Hao Lv 0007, Yang Zhang 0125, Hao Lin 0001, Lin Ning 0002
Briefings Bioinform.7
2022 iLoc-miRNA: extracellular/intracellular miRNA prediction using deep BiLSTM with attention mechanism
abstract
The location of microRNAs (miRNAs) in cells determines their function in regulation activity. Studies have shown that miRNAs are stable in the extracellular environment that mediates cell-to-cell communication and are located in the intracellular region that responds to cellular stress and environmental stimuli. Though in situ detection techniques of miRNAs have made great contributions to the study of the localization and distribution of miRNAs, miRNA subcellular localization and their role are still in progress. Recently, some machine learning-based algorithms have been designed for miRNA subcellular location prediction, but their performance is still far from satisfactory. Here, we present a new data partitioning strategy that categorizes functionally similar locations for the precise and instructive prediction of miRNA subcellular location in Homo sapiens. To characterize the localization signals, we adopted one-hot encoding with post padding to represent the whole miRNA sequences, and proposed a deep bidirectional long short-term memory with the multi-head self-attention algorithm to model. The algorithm showed high selectivity in distinguishing extracellular miRNAs from intracellular miRNAs. Moreover, a series of motif analyses were performed to explore the mechanism of miRNA subcellular localization. To improve the convenience of the model, a user-friendly web server named iLoc-miRNA was established (http://iLoc-miRNA.lin-group.cn/).
Zhao-Yue Zhang 0002, Lin Ning 0002, Xiucai Ye, Yasunori Futamura, Tetsuya Sakurai, Hao Lin 0001
Briefings Bioinform.7
2022 A deep learning model to identify gene expression level using cobinding transcription factor signals
abstract
Gene expression is directly controlled by transcription factors (TFs) in a complex combination manner. It remains a challenging task to systematically infer how the cooperative binding of TFs drives gene activity. Here, we quantitatively analyzed the correlation between TFs and surveyed the TF interaction networks associated with gene expression in GM12878 and K562 cell lines. We identified six TF modules associated with gene expression in each cell line. Furthermore, according to the enrichment characteristics of TFs in these TF modules around a target gene, a convolutional neural network model, called TFCNN, was constructed to identify gene expression level. Results showed that the TFCNN model achieved a good prediction performance for gene expression. The average of the area under receiver operating characteristics curve (AUC) can reach up to 0.975 and 0.976, respectively in GM12878 and K562 cell lines. By comparison, we found that the TFCNN model outperformed the prediction models based on SVM and LDA. This is due to the TFCNN model could better extract the combinatorial interaction among TFs. Further analysis indicated that the abundant binding of regulatory TFs dominates expression of target genes, while the cooperative interaction between TFs has a subtle regulatory effects. And gene expression could be regulated by different TF combinations in a nonlinear way. These results are helpful for deciphering the mechanism of TF combination regulating gene expression.
Lirong Zhang, Lu Chai, Qianzhong Li, Hao Lin 0001
Briefings Bioinform.6
2022 Towards a better prediction of subcellular location of long non-coding RNA
Zhao-Yue Zhang 0002, Hao Lin 0001
Frontiers Comput. Sci.4
2021 iDHS-Deep: an integrated tool for predicting DNase I hypersensitive sites by deep neural network
abstract
DNase I hypersensitive site (DHS) refers to the hypersensitive region of chromatin for the DNase I enzyme. It is an important part of the noncoding region and contains a variety of regulatory elements, such as promoter, enhancer, and transcription factor-binding site, etc. Moreover, the related locus of disease (or trait) are usually enriched in the DHS regions. Therefore, the detection of DHS region is of great significance. In this study, we develop a deep learning-based algorithm to identify whether an unknown sequence region would be potential DHS. The proposed method showed high prediction performance on both training datasets and independent datasets in different cell types and developmental stages, demonstrating that the method has excellent superiority in the identification of DHSs. Furthermore, for the convenience of related wet-experimental researchers, the user-friendly web-server iDHS-Deep was established at http://lin-group.cn/server/iDHS-Deep/, by which users can easily distinguish DHS and non-DHS and obtain the corresponding developmental stage ofDHS.
Fu-Ying Dao, Hao Lv 0007, Hao Lin 0001
Briefings Bioinform.6
2021 A computational platform to identify origins of replication sites in eukaryotes
abstract
The locations of the initiation of genomic DNA replication are defined as origins of replication sites (ORIs), which regulate the onset of DNA replication and play significant roles in the DNA replication process. The study of ORIs is essential for understanding the cell-division cycle and gene expression regulation. Accurate identification of ORIs will provide important clues for DNA replication research and drug development by developing computational methods. In this paper, the first integrated predictor named iORI-Euk was built to identify ORIs in multiple eukaryotes and multiple cell types. In the predictor, seven eukaryotic (Homo sapiens, Mus musculus, Drosophila melanogaster, Arabidopsis thaliana, Pichia pastoris, Schizosaccharomyces pombe and Kluyveromyces lactis) ORI data was collected from public database to construct benchmark datasets. Subsequently, three feature extraction strategies which are k-mer, binary encoding and combination of k-mer and binary were used to formulate DNA sequence samples. We also compared the different classification algorithms' performance. As a result, the best results were obtained by using support vector machine in 5-fold cross-validation test and independent dataset test. Based on the optimal model, an online web server called iORI-Euk (http://lin-group.cn/server/iORI-Euk/) was established for the novel ORI identification.
Fu-Ying Dao, Hao Lv 0007, Hasan Zulfiqar, Hui Ding 0005, Hao Lin 0001
Briefings Bioinform.8
2021 DeepYY1: a deep learning approach to identify YY1-mediated chromatin loops
abstract
The protein Yin Yang 1 (YY1) could form dimers that facilitate the interaction between active enhancers and promoter-proximal elements. YY1-mediated enhancer-promoter interaction is the general feature of mammalian gene control. Recently, some computational methods have been developed to characterize the interactions between DNA elements by elucidating important features of chromatin folding; however, no computational methods have been developed for identifying the YY1-mediated chromatin loops. In this study, we developed a deep learning algorithm named DeepYY1 based on word2vec to determine whether a pair of YY1 motifs would form a loop. The proposed models showed a high prediction performance (AUCs$\ge$0.93) on both training datasets and testing datasets in different cell types, demonstrating that DeepYY1 has an excellent performance in the identification of the YY1-mediated chromatin loops. Our study also suggested that sequences play an important role in the formation of YY1-mediated chromatin loops. Furthermore, we briefly discussed the distribution of the replication origin site in the loops. Finally, a user-friendly web server was established, and it can be freely accessed at http://lin-group.cn/server/DeepYY1.
Fu-Ying Dao, Hao Lv 0007, Dan Zhang 0006, Zi-Mei Zhang, Hao Lin 0001
Briefings Bioinform.6
2021 Deep-Kcr: accurate detection of lysine crotonylation sites using deep learning method
abstract
As a newly discovered protein posttranslational modification, histone lysine crotonylation (Kcr) involved in cellular regulation and human diseases. Various proteomics technologies have been developed to detect Kcr sites. However, experimental approaches for identifying Kcr sites are often time-consuming and labor-intensive, which is difficult to widely popularize in large-scale species. Computational approaches are cost-effective and can be used in a high-throughput manner to generate relatively precise identification. In this study, we develop a deep learning-based method termed as Deep-Kcr for Kcr sites prediction by combining sequence-based features, physicochemical property-based features and numerical space-derived information with information gain feature selection. We investigate the performances of convolutional neural network (CNN) and five commonly used classifiers (long short-term memory network, random forest, LogitBoost, naive Bayes and logistic regression) using 10-fold cross-validation and independent set test. Results show that CNN could always display the best performance with high computational efficiency on large dataset. We also compare the Deep-Kcr with other existing tools to demonstrate the excellent predictive power and robustness of our method. Based on the proposed model, a webserver called Deep-Kcr was established and is freely accessible at http://lin-group.cn/server/Deep-Kcr.
Hao Lv 0007, Fu-Ying Dao, Zheng-Xing Guan, Yan-Wen Li, Hao Lin 0001
Briefings Bioinform.6
2021 Design powerful predictor for mRNA subcellular location prediction in Homo sapiens
abstract
Messenger RNAs (mRNAs) shoulder special responsibilities that transmit genetic code from DNA to discrete locations in the cytoplasm. The locating process of mRNA might provide spatial and temporal regulation of mRNA and protein functions. The situ hybridization and quantitative transcriptomics analysis could provide detail information about mRNA subcellular localization; however, they are time consuming and expensive. It is highly desired to develop computational tools for timely and effectively predicting mRNA subcellular location. In this work, by using binomial distribution and one-way analysis of variance, the optimal nonamer composition was obtained to represent mRNA sequences. Subsequently, a predictor based on support vector machine was developed to identify the mRNA subcellular localization. In 5-fold cross-validation, results showed that the accuracy is 90.12% for Homo sapiens (H. sapiens). The predictor may provide a reference for the study of mRNA localization mechanisms and mRNA translocation strategies. An online web server was established based on our models, which is available at http://lin-group.cn/server/iLoc-mRNA/.
Zhao-Yue Zhang 0002, Hui Ding 0005, Dong Wang 0011, Wei Chen 0064, Hao Lin 0001
Briefings Bioinform.6
2021 iCarPS: a computational tool for identifying protein carbonylation sites by novel encoded features
abstract
MOTIVATION: Protein carbonylation is one of the most important oxidative stress-induced post-translational modifications, which is generally characterized as stability, irreversibility and relative early formation. It plays a significant role in orchestrating various biological processes and has been already demonstrated to be related to many diseases. However, the experimental technologies for carbonylation sites identification are not only costly and time consuming, but also unable of processing a large number of proteins at a time. Thus, rapidly and effectively identifying carbonylation sites by computational methods will provide key clues for the analysis of occurrence and development of diseases. RESULTS: In this study, we developed a predictor called iCarPS to identify carbonylation sites based on sequence information. A novel feature encoding scheme called residues conical coordinates combined with their physicochemical properties was proposed to formulate carbonylated protein and non-carbonylated protein samples. To remove potential redundant features and improve the prediction performance, a feature selection technique was used. The accuracy and robustness of iCarPS were proved by experiments on training and independent datasets. Comparison with other published methods demonstrated that the proposed method is powerful and could provide powerful performance for carbonylation sites identification. AVAILABILITY AND IMPLEMENTATION: Based on the proposed model, a user-friendly webserver and a software package were constructed, which can be freely accessed at http://lin-group.cn/server/iCarPS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dan Zhang 0006, Hao Lv 0007, Hao Lin 0001
Bioinform.7
2020 Evaluation of different computational methods on 5-methylcytosine sites identification
abstract
5-Methylcytosine (m5C) plays an extremely important role in the basic biochemical process. With the great increase of identified m5C sites in a wide variety of organisms, their epigenetic roles become largely unknown. Hence, accurate identification of m5C site is a key step in understanding its biological functions. Over the past several years, more attentions have been paid on the identification of m5C sites in multiple species. In this work, we firstly summarized the current progresses in computational prediction of m5C sites and then constructed a more powerful and reliable model for identifying m5C sites. To train the model, we collected experimentally confirmed m5C data from Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Arabidopsis thaliana, and compared the performances of different feature extraction methods and classification algorithms for optimizing prediction model. Based on the optimal model, a novel predictor called iRNA-m5C was developed for the recognition of m5C sites. Finally, we critically evaluated the performance of iRNA-m5C and compared it with existing methods. The result showed that iRNA-m5C could produce the best prediction performance. We hope that this paper could provide a guide on the computational identification of m5C site and also anticipate that the proposed iRNA-m5C will become a powerful tool for large scale identification of m5C sites.
Hao Lv 0007, Zi-Mei Zhang, Jiu-Xin Tan, Wei Chen 0064, Hao Lin 0001
Briefings Bioinform.6
2020 A comparison and assessment of computational method for identifying recombination hotspots in Saccharomyces cerevisiae
abstract
Meiotic recombination is one of the most important driving forces of biological evolution, which is initiated by double-strand DNA breaks. Recombination has important roles in genome diversity and evolution. This review firstly provides a comprehensive survey of the 15 computational methods developed for identifying recombination hotspots in Saccharomyces cerevisiae. These computational methods were discussed and compared in terms of underlying algorithms, extracted features, predictive capability and practical utility. Subsequently, a more objective benchmark data set was constructed to develop a new predictor iRSpot-Pse6NC2.0 (http://lin-group.cn/server/iRSpot-Pse6NC2.0). To further demonstrate the generalization ability of these methods, we compared iRSpot-Pse6NC2.0 with existing methods on the chromosome XVI of S. cerevisiae. The results of the independent data set test demonstrated that the new predictor is superior to existing tools in the identification of recombination hotspots. The iRSpot-Pse6NC2.0 will become an important tool for identifying recombination hotspot.
Wuritu Yang, Fu-Ying Dao, Hao Lv 0007, Hui Ding 0005, Wei Chen 0064, Hao Lin 0001
Briefings Bioinform.7
2020 DNA4mC-LIP: a linear integration method to identify N4-methylcytosine site in multiple species
abstract
MOTIVATION: DNA N4-methylcytosine (4mC) is a crucial epigenetic modification. However, the knowledge about its biological functions is limited. Effective and accurate identification of 4mC sites will be helpful to reveal its biological functions and mechanisms. Since experimental methods are cost and ineffective, a number of machine learning-based approaches have been proposed to detect 4mC sites. Although these methods yielded acceptable accuracy, there is still room for the improvement of the prediction performance and the stability of existing methods in practical applications. RESULTS: In this work, we first systematically assessed the existing methods based on an independent dataset. And then, we proposed DNA4mC-LIP, a linear integration method by combining existing predictors to identify 4mC sites in multiple species. The results obtained from independent dataset demonstrated that DNA4mC-LIP outperformed existing methods for identifying 4mC sites. To facilitate the scientific community, a web server for DNA4mC-LIP was developed. We anticipated that DNA4mC-LIP could serve as a powerful computational technique for identifying 4mC sites and facilitate the interpretation of 4mC mechanism. AVAILABILITY AND IMPLEMENTATION: http://i.uestc.edu.cn/DNA4mC-LIP/. CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qiang Tang 0013, Juanjuan Kang, Jiaqing Yuan, Hua Tang, Xianhai Li, Hao Lin 0001, Jian Huang 0004, Wei Chen 0064
Bioinform.6
2020 VisFeature: a stand-alone program for visualizing and analyzing statistical features of biological sequences
abstract
SUMMARY: Many efforts have been made in developing bioinformatics algorithms to predict functional attributes of genes and proteins from their primary sequences. One challenge in this process is to intuitively analyze and to understand the statistical features that have been selected by heuristic or iterative methods. In this paper, we developed VisFeature, which aims to be a helpful software tool that allows the users to intuitively visualize and analyze statistical features of all types of biological sequence, including DNA, RNA and proteins. VisFeature also integrates sequence data retrieval, multiple sequence alignments and statistical feature generation functions. AVAILABILITY AND IMPLEMENTATION: VisFeature is a desktop application that is implemented using JavaScript/Electron and R. The source codes of VisFeature are freely accessible from the GitHub repository (https://github.com/wangjun1996/VisFeature). The binary release, which includes an example dataset, can be freely downloaded from the same GitHub repository (https://github.com/wangjun1996/VisFeature/releases). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jun Wang 0081, Pu-Feng Du, Xin-Yu Xue, Guang-Ping Li 0005, Yuan-Ke Zhou, Hao Lin 0001, Wei Chen 0064
Bioinform.7
2019 i6mA-Pred: identifying DNA N6-methyladenine sites in the rice genome
abstract
MOTIVATION: DNA N6-methyladenine (6mA) is associated with a wide range of biological processes. Since the distribution of 6mA site in the genome is non-random, accurate identification of 6mA sites is crucial for understanding its biological functions. Although experimental methods have been proposed for this regard, they are still cost-ineffective for detecting 6mA site in genome-wide scope. Therefore, it is desirable to develop computational methods to facilitate the identification of 6mA site. RESULTS: In this study, a computational method called i6mA-Pred was developed to identify 6mA sites in the rice genome, in which the optimal nucleotide chemical properties obtained by the using feature selection technique were used to encode the DNA sequences. It was observed that the i6mA-Pred yielded an accuracy of 83.13% in the jackknife test. Meanwhile, the performance of i6mA-Pred was also superior to other methods. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server, i6mA-Pred is freely accessible at http://lin-group.cn/server/i6mA-Pred.
Wei Chen 0064, Hao Lv 0007, Fulei Nie, Hao Lin 0001
Bioinform.4
2019 Identify origin of replication in Saccharomyces cerevisiae using two-step feature selection technique
abstract
MOTIVATION: DNA replication is a key step to maintain the continuity of genetic information between parental generation and offspring. The initiation site of DNA replication, also called origin of replication (ORI), plays an extremely important role in the basic biochemical process. Thus, rapidly and effectively identifying the location of ORI in genome will provide key clues for genome analysis. Although biochemical experiments could provide detailed information for ORI, it requires high experimental cost and long experimental period. As good complements to experimental techniques, computational methods could overcome these disadvantages. RESULTS: Thus, in this study, we developed a predictor called iORI-PseKNC2.0 to identify ORIs in the Saccharomyces cerevisiae genome based on sequence information. The PseKNC including 90 physicochemical properties was proposed to formulate ORI and non-ORI samples. In order to improve the accuracy, a two-step feature selection was proposed to exclude redundant and noise information. As a result, the overall success rate of 88.53% was achieved in the 5-fold cross-validation test by using support vector machine. AVAILABILITY AND IMPLEMENTATION: Based on the proposed model, a user-friendly webserver was established and can be freely accessed at http://lin-group.cn/server/iORI-PseKNC2.0. The webserver will provide more convenience to most of wet-experimental scholars.
Fu-Ying Dao, Hao Lv 0007, Chao-Qin Feng, Hui Ding 0005, Wei Chen 0064, Hao Lin 0001
Bioinform.7
2019 iTerm-PseKNC: a sequence-based tool for predicting bacterial transcriptional terminators
abstract
MOTIVATION: Transcription termination is an important regulatory step of gene expression. If there is no terminator in gene, transcription could not stop, which will result in abnormal gene expression. Detecting such terminators can determine the operon structure in bacterial organisms and improve genome annotation. Thus, accurate identification of transcriptional terminators is essential and extremely important in the research of transcription regulations. RESULTS: In this study, we developed a new predictor called 'iTerm-PseKNC' based on support vector machine to identify transcription terminators. The binomial distribution approach was used to pick out the optimal feature subset derived from pseudo k-tuple nucleotide composition (PseKNC). The 5-fold cross-validation test results showed that our proposed method achieved an accuracy of 95%. To further evaluate the generalization ability of 'iTerm-PseKNC', the model was examined on independent datasets which are experimentally confirmed Rho-independent terminators in Escherichia coli and Bacillus subtilis genomes. As a result, all the terminators in E. coli and 87.5% of the terminators in B. subtilis were correctly identified, suggesting that the proposed model could become a powerful tool for bacterial terminator recognition. AVAILABILITY AND IMPLEMENTATION: For the convenience of most of wet-experimental researchers, the web-server for 'iTerm-PseKNC' was established at http://lin-group.cn/server/iTerm-PseKNC/, by which users can easily obtain their desired result without the need to go through the detailed mathematical equations involved.
Chao-Qin Feng, Zhao-Yue Zhang 0002, Xiao-Juan Zhu, Wei Chen 0064, Hua Tang, Hao Lin 0001
Bioinform.7
2019 iRNAD: a computational tool for identifying D modification sites in RNA sequence
abstract
MOTIVATION: Dihydrouridine (D) is a common RNA post-transcriptional modification found in eukaryotes, bacteria and a few archaea. The modification can promote the conformational flexibility of individual nucleotide bases. And its levels are increased in cancerous tissues. Therefore, it is necessary to detect D in RNA for further understanding its functional roles. Since wet-experimental techniques for the aim are time-consuming and laborious, it is urgent to develop computational models to identify D modification sites in RNA. RESULTS: We constructed a predictor, called iRNAD, for identifying D modification sites in RNA sequence. In this predictor, the RNA samples derived from five species were encoded by nucleotide chemical property and nucleotide density. Support vector machine was utilized to perform the classification. The final model could produce the overall accuracy of 96.18% with the area under the receiver operating characteristic curve of 0.9839 in jackknife cross-validation test. Furthermore, we performed a series of validations from several aspects and demonstrated the robustness and reliability of the proposed model. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server called iRNAD can be freely accessible at http://lin-group.cn/server/iRNAD, which will provide convenience and guide to users for further studying D modification.
Peng-Mian Feng, Wangren Qiu, Wei Chen 0064, Hao Lin 0001
Bioinform.6
2019 RIscoper: a tool for RNA-RNA interaction extraction from the literature
abstract
MOTIVATION: Numerous experimental and computational studies in the biomedical literature have provided considerable amounts of data on diverse RNA-RNA interactions (RRIs). However, few text mining systems for RRIs information extraction are available. RESULTS: RNA Interactome Scoper (RIscoper) represents the first tool for full-scale RNA interactome scanning and was developed for extracting RRIs from the literature based on the N-gram model. Notably, a reliable RRI corpus was integrated in RIscoper, and more than 13 300 manually curated sentences with RRI information were recruited. RIscoper allows users to upload full texts or abstracts, and provides an online search tool that is connected with PubMed (PMID and keyword input), and these capabilities are useful for biologists. RIscoper has a strong performance (90.4% precision and 93.9% recall), integrates natural language processing techniques and has a reliable RRI corpus. AVAILABILITY AND IMPLEMENTATION: The standalone software and web server of RIscoper are freely available at www.rna-society.org/riscoper/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Zhang 0125, Jinxurong Yang, Jiayi Yin, Yuncong Zhang, Zhixi Yun, Lin Ning 0002, Feng-Biao Guo, Yongshuai Jiang, Hao Lin 0001, Dong Wang 0011, Jian Huang 0004
Bioinform.12
2019 Predicting protein structural classes for low-similarity sequences by evaluating different features
Xiao-Juan Zhu, Chao-Qin Feng, Hong-Yan Lai, Wei Chen 0064, Hao Lin 0001
Knowl. Based Syst.5
2019 Identifying Sigma70 Promoters with Novel Pseudo Nucleotide Composition
abstract
Promoters are DNA regulatory elements located directly upstream or at the 5' end of the transcription initiation site (TSS), which are in charge of gene transcription initiation. With the completion of a large number of microorganism genomics, it is urgent to predict promoters accurately in bacteria by using the computational method. In this work, a sequence-based predictor named "iPro70-PseZNC" was designed for identifying sigma70 promoters in prokaryote. In the predictor, the samples of DNA sequences are formulated by a novel pseudo nucleotide composition, called PseZNC, into which the multi-window Z-curve composition and six local DNA structural properties are incorporated. In the 5-fold cross-validation, the area under the curve of receiver operating characteristic of 0.909 was obtained on our benchmark dataset, indicating that the proposed predictor is promising and will provide an important guide in this area. Further studies showed that the performance of PseZNC is better than it of multi-window Z-curve composition. For the sake of convenience for researchers, a user-friendly online service was established and can be freely accessible at http://lin.uestc.edu.cn/server/iPro70-PseZNC. The PseZNC approach can be also extended to other DNA-related problems.
Hao Lin 0001, Zhi-Yong Liang, Hua Tang, Wei Chen 0064
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNC
abstract
Motivation: Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. They have important functions in cell development and metabolism, such as genetic markers, genome rearrangements, chromatin modifications, cell cycle regulation, transcription and translation. Their functions are generally closely related to their localization in the cell. Therefore, knowledge about their subcellular locations can provide very useful clues or preliminary insight into their biological functions. Although biochemical experiments could determine the localization of lncRNAs in a cell, they are both time-consuming and expensive. Therefore, it is highly desirable to develop bioinformatics tools for fast and effective identification of their subcellular locations. Results: We developed a sequence-based bioinformatics tool called 'iLoc-lncRNA' to predict the subcellular locations of LncRNAs by incorporating the 8-tuple nucleotide features into the general PseKNC (Pseudo K-tuple Nucleotide Composition) via the binomial distribution approach. Rigorous jackknife tests have shown that the overall accuracy achieved by the new predictor on a stringent benchmark dataset is 86.72%, which is over 20% higher than that by the existing state-of-the-art predictor evaluated on the same tests. Availability and implementation: A user-friendly webserver has been established at http://lin-group.cn/server/iLoc-LncRNA, by which users can easily obtain their desired results. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhen-Dong Su 0001, Zhao-Yue Zhang 0002, Ya-Wei Zhao, Dong Wang 0011, Wei Chen 0064, Kuo-Chen Chou, Hao Lin 0001
Bioinform.8
2017 Identify and analysis crotonylation sites in histone by using support vector machines
Wangren Qiu, Bi-Qian Sun, Hua Tang, Jian Huang 0004, Hao Lin 0001
Artif. Intell. Medicine5
2017 iDNA4mC: identifying DNA N4-methylcytosine sites based on nucleotide chemical properties
abstract
MOTIVATION: DNA N4-methylcytosine (4mC) is an epigenetic modification. The knowledge about the distribution of 4mC is helpful for understanding its biological functions. Although experimental methods have been proposed to detect 4mC sites, they are expensive for performing genome-wide detections. Thus, it is necessary to develop computational methods for predicting 4mC sites. RESULTS: In this work, we developed iDNA4mC, the first webserver to identify 4mC sites, in which DNA sequences are encoded with both nucleotide chemical properties and nucleotide frequency. The predictive results of the rigorous jackknife test and cross species test demonstrated that the performance of iDNA4mC is quite promising and holds high potential to become a useful tool for identifying 4mC sites. AVAILABILITY AND IMPLEMENTATION: The user-friendly web-server, iDNA4mC, is freely accessible at http://lin.uestc.edu.cn/server/iDNA4mC. CONTACT: [email protected] or [email protected].
Wei Chen 0064, Peng-Mian Feng, Hui Ding 0005, Hao Lin 0001
Bioinform.5
2017 Pro54DB: a database for experimentally verified sigma-54 promoters
abstract
Summary: In prokaryotes, the σ54 promoters are unique regulatory elements and have attracted much attention because they are in charge of the transcription of carbon and nitrogen-related genes and participate in numerous ancillary processes and environmental responses. All findings on σ54 promoters are favorable for a better understanding of their regulatory mechanisms in gene transcription and an accurate discovery of genes missed by the wet experimental evidences. In order to provide an up-to-date, interactive and extensible database for σ54 promoter, a free and easy accessed database called Pro54DB (σ54 promoter database) was built to collect information of σ54 promoter. In the current version, it has stored 210 experimental-confirmed σ54 promoters with 297 regulated genes in 43 species manually extracted from 133 publications, which is helpful for researchers in fields of bioinformatics and molecular biology. Availability and Implementation: Pro54DB is freely available on the web at http://lin.uestc.edu.cn/database/pro54db with all major browsers supported. Contacts: [email protected] or [email protected]
Zhi-Yong Liang, Hong-Yan Lai, Huan-Huan Wei, Xin-Xin Chen, Ya-Wei Zhao, Zhen-Dong Su 0001, En-Ze Deng, Hua Tang, Wei Chen 0064, Hao Lin 0001
Bioinform.14
2015 PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositions
abstract
SUMMARY: The avalanche of genomic sequences generated in the post-genomic age requires efficient computational methods for rapidly and accurately identifying biological features from sequence information. Towards this goal, we developed a freely available and open-source package, called PseKNC-General (the general form of pseudo k-tuple nucleotide composition), that allows for fast and accurate computation of all the widely used nucleotide structural and physicochemical properties of both DNA and RNA sequences. PseKNC-General can generate several modes of pseudo nucleotide compositions, including conventional k-tuple nucleotide compositions, Moreau-Broto autocorrelation coefficient, Moran autocorrelation coefficient, Geary autocorrelation coefficient, Type I PseKNC and Type II PseKNC. In every mode, >100 physicochemical properties are available for choosing. Moreover, it is flexible enough to allow the users to calculate PseKNC with user-defined properties. The package can be run on Linux, Mac and Windows systems and also provides a graphical user interface. AVAILABILITY AND IMPLEMENTATION: The package is freely available at: http://lin.uestc.edu.cn/server/pseknc.
Wei Chen 0064, Xitong Zhang, Jordan Brooker, Hao Lin 0001, Liqing Zhang 0002, Kuo-Chen Chou
Bioinform.4
2014 iNuc-PseKNC: a sequence-based predictor for predicting nucleosome positioning in genomes with pseudo k-tuple nucleotide composition
abstract
MOTIVATION: Nucleosome positioning participates in many cellular activities and plays significant roles in regulating cellular processes. With the avalanche of genome sequences generated in the post-genomic age, it is highly desired to develop automated methods for rapidly and effectively identifying nucleosome positioning. Although some computational methods were proposed, most of them were species specific and neglected the intrinsic local structural properties that might play important roles in determining the nucleosome positioning on a DNA sequence. RESULTS: Here a predictor called 'iNuc-PseKNC' was developed for predicting nucleosome positioning in Homo sapiens, Caenorhabditis elegans and Drosophila melanogaster genomes, respectively. In the new predictor, the samples of DNA sequences were formulated by a novel feature-vector called 'pseudo k-tuple nucleotide composition', into which six DNA local structural properties were incorporated. It was observed by the rigorous cross-validation tests on the three stringent benchmark datasets that the overall success rates achieved by iNuc-PseKNC in predicting the nucleosome positioning of the aforementioned three genomes were 86.27%, 86.90% and 79.97%, respectively. Meanwhile, the results obtained by iNuc-PseKNC on various benchmark datasets used by the previous investigators for different genomes also indicated that the current predictor remarkably outperformed its counterparts. AVAILABILITY: A user-friendly web-server, iNuc-PseKNC is freely accessible at http://lin.uestc.edu.cn/server/iNuc-PseKNC.
Shou-Hui Guo, En-Ze Deng, Liqin Xu, Hui Ding 0005, Hao Lin 0001, Wei Chen 0064, Kuo-Chen Chou
Bioinform.5
2014 Training sparse SVM on the core sets of fitting-planes
Mingtian Zhou, Lijia Xu, Hao Lin 0001, Haibo Pu
Neurocomputing4