Fei-Liao Lai

dblp:291/1794 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-7328-1585ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Nmix: a hybrid deep learning model for precise prediction of 2'-O-methylation sites based on multi-feature fusion and ensemble learning
abstract
RNA 2'-O-methylation (Nm) is a crucial post-transcriptional modification with significant biological implications. However, experimental identification of Nm sites is challenging and resource-intensive. While multiple computational tools have been developed to identify Nm sites, their predictive performance, particularly in terms of precision and generalization capability, remains deficient. We introduced Nmix, an advanced computational tool for precise prediction of Nm sites in human RNA. We constructed the largest, low-redundancy dataset of experimentally verified Nm sites and employed an innovative multi-feature fusion approach, combining one-hot, Z-curve and RNA secondary structure encoding. Nmix utilizes a meticulously designed hybrid deep learning architecture, integrating 1D/2D convolutional neural networks, self-attention mechanism and residual connection. We implemented asymmetric loss function and Bayesian optimization-based ensemble learning, substantially improving predictive performance on imbalanced datasets. Rigorous testing on two benchmark datasets revealed that Nmix significantly outperforms existing state-of-the-art methods across various metrics, particularly in precision, with average improvements of 33.1% and 60.0%, and Matthews correlation coefficient, with average improvements of 24.7% and 51.1%. Notably, Nmix demonstrated exceptional cross-species generalization capability, accurately predicting 93.8% of experimentally verified Nm sites in rat RNA. We also developed a user-friendly web server (https://tubic.org/Nm) and provided standalone prediction scripts to facilitate widespread adoption. We hope that by providing a more accurate and robust tool for Nm site prediction, we can contribute to advancing our understanding of Nm mechanisms and potentially benefit the prediction of other RNA modification sites.
Yu-Qing Geng, Fei-Liao Lai, Hao Luo 0002, Feng Gao 0001
Briefings Bioinform.2
2024 Unveiling human origins of replication using deep learning: accurate prediction and comprehensive analysis
abstract
Accurate identification of replication origins (ORIs) is crucial for a comprehensive investigation into the progression of human cell growth and cancer therapy. Here, we proposed a computational approach Ori-FinderH, which can efficiently and precisely predict the human ORIs of various lengths by combining the Z-curve method with deep learning approach. Compared with existing methods, Ori-FinderH exhibits superior performance, achieving an area under the receiver operating characteristic curve (AUC) of 0.9616 for K562 cell line in 10-fold cross-validation. In addition, we also established a cross-cell-line predictive model, which yielded a further improved AUC of 0.9706. The model was subsequently employed as a fitness function to support genetic algorithm for generating artificial ORIs. Sequence analysis through iORI-Euk revealed that a vast majority of the created sequences, specifically 98% or more, incorporate at least one ORI for three cell lines (Hela, MCF7 and K562). This innovative approach could provide more efficient, accurate and comprehensive information for experimental investigation, thereby further advancing the development of this field.
Zhen-Ning Yin, Fei-Liao Lai, Feng Gao 0001
Briefings Bioinform.2
2023 Auto-Kla: a novel web server to discriminate lysine lactylation sites using automated machine learning
abstract
Recently, lysine lactylation (Kla), a novel post-translational modification (PTM), which can be stimulated by lactate, has been found to regulate gene expression and life activities. Therefore, it is imperative to accurately identify Kla sites. Currently, mass spectrometry is the fundamental method for identifying PTM sites. However, it is expensive and time-consuming to achieve this through experiments alone. Herein, we proposed a novel computational model, Auto-Kla, to quickly and accurately predict Kla sites in gastric cancer cells based on automated machine learning (AutoML). With stable and reliable performance, our model outperforms the recently published model in the 10-fold cross-validation. To investigate the generalizability and transferability of our approach, we evaluated the performance of our models trained on two other widely studied types of PTM, including phosphorylation sites in host cells infected with SARS-CoV-2 and lysine crotonylation sites in HeLa cells. The results show that our models achieve comparable or better performance than current outstanding models. We believe that this method will become a useful analytical tool for PTM prediction and provide a reference for the future development of related models. The web server and source code are available at http://tubic.org/Kla and https://github.com/tubic/Auto-Kla, respectively.
Fei-Liao Lai, Feng Gao 0001
Briefings Bioinform.1
2022 GC-Profile 2.0: an extended web server for the prediction and visualization of CpG islands
abstract
SUMMARY: Due to the spontaneous deamination of 5'-methylcytosine into thymine, the number of CpG dinucleotides is less than expected in vertebrate genomes. Exceptionally, there are a large number of CpG dinucleotides clustered at certain genomic loci, known as CpG islands (CGIs), where CpG dinucleotides are free from methylation. Identification of CGIs is of great significance in the field of genomics and epigenetics because they can serve as important gene markers or regulatory elements. Here, GC-Profile 2.0 has been presented as a newly extended application for CGIs detection and visualization. Based on a benchmark test of assembled sequences, GC-Profile 2.0 has shown better overall performance compared with other four popular methods. In addition, cumulative CpG profile, a visualization tool of CpG content variation, is also proposed to intuitively display the change trend of CpG content. AVAILABILITY AND IMPLEMENTATION: GC-Profile 2.0 is freely available at http://tubic.org/GC-Profile2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fei-Liao Lai, Feng Gao 0001
Bioinform.1
2021 Ori-Finder 3: a web server for genome-wide prediction of replication origins in Saccharomyces cerevisiae
abstract
DNA replication is a fundamental process in all organisms; this event initiates at sites termed origins of replication. The characteristics of eukaryotic replication origins are best understood in Saccharomyces cerevisiae. For this species, origin prediction algorithms or web servers have been developed based on the sequence features of autonomously replicating sequences (ARSs). However, their performances are far from satisfactory. By utilizing the Z-curve methodology, we present a novel pipeline, Ori-Finder 3, for the computational prediction of replication origins in S. cerevisiae at the genome-wide level based solely on DNA sequences. The ARS exhibiting both an AT-rich stretch and ARS consensus sequence element can be predicted at the single-nucleotide level. For the identified ARSs in the S. cerevisiae reference genome, 83 and 60% of the top 100 and top 300 predictions matched the known ARS records, respectively. Based on Ori-Finder 3, we subsequently built a database of the predicted ARSs identified in more than a hundred S. cerevisiae genomes. Consequently, we developed a user-friendly web server including the ARS prediction pipeline and the predicted ARSs database, which can be freely accessed at http://tubic.tju.edu.cn/Ori-Finder3.
Dan Wang 0028, Fei-Liao Lai, Feng Gao 0001
Briefings Bioinform.2