Yan Liu 0038

dblp:150/4295-38 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-5331-3655ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Dualmark: A novel dual watermarking approach for large language models
Zihao Qiang, Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Chaowei Zhang 0001, Yan Liu 0038, Wei Li 0121
Inf. Process. Manag.6
2026 CellPredX, a computational framework for cross-data type, cross-sample, and cross-protocol cell type annotation through domain adaptation and deep metric learning
abstract
Accurate cell type annotation is fundamental to single-cell analysis, yet remains challenging across heterogeneous datasets and modalities. In particular, transferring labels between scRNA-seq and scATAC-seq data poses unique difficulties due to discrepancies in sequencing protocols and feature spaces. Existing methods typically handle only a subset of these challenges, often requiring scenario-specific adjustments and offering limited interpretability. Here, we present CellPredX, a structurally unified but adaptively parameterized, semi-supervised cross-modality framework for label transfer across scRNA-seq, scATAC-seq, and cross-protocol datasets. While maintaining a unified model architecture and optimization strategy, CellPredX allows adaptive tuning of loss-weight hyperparameters to account for the varying degree of similarity or discrepancy between different reference-query dataset pairs. CellPredX integrates domain adaptation and deep metric learning to align heterogeneous embeddings, and introduces a sparse center loss with an attention mechanism to enhance discriminative representations while suppressing noise. Moreover, an integrated interpreter module based on gradient attribution enables biological interpretability by identifying key markers and feature dimensions driving model predictions. Through extensive benchmarking across scRNA to scATAC, scATAC to scATAC, and scRNA to scRNA transfers, CellPredX consistently outperforms state-of-the-art annotation methods in both accuracy and robustness. The interpreter module further reveals biologically meaningful marker patterns that are consistent with known cell hierarchies. Together, these results demonstrate that CellPredX provides an interpretable and scalable solution for cross-modality cell type annotation in single-cell multi-omic integration.
Yan Liu 0038, Long-Chen Shen, Jipeng Qiang
PLoS Comput. Biol.1
2025 Supervised contrastive learning enhances MHC-II peptide binding affinity prediction
Long-Chen Shen, Yan Liu 0038, Zi Liu, Zhikang Wang, Yuming Guo 0001, Jamie Rossjohn, Jiangning Song, Dongjun Yu
Expert Syst. Appl.2
2025 MUSIC-GCN: A Novel Multi-Tasking Pipeline for Analyzing Single-Cell Transcriptomic Data Using Residual Graph Convolution Network
abstract
Single-cell transcriptomics is a powerful approach for characterizing gene transcription at cellular resolution. This approach requires efficient computational pipelines to undertake essential tasks, including clustering, dimensionality reduction, imputation, and denoising. Currently, most such pipelines undertake these computational tasks separately without considering the interdependence among these tasks. Here, we present an advanced pipeline, MUSIC-GCN, by employing a graph convolutional neural (GCN) network and autoencoder to perform multi-task single-cell RNA-sequencing (scRNA-seq) data analysis. The rationale is that multiple related tasks can be carried out simultaneously to enable enhanced learning and more effective representations through the 'sharing of knowledge' regarding individual tasks. Benchmarking experiments using various scRNA-seq datasets show that MUSIC-GCN can achieve a competitive performance on multi-tasks when benchmarked with state-of-the-art approaches.
Yan Liu 0038, Chen Li 0021, Long-Chen Shen, Robin B. Gasser, Jiangning Song, Dijun Chen, Dongjun Yu
IEEE Trans. Comput. Biol. Bioinform.1
2025 Integrating Graph Convolutional Networks for Missing Gene Expression Imputation
abstract
Single-cell RNA sequencing (scRNA-seq) techniques are emerging to revolutionize modern biomedical sciences by providing a detailed landscape of individual cells. However, these methods often lack crucial spatial localization information. To address this gap, spatial transcriptomic technologies have developed, enabling gene expression profiling while mapping cells spatial information. Yet, the gene throughput in spatial transcriptomic technologies makes it challenging to characterize whole-transcriptome-level data for single cells in space. In this context, approaches for predicting the spatial distribution of genes are still under development. Here, we present GCNgene, a novel method to predict the spatial distribution of the undetected RNA transcripts, through integrating spatial and scRNA-seq datasets. GCNgene leverages a graph convolutional network to embed spatial transcriptomics data and then applies a learned rule to reconstruct gene expression by combining the reference single-cell data with the calculated cell-type proportions. Ultimately, this learned paradigm enables accurate predictions of gene expression levels.
Ying Zhang 0053, Hong-Jin Yu, Zihao Yan, Tong Pan, Yan Liu 0038, Shanshan Li 0008, Yuming Guo 0001, Jiangning Song, Dongjun Yu
IEEE Trans. Comput. Biol. Bioinform.6
2025 Identification of Protein-Nucleotide Binding Residues With Deep Multi-Task and Multi-Scale Learning
abstract
Accurate identification of protein-nucleotide binding residues is essential for protein functional annotation and drug discovery. Advancements in computational methods for predicting binding residues from protein sequences have significantly improved predictive accuracy. However, it remains a challenge for current methodologies to extract discriminative features and assimilate heterogeneous data from different nucleotide binding residues. To address this, we introduce NucMoMTL, a novel predictor specifically designed for identifying protein-nucleotide binding residues. Specifically, NucMoMTL leverages a pre-trained language model for robust sequence embedding and utilizes deep multi-task and multi-scale learning within parameter-based orthogonal constraints to extract shared representations, capitalizing on auxiliary information from diverse nucleotides binding residues. Evaluation of NucMoMTL on the benchmark datasets demonstrates that it outperforms state-of-the-art methods, achieving an average AUROC and AUPRC of 0.961 and 0.566, respectively. NucMoMTL can be explored as a reliable computational tool for identifying protein-nucleotide binding residues and facilitating drug discovery.
Fang Ge, Shanruo Xu, Yan Liu 0038, Jiangning Song, Dongjun Yu
IEEE J. Biomed. Health Informatics4
2024 GMFGRN: a matrix factorization and graph neural network approach for gene regulatory network inference
abstract
The recent advances of single-cell RNA sequencing (scRNA-seq) have enabled reliable profiling of gene expression at the single-cell level, providing opportunities for accurate inference of gene regulatory networks (GRNs) on scRNA-seq data. Most methods for inferring GRNs suffer from the inability to eliminate transitive interactions or necessitate expensive computational resources. To address these, we present a novel method, termed GMFGRN, for accurate graph neural network (GNN)-based GRN inference from scRNA-seq data. GMFGRN employs GNN for matrix factorization and learns representative embeddings for genes. For transcription factor-gene pairs, it utilizes the learned embeddings to determine whether they interact with each other. The extensive suite of benchmarking experiments encompassing eight static scRNA-seq datasets alongside several state-of-the-art methods demonstrated mean improvements of 1.9 and 2.5% over the runner-up in area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). In addition, across four time-series datasets, maximum enhancements of 2.4 and 1.3% in AUROC and AUPRC were observed in comparison to the runner-up. Moreover, GMFGRN requires significantly less training time and memory consumption, with time and memory consumed <10% compared to the second-best method. These findings underscore the substantial potential of GMFGRN in the inference of GRNs. It is publicly available at https://github.com/Lishuoyy/GMFGRN.
Yan Liu 0038, Long-Chen Shen, Jiangning Song, Dongjun Yu
Briefings Bioinform.2
2024 ULDNA: integrating unsupervised multi-source language models with LSTM-attention network for high-accuracy protein-DNA binding site prediction
abstract
Efficient and accurate recognition of protein-DNA interactions is vital for understanding the molecular mechanisms of related biological processes and further guiding drug discovery. Although the current experimental protocols are the most precise way to determine protein-DNA binding sites, they tend to be labor-intensive and time-consuming. There is an immediate need to design efficient computational approaches for predicting DNA-binding sites. Here, we proposed ULDNA, a new deep-learning model, to deduce DNA-binding sites from protein sequences. This model leverages an LSTM-attention architecture, embedded with three unsupervised language models that are pre-trained on large-scale sequences from multiple database sources. To prove its effectiveness, ULDNA was tested on 229 protein chains with experimental annotation of DNA-binding sites. Results from computational experiments revealed that ULDNA significantly improves the accuracy of DNA-binding site prediction in comparison with 17 state-of-the-art methods. In-depth data analyses showed that the major strength of ULDNA stems from employing three transformer language models. Specifically, these language models capture complementary feature embeddings with evolution diversity, in which the complex DNA-binding patterns are buried. Meanwhile, the specially crafted LSTM-attention network effectively decodes evolution diversity-based embeddings as DNA-binding results at the residue level. Our findings demonstrated a new pipeline for predicting DNA-binding sites on a large scale with high accuracy from protein sequence alone.
Yiheng Zhu 0001, Zi Liu, Yan Liu 0038, Zhiwei Ji, Dongjun Yu
Briefings Bioinform.3
2024 CTISL: a dynamic stacking multi-class classification approach for identifying cell types from single-cell RNA-seq data
abstract
MOTIVATION: Effective identification of cell types is of critical importance in single-cell RNA-sequencing (scRNA-seq) data analysis. To date, many supervised machine learning-based predictors have been implemented to identify cell types from scRNA-seq datasets. Despite the technical advances of these state-of-the-art tools, most existing predictors were single classifiers, of which the performances can still be significantly improved. It is therefore highly desirable to employ the ensemble learning strategy to develop more accurate computational models for robust and comprehensive identification of cell types on scRNA-seq datasets. RESULTS: We propose a two-layer stacking model, termed CTISL (Cell Type Identification by Stacking ensemble Learning), which integrates multiple classifiers to identify cell types. In the first layer, given a reference scRNA-seq dataset with known cell types, CTISL dynamically combines multiple cell-type-specific classifiers (i.e. support-vector machine and logistic regression) as the base learners to deliver the outcomes for the input of a meta-classifier in the second layer. We conducted a total of 24 benchmarking experiments on 17 human and mouse scRNA-seq datasets to evaluate and compare the prediction performance of CTISL and other state-of-the-art predictors. The experiment results demonstrate that CTISL achieves superior or competitive performance compared to these state-of-the-art approaches. We anticipate that CTISL can serve as a useful and reliable tool for cost-effective identification of cell types from scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: The webserver and source code are freely available at http://bigdata.biocie.cn/CTISLweb/home and https://zenodo.org/records/10568906, respectively.
Ziyi Chai, Yan Liu 0038, Chen Li 0021, Yu Jiang 0014, Quanzhong Liu
Bioinform.4
2024 Robust GEPSVM classifier: An efficient iterative optimization framework
Yan Liu 0038, Yanmeng Li, Qiaolin Ye, Dongjun Yu, Yong Qi 0002
Inf. Sci.2
2024 Improving Antifreeze Proteins Prediction With Protein Language Models and Hybrid Feature Extraction Networks
abstract
Accurate identification of antifreeze proteins (AFPs) is crucial in developing biomimetic synthetic anti-icing materials and low-temperature organ preservation materials. Although numerous machine learning-based methods have been proposed for AFPs prediction, the complex and diverse nature of AFPs limits the prediction performance of existing methods. In this study, we propose AFP-Deep, a new deep learning method to predict antifreeze proteins by integrating embedding from protein sequences with pre-trained protein language models and evolutionary contexts with hybrid feature extraction networks. The experimental results demonstrated that the main advantage of AFP-Deep is its utilization of pre-trained protein language models, which can extract discriminative global contextual features from protein sequences. Additionally, the hybrid deep neural networks designed for protein language models and evolutionary context feature extraction enhance the correlation between embeddings and antifreeze pattern. The performance evaluation results show that AFP-Deep achieves superior performance compared to state-of-the-art models on benchmark datasets, achieving an AUPRC of 0.724 and 0.924, respectively.
Yan Liu 0038, Yiheng Zhu 0001, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 TripletCell: a deep metric learning framework for accurate annotation of cell types at the single-cell level
abstract
Single-cell RNA sequencing (scRNA-seq) has significantly accelerated the experimental characterization of distinct cell lineages and types in complex tissues and organisms. Cell-type annotation is of great importance in most of the scRNA-seq analysis pipelines. However, manual cell-type annotation heavily relies on the quality of scRNA-seq data and marker genes, and therefore can be laborious and time-consuming. Furthermore, the heterogeneity of scRNA-seq datasets poses another challenge for accurate cell-type annotation, such as the batch effect induced by different scRNA-seq protocols and samples. To overcome these limitations, here we propose a novel pipeline, termed TripletCell, for cross-species, cross-protocol and cross-sample cell-type annotation. We developed a cell embedding and dimension-reduction module for the feature extraction (FE) in TripletCell, namely TripletCell-FE, to leverage the deep metric learning-based algorithm for the relationships between the reference gene expression matrix and the query cells. Our experimental studies on 21 datasets (covering nine scRNA-seq protocols, two species and three tissues) demonstrate that TripletCell outperformed state-of-the-art approaches for cell-type annotation. More importantly, regardless of protocols or species, TripletCell can deliver outstanding and robust performance in annotating different types of cells. TripletCell is freely available at https://github.com/liuyan3056/TripletCell. We believe that TripletCell is a reliable computational tool for accurately annotating various cell types using scRNA-seq data and will be instrumental in assisting the generation of novel biological hypotheses in cell biology.
Yan Liu 0038, Chen Li 0021, Long-Chen Shen, Robin B. Gasser, Jiangning Song, Dijun Chen, Dongjun Yu
Briefings Bioinform.1
2021 Leveraging the attention mechanism to improve the identification of DNA N6-methyladenine sites
abstract
DNA N6-methyladenine is an important type of DNA modification that plays important roles in multiple biological processes. Despite the recent progress in developing DNA 6mA site prediction methods, several challenges remain to be addressed. For example, although the hand-crafted features are interpretable, they contain redundant information that may bias the model training and have a negative impact on the trained model. Furthermore, although deep learning (DL)-based models can perform feature extraction and classification automatically, they lack the interpretability of the crucial features learned by those models. As such, considerable research efforts have been focused on achieving the trade-off between the interpretability and straightforwardness of DL neural networks. In this study, we develop two new DL-based models for improving the prediction of N6-methyladenine sites, termed LA6mA and AL6mA, which use bidirectional long short-term memory to respectively capture the long-range information and self-attention mechanism to extract the key position information from DNA sequences. The performance of the two proposed methods is benchmarked and evaluated on the two model organisms Arabidopsis thaliana and Drosophila melanogaster. On the two benchmark datasets, LA6mA achieves an area under the receiver operating characteristic curve (AUROC) value of 0.962 and 0.966, whereas AL6mA achieves an AUROC value of 0.945 and 0.941, respectively. Moreover, an in-depth analysis of the attention matrix is conducted to interpret the important information, which is hidden in the sequence and relevant for 6mA site prediction. The two novel pipelines developed for DNA 6mA site prediction in this work will facilitate a better understanding of the underlying principle of DL-based DNA methylation site prediction and its future applications.
Ying Zhang 0053, Yan Liu 0038, Jian Xu 0009, Xiaoyu Wang 0016, Xinxin Peng, Jiangning Song, Dongjun Yu
Briefings Bioinform.2
2021 Improving protein fold recognition using triplet network and ensemble deep learning
abstract
Protein fold recognition is a critical step toward protein structure and function prediction, aiming at providing the most likely fold type of the query protein. In recent years, the development of deep learning (DL) technique has led to massive advances in this important field, and accordingly, the sensitivity of protein fold recognition has been dramatically improved. Most DL-based methods take an intermediate bottleneck layer as the feature representation of proteins with new fold types. However, this strategy is indirect, inefficient and conditional on the hypothesis that the bottleneck layer's representation is assumed as a good representation of proteins with new fold types. To address the above problem, in this work, we develop a new computational framework by combining triplet network and ensemble DL. We first train a DL-based model, termed FoldNet, which employs triplet loss to train the deep convolutional network. FoldNet directly optimizes the protein fold embedding itself, making the proteins with the same fold types be closer to each other than those with different fold types in the new protein embedding space. Subsequently, using the trained FoldNet, we implement a new residue-residue contact-assisted predictor, termed FoldTR, which improves protein fold recognition. Furthermore, we propose a new ensemble DL method, termed FSD_XGBoost, which combines protein fold embedding with the other two discriminative fold-specific features extracted by two DL-based methods SSAfold and DeepFR. The Top 1 sensitivity of FSD_XGBoost increases to 74.8% at the fold level, which is ~9% higher than that of the state-of-the-art method. Together, the results suggest that fold-specific features extracted by different DL methods complement with each other, and their combination can further improve fold recognition at the fold level. The implemented web server of FoldTR and benchmark datasets are publicly available at http://csbio.njust.edu.cn/bioinf/foldtr/.
Yan Liu 0038, Yiheng Zhu 0001, Ying Zhang 0053, Long-Chen Shen, Jiangning Song, Dongjun Yu
Briefings Bioinform.1
2021 Why can deep convolutional neural networks improve protein fold recognition? A visual explanation by interpretation
abstract
As an essential task in protein structure and function prediction, protein fold recognition has attracted increasing attention. The majority of the existing machine learning-based protein fold recognition approaches strongly rely on handcrafted features, which depict the characteristics of different protein folds; however, effective feature extraction methods still represent the bottleneck for further performance improvement of protein fold recognition. As a powerful feature extractor, deep convolutional neural network (DCNN) can automatically extract discriminative features for fold recognition without human intervention, which has demonstrated an impressive performance on protein fold recognition. Despite the encouraging progress, DCNN often acts as a black box, and as such, it is challenging for users to understand what really happens in DCNN and why it works well for protein fold recognition. In this study, we explore the intrinsic mechanism of DCNN and explain why it works for protein fold recognition using a visual explanation technique. More specifically, we first trained a VGGNet-based DCNN model, termed VGGNet-FE, which can extract fold-specific features from the predicted protein residue-residue contact map for protein fold recognition. Subsequently, based on the trained VGGNet-FE, we implemented a new contact-assisted predictor, termed VGGfold, for protein fold recognition; we then visualized what features were extracted by each of the convolutional layers in VGGNet-FE using a deconvolution technique. Furthermore, we visualized the high-level semantic information, termed fold-discriminative region, of a predicted contact map from the localization map obtained from the last convolutional layer of VGGNet-FE. It is visually confirmed that VGGNet-FE could effectively extract distinct fold-discriminative regions for different types of protein folds, thereby accounting for the improved performance of VGGfold for protein fold recognition. In summary, this study is of great significance for both understanding the working principle of DCNNs in protein fold recognition and exploring the relationship between the predicted protein contact map and protein tertiary structure. This proposed visualization method is flexible and applicable to address other DCNN-based bioinformatics and computational biology questions. The online web server of VGGfold is freely available at http://csbio.njust.edu.cn/bioinf/vggfold/.
Yan Liu 0038, Yiheng Zhu 0001, Xiaoning Song, Jiangning Song, Dongjun Yu
Briefings Bioinform.1
2021 SAResNet: self-attention residual network for predicting DNA-protein binding
abstract
Knowledge of the specificity of DNA-protein binding is crucial for understanding the mechanisms of gene expression, regulation and gene therapy. In recent years, deep-learning-based methods for predicting DNA-protein binding from sequence data have achieved significant success. Nevertheless, the current state-of-the-art computational methods have some drawbacks associated with the use of limited datasets with insufficient experimental data. To address this, we propose a novel transfer learning-based method, termed SAResNet, which combines the self-attention mechanism and residual network structure. More specifically, the attention-driven module captures the position information of the sequence, while the residual network structure guarantees that the high-level features of the binding site can be extracted. Meanwhile, the pre-training strategy used by SAResNet improves the learning ability of the network and accelerates the convergence speed of the network during transfer learning. The performance of SAResNet is extensively tested on 690 datasets from the ChIP-seq experiments with an average AUC of 92.0%, which is 4.4% higher than that of the best state-of-the-art method currently available. When tested on smaller datasets, the predictive performance is more clearly improved. Overall, we demonstrate that the superior performance of DNA-protein binding prediction on DNA sequences can be achieved by combining the attention mechanism and residual structure, and a novel pipeline is accordingly developed. The proposed methodology is generally applicable and can be used to address any other sequence classification problems.
Long-Chen Shen, Yan Liu 0038, Jiangning Song, Dongjun Yu
Briefings Bioinform.2
2018 A Complete Canonical Correlation Analysis for Multiview Learning
abstract
Canonical correlation analysis (CCA) is an effective feature learning method, which has wide applications in pattern recognition and computer vision. However, CCA considers the correlation only between the one-to-one aligned samples in two views, ignoring the correlation between all the samples sharing the same label. In this paper, we propose a deep complete canonical correlation analysis (Deep Complete-CCA), which learns the relationships between all pairwise correspondences of sample points in the same classes. Unlike CCA, our method can learn discriminant representations that maximize the correlation between the two views while segregating the different classes on the learned space. We test Deep Complete-CCA on handwriting recognition and speech based emotion recognition using two popular MNIST and RAVDESS datasets. Experimental results show that our proposed method can obtain better performances than several related algorithms.
Yan Liu 0038, Yun Li 0010, Yun-Hao Yuan 0001
ICIP1
2017 Supervised Deep Canonical Correlation Analysis for Multiview Feature Learning
Yan Liu 0038, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Min Ruan, Zhao Zhang 0018
ICONIP (6)1