Quanzhong Liu

dblp:96/6112 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 scMoSAT: A Semi-supervised Transformer Framework for Cross-Modal Cell Type Annotation
Miaomiao Xie, Quanzhong Liu
ICIC (6)4
2026 AGEP_TWAS: A Deep Learning-Based Framework for Predicting Gene Expression Levels in Tissues
abstract
Accurate prediction of gene expression levels across different tissues is of great significance in understanding the functional roles of genes in various biological processes and assisting in transcriptome-wide association studies (TWAS). Traditional methods rely on the construction of a prediction model for each gene in a specific tissue, which is time-consuming and inefficient when dealing with numerous genes or tissues. In addition, current approaches do not consider missing single nucleotide polymorphisms (SNPs) in the training population. These SNPs significantly affect gene expression levels and in turn limit the predictive capability of these approaches in new samples. Recent research indicates that by identifying a specific group of core genes (known as landmark genes) that accurately reflect the cellular states of samples across different experimental conditions, it is possible to predict the expression levels of other genes in the genome. In light of this, we propose AGEP_TWAS (Adaptive Gene Expression Predictor for TWAS), a gene expression prediction method that utilizes a dense connection network, adaptive activation functions, and parameter pruning strategies within a nonlinear feature extraction computational framework. AGEP_TWAS leverages landmark genes within a tissue to predict the expression levels of other genes that are challenging to predict using traditional methods. Results on the human GEO expression dataset demonstrate that AGEP_TWAS achieved a mean squared error (MSE) of 0.1821 and a Pearson correlation coefficient (PCC) of 0.9004, outperforming existing state-of-the-art prediction models. Additionally, when applied to the CattleGTEx dataset to infer gene expression levels across different tissues in cattle, AGEP_TWAS exhibited superior predictive performance compared to existing methods. A TWAS on milk production traits in cattle highlights the practical utility of AGEP_TWAS, with six identified significant genes already reported in scientific literature to be associated with milk production traits.
Chen Li 0021, Yudong Cai 0003, Zhuangbiao Zhang, Zhannur Niyazbekova, Yu Jiang 0014, Quanzhong Liu
IEEE Trans. Comput. Biol. Bioinform.10
2025 SCANER: robust and sensitive identification of malignant cells from the scRNA-seq profiled tumor ecosystem
abstract
Single-cell RNA sequencing (scRNA-seq) has enabled the dissection of complex tumor ecosystems. Recognition of malignant cells as an essential step has a profound impact on downstream interpretation. However, most existing computational strategies are based on prior knowledge of canonical cell-type markers. We have developed a marker-free approach, the Seed-Cluster based Approach for NEoplastic cells Recognition (SCANER), to identify malignant cells based on significant gene expression variations caused by genomic instability. Upon analyzing different cancer types, SCANER achieved superior accuracy and robustness in identifying malignant cells, effectively addressing dropout events and tumor purity variations. Besides, SCANER can significantly detect copy number variations (CNVs) in malignant cells compared to nonmalignant cells, which is further confirmed through the paired whole exome sequencing data. In conclusion, SCANER has the potential to facilitate the biological exploration of the tumor ecosystem by accurately identifying malignant cells and it is applicable across various solid cancer types regardless of prior knowledge. SCANER is available at https://github.com/woolingxiang/SCANER.
Quanzhong Liu, Mengyan Zhu, Miao Yu 0010, Kening Li, Lingxiang Wu, Qianghu Wang
Briefings Bioinform.3
2024 CTISL: a dynamic stacking multi-class classification approach for identifying cell types from single-cell RNA-seq data
abstract
MOTIVATION: Effective identification of cell types is of critical importance in single-cell RNA-sequencing (scRNA-seq) data analysis. To date, many supervised machine learning-based predictors have been implemented to identify cell types from scRNA-seq datasets. Despite the technical advances of these state-of-the-art tools, most existing predictors were single classifiers, of which the performances can still be significantly improved. It is therefore highly desirable to employ the ensemble learning strategy to develop more accurate computational models for robust and comprehensive identification of cell types on scRNA-seq datasets. RESULTS: We propose a two-layer stacking model, termed CTISL (Cell Type Identification by Stacking ensemble Learning), which integrates multiple classifiers to identify cell types. In the first layer, given a reference scRNA-seq dataset with known cell types, CTISL dynamically combines multiple cell-type-specific classifiers (i.e. support-vector machine and logistic regression) as the base learners to deliver the outcomes for the input of a meta-classifier in the second layer. We conducted a total of 24 benchmarking experiments on 17 human and mouse scRNA-seq datasets to evaluate and compare the prediction performance of CTISL and other state-of-the-art predictors. The experiment results demonstrate that CTISL achieves superior or competitive performance compared to these state-of-the-art approaches. We anticipate that CTISL can serve as a useful and reliable tool for cost-effective identification of cell types from scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: The webserver and source code are freely available at http://bigdata.biocie.cn/CTISLweb/home and https://zenodo.org/records/10568906, respectively.
Ziyi Chai, Yan Liu 0038, Chen Li 0021, Yu Jiang 0014, Quanzhong Liu
Bioinform.7
2022 scHiCStackL: a stacking ensemble learning-based method for single-cell Hi-C classification using cell embedding
abstract
Single-cell Hi-C data are a common data source for studying the differences in the three-dimensional structure of cell chromosomes. The development of single-cell Hi-C technology makes it possible to obtain batches of single-cell Hi-C data. How to quickly and effectively discriminate cell types has become one hot research field. However, the existing computational methods to predict cell types based on Hi-C data are found to be low in accuracy. Therefore, we propose a high accuracy cell classification algorithm, called scHiCStackL, based on single-cell Hi-C data. In our work, we first improve the existing data preprocessing method for single-cell Hi-C data, which allows the generated cell embedding better to represent cells. Then, we construct a two-layer stacking ensemble model for classifying cells. Experimental results show that the cell embedding generated by our data preprocessing method increases by 0.23, 1.22, 1.46 and 1.61$\%$ comparing with the cell embedding generated by the previously published method scHiCluster, in terms of the Acc, MCC, F1 and Precision confidence intervals, respectively, on the task of classifying human cells in the ML1 and ML3 datasets. When using the two-layer stacking ensemble framework with the cell embedding, scHiCStackL improves by 13.33, 19, 19.27 and 14.5 over the scHiCluster, in terms of the Acc, ARI, NMI and F1 confidence intervals, respectively. In summary, scHiCStackL achieves superior performance in predicting cell types using the single-cell Hi-C data. The webserver and source code of scHiCStackL are freely available at http://hww.sdu.edu.cn:8002/scHiCStackL/ and https://github.com/HaoWuLab-Bioinformatics/scHiCStackL, respectively.
Hao Wu 0062, Yingfu Wu, Haoru Zhou, Zhongli Chen, Yi Xiong 0002, Quanzhong Liu, Hongming Zhang 0002
Briefings Bioinform.8
2022 DeepGenGrep: a general deep learning-based predictor for multiple genomic signals and regions
abstract
MOTIVATION: Accurate annotation of different genomic signals and regions (GSRs) from DNA sequences is fundamentally important for understanding gene structure, regulation and function. Numerous efforts have been made to develop machine learning-based predictors for in silico identification of GSRs. However, it remains a great challenge to identify GSRs as the performance of most existing approaches is unsatisfactory. As such, it is highly desirable to develop more accurate computational methods for GSRs prediction. RESULTS: In this study, we propose a general deep learning framework termed DeepGenGrep, a general predictor for the systematic identification of multiple different GSRs from genomic DNA sequences. DeepGenGrep leverages the power of hybrid neural networks comprising a three-layer convolutional neural network and a two-layer long short-term memory to effectively learn useful feature representations from sequences. Benchmarking experiments demonstrate that DeepGenGrep outperforms several state-of-the-art approaches on identifying polyadenylation signals, translation initiation sites and splice sites across four eukaryotic species including Homo sapiens, Mus musculus, Bos taurus and Drosophila melanogaster. Overall, DeepGenGrep represents a useful tool for the high-throughput and cost-effective identification of potential GSRs in eukaryotic genomes. AVAILABILITY AND IMPLEMENTATION: The webserver and source code are freely available at http://bigdata.biocie.cn/deepgengrep/home and Github (https://github.com/wx-cie/DeepGenGrep/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Quanzhong Liu, Honglin Fang, Lachlan James M. Coin, Fuyi Li, Jiangning Song
Bioinform.1
2021 Large-scale comparative review and assessment of computational methods for anti-cancer peptide identification
abstract
Anti-cancer peptides (ACPs) are known as potential therapeutics for cancer. Due to their unique ability to target cancer cells without affecting healthy cells directly, they have been extensively studied. Many peptide-based drugs are currently evaluated in the preclinical and clinical trials. Accurate identification of ACPs has received considerable attention in recent years; as such, a number of machine learning-based methods for in silico identification of ACPs have been developed. These methods promote the research on the mechanism of ACPs therapeutics against cancer to some extent. There is a vast difference in these methods in terms of their training/testing datasets, machine learning algorithms, feature encoding schemes, feature selection methods and evaluation strategies used. Therefore, it is desirable to summarize the advantages and disadvantages of the existing methods, provide useful insights and suggestions for the development and improvement of novel computational tools to characterize and identify ACPs. With this in mind, we firstly comprehensively investigate 16 state-of-the-art predictors for ACPs in terms of their core algorithms, feature encoding schemes, performance evaluation metrics and webserver/software usability. Then, comprehensive performance assessment is conducted to evaluate the robustness and scalability of the existing predictors using a well-prepared benchmark dataset. We provide potential strategies for the model performance improvement. Moreover, we propose a novel ensemble learning framework, termed ACPredStackL, for the accurate identification of ACPs. ACPredStackL is developed based on the stacking ensemble strategy combined with SVM, Naïve Bayesian, lightGBM and KNN. Empirical benchmarking experiments against the state-of-the-art methods demonstrate that ACPredStackL achieves a comparative performance for predicting ACPs. The webserver and source code of ACPredStackL is freely available at http://bigdata.biocie.cn/ACPredStackL/ and https://github.com/liangxiaoq/ACPredStackL, respectively.
Fuyi Li, Hao Wu 0062, Jiangning Song, Quanzhong Liu
Briefings Bioinform.8
2021 DeepTorrent: a deep learning-based approach for predicting DNA N4-methylcytosine sites
abstract
DNA N4-methylcytosine (4mC) is an important epigenetic modification that plays a vital role in regulating DNA replication and expression. However, it is challenging to detect 4mC sites through experimental methods, which are time-consuming and costly. Thus, computational tools that can identify 4mC sites would be very useful for understanding the mechanism of this important type of DNA modification. Several machine learning-based 4mC predictors have been proposed in the past 3 years, although their performance is unsatisfactory. Deep learning is a promising technique for the development of more accurate 4mC site predictions. In this work, we propose a deep learning-based approach, called DeepTorrent, for improved prediction of 4mC sites from DNA sequences. It combines four different feature encoding schemes to encode raw DNA sequences and employs multi-layer convolutional neural networks with an inception module integrated with bidirectional long short-term memory to effectively learn the higher-order feature representations. Dimension reduction and concatenated feature maps from the filters of different sizes are then applied to the inception module. In addition, an attention mechanism and transfer learning techniques are also employed to train the robust predictor. Extensive benchmarking experiments demonstrate that DeepTorrent significantly improves the performance of 4mC site prediction compared with several state-of-the-art methods.
Quanzhong Liu, Cangzhi Jia, Jiangning Song, Fuyi Li
Briefings Bioinform.1
2020 PRISMOID: a comprehensive 3D structure database for post-translational modifications and mutations with functional impact
abstract
Post-translational modifications (PTMs) play very important roles in various cell signaling pathways and biological process. Due to PTMs' extremely important roles, many major PTMs have been studied, while the functional and mechanical characterization of major PTMs is well documented in several databases. However, most currently available databases mainly focus on protein sequences, while the real 3D structures of PTMs have been largely ignored. Therefore, studies of PTMs 3D structural signatures have been severely limited by the deficiency of the data. Here, we develop PRISMOID, a novel publicly available and free 3D structure database for a wide range of PTMs. PRISMOID represents an up-to-date and interactive online knowledge base with specific focus on 3D structural contexts of PTMs sites and mutations that occur on PTMs and in the close proximity of PTM sites with functional impact. The first version of PRISMOID encompasses 17 145 non-redundant modification sites on 3919 related protein 3D structure entries pertaining to 37 different types of PTMs. Our entry web page is organized in a comprehensive manner, including detailed PTM annotation on the 3D structure and biological information in terms of mutations affecting PTMs, secondary structure features and per-residue solvent accessibility features of PTM sites, domain context, predicted natively disordered regions and sequence alignments. In addition, high-definition JavaScript packages are employed to enhance information visualization in PRISMOID. PRISMOID equips a variety of interactive and customizable search options and data browsing functions; these capabilities allow users to access data via keyword, ID and advanced options combination search in an efficient and user-friendly way. A download page is also provided to enable users to download the SQL file, computational structural features and PTM sites' data. We anticipate PRISMOID will swiftly become an invaluable online resource, assisting both biologists and bioinformaticians to conduct experiments and develop applications supporting discovery efforts in the sequence-structural-functional relationship of PTMs and providing important insight into mutations and PTM sites interaction mechanisms. The PRISMOID database is freely accessible at http://prismoid.erc.monash.edu/. The database and web interface are implemented in MySQL, JSP, JavaScript and HTML with all major browsers supported.
Fuyi Li, Cunshuo Fan, Tatiana T. Marquez-Lago, André Leier, Jerico Revote, Cangzhi Jia, Yan Zhu 0006, Alexander Ian Smith, Geoffrey I. Webb, Quanzhong Liu, Leyi Wei, Jian Li 0052, Jiangning Song
Briefings Bioinform.10
2020 DeepCleave: a deep learning predictor for caspase and matrix metalloprotease substrates and cleavage sites
abstract
MOTIVATION: Proteases are enzymes that cleave target substrate proteins by catalyzing the hydrolysis of peptide bonds between specific amino acids. While the functional proteolysis regulated by proteases plays a central role in the 'life and death' cellular processes, many of the corresponding substrates and their cleavage sites were not found yet. Availability of accurate predictors of the substrates and cleavage sites would facilitate understanding of proteases' functions and physiological roles. Deep learning is a promising approach for the development of accurate predictors of substrate cleavage events. RESULTS: We propose DeepCleave, the first deep learning-based predictor of protease-specific substrates and cleavage sites. DeepCleave uses protein substrate sequence data as input and employs convolutional neural networks with transfer learning to train accurate predictive models. High predictive performance of our models stems from the use of high-quality cleavage site features extracted from the substrate sequences through the deep learning process, and the application of transfer learning, multiple kernels and attention layer in the design of the deep network. Empirical tests against several related state-of-the-art methods demonstrate that DeepCleave outperforms these methods in predicting caspase and matrix metalloprotease substrate-cleavage sites. AVAILABILITY AND IMPLEMENTATION: The DeepCleave webserver and source code are freely available at http://deepcleave.erc.monash.edu/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fuyi Li, André Leier, Tatiana T. Marquez-Lago, Quanzhong Liu, Jerico Revote, Alexander Ian Smith, Tatsuya Akutsu, Geoffrey I. Webb, Lukasz A. Kurgan, Jiangning Song
Bioinform.5
2008 Extracting Decision Rules from Sigmoid Kernel
Quanzhong Liu, Yang Zhang 0010, Zhengguo Hu
ADMA1