Wei Zhang 0079

dblp:10/4661-79 · DBLP profile ↗
← Back
12ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0002-8884-7956ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 9 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 scVGAMF: a novel imputation method for scRNA-seq data by integrating linear and non-linear features
abstract
Single-cell RNA sequencing (scRNA-seq) is crucial for elucidating gene expression dynamics and cellular heterogeneity at the individual cell level, thereby advancing our understanding of transcriptional regulation across distinct cell populations. However, a significant challenge in scRNA-seq data analysis is the prevalence of dropout events, which complicate downstream analyses. Most existing imputation tools either rely solely on linear assumptions or overlook the non-linear regulatory relationships embedded in the data. To address this issue, we propose single-cell variational graph autoencoder and matrix factorization (scVGAMF), a novel imputation method that integrates both linear and non-linear features. Specifically, scVGAMF first identifies highly variable genes and partitions them into groups. Cells are then clustered by applying spectral clustering to the principal component analysis results of the representative groups. Based on the resulting submatrices, along with the gene similarity and cell-cell similarity matrices, scVGAMF employs non-negative matrix factorization to extract underlying linear features while utilizing two variational graph autoencoders to capture non-linear features. A fully connected neural network then integrates these features to predict missing values. Extensive experimental evaluations on simulated dropout datasets and real scRNA-seq data demonstrate that scVGAMF outperforms existing methods in gene expression recovery, cell clustering accuracy, differential gene identification, and pseudo-trajectory analysis. Furthermore, ablation studies confirm that the integration of both linear and non-linear features significantly enhances overall data imputation performance.
Wei Zhang 0079
Briefings Bioinform.2
2025 AcImpute: a constraint-enhancing smooth-based approach for imputing single-cell RNA sequencing data
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) provides a powerful tool for studying cellular heterogeneity and complexity. However, dropout events in single-cell RNA-seq data severely hinder the effectiveness and accuracy of downstream analysis. Therefore, data preprocessing with imputation methods is crucial to scRNA-seq analysis. RESULTS: To address the issue of oversmoothing in smoothing-based imputation methods, the presented AcImpute, an unsupervised method that enhances imputation accuracy by constraining the smoothing weights among cells for genes with different expression levels. Compared with nine other imputation methods in cluster analysis and trajectory inference, the experimental results can demonstrate that AcImpute effectively restores gene expression, preserves inter-cell variability, preventing oversmoothing and improving clustering and trajectory inference performance. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Liutto/AcImpute.
Wei Zhang 0079
Bioinform.1
2025 Low-rank kernel consistent multi-view subspace clustering
Wei Zhang 0079, Shiqi Wang 0001, Zizhu Fan
Neurocomputing1
2024 Identifying cell types by lasso-constraint regularized Gaussian graphical model based on weighted distance penalty
abstract
Single-cell RNA sequencing (scRNA-seq) technology is one of the most cost-effective and efficacious methods for revealing cellular heterogeneity and diversity. Precise identification of cell types is essential for establishing a robust foundation for downstream analyses and is a prerequisite for understanding heterogeneous mechanisms. However, the accuracy of existing methods warrants improvement, and highly accurate methods often impose stringent equipment requirements. Moreover, most unsupervised learning-based approaches are constrained by the need to input the number of cell types a prior, which limits their widespread application. In this paper, we propose a novel algorithm framework named WLGG. Initially, to capture the underlying nonlinear information, we introduce a weighted distance penalty term utilizing the Gaussian kernel function, which maps data from a low-dimensional nonlinear space to a high-dimensional linear space. We subsequently impose a Lasso constraint on the regularized Gaussian graphical model to enhance its ability to capture linear data characteristics. Additionally, we utilize the Eigengap strategy to predict the number of cell types and obtain predicted labels via spectral clustering. The experimental results on 14 test datasets demonstrate the superior clustering accuracy of the WLGG algorithm over 16 alternative methods. Furthermore, downstream analysis, including marker gene identification, pseudotime inference, and functional enrichment analysis based on the similarity matrix and predicted labels from the WLGG algorithm, substantiates the reliability of WLGG and offers valuable insights into biological dynamic biological processes and regulatory mechanisms.
Wei Zhang 0079, Yaxin Xu
Briefings Bioinform.1
2022 NMFLRR: Clustering scRNA-Seq Data by Integrating Nonnegative Matrix Factorization With Low Rank Representation
abstract
Fast-developing single-cell technologies create unprecedented opportunities to reveal cell heterogeneity and diversity. Accurate classification of single cells is a critical prerequisite for recovering the mechanisms of heterogeneity. However, the scRNA-seq profiles we obtained at present have high dimensionality, sparsity, and noise, which pose challenges for existing clustering methods in grouping cells that belong to the same subpopulation based on transcriptomic profiles. Although many computational methods have been proposed developing novel and effective computational methods to accurately identify cell types remains a considerable challenge. We present a new computational framework to identify cell types by integrating low-rank representation (LRR) and nonnegative matrix factorization (NMF); this framework is named NMFLRR. The LRR captures the global properties of original data by using nuclear norms, and a locality constrained graph regularization term is introduced to characterize the data's local geometric information. The similarity matrix and low-dimensional features of data can be simultaneously obtained by applying the alternating direction method of multipliers (ADMM) algorithm to handle each variable alternatively in an iterative way. We finally obtained the predicted cell types by using a spectral algorithm based on the optimized similarity matrix. Nine real scRNA-seq datasets were used to test the performance of NMFLRR and fifteen other competitive methods, and the accuracy and robustness of the simulation results suggest the NMFLRR is a promising algorithm for the classification of single cells. The simulation code is freely available at: https://github.com/wzhangwhu/NMFLRR_code.
Wei Zhang 0079, Xiaoli Xue, Zizhu Fan
IEEE J. Biomed. Health Informatics1
2021 SCCLRR: A Robust Computational Method for Accurate Clustering Single Cell RNA-Seq Data
abstract
Single-cell RNA transcriptome data present a tremendous opportunity for studying the cellular heterogeneity. Identifying subpopulations based on scRNA-seq data is a hot topic in recent years, although many researchers have been focused on designing elegant computational methods for identifying new cell types; however, the performance of these methods is still unsatisfactory due to the high dimensionality, sparsity and noise of scRNA-seq data. In this study, we propose a new cell type detection method by learning a robust and accurate similarity matrix, named SCCLRR. The method simultaneously captures both global and local intrinsic properties of data based on a low rank representation (LRR) framework mathematical model. The integrated normalized Euclidean distance and cosine similarity are used to balance the intrinsic linear and nonlinear manifold of data in the local regularization term. To solve the non-convex optimization model, we present an iterative optimization procedure using the alternating direction method of multipliers (ADMM) algorithm. We evaluate the performance of the SCCLRR method on nine real scRNA-seq datasets and compare it with seven state-of-the-art methods. The simulation results show that the SCCLRR outperforms other methods and is robust and effective for clustering scRNA-seq data. (The code of SCCLRR is free available for academic https://github.com/wzhangwhu/SCCLRR).
Wei Zhang 0079, Xiu-Fen Zou
IEEE J. Biomed. Health Informatics1
2020 Comparative analysis of similarity measurements in miRNAs with applications to miRNA-disease association predictions
abstract
BACKGROUND: As regulators of gene expression, microRNAs (miRNAs) are increasingly recognized as critical biomarkers of human diseases. Till now, a series of computational methods have been proposed to predict new miRNA-disease associations based on similarity measurements. Different categories of features in miRNAs are applied in these methods for miRNA-miRNA similarity calculation. Benchmarking tests on these miRNA similarity measures are warranted to assess their effectiveness and robustness. RESULTS: In this study, 5 categories of features, i.e. miRNA sequences, miRNA expression profiles in cell-lines, miRNA expression profiles in tissues, gene ontology (GO) annotations of miRNA target genes and Medical Subject Heading (MeSH) terms of miRNA-associated diseases, are collected and similarity values between miRNAs are quantified based on these feature spaces, respectively. We systematically compare the 5 similarities from multi-statistical views. Furthermore, we adopt a rule-based inference method to test their performance on miRNA-disease association predictions with the similarity measurements. Comprehensive comparison is made based on leave-one-out cross-validations and a case study. Experimental results demonstrate that the similarity measurement using MeSH terms performs best among the 5 measurements. It should be noted that the other 4 measurements can also achieve reliable prediction performance. The best-performed similarity measurement is used for new miRNA-disease association predictions and the inferred results are released for further biomedical screening. CONCLUSIONS: Our study suggests that all the 5 features, even though some are restricted by data availability, are useful information for inferring novel miRNA-disease associations. However, biased prediction results might be produced in GO- and MeSH-based similarity measurements due to incomplete feature spaces. Similarity fusion may help produce more reliable prediction results. We expect that future studies will provide more detailed information into the 5 feature spaces and widen our understanding about disease pathogenesis.
Hailin Chen, Ruiyu Guo, Guanghui Li 0003, Wei Zhang 0079, Zuping Zhang 0001
BMC Bioinform.4
2020 Predicting Essential Proteins by Integrating Network Topology, Subcellular Localization Information, Gene Expression Profile and GO Annotation Data
abstract
Essential proteins are indispensable for maintaining normal cellular functions. Identification of essential proteins from Protein-protein interaction (PPI) networks has become a hot topic in recent years. Traditionally biological experimental based approaches are time-consuming and expensive, although lots of computational based methods have been developed in the past years; however, the prediction accuracy is still unsatisfied. In this research, by introducing the protein sub-cellular localization information, we define a new measurement for characterizing the protein's subcellular localization essentiality, and a new data fusion based method is developed for identifying essential proteins, named TEGS, based on integrating network topology, gene expression profile, GO annotation information, and protein subcellular localization information. To demonstrate the efficiency of the proposed method TEGS, we evaluate its performance on two Saccharomyces cerevisiae datasets and compare with other seven state-of-the-art methods (DC, BC, NC, PeC, WDC, SON, and TEO) in terms of true predicted number, jackknife curve, and precision-recall curve. Simulation results show that the TEGS outperforms the other compared methods in identifying essential proteins. The source code of TEGS is freely available at https://github.com/wzhangwhu/TEGS.
Wei Zhang 0079, Jia Xu 0002, Xiu-Fen Zou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 Correction to "Detecting Essential Proteins Based on Network Topology, Gene Expression Data, and Gene Ontology Information"
abstract
Presents corrections to author information for the paper, W. Zhang, J. Xu, Y. Li, and X. Zou, "Detecting essential proteins based on network topology, gene expression data, and gene ontology information,", IEEE/ACMTrans. Comput. Biol. Bioinf., vol. 15, no. 1, pp. 109-116, Jan./Feb. 2018.
Wei Zhang 0079, Jia Xu 0002, Xiu-Fen Zou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 Detecting Essential Proteins Based on Network Topology, Gene Expression Data, and Gene Ontology Information
abstract
The identification of essential proteins in protein-protein interaction (PPI) networks is of great significance for understanding cellular processes. With the increasing availability of large-scale PPI data, numerous centrality measures based on network topology have been proposed to detect essential proteins from PPI networks. However, most of the current approaches focus mainly on the topological structure of PPI networks, and largely ignore the gene ontology annotation information. In this paper, we propose a novel centrality measure, called TEO, for identifying essential proteins by combining network topology, gene expression profiles, and GO information. To evaluate the performance of the TEO method, we compare it with five other methods (degree, betweenness, NC, Pec, and CowEWC) in detecting essential proteins from two different yeast PPI datasets. The simulation results show that adding GO information can effectively improve the predicted precision and that our method outperforms the others in predicting essential proteins.
Wei Zhang 0079, Jia Xu 0002, Xiu-Fen Zou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 A New Method for Detecting Protein Complexes based on the Three Node Cliques
abstract
The identification of protein complexes in protein-protein interaction (PPI) networks is fundamental for understanding biological processes and cellular molecular mechanisms. Many graph computational algorithms have been proposed to identify protein complexes from PPI networks by detecting densely connected groups of proteins. These algorithms assess the density of subgraphs through evaluation of the sum of individual edges or nodes; thus, incomplete and inaccurate measures may miss meaningful biological protein complexes with functional significance. In this study, we propose a novel method for assessing the compactness of local subnetworks by measuring the number of three node cliques. The present method detects each optimal cluster by growing a seed and maximizing the compactness function. To demonstrate the efficacy of the new proposed method, we evaluate its performance using five PPI networks on three reference sets of yeast protein complexes with five different measurements and compare the performance of the proposed method with four state-of-the-art methods. The results show that the protein complexes generated by the proposed method are of better quality than those generated by four classic methods. Therefore, the new proposed method is effective and useful for detecting protein complexes in PPI networks.
Wei Zhang 0079, Xiu-Fen Zou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2013 Systematic Analysis of the Mechanisms of Virus-Triggered Type I IFN Signaling Pathways through Mathematical Modeling
abstract
Based on biological experimental data, we developed a mathematical model of the virus-triggered signaling pathways that lead to induction of type I IFNs and systematically analyzed the mechanisms of the cellular antiviral innate immune responses, including the negative feedback regulation of ISG56 and the positive feedback regulation of IFNs. We found that the time between 5 and 48 hours after viral infection is vital for the control and/or elimination of the virus from the host cells and demonstrated that the ISG56-induced inhibition of MITA activation is stronger than the ISG56-induced inhibition of TBK1 activation. The global parameter sensitivity analysis suggests that the positive feedback regulation of IFNs is very important in the innate antiviral system. Furthermore, the robustness of the innate immune signaling network was demonstrated using a new robustness index. These results can help us understand the mechanisms of the virus-induced innate immune response at a system level and provide instruction for further biological experiments.
Wei Zhang 0079, Xiu-Fen Zou
IEEE ACM Trans. Comput. Biol. Bioinform.1