Chuan-Yuan Wang

dblp:283/3885 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0002-3206-1372ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 7 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Network Activity Evaluation Reveals Significant Gene Regulatory Architectures During SARS-CoV-2 Viral Infection From Dynamic scRNA-seq Data
abstract
The key to understand COVID-19 caused by SARS-CoV-2, which has caused massive deaths worldwide, is to reveal the gene activities at molecular level. Single-cell RNA-sequencing (scRNA-seq) technology allows us to capture gene expression at high resolution, thereby delineating cell-specific gene regulatory network (GRN). Network activity refers to the degree of consistency between GRN architectures and gene expression profiles in a specific condition or cellular microenvironment. Currently, numerous experimentally determined molecular interactions, including regulatory relationships closely related to SARS-CoV-2 infection, are documented in knowledge-bases. However, GRN activity is closely related to the cell dynamic environment and the heterogeneity of cell clusters. Therefore, to evaluate the consistency of GRN with gene expression profiles, we propose a single-cell Network Activity Evaluation framework, called scNAE. First, scNAE performs ODE modeling of time-course gene expression data. Then, the loss function with regularization penalty terms is constructed for formulating GRN inference rules from transcriptomic data. Furthermore, we have devised a rapid-convergence alternating direction method of multipliers to solve the regularized and constrained programs. Finally, an empirical P-value is derived based on a permutation statistical testing procedure to quantify the likelihood significance of the network matching with the data. The efficiency and advantage of scNAE have also been demonstrated by extensive numerical experiments, which can clearly depict the dynamic responses underlying GRN architectures triggered by the infection of SARS-CoV-2 in cells. The code and data of scNAE are available at https://github.com/zpliulab/scNAE.
Chuan-Yuan Wang, Zhi-Ping Liu
IEEE J. Biomed. Health Informatics1
2023 Multi-objective Optimization-Based Approach for Detection of Breast Cancer Biomarkers
Chuan-Yuan Wang, Duanchen Sun, Zhi-Ping Liu
ICIC (3)2
2023 ActivePPI: quantifying protein-protein interaction network activity with Markov random fields
abstract
MOTIVATION: Protein-protein interactions (PPI) are crucial components of the biomolecular networks that enable cells to function. Biological experiments have identified a large number of PPI, and these interactions are stored in knowledge bases. However, these interactions are often restricted to specific cellular environments and conditions. Network activity can be characterized as the extent of agreement between a PPI network (PPIN) and a distinct cellular environment measured by protein mass spectrometry, and it can also be quantified as a statistical significance score. Without knowing the activity of these PPI in the cellular environments or specific phenotypes, it is impossible to reveal how these PPI perform and affect cellular functioning. RESULTS: To calculate the activity of PPIN in different cellular conditions, we proposed a PPIN activity evaluation framework named ActivePPI to measure the consistency between network architecture and protein measurement data. ActivePPI estimates the probability density of protein mass spectrometry abundance and models PPIN using a Markov-random-field-based method. Furthermore, empirical P-value is derived based on a nonparametric permutation test to quantify the likelihood significance of the match between PPIN structure and protein abundance data. Extensive numerical experiments demonstrate the superior performance of ActivePPI and result in network activity evaluation, pathway activity assessment, and optimal network architecture tuning tasks. To summarize it succinctly, ActivePPI is a versatile tool for evaluating PPI network that can uncover the functional significance of protein interactions in crucial cellular biological processes and offer further insights into physiological phenomena. AVAILABILITY AND IMPLEMENTATION: All source code and data are freely available at https://github.com/zpliulab/ActivePPI.
Chuan-Yuan Wang, Duanchen Sun, Zhi-Ping Liu
Bioinform.1
2022 Single-Cell RNA Sequencing Data Clustering by Low-Rank Subspace Ensemble Framework
abstract
The rapid development of single-cell RNA sequencing (scRNA-seq)technology reveals the gene expression status and gene structure of individual cells, reflecting the heterogeneity and diversity of cells. The traditional methods of scRNA-seq data analysis treat data as the same subspace, and hide structural information in other subspaces. In this paper, we propose a low-rank subspace ensemble clustering framework (LRSEC)to analyze scRNA-seq data. Assuming that the scRNA-seq data exist in multiple subspaces, the low-rank model is used to find the lowest rank representation of the data in the subspace. It is worth noting that the penalty factor of the low-rank kernel function is uncertain, and different penalty factors correspond to different low-rank structures. Moreover, the single cluster model is difficult to find the cellular structure of all datasets. To strengthen the correlation between model solutions, we construct a new ensemble clustering framework LRSEC by using the low-rank model as the basic learner. The LRSEC framework captures the global structure of data through low-rank subspaces, which has better clustering performance than a single clustering model. We validate the performance of the LRSEC framework on seven small datasets and one large dataset and obtain satisfactory results.
Chuan-Yuan Wang, Ying-Lian Gao, Jin-Xing Liu 0001, Xiang-Zhen Kong, Chun-Hou Zheng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 Unsupervised Cluster Analysis and Gene Marker Extraction of scRNA-seq Data Based On Non-Negative Matrix Factorization
abstract
The development of single-cell RNA sequencing (scRNA-seq) technology has made it possible to measure gene expression levels at the resolution of a single cell, which further reveals the complex growth processes of cells such as mutation and differentiation. Recognizing cell heterogeneity is one of the most critical tasks in scRNA-seq research. To solve it, we propose a non-negative matrix factorization framework based on multi-subspace cell similarity learning for unsupervised scRNA-seq data analysis (MscNMF). MscNMF includes three parts: data decomposition, similarity learning, and similarity fusion. The three work together to complete the data similarity learning task. MscNMF can learn the gene features and cell features of different subspaces, and the correlation and heterogeneity between cells will be more prominent in multi-subspaces. The redundant information and noise in each low-dimensional feature space are eliminated, and its gene weight information can be further analyzed to calculate the optimal number of subpopulations. The final cell similarity learning will be more satisfactory due to the fusion of cell similarity information in different subspaces. The advantage of MscNMF is that it can calculate the number of cell types and the rank of Non-negative matrix factorization (NMF) reasonably. Experiments on eight real scRNA-seq datasets show that MscNMF can effectively perform clustering tasks and extract useful genetic markers. To verify its clustering performance, the framework is compared with other latest clustering algorithms and satisfactory results are obtained. The code of MscNMF is free available for academic (https://github.com/wangchuanyuan1/project-MscNMF).
Chuan-Yuan Wang, Ying-Lian Gao, Xiang-Zhen Kong, Jin-Xing Liu 0001, Chun-Hou Zheng 0001
IEEE J. Biomed. Health Informatics1
2022 Evaluating Gene Regulatory Network Activity From Dynamic Expression Data by Regularized Constraint Programming
abstract
By extracting molecular interactions identified by experiments, gene regulatory networks or gene circuits have documented in a large number of knowledge-based repositories. They provide systematic information and guidance of the functional connections between regulators, e.g., transcription factor proteins and miRNAs, and target genes. Network activity is defined as the degree of consistency between a regulatory network architecture and a specific cellular context of gene expression and can also be measured as a score of statistical significance. The gene network activities are closely related to the dynamics of cell states. To evaluate the activity of regulatory events in the form of network, we propose a network activity evaluation (NAE) framework by measuring the consistency between network architecture and gene expression data across specific states based on mathematical programming. NAE firstly employs the dynamic Bayesian network model to formulate the network structure with time series profiling data. For the constraints of prior knowledge about gene regulatory network, NAE introduces an interpretable general loss function with regularization penalties to calculate the degree of consistency between gene network and gene expression data. Moreover, we design a fast and convergent alternating direction method of multipliers algorithm to optimize the regularized constraint programming. The efficiency and advantage of the NAE framework is deduced through numerous experiments and comparison studies. It reflects the possibility and potential of the match between network and data, thereby helping to reveal the network activity and to explain the dynamic responds underlying the network structure caused by changes in molecular environment of living cells. The code of NAE is freely available for academic use (https://github.com/zpliulab/NAE).
Chuan-Yuan Wang, Zhi-Ping Liu
IEEE J. Biomed. Health Informatics1
2021 Dual Hyper-Graph Regularized Supervised NMF for Selecting Differentially Expressed Genes and Tumor Classification
abstract
Non-negative matrix factorization (NMF) is a dimensionality reduction technique based on high-dimensional mapping. It can learn part-based representations effectively. In this paper, we propose a method called Dual Hyper-graph Regularized Supervised Non-negative Matrix Factorization (HSNMF). To encode the geometric information of the data, the hyper-graph is introduced into the model as a regularization term. The advantage of hyper-graph learning is to find higher order data relationship to enhance data relevance. This method constructs the data hyper-graph and the feature hyper-graph to find the data manifold and the feature manifold simultaneously. The application of hyper-graph theory in cancer datasets can effectively find pathogenic genes. The discrimination information is further introduced into the objective function to obtain more information about the data. Supervised learning with label information greatly improves the classification effect. Furthermore, the real datasets of cancer usually contain sparse noise, so the$L_{2,1}$-norm is applied to enhance the robustness of HSNMF algorithm. Experiments under The Cancer Genome Atlas (TCGA) datasets verify the feasibility of the HSNMF method.
Chuan-Yuan Wang, Na Yu 0004, Ming-Juan Wu, Ying-Lian Gao, Jin-Xing Liu 0001, Juan Wang 0003
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Locally Manifold Non-negative Matrix Factorization Based on Centroid for scRNA-seq Data Analysis
abstract
The rapid development of single cell RNA sequencing (scRNA-seq) has made it possible to study the association between cells and genes at molecular resolution. When the follow-up analysis is carried out, it is often difficult to extract the cell information in high-dimensional space because of the high gene dimension in single-cell sequencing, which leads to inaccurate results in the follow-up analysis. To solve the problem, we propose a method called locally manifold non-negative matrix factorization based on centroid for scRNA-seq data analysis (MNMFC). MNMFC is a similarity modeling scheme based on locally manifold, which can map cell association in high dimensional space. Through similarity learning based on locally manifold and non-negative matrix decomposition (NMF) algorithm, the data in high-dimensional space can be mapped to low-dimensional space, which provides help for downstream clustering analysis. The performance of the model was validated experimentally on 10 scRNA-seq datasets. Compared with other nine advanced single-cell clustering methods, whether it is a comprehensive analysis or an individual analysis of the dataset, MNMFC has achieved encouraging results.
Chuan-Yuan Wang, Ying-Lian Gao, Cui-Na Jiao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Xiang-Zhen Kong
BIBM1