EDBT 2026 Demo / reviewers in the wild / expert
Xiang-Zhen Kong
dblp:121/5758
· DBLP profile ↗
19ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0002-1315-3570ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Scalable and Unified Hierarchical Coarsening Hypergraph Framework for Single-Cell Multi-Omics Data AnalysisabstractNowadays, technological advancements enable the simultaneous profiling of the epigenome, transcriptome, and proteome at the single-cell level, and an urgent need for tools to integrate such complex tri-modal data is demanded. In this paper, we propose a novel method named scHCHF, a scalable and unified Hierarchical Coarsening Hypergraph Framework. scHCHF first employs a hierarchical coarsening process to groups cells into hypervertices. A unified multiomics hypergraph is then constructed upon these hypervertices, which substantially reduces computational complexity and memory footprint. Subsequently, a hypergraph attention network captures high-order relationships and learns discriminative cell representations, guided by a weakly supervised prototypical contrastive loss. Building upon the pre-trained network, scHCHF further employs a fine-tuning paradigm to achieve accurate cell type annotation using a minimal number of labeled cells. Experiments on real tri-modal datasets demonstrate that scHCHF outperforms other state-of-the-art methods in both cell clustering and cell type annotation tasks. Ke-Ze Yu, Xiang-Zhen Kong, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Junliang Shang |
BIBM | 2 |
| 2023 | Identify Complex Higher-Order Associations Between Alzheimer's Disease Genes and Imaging Markers Through Improved Adaptive Sparse Multi-view Canonical Correlation Analysis
Xiang-Zhen Kong, Boxin Guan, Chun-Hou Zheng 0001, Ying-Lian Gao |
ICIC (3) | 2 |
| 2023 | CHLPCA: Correntropy-Based Hypergraph Regularized Sparse PCA for Single-Cell Type Identification
Tai-Ge Wang, Xiang-Zhen Kong, Shengjun Li, Juan Wang 0003 |
ISBRA | 2 |
| 2023 | A New Binary Biclustering Algorithm Based on Weight Adjacency Difference Matrix for Analyzing Gene Expression DataabstractBiclustering algorithms are essential for processing gene expression data. However, to process the dataset, most biclustering algorithms require preprocessing the data matrix into a binary matrix. Regrettably, this type of preprocessing may introduce noise or cause information loss in the binary matrix, which would reduce the biclustering algorithm's ability to effectively obtain the optimal biclusters. In this paper, we propose a new preprocessing method named Mean-Standard Deviation (MSD) to resolve the problem. Additionally, we introduce a new biclustering algorithm called Weight Adjacency Difference Matrix Binary Biclustering (W-AMBB) to effectively process datasets containing overlapping biclusters. The basic idea is to create a weighted adjacency difference matrix by applying weights to a binary matrix that is derived from the data matrix. This allows us to identify genes with significant associations in sample data by efficiently identifying similar genes that respond to specific conditions. Furthermore, the performance of the W-AMBB algorithm was tested on both synthetic and real datasets and compared with other classical biclustering methods. The experiment results demonstrate that the W-AMBB algorithm is significantly more robust than the compared biclustering methods on the synthetic dataset. Additionally, the results of the GO enrichment analysis show that the W-AMBB method possesses biological significance on real datasets. He-Ming Chu, Xiang-Zhen Kong, Jin-Xing Liu 0001, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | A binary biclustering algorithm based on the adjacency difference matrix for gene expression data analysisabstractBiclustering algorithm is an effective tool for processing gene expression datasets. There are two kinds of data matrices, binary data and non-binary data, which are processed by biclustering method. A binary matrix is usually converted from pre-processed gene expression data, which can effectively reduce the interference from noise and abnormal data, and is then processed using a biclustering algorithm. However, biclustering algorithms of dealing with binary data have a poor balance between running time and performance. In this paper, we propose a new biclustering algorithm called the Adjacency Difference Matrix Binary Biclustering algorithm (AMBB) for dealing with binary data to address the drawback. The AMBB algorithm constructs the adjacency matrix based on the adjacency difference values, and the submatrix obtained by continuously updating the adjacency difference matrix is called a bicluster. The adjacency matrix allows for clustering of gene that undergo similar reactions under different conditions into clusters, which is important for subsequent genes analysis. Meanwhile, experiments on synthetic and real datasets visually demonstrate that the AMBB algorithm has high practicability. He-Ming Chu, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Juan Wang 0003, Xiang-Zhen Kong |
BMC Bioinform. | 6 |
| 2022 | Single-Cell RNA Sequencing Data Clustering by Low-Rank Subspace Ensemble FrameworkabstractThe rapid development of single-cell RNA sequencing (scRNA-seq)technology reveals the gene expression status and gene structure of individual cells, reflecting the heterogeneity and diversity of cells. The traditional methods of scRNA-seq data analysis treat data as the same subspace, and hide structural information in other subspaces. In this paper, we propose a low-rank subspace ensemble clustering framework (LRSEC)to analyze scRNA-seq data. Assuming that the scRNA-seq data exist in multiple subspaces, the low-rank model is used to find the lowest rank representation of the data in the subspace. It is worth noting that the penalty factor of the low-rank kernel function is uncertain, and different penalty factors correspond to different low-rank structures. Moreover, the single cluster model is difficult to find the cellular structure of all datasets. To strengthen the correlation between model solutions, we construct a new ensemble clustering framework LRSEC by using the low-rank model as the basic learner. The LRSEC framework captures the global structure of data through low-rank subspaces, which has better clustering performance than a single clustering model. We validate the performance of the LRSEC framework on seven small datasets and one large dataset and obtain satisfactory results. Chuan-Yuan Wang, Ying-Lian Gao, Jin-Xing Liu 0001, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | NCPLP: A Novel Approach for Predicting Microbe-Associated Diseases With Network Consistency Projection and Label PropagationabstractA growing number of clinical studies have provided substantial evidence of a close relationship between the microbe and the disease. Thus, it is necessary to infer potential microbe-disease associations. But traditional approaches use experiments to validate these associations that often spend a lot of materials and time. Hence, more reliable computational methods are expected to be applied to predict disease-associated microbes. In this article, an innovative mean for predicting microbe-disease associations is proposed, which is based on network consistency projection and label propagation (NCPLP). Given that most existing algorithms use the Gaussian interaction profile (GIP) kernel similarity as the similarity criterion between microbe pairs and disease pairs, in this model, Medical Subject Headings descriptors are considered to calculate disease semantic similarity. In addition, 16S rRNA gene sequences are borrowed for the calculation of microbe functional similarity. In view of the gene-based sequence information, we use two conventional methods (BLAST+ and MEGA7) to assess the similarity between each pair of microbes from different perspectives. Especially, network consistency projection is added to obtain network projection scores from the microbe space and the disease space. Ultimately, label propagation is utilized to reliably predict microbes related to diseases. NCPLP achieves better performance in various evaluation indicators and discovers a greater number of potential associations between microbes and diseases. Also, case studies further confirm the reliable prediction performance of NCPLP. To conclude, our algorithm NCPLP has the ability to discover these underlying microbe-disease associations and can provide help for biological study. Meng-Meng Yin, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Multi-View Random-Walk Graph Regularization Low-Rank Representation for Cancer Clustering and Differentially Expressed Gene SelectionabstractCancer genome data generally consists of multiple views from different sources. These views provide different levels of information about gene activity, as well as more comprehensive cancer information. The low-rank representation (LRR) method, as a powerful subspace clustering method, has been extended and applied in cancer data research. Although the multi-view learning methods based on low rank representation have achieved good results in cancer multi-omics analysis because they fully consider the consistency and complementarity between views, these methods have some shortcomings in mining the potential local geometry of data. In view of this, this paper proposes a new method named Multi-view Random-walk Graph regularization Low-Rank Representation (MRGLRR) to comprehensively analyze multi-view genomics data. This method uses multi-view model to find the common centroid of view. By constructing a joint affinity matrix to learn the low-rank subspace representation of multiple sets of data, the hidden information of each view is fully obtained. In addition, this method introduces random walk graph regularization constraint to obtain more accurate similarity between samples. Different from the traditional graph regularization constraint, after constructing the KNN graph, we use the random walk algorithm to obtain the weight matrix. The random walk algorithm can retain more local geometric information and better learn the topological structure of the data. What's more, a feature gene selection strategy suitable for multi-view model is proposed to find more differentially expressed genes with research value. Experimental results show that our method is better than other representative methods in terms of clustering and feature gene selection for cancer multi-omics data. Juan Wang 0003, Li-Hong Wang, Jin-Xing Liu 0001, Xiang-Zhen Kong, Shengjun Li |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Unsupervised Cluster Analysis and Gene Marker Extraction of scRNA-seq Data Based On Non-Negative Matrix FactorizationabstractThe development of single-cell RNA sequencing (scRNA-seq) technology has made it possible to measure gene expression levels at the resolution of a single cell, which further reveals the complex growth processes of cells such as mutation and differentiation. Recognizing cell heterogeneity is one of the most critical tasks in scRNA-seq research. To solve it, we propose a non-negative matrix factorization framework based on multi-subspace cell similarity learning for unsupervised scRNA-seq data analysis (MscNMF). MscNMF includes three parts: data decomposition, similarity learning, and similarity fusion. The three work together to complete the data similarity learning task. MscNMF can learn the gene features and cell features of different subspaces, and the correlation and heterogeneity between cells will be more prominent in multi-subspaces. The redundant information and noise in each low-dimensional feature space are eliminated, and its gene weight information can be further analyzed to calculate the optimal number of subpopulations. The final cell similarity learning will be more satisfactory due to the fusion of cell similarity information in different subspaces. The advantage of MscNMF is that it can calculate the number of cell types and the rank of Non-negative matrix factorization (NMF) reasonably. Experiments on eight real scRNA-seq datasets show that MscNMF can effectively perform clustering tasks and extract useful genetic markers. To verify its clustering performance, the framework is compared with other latest clustering algorithms and satisfactory results are obtained. The code of MscNMF is free available for academic (https://github.com/wangchuanyuan1/project-MscNMF). Chuan-Yuan Wang, Ying-Lian Gao, Xiang-Zhen Kong, Jin-Xing Liu 0001, Chun-Hou Zheng 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Sparse Hyper-graph Non-negative Matrix Factorization by Maximizing CorrentropyabstractNon-negative Matrix Factorization (NMF) as a powerful dimension reduction tool, which is widely used in the bioinformatics field. However, the loss function of conventional NMF is sensitive to non-Gaussian noise and outliers. In addition, NMF-based algorithm overlooks the geometric structure of high dimensional data. To improve the robustness of NMF, we propose a novel method called Sparse Hyper-graph regularized Non-negative Matrix Factorization by Maximizing Correntropy (SHNMF-MCC) in this paper. Specifically, the maximum correntropy criterion replaces the Euclidean distance in the loss term of SHNMF-MCC, which can filter out the noise with large outliers. Moreover, the high-order geometric structure in more sample points is completely preserved in the low-dimensional manifold through the hyper-graph regularization. Meanwhile, the sparse constraint is applied to the loss function to reduce matrix complexity and analysis difficulty. Then, the complex optimization problem can be solved by a half-quadratic (HQ) optimization approach. Before carrying out experiments, we analyze the convergence of SHNMF-MCC. Sample clustering experiments on The Cancer Genome Atlas (TCGA) data and single cell RNA-sequencing (scRNA-seq) data verify that the proposed method is more robust and effective than other similar robust approaches. Cui-Na Jiao, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001, Xianzi Yu |
BIBM | 4 |
| 2021 | Extreme Learning Machine Based on Double Kernel Risk-Sensitive Loss for Cancer Samples Classification
Zhen-Xin Niu, Liangrui Ren, Xiang-Zhen Kong, Ying-Lian Gao, Jin-Xing Liu 0001 |
ICIC (2) | 4 |
| 2021 | Kernel Risk-Sensitive Loss based Hyper-graph Regularized Robust Extreme Learning Machine and Its Semi-supervised Extension for Classification
Liangrui Ren, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
Knowl. Based Syst. | 4 |
| 2021 | WGRCMF: A Weighted Graph Regularized Collaborative Matrix Factorization Method for Predicting Novel LncRNA-Disease AssociationsabstractIn recent years, many human diseases have been determined to be associated with certain lncRNAs. Only a small percentage of all lncRNA-disease associations (LDAs) have been discovered by researchers. Predicting novel LDAs is time-consuming and costly. It is crucial to propose a method that can effectively identify potential LDAs to solve this problem based on the available datasets. Although some current methods can effectively predict potential LDAs, the prediction accuracy needs to be improved, and there are few known associations. Moreover, there are notable errors in the method of constructing the network and the bipartite graph, which interfere with the final results. A weighted graph regularized collaborative matrix factorization (WGRCMF) method is proposed to predict novel LDAs. We introduce the graph regularization terms into the collaborative matrix factorization. Considering that manifold learning can recover low-dimensional manifold structures from high-dimensional sampled data, we can find low-dimensional manifolds in high-dimensional space. In addition, a weight matrix is also introduced into the method, the significance of which is to prevent unknown associations from contributing to the final prediction matrix. Finally, the prediction accuracy of this method is better than those of other methods. In several cancer cases, we implemented the corresponding simulation experiments. According to the experimental results, the proposed method is feasible and effective. Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Locally Manifold Non-negative Matrix Factorization Based on Centroid for scRNA-seq Data AnalysisabstractThe rapid development of single cell RNA sequencing (scRNA-seq) has made it possible to study the association between cells and genes at molecular resolution. When the follow-up analysis is carried out, it is often difficult to extract the cell information in high-dimensional space because of the high gene dimension in single-cell sequencing, which leads to inaccurate results in the follow-up analysis. To solve the problem, we propose a method called locally manifold non-negative matrix factorization based on centroid for scRNA-seq data analysis (MNMFC). MNMFC is a similarity modeling scheme based on locally manifold, which can map cell association in high dimensional space. Through similarity learning based on locally manifold and non-negative matrix decomposition (NMF) algorithm, the data in high-dimensional space can be mapped to low-dimensional space, which provides help for downstream clustering analysis. The performance of the model was validated experimentally on 10 scRNA-seq datasets. Compared with other nine advanced single-cell clustering methods, whether it is a comprehensive analysis or an individual analysis of the dataset, MNMFC has achieved encouraging results. Chuan-Yuan Wang, Ying-Lian Gao, Cui-Na Jiao, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Xiang-Zhen Kong |
BIBM | 6 |
| 2020 | Sparse Regularization Tensor Robust PCA Based on t-product and Its Application in Cancer Genomic DataabstractGenetic information becomes more and more important in the process of biological research. Gene analysis is an effective mean in biological research, especially the analysis of differentially expressed genes. Robust principal component analysis (RPCA) is an effective method to identify differentially expressed genes. But tensor robust principal component analysis (TRPCA) performs better than RPCA when processing multi-dimensional data. The traditional TRPCA method also has limitations in restoring low-rank sparse components. To further improve the accuracy of the TRPCA method in restoring low-rank components and sparse components, we propose a novel TRPCA method to obtain high-order correlations information of multi-dimensional data. It uses a new nuclear norm based on t-product operator to approximate the rank function. The L2,1-norm is used to improve the sparsity of tensors and reduce the negative effects caused by noises and outliers. At the same time, the introduction of L2,1-norm enhances the sparsity of error components, and improves the accuracy of low-rank component recovery. The low-rank sparse components are obtained by solving the convex problem of the new tensor nuclear norm. It can well preserve the spatial structure and make full use of complementary information to improve the clustering effect. Alternating direction method of multiplier (ADMM) is used to solve the optimization problem of this method. Experimental results on different cancer genomic datasets indicate that our method is superior to other methods. Hang-Jin Yang, Yu-Ying Zhao, Jin-Xing Liu 0001, Yuxia Lei, Junliang Shang, Xiang-Zhen Kong |
BIBM | 6 |
| 2020 | Tensor Robust Principal Component Analysis with Low-Rank Weight Constraints for Sample ClusteringabstractWith the rapid development of the next-generation sequencing technology, a large amount of genomics information has been obtained. The scale of biological sequencing data is particularly large and complex. The tensor robust principal component analysis (TRPCA) method can effectively preserve the spatial structure of tensor data, so it has received extensive attention. However, the low-rank tensor obtained by TRPCA may be damaged to a certain extent. To solve this problem, this paper proposes a model for weighting low-rank data based on the method of TRPCA. This model has an additional constraint penalty term that can repair corrupted low-rank data and the effective information in it can be fully utilized. In addition, the norm is used to constrain the sparse tensor to make the sparse effect better. In the experimental part, TRPCA model clusters samples by low-rank tensor. The experimental results on cancer omics data show that our method is superior to other methods. Yu-Ying Zhao, Maoli Wang, Juan Wang 0003, Shasha Yuan, Jin-Xing Liu 0001, Xiang-Zhen Kong |
BIBM | 6 |
| 2020 | Robust Graph Regularized Extreme Learning Machine Auto Encoder and Its Application to Single-Cell Samples Classification
Liangrui Ren, Jin-Xing Liu 0001, Ying-Lian Gao, Xiang-Zhen Kong, Chun-Hou Zheng 0001 |
ICIC (2) | 4 |
| 2020 | MCCMF: collaborative matrix factorization based on matrix completion for predicting miRNA-disease associationsabstractBACKGROUND: MicroRNAs (miRNAs) are non-coding RNAs with regulatory functions. Many studies have shown that miRNAs are closely associated with human diseases. Among the methods to explore the relationship between the miRNA and the disease, traditional methods are time-consuming and the accuracy needs to be improved. In view of the shortcoming of previous models, a method, collaborative matrix factorization based on matrix completion (MCCMF) is proposed to predict the unknown miRNA-disease associations. RESULTS: The complete matrix of the miRNA and the disease is obtained by matrix completion. Moreover, Gaussian Interaction Profile kernel is added to the miRNA functional similarity matrix and the disease semantic similarity matrix. Then the Weight K Nearest Known Neighbors method is used to pretreat the association matrix, so the model is close to the reality. Finally, collaborative matrix factorization method is applied to obtain the prediction results. Therefore, the MCCMF obtains a satisfactory result in the fivefold cross-validation, with an AUC of 0.9569 (0.0005). CONCLUSIONS: The AUC value of MCCMF is higher than other advanced methods in the fivefold cross validation experiment. In order to comprehensively evaluate the performance of MCCMF, accuracy, precision, recall and f-measure are also added. The final experimental results demonstrate that MCCMF outperforms other methods in predicting miRNA-disease associations. In the end, the effectiveness and practicability of MCCMF are further verified by researching three specific diseases. Tian-Ru Wu, Meng-Meng Yin, Cui-Na Jiao, Ying-Lian Gao, Xiang-Zhen Kong, Jin-Xing Liu 0001 |
BMC Bioinform. | 5 |
| 2019 | A Mixed-Norm Laplacian Regularized Low-Rank Representation Method for Tumor Samples ClusteringabstractTumor samples clustering based on biomolecular data is a hot issue of cancer classifications discovery. How to extract the valuable information from high dimensional genomic data is becoming an urgent problem in tumor samples clustering. In this paper, we introduce manifold regularization into low-rank representation model and present a novel method named Mixed-norm Laplacian regularized Low-Rank Representation (MLLRR) to identify the differentially expressed genes for tumor clustering based on gene expression data. Then, in order to advance the accuracy and stability of tumor clustering, we establish the clustering model based on Penalized Matrix Decomposition (PMD) and propose a novel cluster method named MLLRR-PMD. In this method, the cancer clustering research includes three steps. First, the matrix of gene expression data is decomposed into a low rank representation matrix and a sparse matrix by MLLRR. Second, the differentially expressed genes are identified based on the sparse matrix. Finally, the PMD is applied to cluster the samples based on the differentially expressed genes. The experiment results on simulation data and real genomic data illustrate that MLLRR method enhances the robustness to outliers and achieves remarkable performance in the extraction of differentially expressed genes. Juan Wang 0003, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Yaxuan Wang, Xiang-Zhen Kong, Chang-Gang Wen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |