VLDB 2026 Research / reviewers in the wild / expert
Zhu Wang 0007
dblp:03/6588-7
· DBLP profile ↗
11ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-7752-3606ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Distribution Learning for Graph ClassificationabstractLeveraging the diversity and quantity of data provided by various graph-structured data augmentations while preserving intrinsic semantic information is challenging. Additionally, successive layers in graph neural network (GNN) tend to produce more similar node embeddings, while graph contrastive learning aims to increase the dissimilarity between negative pairs of node embeddings. This inevitably results in a conflict between the message-passing mechanism (MPM) of GNNs and the contrastive learning (CL) of negative pairs via intraviews. In this paper, we propose a conditional distribution learning (CDL) method that learns graph representations from graph-structured data for semisupervised graph classification. Specifically, we present an end-to-end graph representation learning model to align the conditional distributions of weakly and strongly augmented features over the original features. This alignment enables the CDL model to effectively preserve intrinsic semantic information when both weak and strong augmentations are applied to graph-structured data. To avoid the conflict between the MPM and the CL of negative pairs, positive pairs of node representations are retained for measuring the similarity between the original features and the corresponding weakly augmented features. Extensive experiments with several benchmark graph datasets demonstrate the effectiveness of the proposed CDL method. Jie Chen 0065, Hua Mao 0001, Chuanbin Liu 0003, Zhu Wang 0007, Xi Peng 0001 |
AAAI | 4 |
| 2025 | One-Step Adaptive Graph Learning for Incomplete Multiview Subspace ClusteringabstractIncomplete multiview clustering (IMVC) optimally integrates complementary information within incomplete multiview data to improve clustering performance. Several one-step graph-based methods show great potential for IMVC. However, the low-rank structures of similarity graphs are neglected at the initialization stage of similarity graph construction. Moreover, further investigation into complementary information integration across incomplete multiple views is needed, particularly when considering the low-rank structures implied in high-dimensional multiview data. In this paper, we present one-step adaptive graph learning (OAGL) that adaptively performs spectral embedding fusion to achieve clustering assignments at the clustering indicator level. We first initiate affinity matrices corresponding to incomplete multiple views using spare representation under two constraints, i.e., the sparsity constraint on each affinity matrix corresponding to an incomplete view and the degree matrix of the affinity matrix approximating an identity matrix. This approach promotes exploring complementary information across incomplete multiple views. Subsequently, we perform an alignment of the spectral block-diagonal matrices among incomplete multiple views using low-rank tensor learning theory. This facilitates consistency information exploration across incomplete multiple views. Furthermore, we present an effective alternating iterative algorithm to solve the resulting optimization problem. Extensive experiments on benchmark datasets demonstrate that the proposed OAGL method outperforms several state-of-the-art approaches. Jie Chen 0065, Hua Mao 0001, Wai Lok Woo, Chuanbin Liu 0003, Zhu Wang 0007, Xi Peng 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Online Sparse Representation Clustering for Evolving Data StreamsabstractData stream clustering can be performed to discover the patterns underlying continuously arriving sequences of data. A number of data stream clustering algorithms for finding clusters in arbitrary shapes and handling outliers, such as density-based clustering algorithms, have been proposed. However, these algorithms are often limited in their ability to construct and merge microclusters by measuring the Euclidean distances between high-dimensional data objects, e.g., transferring valuable knowledge from historical landmark windows to the current landmark window, and exploiting evolving subspace structures adaptively. We propose an online sparse representation clustering (OSRC) method to learn an affinity matrix for evaluating the relationships among high-dimensional data objects in evolving data streams. We first introduce a low-dimensional projection (LDP) into sparse representation to adaptively reduce the potential negative influence associated with the noise and redundancy contained in high-dimensional data. Then, we take advantage of the -norm optimization technique to choose the appropriate number of representative data objects and form a specific dictionary for sparse representation. The specific dictionary is integrated into sparse representation to adaptively exploit the evolving subspace structures of the high-dimensional data objects. Moreover, the data object representatives from the current landmark window can transfer valuable knowledge to the next landmark window. The experimental results based on a synthetic dataset and six benchmark datasets validate the effectiveness of the proposed method compared to that of state-of-the-art methods for data stream clustering. Jie Chen 0065, Shengxiang Yang, Conor Fahy, Zhu Wang 0007, Yinan Guo 0001, Yingke Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Spectral Embedding Fusion for Incomplete Multiview ClusteringabstractIncomplete multiview clustering (IMVC) aims to reveal the underlying structure of incomplete multiview data by partitioning data samples into clusters. Several graph-based methods exhibit a strong ability to explore high-order information among multiple views using low-rank tensor learning. However, spectral embedding fusion of multiple views is ignored in low-rank tensor learning. In addition, addressing missing instances or features is still an intractable problem for most existing IMVC methods. In this paper, we present a unified spectral embedding tensor learning (USETL) framework that integrates the spectral embedding fusion of multiple similarity graphs and spectral embedding tensor learning for IMVC. To remove redundant information from the original incomplete multiview data, spectral embedding fusion is performed by introducing spectral rotations at two different data levels, i.e., the spectral embedding feature level and the clustering indicator level. The aim of introducing spectral embedding tensor learning is to capture consistent and complementary information by seeking high-order correlations among multiple views. The strategy of removing missing instances is adopted to construct multiple similarity graphs for incomplete multiple views. Consequently, this strategy provides an intuitive and feasible way to construct multiple similarity graphs. Extensive experimental results on multiview datasets demonstrate the effectiveness of the two spectral embedding fusion methods within the USETL framework. Jie Chen 0065, Yingke Chen, Zhu Wang 0007, Haixian Zhang, Xi Peng 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Augmented Sparse Representation for Incomplete Multiview ClusteringabstractIncomplete multiview data are collected from multiple sources or characterized by multiple modalities, where the features of some samples or some views may be missing. Incomplete multiview clustering (IMVC) aims to partition the data into different groups by taking full advantage of the complementary information from multiple incomplete views. Most existing methods based on matrix factorization or subspace learning attempt to recover the missing views or perform imputation of the missing features to improve clustering performance. However, this problem is intractable due to a lack of prior knowledge, e.g., label information or data distribution, especially when the missing views or features are completely damaged. In this article, we proposed an augmented sparse representation (ASR) method for IMVC. We first introduce a discriminative sparse representation learning (DSRL) model, which learns the sparse representations of multiple views as applied to measure the similarity of the existing features. The DSRL model explores complementary and consistent information by integrating the sparse regularization item and a consensus regularization item, respectively. Simultaneously, it learns a discriminative dictionary from the original samples. The sparsity constrained optimization problem in the DSRL model can be efficiently solved by the alternating direction method of multipliers (ADMM). Then, we present a similarity fusion scheme, namely, a sparsity augmented fusion of sparse representations, to obtain a sparsity augmented similarity matrix across different views for spectral clustering. Experimental results on several datasets demonstrate the effectiveness of the proposed ASR method for IMVC. Jie Chen 0065, Shengxiang Yang, Xi Peng 0001, Dezhong Peng, Zhu Wang 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Two-Stage Sparse Representation Clustering for Dynamic Data StreamsabstractData streams are a potentially unbounded sequence of data objects, and the clustering of such data is an effective way of identifying their underlying patterns. Existing data stream clustering algorithms face two critical issues: 1) evaluating the relationship among data objects with individual landmark windows of fixed size and 2) passing useful knowledge from previous landmark windows to the current landmark window. Based on sparse representation techniques, this article proposes a two-stage sparse representation clustering (TSSRC) method. The novelty of the proposed TSSRC algorithm comes from evaluating the effective relationship among data objects in the landmark windows with an accurate number of clusters. First, the proposed algorithm evaluates the relationship among data objects using sparse representation techniques. The dictionary and sparse representations are iteratively updated by solving a convex optimization problem. Second, the proposed TSSRC algorithm presents a dictionary initialization strategy that seeks representative data objects by making full use of the sparse representation results. This efficiently passes previously learned knowledge to the current landmark window over time. Moreover, the convergence and sparse stability of TSSRC can be theoretically guaranteed in continuous landmark windows under certain conditions. Experimental results on benchmark datasets demonstrate the effectiveness and robustness of TSSRC. Jie Chen 0065, Zhu Wang 0007, Shengxiang Yang, Hua Mao 0001 |
IEEE Trans. Cybern. | 2 |
| 2023 | Low-Rank Tensor Learning for Incomplete Multiview ClusteringabstractIncomplete multiview clustering (IMVC) is an effective way to identify the underlying structure of incomplete multiview data. Most existing algorithms based on matrix factorization, graph learning or subspace learning have at least one of the following limitations: (1) the global and local structures of high-dimensional data are not effectively explored simultaneously; (2) the high-order correlations among multiple views are ignored. In this article, we propose a low-rank tensor learning (LRTL) method that learns a consensus low-dimensional embedding matrix for IMVC. We first take advantage of the self-expressiveness property of high-dimensional data to construct sparse similarity matrices for individual views under low-rank and sparsity constraints. Individual low-dimensional embedding matrices can be obtained from the sparse similarity matrices using spectral embedding techniques. This approach simultaneously explores the global and local structures of incomplete multiview data. Then, we present a multiview embedding matrix fusion model that incorporates individual low-dimensional embedding matrices into a third-norm tensor to achieve a consensus low-dimensional embedding matrix. The fusion model exploits complementary information by finding the high-order correlations among multiple views. In addition, the computational cost of an improved fusion strategy is dramatically reduced. Extensive experimental results demonstrate that the proposed LRTL method outperforms several state-of-the-art approaches. Jie Chen 0065, Zhu Wang 0007, Hua Mao 0001, Xi Peng 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Efficient Sparse Representation for Learning With High-Dimensional DataabstractDue to the capability of effectively learning intrinsic structures from high-dimensional data, techniques based on sparse representation have begun to display an impressive impact on several fields, such as image processing, computer vision, and pattern recognition. Learning sparse representations isoften computationally expensive due to the iterative computations needed to solve convex optimization problems in which the number of iterations is unknown before convergence. Moreover, most sparse representation algorithms focus only on determining the final sparse representation results and ignore the changes in the sparsity ratio (SR) during iterative computations. In this article, two algorithms are proposed to learn sparse representations based on locality-constrained linear representation learning with probabilistic simplex constraints. Specifically, the first algorithm, called approximated local linear representation (ALLR), obtains a closed-form solution from individual locality-constrained sparse representations. The second algorithm, called ALLR with symmetric constraints (ALLRSC), further obtains a symmetric sparse representation result with a limited number of computations; notably, the sparsity and convergence of sparse representations can be guaranteed based on theoretical analysis. The steady decline in the SR during iterative computations is a critical factor in practical applications. Experimental results based on public datasets demonstrate that the proposed algorithms perform better than several state-of-the-art algorithms for learning with high-dimensional data. Jie Chen 0065, Shengxiang Yang, Zhu Wang 0007, Hua Mao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Multi-view representation learning for data stream clusteringabstractData stream clustering provides valuable insights into the evolving patterns of long sequences of continuously generated data objects. Most existing clustering methods focus on single-view data streams. In this paper, we propose a multi-view representation learning (MVRL) method for multi-view clustering of data streams. We first introduce an integrated representation learning model to learn a fused sparse affinity matrix across multiple views for spectral clustering. Motivated by the optimization procedure of the integrated representation learning model, we propose three consecutive stages: collaborative representation, the construction of individual global affinity matrices using a mapping function, and the calculation of a fused sparse affinity matrix using Euclidean projection. These stages allow the effective capture of the global and local structures of high-dimensional data objects. Moreover, each stage has a closed-form solution, which determines the upper bound of the computational cost and memory consumption. We then employ the construction residuals of the collaborative representation to adaptively update a dynamic set, which is used to preserve the representative data objects. The dynamic set efficiently transfers previously learned useful knowledge to the arriving data objects. Extensive experimental results on multi-view data stream datasets demonstrate the effectiveness of the proposed MVRL method. Jie Chen 0065, Shengxiang Yang, Zhu Wang 0007 |
Inf. Sci. | 3 |
| 2021 | Low-rank representation with adaptive dictionary learning for subspace clustering
Jie Chen 0065, Hua Mao 0001, Zhu Wang 0007, Xinpei Zhang |
Knowl. Based Syst. | 3 |
| 2018 | K-Means Clustering for Controversial Issues Merging in Chinese Legal TextsabstractIn the fact of growing number of cases, Chinese courts have gradually formed a trial mode to improve the efficiency of trials by conducting trials around the controversial issues. However, identifying the controversy issue in specific cases is not only affected by the uncertainty of facts and laws, but also by the discretion of the judges and extra-case factors, and cannot be expressed as a standard format, which lead to the controversial issues based case retrieval a challenge problem. In this paper, we propose a controversial issues merging algorithm based on K-means clustering for Chinese legal texts. The proposed algorithm can determine the number of clusters of the given cause of action automatically and merge the controversial issues semantically, which makes the case information retrieval more accurate and effective. Xin Tian 0011, Yin Fang, Yang Weng, Yawen Luo, Huifang Cheng, Zhu Wang 0007 |
JURIX | 6 |