VLDB 2026 Research / reviewers in the wild / expert
Jie Chen 0065
dblp:92/6289-65
· DBLP profile ↗
21ranked-venue papers
19as first author
17since 2021 · last 2026
0000-0003-0827-8819ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 13 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Distribution Learning for Graph ClassificationabstractLeveraging the diversity and quantity of data provided by various graph-structured data augmentations while preserving intrinsic semantic information is challenging. Additionally, successive layers in graph neural network (GNN) tend to produce more similar node embeddings, while graph contrastive learning aims to increase the dissimilarity between negative pairs of node embeddings. This inevitably results in a conflict between the message-passing mechanism (MPM) of GNNs and the contrastive learning (CL) of negative pairs via intraviews. In this paper, we propose a conditional distribution learning (CDL) method that learns graph representations from graph-structured data for semisupervised graph classification. Specifically, we present an end-to-end graph representation learning model to align the conditional distributions of weakly and strongly augmented features over the original features. This alignment enables the CDL model to effectively preserve intrinsic semantic information when both weak and strong augmentations are applied to graph-structured data. To avoid the conflict between the MPM and the CL of negative pairs, positive pairs of node representations are retained for measuring the similarity between the original features and the corresponding weakly augmented features. Extensive experiments with several benchmark graph datasets demonstrate the effectiveness of the proposed CDL method. Jie Chen 0065, Hua Mao 0001, Chuanbin Liu 0003, Zhu Wang 0007, Xi Peng 0001 |
AAAI | 1 |
| 2026 | Homophilic-aware graph contrastive learning
Hua Mao 0001, Wai Lok Woo, Jie Chen 0065 |
Pattern Recognit. | 4 |
| 2025 | Cross-View Graph Consistency Learning for Invariant Graph RepresentationsabstractGraph representation learning is fundamental for analyzing graph-structured data. Exploring invariant graph representations remains a challenge for most existing graph representation learning methods. In this paper, we propose a cross-view graph consistency learning (CGCL) method that learns invariant graph representations for link prediction. First, two complementary augmented views are derived from an incomplete graph structure through a coupled graph structure augmentation scheme. This augmentation scheme mitigates the potential information loss that is commonly associated with various data augmentation techniques involving raw graph data, such as edge perturbation, node removal, and attribute masking. Second, we propose a CGCL model that can learn invariant graph representations. A cross-view training scheme is proposed to train the proposed CGCL model. This scheme attempts to maximize the consistency information between one augmented view and the graph structure reconstructed from the other augmented view. Furthermore, we offer a comprehensive theoretical CGCL analysis. This paper empirically and experimentally demonstrates the effectiveness of the proposed CGCL method, achieving competitive results on graph datasets in comparisons with several state-of-the-art algorithms. Jie Chen 0065, Hua Mao 0001, Wai Lok Woo, Chuanbin Liu 0003, Xi Peng 0001 |
AAAI | 1 |
| 2025 | Progressive low-confidence pseudolabeling for semisupervised node classification
Hua Mao 0001, Jie Chen 0065 |
Neurocomputing | 4 |
| 2025 | One-Step Adaptive Graph Learning for Incomplete Multiview Subspace ClusteringabstractIncomplete multiview clustering (IMVC) optimally integrates complementary information within incomplete multiview data to improve clustering performance. Several one-step graph-based methods show great potential for IMVC. However, the low-rank structures of similarity graphs are neglected at the initialization stage of similarity graph construction. Moreover, further investigation into complementary information integration across incomplete multiple views is needed, particularly when considering the low-rank structures implied in high-dimensional multiview data. In this paper, we present one-step adaptive graph learning (OAGL) that adaptively performs spectral embedding fusion to achieve clustering assignments at the clustering indicator level. We first initiate affinity matrices corresponding to incomplete multiple views using spare representation under two constraints, i.e., the sparsity constraint on each affinity matrix corresponding to an incomplete view and the degree matrix of the affinity matrix approximating an identity matrix. This approach promotes exploring complementary information across incomplete multiple views. Subsequently, we perform an alignment of the spectral block-diagonal matrices among incomplete multiple views using low-rank tensor learning theory. This facilitates consistency information exploration across incomplete multiple views. Furthermore, we present an effective alternating iterative algorithm to solve the resulting optimization problem. Extensive experiments on benchmark datasets demonstrate that the proposed OAGL method outperforms several state-of-the-art approaches. Jie Chen 0065, Hua Mao 0001, Wai Lok Woo, Chuanbin Liu 0003, Zhu Wang 0007, Xi Peng 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Hierarchical Sparse Representation Clustering for High-Dimensional Data StreamsabstractData stream clustering reveals patterns within continuously arriving, potentially unbounded data sequences. Numerous data stream algorithms have been proposed to cluster data streams. The existing data stream clustering algorithms still face significant challenges when addressing high-dimensional data streams. First, it is intractable to measure the similarities among high-dimensional data objects via Euclidean distances when constructing and merging microclusters. Second, these algorithms are highly sensitive to the noise contained in high-dimensional data streams. In this article, we propose a hierarchical sparse representation clustering (HSRC) framework for clustering high-dimensional data streams. HSRC first employs a sparse representation-based technique to learn an affinity matrix for data objects in individual landmark windows with a fixed size, where the number of neighboring data objects is automatically selected. The sparse representation-based technique ensures that highly correlated data samples within clusters are grouped together. Then, HSRC applies a spectral clustering technique to the affinity matrix to generate microclusters. These microclusters are subsequently merged into macroclusters based on their sparse similarity degrees (SSDs). In addition, HSRC introduces sparsity residual values (SRVs) to adaptively select representative data objects from the current landmark window. These representatives serve as dictionary samples for the next landmark window. Finally, HSRC refines each macrocluster through fine-tuning. In particular, HSRC enables the detection of outliers in high-dimensional data streams via the associated SRVs. The experimental results obtained on several benchmark datasets demonstrate the effectiveness and robustness of the proposed HSRC framework. Jie Chen 0065, Hua Mao 0001, Yuanbiao Gou, Xi Peng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Online Sparse Representation Clustering for Evolving Data StreamsabstractData stream clustering can be performed to discover the patterns underlying continuously arriving sequences of data. A number of data stream clustering algorithms for finding clusters in arbitrary shapes and handling outliers, such as density-based clustering algorithms, have been proposed. However, these algorithms are often limited in their ability to construct and merge microclusters by measuring the Euclidean distances between high-dimensional data objects, e.g., transferring valuable knowledge from historical landmark windows to the current landmark window, and exploiting evolving subspace structures adaptively. We propose an online sparse representation clustering (OSRC) method to learn an affinity matrix for evaluating the relationships among high-dimensional data objects in evolving data streams. We first introduce a low-dimensional projection (LDP) into sparse representation to adaptively reduce the potential negative influence associated with the noise and redundancy contained in high-dimensional data. Then, we take advantage of the -norm optimization technique to choose the appropriate number of representative data objects and form a specific dictionary for sparse representation. The specific dictionary is integrated into sparse representation to adaptively exploit the evolving subspace structures of the high-dimensional data objects. Moreover, the data object representatives from the current landmark window can transfer valuable knowledge to the next landmark window. The experimental results based on a synthetic dataset and six benchmark datasets validate the effectiveness of the proposed method compared to that of state-of-the-art methods for data stream clustering. Jie Chen 0065, Shengxiang Yang, Conor Fahy, Zhu Wang 0007, Yinan Guo 0001, Yingke Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Spectral Embedding Fusion for Incomplete Multiview ClusteringabstractIncomplete multiview clustering (IMVC) aims to reveal the underlying structure of incomplete multiview data by partitioning data samples into clusters. Several graph-based methods exhibit a strong ability to explore high-order information among multiple views using low-rank tensor learning. However, spectral embedding fusion of multiple views is ignored in low-rank tensor learning. In addition, addressing missing instances or features is still an intractable problem for most existing IMVC methods. In this paper, we present a unified spectral embedding tensor learning (USETL) framework that integrates the spectral embedding fusion of multiple similarity graphs and spectral embedding tensor learning for IMVC. To remove redundant information from the original incomplete multiview data, spectral embedding fusion is performed by introducing spectral rotations at two different data levels, i.e., the spectral embedding feature level and the clustering indicator level. The aim of introducing spectral embedding tensor learning is to capture consistent and complementary information by seeking high-order correlations among multiple views. The strategy of removing missing instances is adopted to construct multiple similarity graphs for incomplete multiple views. Consequently, this strategy provides an intuitive and feasible way to construct multiple similarity graphs. Extensive experimental results on multiview datasets demonstrate the effectiveness of the two spectral embedding fusion methods within the USETL framework. Jie Chen 0065, Yingke Chen, Zhu Wang 0007, Haixian Zhang, Xi Peng 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Augmented Sparse Representation for Incomplete Multiview ClusteringabstractIncomplete multiview data are collected from multiple sources or characterized by multiple modalities, where the features of some samples or some views may be missing. Incomplete multiview clustering (IMVC) aims to partition the data into different groups by taking full advantage of the complementary information from multiple incomplete views. Most existing methods based on matrix factorization or subspace learning attempt to recover the missing views or perform imputation of the missing features to improve clustering performance. However, this problem is intractable due to a lack of prior knowledge, e.g., label information or data distribution, especially when the missing views or features are completely damaged. In this article, we proposed an augmented sparse representation (ASR) method for IMVC. We first introduce a discriminative sparse representation learning (DSRL) model, which learns the sparse representations of multiple views as applied to measure the similarity of the existing features. The DSRL model explores complementary and consistent information by integrating the sparse regularization item and a consensus regularization item, respectively. Simultaneously, it learns a discriminative dictionary from the original samples. The sparsity constrained optimization problem in the DSRL model can be efficiently solved by the alternating direction method of multipliers (ADMM). Then, we present a similarity fusion scheme, namely, a sparsity augmented fusion of sparse representations, to obtain a sparsity augmented similarity matrix across different views for spectral clustering. Experimental results on several datasets demonstrate the effectiveness of the proposed ASR method for IMVC. Jie Chen 0065, Shengxiang Yang, Xi Peng 0001, Dezhong Peng, Zhu Wang 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Deep Multiview Clustering by Contrasting Cluster AssignmentsabstractMultiview clustering (MVC) aims to reveal the underlying structure of multiview data by categorizing data samples into clusters. Deep learning-based methods exhibit strong feature learning capabilities on large-scale datasets. For most existing deep MVC methods, exploring the invariant representations of multiple views is still an intractable problem. In this paper, we propose a cross-view contrastive learning (CVCL) method that learns view-invariant representations and produces clustering results by contrasting the cluster assignments among multiple views. Specifically, we first employ deep autoencoders to extract view-dependent features in the pretraining stage. Then, a cluster-level CVCL strategy is presented to explore consistent semantic label information among the multiple views in the fine-tuning stage. Thus, the proposed CVCL method is able to produce more discriminative cluster assignments by virtue of this learning strategy. Moreover, we provide a theoretical analysis of soft cluster assignment alignment. The extensive experimental results obtained on several datasets demonstrate that the proposed CVCL method outperforms several state-of-the-art approaches. Jie Chen 0065, Hua Mao 0001, Wai Lok Woo, Xi Peng 0001 |
ICCV | 1 |
| 2023 | Two-Stage Sparse Representation Clustering for Dynamic Data StreamsabstractData streams are a potentially unbounded sequence of data objects, and the clustering of such data is an effective way of identifying their underlying patterns. Existing data stream clustering algorithms face two critical issues: 1) evaluating the relationship among data objects with individual landmark windows of fixed size and 2) passing useful knowledge from previous landmark windows to the current landmark window. Based on sparse representation techniques, this article proposes a two-stage sparse representation clustering (TSSRC) method. The novelty of the proposed TSSRC algorithm comes from evaluating the effective relationship among data objects in the landmark windows with an accurate number of clusters. First, the proposed algorithm evaluates the relationship among data objects using sparse representation techniques. The dictionary and sparse representations are iteratively updated by solving a convex optimization problem. Second, the proposed TSSRC algorithm presents a dictionary initialization strategy that seeks representative data objects by making full use of the sparse representation results. This efficiently passes previously learned knowledge to the current landmark window over time. Moreover, the convergence and sparse stability of TSSRC can be theoretically guaranteed in continuous landmark windows under certain conditions. Experimental results on benchmark datasets demonstrate the effectiveness and robustness of TSSRC. Jie Chen 0065, Zhu Wang 0007, Shengxiang Yang, Hua Mao 0001 |
IEEE Trans. Cybern. | 1 |
| 2023 | Multiview Clustering by Consensus Spectral Rotation FusionabstractMultiview clustering (MVC) aims to partition data into different groups by taking full advantage of the complementary information from multiple views. Most existing MVC methods fuse information of multiple views at the raw data level. They may suffer from performance degradation due to the redundant information contained in the raw data. Graph learning-based methods often heavily depend on one specific graph construction, which limits their practical applications. Moreover, they often require a computational complexity ofO(n3) because of matrix inversion or eigenvalue decomposition for each iterative computation. In this paper, we propose a consensus spectral rotation fusion (CSRF) method to learn a fused affinity matrix for MVC at the spectral embedding feature level. Specifically, we first introduce a CSRF model to learn a consensus low-dimensional embedding, which explores the complementary and consistent information across multiple views. We develop an alternating iterative optimization algorithm to solve the CSRF optimization problem, where a computational complexity ofO(n2) is required during each iterative computation. Then, the sparsity policy is introduced to design two different graph construction schemes, which are effectively integrated with the CSRF model. Finally, a multiview fused affinity matrix is constructed from the consensus low-dimensional embedding in spectral embedding space. We analyze the convergence of the alternating iterative optimization algorithm and provide an extension of CSRF for incomplete MVC. Extensive experiments on multiview datasets demonstrate the effectiveness and efficiency of the proposed CSRF method. Jie Chen 0065, Hua Mao 0001, Dezhong Peng, Changqing Zhang 0002, Xi Peng 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Low-Rank Tensor Learning for Incomplete Multiview ClusteringabstractIncomplete multiview clustering (IMVC) is an effective way to identify the underlying structure of incomplete multiview data. Most existing algorithms based on matrix factorization, graph learning or subspace learning have at least one of the following limitations: (1) the global and local structures of high-dimensional data are not effectively explored simultaneously; (2) the high-order correlations among multiple views are ignored. In this article, we propose a low-rank tensor learning (LRTL) method that learns a consensus low-dimensional embedding matrix for IMVC. We first take advantage of the self-expressiveness property of high-dimensional data to construct sparse similarity matrices for individual views under low-rank and sparsity constraints. Individual low-dimensional embedding matrices can be obtained from the sparse similarity matrices using spectral embedding techniques. This approach simultaneously explores the global and local structures of incomplete multiview data. Then, we present a multiview embedding matrix fusion model that incorporates individual low-dimensional embedding matrices into a third-norm tensor to achieve a consensus low-dimensional embedding matrix. The fusion model exploits complementary information by finding the high-order correlations among multiple views. In addition, the computational cost of an improved fusion strategy is dramatically reduced. Extensive experimental results demonstrate that the proposed LRTL method outperforms several state-of-the-art approaches. Jie Chen 0065, Zhu Wang 0007, Hua Mao 0001, Xi Peng 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Efficient Sparse Representation for Learning With High-Dimensional DataabstractDue to the capability of effectively learning intrinsic structures from high-dimensional data, techniques based on sparse representation have begun to display an impressive impact on several fields, such as image processing, computer vision, and pattern recognition. Learning sparse representations isoften computationally expensive due to the iterative computations needed to solve convex optimization problems in which the number of iterations is unknown before convergence. Moreover, most sparse representation algorithms focus only on determining the final sparse representation results and ignore the changes in the sparsity ratio (SR) during iterative computations. In this article, two algorithms are proposed to learn sparse representations based on locality-constrained linear representation learning with probabilistic simplex constraints. Specifically, the first algorithm, called approximated local linear representation (ALLR), obtains a closed-form solution from individual locality-constrained sparse representations. The second algorithm, called ALLR with symmetric constraints (ALLRSC), further obtains a symmetric sparse representation result with a limited number of computations; notably, the sparsity and convergence of sparse representations can be guaranteed based on theoretical analysis. The steady decline in the SR during iterative computations is a critical factor in practical applications. Experimental results based on public datasets demonstrate that the proposed algorithms perform better than several state-of-the-art algorithms for learning with high-dimensional data. Jie Chen 0065, Shengxiang Yang, Zhu Wang 0007, Hua Mao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Multi-view representation learning for data stream clusteringabstractData stream clustering provides valuable insights into the evolving patterns of long sequences of continuously generated data objects. Most existing clustering methods focus on single-view data streams. In this paper, we propose a multi-view representation learning (MVRL) method for multi-view clustering of data streams. We first introduce an integrated representation learning model to learn a fused sparse affinity matrix across multiple views for spectral clustering. Motivated by the optimization procedure of the integrated representation learning model, we propose three consecutive stages: collaborative representation, the construction of individual global affinity matrices using a mapping function, and the calculation of a fused sparse affinity matrix using Euclidean projection. These stages allow the effective capture of the global and local structures of high-dimensional data objects. Moreover, each stage has a closed-form solution, which determines the upper bound of the computational cost and memory consumption. We then employ the construction residuals of the collaborative representation to adaptively update a dynamic set, which is used to preserve the representative data objects. The dynamic set efficiently transfers previously learned useful knowledge to the arriving data objects. Extensive experimental results on multi-view data stream datasets demonstrate the effectiveness of the proposed MVRL method. Jie Chen 0065, Shengxiang Yang, Zhu Wang 0007 |
Inf. Sci. | 1 |
| 2022 | Multiview Subspace Clustering Using Low-Rank RepresentationabstractMultiview subspace clustering is one of the most widely used methods for exploiting the internal structures of multiview data. Most previous studies have performed the task of learning multiview representations by individually constructing an affinity matrix for each view without simultaneously exploiting the intrinsic characteristics of multiview data. In this article, we propose a multiview low-rank representation (MLRR) method to comprehensively discover the correlation of multiview data for multiview subspace clustering. MLRR considers symmetric low-rank representations (LRRs) to be an approximately linear spatial transformation under the new base, that is, the multiview data themselves, to fully exploit the angular information of the principal directions of LRRs, which is adopted to construct an affinity matrix for multiview subspace clustering, under a symmetric condition. MLRR takes full advantage of LRR techniques and a diversity regularization term to exploit the diversity and consistency of multiple views, respectively, and this method simultaneously imposes a symmetry constraint on LRRs. Hence, the angular information of the principal directions of rows is consistent with that of columns in symmetric LRRs. The MLRR model can be efficiently calculated by solving a convex optimization problem. Moreover, we present an intuitive fusion strategy for symmetric LRRs from the perspective of spectral clustering to obtain a compact representation, which can be shared by multiple views and comprehensively represents the intrinsic features of multiview data. Finally, the experimental results based on benchmark datasets demonstrate the effectiveness and robustness of MLRR compared with several state-of-the-art multiview subspace clustering algorithms. Jie Chen 0065, Shengxiang Yang, Hua Mao 0001, Conor Fahy |
IEEE Trans. Cybern. | 1 |
| 2021 | Low-rank representation with adaptive dictionary learning for subspace clustering
Jie Chen 0065, Hua Mao 0001, Zhu Wang 0007, Xinpei Zhang |
Knowl. Based Syst. | 1 |
| 2018 | Symmetric low-rank preserving projections for subspace learningabstractGraph construction plays an important role in graph-oriented subspace learning. However, most existing approaches cannot simultaneously consider the global and local structures of high-dimensional data. In order to solve this deficiency, we propose a symmetric low-rank preserving projection (SLPP) framework incorporating a symmetric constraint and a local regularization into low-rank representation learning for subspace learning. Under this framework, SLPP-M is incorporated with manifold regularization as its local regularization while SLPP-S uses sparsity regularization. Besides characterizing the global structure of high-dimensional data by a symmetric low-rank representation, both SLPP-M and SLPP-S effectively exploit the local manifold and geometric structure by incorporating manifold and sparsity regularization, respectively. The similarity matrix is successfully learned by solving the nuclear-norm minimization optimization problem . Combined with graph embedding techniques, a transformation matrix effectively preserves the low-dimensional structure features of high-dimensional data. In order to facilitate classification by exploiting available labels of training samples , we also develop a supervised version of SLPP-M and SLPP-S under the SLPP framework, named S-SLPP-M and S-SLPP-S, respectively. Experimental results in face, handwriting and object recognition applications demonstrate the efficiency of the proposed algorithm for subspace learning. Jie Chen 0065, Hua Mao 0001, Haixian Zhang, Zhang Yi 0001 |
Neurocomputing | 1 |
| 2017 | Subspace clustering using a symmetric low-rank representation
Jie Chen 0065, Hua Mao 0001, Yongsheng Sang, Zhang Yi 0001 |
Knowl. Based Syst. | 1 |
| 2016 | Symmetric low-rank representation for subspace clustering
Jie Chen 0065, Haixian Zhang, Hua Mao 0001, Yongsheng Sang, Zhang Yi 0001 |
Neurocomputing | 1 |
| 2014 | Sparse representation for face recognition by discriminative low-rank matrix recovery
Jie Chen 0065, Zhang Yi 0001 |
J. Vis. Commun. Image Represent. | 1 |