VLDB 2026 Research / reviewers in the wild / expert
Junpu Zhang
dblp:276/4523
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-7455-1989ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JetBGC: Joint Robust Embedding and Structural Fusion Bipartite Graph Clustering (Extended Abstract)
Yuangang Pan, Junpu Zhang, Pei Zhang 0008, Xinwang Liu 0002, Kenli Li 0001, Ivor W. Tsang, Keqin Li 0001 |
ICDE | 3 |
| 2025 | Learning the Anchors with Similar Distributions to Original Data for Multi-view ClusteringabstractIn multi-view clustering (MVC), anchor technique is generally hailed as an effective means for filtering noise and improving computation efficiency. However, existing methods usually construct anchors via heuristic strategy, random sampling, or orthogonal learning, which overlook the distribution differences between anchors and original data, leading to anchors lacking structural characteristics. To generate the anchors that are with similar distributions to original data, in the paper we carefully devise a LASD algorithm from the perspective of optimal transport (OT). Concretely, we firstly design a Multi-View OT (MVOT) framework through complementary and consensus representation learning. Then, we theoretically demonstrate the convexity of MVOT using the positive semidefiniteness of its Hessian matrix, and accordingly the global optimal solution of each transport plan can be reached. Further, we establish the strong dual condition for MVOT by the relative interior. Based on dual programming, consequently, we successfully obtain the transport plan between anchors and original data for each view within linear computational complexity. Afterwards, the spectral clustering operation is employed on the consensus plan to produce the discrete cluster labels. Abundant experiments underscore that our learned anchors do well reflect the distributions of original data, and the generated clustering results outperform multiple strong MVC competitors, even under large-scale scenarios. The source code is available at https://github.com/junpuzhang/LASD. Junpu Zhang, Shengju Yu, Suyuan Liu, Siwei Wang 0001, Miaomiao Li 0001, Xinwang Liu 0002, En Zhu, Kunlun He |
ACM Multimedia | 1 |
| 2025 | Generalized Probabilistic Graphical Modeling for Multi-View Bipartite Graph ClusteringabstractMulti-view bipartite graph clustering (MVBGC) is an active pipeline in unsupervised learning to tackle the limited scalability issue of traditional graph clustering. Despite improved performance, numerous variants still fall under conventional modeling that plugs additional modules, which however induces increasingly intricate models and fails to reveal the inherent variable relationship. We make the first attempt to introduce probabilistic graphical models for modeling the multi-view bipartite graph clustering task, reformulating it as a maximum likelihood estimation (MLE) problem. Such a setting uncovers the underlying probabilistic correlations among the commonality, view-specific variables, and noisy components. By pruning redundancy and disturbance collectively referred to as noise, we prove that minimizing the total noise is an approximation of the lower bound of MLE for multi-view data observations. We further generalize the MLE setting with clustering-suited constraints, deriving a Generalized Probabilistic Graphical Modeling framework (GProM), achieving an interpretable, concise, and flexible MVBGC framework. Extensive experiments verify the effectiveness of our framework. Furthermore, statistical significance analysis reveals the effectiveness of different distribution assumptions, providing valuable insights for model design. Liang Li 0041, Yuangang Pan, Yinghua Yao, Junpu Zhang, Moyun Liu, Xueling Zhu, Xinwang Liu 0002, Kenli Li 0001, Ivor W. Tsang, Keqin Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | JetBGC: Joint Robust Embedding and Structural Fusion Bipartite Graph ClusteringabstractBipartite graph clustering (BGC) has emerged as a fast-growing research in the clustering community. Despite BGC has achieved promising scalability, most variants still suffer from the following concerns: a) Susceptibility to noisy features. They construct bipartite graphs in the raw feature space, inducing poor robustness to noisy features. b) Inflexible anchor selection strategies. They usually select anchors through heuristic sampling or constrained learning methods, degrading flexibility. c) Partial structure mining. Existing methods are mainly built upon Linear Reconstruction Paradigm (LRP) from subspace clustering or Locally Linear Paradigm (LLP) from manifold learning, which partially exploit linear or locally linear structures, lacking a unified perspective to integrate global complementary structures. To this end, we propose a novel model, termedJoint Robust Embedding and Structural FusionBipartiteGraphClustering (JetBGC), which focuses on three aspects, namely robustness, flexibility, and complementarity. Concretely, we first introduce a robust embedding learning module to extract latent representation that can reduce the impact of noisy features. Then, we optimize anchors via a constraint-free strategy that can flexibly capture data distribution. Furthermore, we revisit the consistency and specificity of LRP and LLP, and design a new unified structural fusion strategy to integrate both linear and locally linear structures from a global perspective. Therefore, JetBGC unifies robust representation learning, flexible anchor optimization, and structural bipartite graph fusion in a framework. Extensive experiments on synthetic and real-world datasets validate our effectiveness against existing baselines. Liang Li 0041, Yuangang Pan, Junpu Zhang, Pei Zhang 0008, Jie Liu 0002, Xinwang Liu 0002, Kenli Li 0001, Ivor W. Tsang, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | TFMKC: Tuning-Free Multiple Kernel Clustering Coupled With Diverse Partition FusionabstractClustering is a popular research pipeline in unsupervised learning to find potential groupings. As a representative paradigm in multiple kernel clustering (MKC), late fusion-based models learn a consistent partition across multiple base kernels. Despite their promising performance, a common concern is the limited representation capacity caused by the inflexible fusion mechanism. Concretely, the representations are constrained by truncated-k Eigen-decomposition (EVD) without fully exploiting potential information. An intuitive idea to alleviate this concern is to generate a set of augmented partitions and then select the optimal partition by fine-tuning. However, this is overlimited by: 1) introducing undesired hyperparameters and dataset-related consequences; 2) neglecting rich information across diverse partitions; and 3) expensive parameter-tuning costs. To address these problems, we propose transforming the challenging problem of directly determining the optimal partition (optimal parameter) into a diverse partition fusion (parameter ensemble) problem. We design a novel flexible fusion mechanism called tuning-free multiple kernel clustering coupled with diverse partition fusion (TFMKC) by reweighting diverse partitions through optimization, achieving an optimal consensus partition by integrating diverse and complementary information rather than traditional fine-tuning, and distinguishing our work from existing methods. Extensive experiments verify that TFMKC achieves competitive effectiveness and efficiency over comparison baselines. The code can be accessed at https://github.com/ZJP/TFMKC. Junpu Zhang, Liang Li 0041, Pei Zhang 0008, Yue Liu 0008, Siwei Wang 0001, Changbao Zhou, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Sample-Level Cross-View Similarity Learning for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering has attracted much attention due to its ability to handle partial multi-view data. Recently, similarity-based methods have been developed to explore the complete relationship among incomplete multi-view data. Although widely applied to partial scenarios, most of the existing approaches are still faced with two limitations. Firstly, fusing similarities constructed individually on each view fails to yield a complete unified similarity. Moreover, incomplete similarity generation may lead to anomalous similarity values with column sum constraints, affecting the final clustering results. To solve the above challenging issues, we propose a Sample-level Cross-view Similarity Learning (SCSL) method for Incomplete Multi-view Clustering. Specifically, we project all samples to the same dimension and simultaneously construct a complete similarity matrix across views based on the inter-view sample relationship and the intra-view sample relationship. In addition, a simultaneously learning consensus representation ensures the validity of the projection, which further enhances the quality of the similarity matrix through the graph Laplacian regularization. Experimental results on six benchmark datasets demonstrate the ability of SCSL in processing incomplete multi-view clustering tasks. Our code is publicly available at https://github.com/Tracesource/SCSL. Suyuan Liu, Junpu Zhang, Yi Wen 0001, Xihong Yang, Siwei Wang 0001, Yi Zhang 0104, En Zhu, Chang Tang, Long Zhao 0002, Xinwang Liu 0002 |
AAAI | 2 |
| 2024 | Reliable Attribute-missing Multi-view Clustering with Instance-level and feature-level Cooperative ImputationabstractMulti-view clustering (MVC) constitutes a distinct approach to data mining within the field of machine learning. Due to limitations in the data collection process, missing attributes are frequently encountered. However, existing MVC methods primarily focus on missing instances, showing limited attention to missing attributes. A small number of studies employ the reconstruction of missing instances to address missing attributes, potentially overlooking the synergistic effects between the instance and feature spaces, which could lead to distorted imputation outcomes. Furthermore, current methods uniformly treat all missing attributes as zero values, thus failing to differentiate between real and technical zeroes, potentially resulting in data over-imputation. To mitigate these challenges, we introduce a novel Reliable Attribute-Missing Multi-View Clustering method (RAM-MVC). Specifically, feature reconstruction is utilized to address missing attributes, while similarity graphs are simultaneously constructed within the instance and feature spaces. By leveraging structural information from both spaces, RAM-MVC learns a high-quality feature reconstruction matrix during the joint optimization process. Additionally, we introduce a reliable imputation guidance module that distinguishes between real and technical attribute-missing events, enabling discriminative imputation. The proposed RAM-MVC method outperforms nine baseline methods, as evidenced by real-world experiments using single-cell multi-view data. Dayu Hu, Suyuan Liu, Jun Wang 0118, Junpu Zhang, Siwei Wang 0001, Xingchen Hu 0001, Xinzhong Zhu, Chang Tang, Xinwang Liu 0002 |
ACM Multimedia | 4 |
| 2024 | Alleviate Anchor-Shift: Explore Blind Spots with Cross-View Reconstruction for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering aims to learn complete correlations among samples by leveraging complementary information across multiple views for clustering. Anchor-based methods further establish sample-level similarities for representative anchor generation, effectively addressing scalability issues in large-scale scenarios. Despite efficiency improvements, existing methods overlook the misguidance in anchors learning induced by partial missing samples, i.e., the absence of samples results in shift of learned anchors, further leading to sub-optimal clustering performance. To conquer the challenges, our solution involves a cross-view reconstruction strategy that not only alleviate the anchor shift problem through a carefully designed cross-view learning process, but also reconstructs missing samples in a way that transcends the limitations imposed by convex combinations. By employing affine combinations, our method explores areas beyond the convex hull defined by anchors, thereby illuminating blind spots in the reconstruction of missing samples. Experimental results on four benchmark datasets and three large-scale datasets validate the effectiveness of our proposed method. Suyuan Liu, Siwei Wang 0001, Ke Liang 0006, Junpu Zhang, Zhibin Dong, Tianrui Liu 0001, En Zhu, Xinwang Liu 0002, Kunlun He |
NeurIPS | 4 |
| 2024 | Symmetric Multi-View Subspace Clustering With Automatic Neighbor DiscoveryabstractMulti-view subspace clustering (MVSC) is a popular area of research that concentrates on partitioning data points from multiple views. It has gained wide attention in recent years due to the ability to handle complex data with diverse features across different views. However, the success of MVSC largely relies on the quality of the learned similarity matrix, and existing methods normally adopt the separate two-step procedures of optimization and symmetrization, which could not guarantee symmetry and adaptive locality of the similarity matrix. To alleviate this issue, in this paper, we propose a novel paradigm called Symmetric Multi-view Subspace Clustering with Automatic Neighbor Discovery (SMSC-AND), which aims at formulating the symmetrization and localization of the ideal similarity matrix into one unified framework. In particular, we theoretically and experimentally demonstrate that SMSC-AND can directly receive the refined symmetric similarity matrix without previous post-processing procedures. Additionally, we propose an automatic neighbor discovery strategy that avoids previous rank constraints or fixed neighbor size, thereby eliminating the requirement for additional hyperparameters. Benefiting from the aforementioned merits, we can directly explore the local structure of the consensus similarity matrix of multi-view data without pre-searching hyperparameters. Comprehensive experimental results on various benchmark datasets have demonstrated the superiority of the proposed algorithm when compared with other MVSC competitors. Siwei Wang 0001, Junpu Zhang, Shengju Yu, Suyuan Liu, Xinwang Liu 0002, Kunlun He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Multi-View Bipartite Graph Clustering With Coupled Noisy Feature FilterabstractUnsupervised bipartite graph learning has been a hotpot in multi-view clustering, to tackle the restricted scalability issue of traditional full graph clustering in large-scale applications. However, the existing bipartite graph clustering paradigm pays little attention to the adverse impact of noisy features on learning process. To further facilitate this part of research, apart from simply reweighting features to depress the noisy ones, we take the first step towards analyzing the induced adverse impact via theoretical and experimental investigations. One crucial finding in this paper is that the existence of noisy features will incur “anchor shift” phenomenon, which deviates the potential representations of anchors and then degrades performance. To this end, we propose a coupled noisy feature filter mechanism with automatically finding feature importance to remedy the anchor shift issue in this paper. Apart from leveraging features, we theoretically analyze the bounds of proposed feature-adaptive bipartite graph's fuzzy membership. Specifically, distinguishing features' discrimination will increase the fuzzy membership to achieve soft partitions against the potential inaccurate absolute relationship. With the afore-mentioned merits, our proposed multi-view bipartite graph clustering with coupled noisy feature filter model (MVBGC-NFF) provides novel and interesting insights on the feature level of anchor shift. The effectiveness and efficiency of MVBGC-NFF are demonstrated on synthetic and real-world datasets with improving clustering performance, increasing fuzzy membership, and filtering noisy features. The code is available onhttps://github.com/liliangnudt/MVBGC-NFF. Liang Li 0041, Junpu Zhang, Siwei Wang 0001, Xinwang Liu 0002, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Multiple Kernel Clustering with Dual Noise MinimizationabstractClustering is a representative unsupervised method widely applied in multi-modal and multi-view scenarios. Multiple kernel clustering (MKC) aims to group data by integrating complementary information from base kernels. As a representative, late fusion MKC first decomposes the kernels into orthogonal partition matrices, then learns a consensus one from them, achieving promising performance recently. However, these methods fail to consider the noise inside the partition matrix, preventing further improvement of clustering performance. We discover that the noise can be disassembled into separable dual parts, i.e. N-noise and C-noise (Null space noise and Column space noise). In this paper, we rigorously define dual noise and propose a novel parameter-free MKC algorithm by minimizing them. To solve the resultant optimization problem, we design an efficient two-step iterative strategy. To our best knowledge, it is the first time to investigate dual noise within the partition in the kernel space. We observe that dual noise will pollute the block diagonal structures and incur the degeneration of clustering performance, and C-noise exhibits stronger destruction than N-noise. Owing to our efficient mechanism to minimize dual noise, the proposed algorithm surpasses the recent methods by large margins. Junpu Zhang, Liang Li 0041, Siwei Wang 0001, Jiyuan Liu 0003, Yue Liu 0008, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 1 |