VLDB 2026 Research / reviewers in the wild / expert
Pei Zhang 0008
dblp:78/5323-8
· DBLP profile ↗
23ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-3018-5080ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JetBGC: Joint Robust Embedding and Structural Fusion Bipartite Graph Clustering (Extended Abstract)
Yuangang Pan, Junpu Zhang, Pei Zhang 0008, Xinwang Liu 0002, Kenli Li 0001, Ivor W. Tsang, Keqin Li 0001 |
ICDE | 4 |
| 2025 | Max-Mahalanobis Anchors Guidance for Multi-View ClusteringabstractAnchor selection or learning has become a critical component in large-scale multi-view clustering. Existing anchor-based methods, which either select-then-fix or initialize-then-optimize with orthogonality, yield promising performance. However, these methods still suffer from instability of initialization or insufficient depiction of data distribution. Moreover, the desired properties of anchors in multi-view clustering remain unspecified. To address these issues, this paper first formalizes the desired characteristics of anchors, namely Diversity, Balance and Compactness. We then devise and mathematically validate anchors that satisfy these properties by maximizing the Mahalanobis distance between anchors. Furthermore, we introduce a novel method called Max-Mahalanobis Anchors Guidance for multi-view Clustering (MAGIC), which guides the cross-view representations to progressively align with our well-defined anchors. This process yields highly discriminative and compact representations, significantly enhancing the performance of multi-view clustering. Experimental results show that our meticulously designed strategy significantly outperforms existing anchor-based methods in enhancing anchor efficacy, leading to substantial improvement in multi-view clustering performance. Pei Zhang 0008, Yuangang Pan, Siwei Wang 0001, Shengju Yu, En Zhu, Xinwang Liu 0002, Ivor W. Tsang |
AAAI | 1 |
| 2025 | Simple yet Effective Incomplete Multi-view Clustering: Similarity-level Imputation and Intra-view Hybrid-group Prototype ConstructionabstractMost of incomplete multi-view clustering (IMVC) methods typically choose to ignore the missing samples and only utilize observed unpaired samples to construct bipartite similarity. Moreover, they employ a single quantity of prototypes to extract the information of $\textbf{all}$ views. To eliminate these drawbacks, we present a simple yet effective IMVC approach, SIIHPC, in this work. It firstly transforms partial bipartition learning into original sample form by virtue of reconstruction concept to split out of observed similarity, and then loosens traditional non-negative constraints via regularizing samples to more freely characterize the similarity. Subsequently,
it learns to recover the incomplete parts by utilizing the connection built between the similarity exclusive on respective view and the consensus graph shared for all views. On this foundation, it further introduces a group of hybrid prototype quantities for each individual view to flexibly extract the data features belonging to each view itself. Accordingly, the resulting graphs are with various scales and describe the overall similarity more comprehensively. It is worth mentioning that these all are optimized in one unified learning framework,
which makes it possible for them to reciprocally promote. Then, to effectively solve the formulated optimization problem, we design an ingenious auxiliary function that is with theoretically proven monotonic-increasing properties. Finally, the clustering results are obtained by implementing spectral grouping action on the eigenvectors of stacked multi-scale consensus similarity. Experimental results confirm the effectiveness of SIIHPC. Shengju Yu, Zhibin Dong, Siwei Wang 0001, Pei Zhang 0008, Yi Zhang 0104, Xinwang Liu 0002, Naiyang Guan, Yiu-Ming Cheung |
ICLR | 4 |
| 2025 | Bit-swapping Oriented Twin-memory Multi-view Clustering in Lifelong Incomplete ScenariosabstractAlthough receiving notable improvements, current multi-view clustering (MVC) techniques generally rely on feature library mechanisms to propagate accumulated knowledge from historical views to newly-arrived data, which overlooks the information pertaining to basis embedding within each view. Moreover, the mapping paradigm inevitably alters the values of learned landmarks and built affinities due to the uninterruption nature, accordingly disarraying the hierarchical cluster structures. To mitigate these two issues, we in the paper provide a named BSTM algorithm. Concretely, we firstly synchronize with the distinct dimensions by introducing a group of specialized projectors, and then establish unified anchors for all views collected so far to capture intrinsic patterns.
Afterwards, departing from per-view architectures, we devise a shared bipartite graph construction via indicators to quantify similarity, which not only avoids redundant data-recalculations but alleviates the representation distortion caused by fusion.
Crucially, there two components are optimized within an integrated framework, and collectively facilitate knowledge transfer upon encountering incoming views. Subsequently, to flexibly do transformation on anchors and meanwhile maintain numerical consistency, we develop a bit-swapping scheme operating exclusively on 0 and 1. It harmonizes anchors on current view and that on previous views through one-hot encoded row and column attributes, and the graph structures are correspondingly reordered to reach a matched configuration. Furthermore, a computationally efficient four-step updating strategy with linear complexity is designed to minimize the associated loss. Extensive experiments organized on publicly-available benchmark datasets with varying missing percentages confirm the superior effectiveness of our BSTM. Shengju Yu, Pei Zhang 0008, Siwei Wang 0001, Suyuan Liu, Xinhang Wan, Zhibin Dong, Xinwang Liu 0002 |
NeurIPS | 2 |
| 2025 | JetBGC: Joint Robust Embedding and Structural Fusion Bipartite Graph ClusteringabstractBipartite graph clustering (BGC) has emerged as a fast-growing research in the clustering community. Despite BGC has achieved promising scalability, most variants still suffer from the following concerns: a) Susceptibility to noisy features. They construct bipartite graphs in the raw feature space, inducing poor robustness to noisy features. b) Inflexible anchor selection strategies. They usually select anchors through heuristic sampling or constrained learning methods, degrading flexibility. c) Partial structure mining. Existing methods are mainly built upon Linear Reconstruction Paradigm (LRP) from subspace clustering or Locally Linear Paradigm (LLP) from manifold learning, which partially exploit linear or locally linear structures, lacking a unified perspective to integrate global complementary structures. To this end, we propose a novel model, termedJoint Robust Embedding and Structural FusionBipartiteGraphClustering (JetBGC), which focuses on three aspects, namely robustness, flexibility, and complementarity. Concretely, we first introduce a robust embedding learning module to extract latent representation that can reduce the impact of noisy features. Then, we optimize anchors via a constraint-free strategy that can flexibly capture data distribution. Furthermore, we revisit the consistency and specificity of LRP and LLP, and design a new unified structural fusion strategy to integrate both linear and locally linear structures from a global perspective. Therefore, JetBGC unifies robust representation learning, flexible anchor optimization, and structural bipartite graph fusion in a framework. Extensive experiments on synthetic and real-world datasets validate our effectiveness against existing baselines. Liang Li 0041, Yuangang Pan, Junpu Zhang, Pei Zhang 0008, Jie Liu 0002, Xinwang Liu 0002, Kenli Li 0001, Ivor W. Tsang, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | TFMKC: Tuning-Free Multiple Kernel Clustering Coupled With Diverse Partition FusionabstractClustering is a popular research pipeline in unsupervised learning to find potential groupings. As a representative paradigm in multiple kernel clustering (MKC), late fusion-based models learn a consistent partition across multiple base kernels. Despite their promising performance, a common concern is the limited representation capacity caused by the inflexible fusion mechanism. Concretely, the representations are constrained by truncated-k Eigen-decomposition (EVD) without fully exploiting potential information. An intuitive idea to alleviate this concern is to generate a set of augmented partitions and then select the optimal partition by fine-tuning. However, this is overlimited by: 1) introducing undesired hyperparameters and dataset-related consequences; 2) neglecting rich information across diverse partitions; and 3) expensive parameter-tuning costs. To address these problems, we propose transforming the challenging problem of directly determining the optimal partition (optimal parameter) into a diverse partition fusion (parameter ensemble) problem. We design a novel flexible fusion mechanism called tuning-free multiple kernel clustering coupled with diverse partition fusion (TFMKC) by reweighting diverse partitions through optimization, achieving an optimal consensus partition by integrating diverse and complementary information rather than traditional fine-tuning, and distinguishing our work from existing methods. Extensive experiments verify that TFMKC achieves competitive effectiveness and efficiency over comparison baselines. The code can be accessed at https://github.com/ZJP/TFMKC. Junpu Zhang, Liang Li 0041, Pei Zhang 0008, Yue Liu 0008, Siwei Wang 0001, Changbao Zhou, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | DVSAI: Diverse View-Shared Anchors Based Incomplete Multi-View ClusteringabstractIn numerous real-world applications, it is quite common that sample information is partially available for some views due to machine breakdown or sensor failure, causing the problem of incomplete multi-view clustering (IMVC). While several IMVC approaches using view-shared anchors have successfully achieved pleasing performance improvement, (1) they generally construct anchors with only one dimension, which could deteriorate the multi-view diversity, bringing about serious information loss; (2) the constructed anchors are typically with a single size, which could not sufficiently characterize the distribution of the whole samples, leading to limited clustering performance. For generating view-shared anchors with multi-dimension and multi-size for IMVC, we design a novel framework called Diverse View-Shared Anchors based Incomplete multi-view clustering (DVSAI). Concretely, we associate each partial view with several potential spaces. In each space, we enable anchors to communicate among views and generate the view-shared anchors with space-specific dimension and size. Consequently, spaces with various scales make the generated view-shared anchors enjoy diverse dimensions and sizes. Subsequently, we devise an integration scheme with linear computational and memory expenditures to integrate the outputted multi-scale unified anchor graphs such that running spectral algorithm generates the spectral embedding. Afterwards, we theoretically demonstrate that DVSAI owns linear time and space costs, thus well-suited for tackling large-size datasets. Finally, comprehensive experiments confirm the effectiveness and advantages of DVSAI. Shengju Yu, Siwei Wang 0001, Pei Zhang 0008, Zhe Liu 0001, Liming Fang 0001, En Zhu, Xinwang Liu 0002 |
AAAI | 3 |
| 2024 | One-Step Late Fusion Multi-View Clustering with Compressed SubspaceabstractLate fusion multi-view clustering (LFMVC) has become a rapidly growing class of methods in the multi-view clustering (MVC) field, owing to its excellent computational speed and clustering performance. One bottleneck faced by existing late fusion methods is that they are usually aligned to the average kernel function, which makes the clustering performance highly dependent on the quality of datasets. Another problem is that they require subsequent k-means clustering after obtaining the consensus partition matrix to get the final discrete labels, and the resulting separation of the label learning and cluster structure optimization processes limits the integrity of these models. To address the above issues, we propose an integrated framework named One-Step Late Fusion Multi-view Clustering with Compressed Subspace (OS-LFMVC-CS). Specifically, we use the consensus subspace to align the partition matrix while optimizing the partition fusion, and utilize the fused partition matrix to guide the learning of discrete labels. A six-step iterative optimization approach with verified convergence is proposed. Sufficient experiments on multiple datasets validate the effectiveness and efficiency of our proposed method. Qiyuan Ou, Pei Zhang 0008, Sihang Zhou 0001, En Zhu |
ICASSP | 2 |
| 2024 | Towards Resource-friendly, Extensible and Stable Incomplete Multi-view ClusteringabstractIncomplete multi-view clustering (IMVC) methods typically encounter three drawbacks: (1) intense time and/or space overheads; (2) intractable hyper-parameters; (3) non-zero variance results. With these concerns in mind, we give a simple yet effective IMVC scheme, termed as ToRES. Concretely, instead of self-expression affinity, we manage to construct prototype-sample affinity for incomplete data so as to decrease the memory requirements. To eliminate hyper-parameters, besides mining complementary features among views by view-wise prototypes, we also attempt to devise cross-view prototypes to capture consensus features for jointly forming high-quality clustering representation. To avoid the variance, we successfully unify representation learning and clustering operation, and directly optimize the discrete cluster indicators from incomplete data. Then, for the resulting objective function, we provide two equivalent solutions from perspectives of feasible region partitioning and objective transformation. Many results suggest that ToRES exhibits advantages against 20 SOTA algorithms, even in scenarios with a higher ratio of incomplete data. Shengju Yu, Zhibin Dong, Siwei Wang 0001, Xinhang Wan, Yue Liu 0008, Weixuan Liang, Pei Zhang 0008, Wenxuan Tu, Xinwang Liu 0002 |
ICML | 7 |
| 2024 | Differentiated Anchor Quantity Assisted Incomplete Multiview Clustering Without Number-TuningabstractIncomplete multiview clustering (IMVC) generally requires the number of anchors to be the same in all views. Also, this number needs to be tuned with extra manual efforts. This not only degenerates the diversity of multiview data but also limits the model's scalability. For generating differentiated numbers of anchors without tuning, in this article we devise a novel framework named DAQINT. To be specific, the most perfect solution is to jointly find the optimal number of anchors that belongs to respective view. Regretfully, it is extremely time consuming. In view of this, we choose to first offer a set of anchor numbers for each view, and then integrate their contributions by adaptive weighting to approximate the optimal number. In particular, these offered numbers are all predefined and do not require any tuning. Through adaptively weighting them, we hold that this equivalently makes each view enjoy a different number of anchors. Accordingly, the bipartite graphs generated on all views are with diverse scales. Besides exploring multiview features more deeply, they also balance the importance between views. Then, to fuse these multiscale bipartite graphs, we design a combination strategy that owns linear computation and storage overheads. Afterward, to solve the resulting optimization problem, we also carefully develop a three-step iterative algorithm with linear complexities and demonstrated convergence. Experiments on the multiple public datasets validate the superiority of DAQINT against several advanced IMVC methods, such as on Mfeat, DAQINT surpasses the competitors like MKC, EEIMVC, FLSD, DSIMVC, IMVC-CBG, and DCP by 36.65%, 6.33%, 48.53%, 22.46%, 15.06%, and 32.04%, respectively, in ACC. Shengju Yu, Pei Zhang 0008, Siwei Wang 0001, Zhibin Dong, Hengfu Yang, En Zhu, Xinwang Liu 0002 |
IEEE Trans. Cybern. | 2 |
| 2023 | Let the Data Choose: Flexible and Diverse Anchor Graph Fusion for Scalable Multi-View ClusteringabstractIn the past few years, numerous multi-view graph clustering algorithms have been proposed to enhance the clustering performance by exploring information from multiple views. Despite the superior performance, the high time and space expenditures limit their scalability. Accordingly, anchor graph learning has been introduced to alleviate the computational complexity. However, existing approaches can be further improved by the following considerations: (i) Existing anchor-based methods share the same number of anchors across views. This strategy violates the diversity and flexibility of multi-view data distribution. (ii) Searching for the optimal anchor number within hyper-parameters takes much extra tuning time, which makes existing methods impractical. (iii) How to flexibly fuse multi-view anchor graphs of diverse sizes has not been well explored in existing literature. To address the above issues, we propose a novel anchor-based method termed Flexible and Diverse Anchor Graph Fusion for Scalable Multi-view Clustering (FDAGF) in this paper. Instead of manually tuning optimal anchor with massive hyper-parameters, we propose to optimize the contribution weights of a group of pre-defined anchor numbers to avoid extra time expenditure among views. Most importantly, we propose a novel hybrid fusion strategy for multi-size anchor graphs with theoretical proof, which allows flexible and diverse anchor graph fusion. Then, an efficient linear optimization algorithm is proposed to solve the resultant problem. Comprehensive experimental results demonstrate the effectiveness and efficiency of our proposed framework. The source code is available at https://github.com/Jeaninezpp/FDAGF. Pei Zhang 0008, Siwei Wang 0001, Liang Li 0041, Changwang Zhang, Xinwang Liu 0002, En Zhu, Zhe Liu 0001, Lu Zhou 0002, Lei Luo 0002 |
AAAI | 1 |
| 2023 | Graph Anomaly Detection via Multi-Scale Contrastive Learning Networks with Augmented ViewabstractGraph anomaly detection (GAD) is a vital task in graph-based machine learning and has been widely applied in many real-world applications. The primary goal of GAD is to capture anomalous nodes from graph datasets, which evidently deviate from the majority of nodes. Recent methods have paid attention to various scales of contrastive strategies for GAD, i.e., node-subgraph and node-node contrasts. However, they neglect the subgraph-subgraph comparison information which the normal and abnormal subgraph pairs behave differently in terms of embeddings and structures in GAD, resulting in sub-optimal task performance. In this paper, we fulfill the above idea in the proposed multi-view multi-scale contrastive learning framework with subgraph-subgraph contrast for the first practice. To be specific, we regard the original input graph as the first view and generate the second view by graph augmentation with edge modifications. With the guidance of maximizing the similarity of the subgraph pairs, the proposed subgraph-subgraph contrast contributes to more robust subgraph embeddings despite of the structure variation. Moreover, the introduced subgraph-subgraph contrast cooperates well with the widely-adopted node-subgraph and node-node contrastive counterparts for mutual GAD performance promotions. Besides, we also conduct sufficient experiments to investigate the impact of different graph augmentation approaches on detection performance. The comprehensive experimental results well demonstrate the superiority of our method compared with the state-of-the-art approaches and the effectiveness of the multi-view subgraph pair contrastive strategy for the GAD task. The source code is released at https://github.com/FelixDJC/GRADATE. Jingcan Duan, Siwei Wang 0001, Pei Zhang 0008, En Zhu, Jingtao Hu, Hu Jin 0005, Yue Liu 0008, Zhibin Dong |
AAAI | 3 |
| 2023 | Consensus One-step Multi-view Subspace Clustering (Extended abstract)abstractMulti-view clustering has attracted increasing attention in data mining communities. Despite superior clustering performance, we observe that existing multi-view subspace clustering methods directly fuse multi-view information in the similarity level by merging noisy affinity matrices; and isolate the processes of affinity learning, multiple information fusion and clustering. Both factors may cause insufficient utilization of multi-view information, leading to unsatisfying clustering performance. This paper proposes a novel consensus one-step multi-view subspace clustering (COMVSC) method to address these issues. Instead of directly fusing affinity matrices, COMVSC optimally integrates discriminative partition-level information, which is helpful in eliminating noise among data. Moreover, the affinity matrices, consensus representation and final clustering labels are learned simultaneously in a unified framework. Extensive experiment results on benchmark datasets demonstrate the superiority of our method over other state-of-the-art approaches. Pei Zhang 0008, Xinwang Liu 0002, Jian Xiong 0002, Sihang Zhou 0001, En Zhu, Zhiping Cai |
ICDE | 1 |
| 2023 | Normality Learning-based Graph Anomaly Detection via Multi-Scale Contrastive LearningabstractGraph anomaly detection (GAD) has attracted increasing attention in machine learning and data mining. Recent works have mainly focused on how to capture richer information to improve the quality of node embeddings for GAD. Despite their significant advances in detection performance, there is still a relative dearth of research on the properties of the task. GAD aims to discern the anomalies that deviate from most nodes. However, the model is prone to learn the pattern of normal samples which make up the majority of samples. Meanwhile, anomalies can be easily detected when their behaviors differ from normality. Therefore, the performance can be further improved by enhancing the ability to learn the normal pattern. To this end, we propose a normality learning-based GAD framework via multi-scale contrastive learning networks (NLGAD for abbreviation). Specifically, we first initialize the model with the contrastive networks on different scales. To provide sufficient and reliable normal nodes for normality learning, we design an effective hybrid strategy for normality selection. Finally, the model is refined with the only input of reliable normal nodes and learns a more accurate estimate of normality so that anomalous nodes can be more easily distinguished. Eventually, extensive experiments on six benchmark graph datasets demonstrate the effectiveness of our normality learning-based scheme on GAD. Notably, the proposed algorithm improves the detection performance (up to 5.89% AUC gain) compared with the state-of-the-art methods. The source code is released at https://github.com/FelixDJC/NLGAD. Jingcan Duan, Pei Zhang 0008, Siwei Wang 0001, Jingtao Hu, Hu Jin 0005, Jiaxin Zhang 0030, Haifang Zhou, Xinwang Liu 0002 |
ACM Multimedia | 2 |
| 2023 | Efficient Multi-View Graph Clustering with Local and Global Structure PreservationabstractAnchor-based multi-view graph clustering (AMVGC) has received abundant attention owing to its high efficiency and the capability to capture complementary structural information across multiple views. Intuitively, a high-quality anchor graph plays an essential role in the success of AMVGC. However, the existing AMVGC methods only consider single-structure information, i.e., local or global structure, which provides insufficient information for the learning task. To be specific, the over-scattered global structure leads to learned anchors failing to depict the cluster partition well. In contrast, the local structure with an improper similarity measure results in potentially inaccurate anchor assignment, ultimately leading to sub-optimal clustering performance. To tackle the issue, we propose a novel anchor-based multi-view graph clustering framework termed Efficient Multi-View Graph Clustering with Local and Global Structure Preservation (EMVGC-LG). Specifically, a unified framework with a theoretical guarantee is designed to capture local and global information. Besides, EMVGC-LG jointly optimizes anchor construction and graph learning to enhance the clustering quality. In addition, EMVGC-LG inherits the linear complexity of existing AMVGC methods respecting the sample number, which is time-economical and scales well with the data size. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method. Yi Wen 0001, Suyuan Liu, Xinhang Wan, Siwei Wang 0001, Ke Liang 0006, Xinwang Liu 0002, Xihong Yang, Pei Zhang 0008 |
ACM Multimedia | 8 |
| 2022 | Efficient One-Pass Multi-View Subspace Clustering with Consensus AnchorsabstractMulti-view subspace clustering (MVSC) optimally integrates multiple graph structure information to improve clustering performance. Recently, many anchor-based variants are proposed to reduce the computational complexity of MVSC. Though achieving considerable acceleration, we observe that most of them adopt fixed anchor points separating from the subsequential anchor graph construction, which may adversely affect the clustering performance. In addition, post-processing is required to generate discrete clustering labels with additional time consumption. To address these issues, we propose a scalable and parameter-free MVSC method to directly output the clustering labels with optimal anchor graph, termed as Efficient One-pass Multi-view Subspace Clustering with Consensus Anchors (EOMSC-CA). Specially, we combine anchor learning and graph construction into a uniform framework to boost clustering performance. Meanwhile, by imposing a graph connectivity constraint, our algorithm directly outputs the clustering labels without any post-processing procedures as previous methods do. Our proposed EOMSC-CA is proven to be linear complexity respecting to the data size. The superiority of our EOMSC-CA over the effectiveness and efficiency is demonstrated by extensive experiments. Our code is publicly available at https://github.com/Tracesource/EOMSC-CA. Suyuan Liu, Siwei Wang 0001, Pei Zhang 0008, Kai Xu 0004, Xinwang Liu 0002, Changwang Zhang |
AAAI | 3 |
| 2022 | Fast Parameter-Free Multi-View Subspace Clustering With Consensus Anchor GuidanceabstractMulti-view subspace clustering has attracted intensive attention to effectively fuse multi-view information by exploring appropriate graph structures. Although existing works have made impressive progress in clustering performance, most of them suffer from the cubic time complexity which could prevent them from being efficiently applied into large-scale applications. To improve the efficiency, anchor sampling mechanism has been proposed to select vital landmarks to represent the whole data. However, existing anchor selecting usually follows the heuristic sampling strategy, e.g. k -means or uniform sampling. As a result, the procedures of anchor selecting and subsequent subspace graph construction are separated from each other which may adversely affect clustering performance. Moreover, the involved hyper-parameters further limit the application of traditional algorithms. To address these issues, we propose a novel subspace clustering method termed Fast Parameter-free Multi-view Subspace Clustering with Consensus Anchor Guidance (FPMVS-CAG). Firstly, we jointly conduct anchor selection and subspace graph construction into a unified optimization formulation. By this way, the two processes can be negotiated with each other to promote clustering quality. Moreover, our proposed FPMVS-CAG is proved to have linear time complexity with respect to the sample number. In addition, FPMVS-CAG can automatically learn an optimal anchor subspace graph without any extra hyper-parameters. Extensive experimental results on various benchmark datasets demonstrate the effectiveness and efficiency of the proposed method against the existing state-of-the-art multi-view subspace clustering competitors. These merits make FPMVS-CAG more suitable for large-scale subspace clustering. The code of FPMVS-CAG is publicly available at https://github.com/wangsiwei2010/FPMVS-CAG. Siwei Wang 0001, Xinwang Liu 0002, Xinzhong Zhu, Pei Zhang 0008, Yi Zhang 0104, En Zhu |
IEEE Trans. Image Process. | 4 |
| 2022 | Consensus One-Step Multi-View Subspace ClusteringabstractMulti-view clustering has attracted increasing attention in multimedia, machine learning and data mining communities. As one kind of the essential multi-view clustering algorithm, multi-view subspace clustering (MVSC) becomes more and more popular due to its strong ability to reveal the intrinsic low dimensional clustering structure hidden across views. Despite superior clustering performance in various applications, we observe that existing MVSC methodsdirectly fuse multi-view information in the similarity level by merging noisy affinity matrices; andisolate the processes of affinity learning, multi-view information fusion and clustering. Both factors may cause insufficient utilization of multi-view information, leading to unsatisfying clustering performance. This paper proposes a novel consensus one-step multi-view subspace clustering (COMVSC) method to address these issues. Instead of directly fusing multiple affinity matrices, COMVSC optimally integrates discriminative partition-level information, which is helpful to eliminate noise among data. Moreover, the affinity matrices, consensus representation and final clustering labels matrix are learned simultaneously in a unified framework. By doing so, the three steps can negotiate with each other to best serve the clustering task, leading to improved performance. Accordingly, we propose an iterative algorithm to solve the resulting optimization problem. Extensive experiment results on benchmark datasets demonstrate the superiority of our method against other state-of-the-art approaches. Pei Zhang 0008, Xinwang Liu 0002, Jian Xiong 0002, Sihang Zhou 0001, En Zhu, Zhiping Cai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Projective Multiple Kernel Subspace ClusteringabstractMultiple kernel subspace clustering (MKSC), as an important extension for handling multi-view non-linear subspace data, has shown notable success in a wide variety of machine learning tasks. The key objective of MKSC is to build a flexible and appropriate graph for clustering from the kernel space. However, existing MKSC methods apply a mechanism utilizing the kernel trick to the traditional self-expressive principle, where the similarity graphs are built on the respective high-dimensional (or even infinite) reproducing kernel Hilbert space (RKHS). Regarding this strategy, we argue that the original high-dimensional spaces usually include noise and unreliable similarity measures and, therefore, output a low-quality graph matrix, which degrades clustering performance. In this paper, inspired by projective clustering, we propose the utilization of a complementary similarity graph by fusing the multiple kernel graphs constructed in the low-dimensional partition space, termed projective multiple kernel subspace clustering (PMKSC). By incorporating intrinsic structures with multi-view data, PMKSC alleviates the noise and redundancy in the original kernel space and obtains high-quality similarity to uncover the underlying clustering structures. Furthermore, we design a three-step alternate algorithm with proven convergence to solve the proposed optimization problem. The experimental results on ten multiple kernel benchmark datasets validate the effectiveness of our proposed PMKSC, compared to the state-of-the-art multiple kernel and kernel subspace clustering methods, by a large margin. Our code is available athttps://github.com/MengjingSun/PMKSC-code. Mengjing Sun, Siwei Wang 0001, Pei Zhang 0008, Xinwang Liu 0002, Xifeng Guo 0001, Sihang Zhou 0001, En Zhu |
IEEE Trans. Multim. | 3 |
| 2021 | Self-Representation Subspace Clustering for Incomplete Multi-view DataabstractIncomplete multi-view clustering is an important research topic in multimedia where partial data entries of one or more views are missing. Current subspace clustering approaches mostly employ matrix factorization on the observed feature matrices to address this issue. Meanwhile, self-representation technique is left unexplored, since it explicitly relies on full data entries to construct the coefficient matrix, which is contradictory to the incomplete data setting. However, it is widely observed that self-representation subspace method enjoys a better clustering performance over the factorization based one. Therefore, we adapt it to incomplete data by jointly performing data imputation and self-representation learning. To the best of our knowledge, this is the first attempt in incomplete multi-view clustering literature. Besides, the proposed method is carefully compared with current advances in experiment with respect to different missing ratios, verifying its effectiveness. Jiyuan Liu 0003, Xinwang Liu 0002, Yi Zhang 0104, Pei Zhang 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Weixuan Liang, Siqi Wang 0001, Yuexiang Yang |
ACM Multimedia | 4 |
| 2021 | Scalable Multi-view Subspace Clustering with Unified AnchorsabstractMulti-view subspace clustering has received widespread attention to effectively fuse multi-view information among multimedia applications. Considering that most existing approaches' cubic time complexity makes it challenging to apply to realistic large-scale scenarios, some researchers have addressed this challenge by sampling anchor points to capture distributions in different views. However, the separation of the heuristic sampling and clustering process leads to weak discriminate anchor points. Moreover, the complementary multi-view information has not been well utilized since the graphs are constructed independently by the anchors from the corresponding views. To address these issues, we propose a Scalable Multi-view Subspace Clustering with Unified Anchors (SMVSC). To be specific, we combine anchor learning and graph construction into a unified optimization framework. Therefore, the learned anchors can represent the actual latent data distribution more accurately, leading to a more discriminative clustering structure. Most importantly, the linear time complexity of our proposed algorithm allows the multi-view subspace clustering approach to be applied to large-scale data. Then, we design a four-step alternative optimization algorithm with proven convergence. Compared with state-of-the-art multi-view subspace clustering methods and large-scale oriented methods, the experimental results on several datasets demonstrate that our SMVSC method achieves comparable or better clustering performance much more efficiently. The code of SMVSC is available at https://github.com/Jeaninezpp/SMVSC. Mengjing Sun, Pei Zhang 0008, Siwei Wang 0001, Sihang Zhou 0001, Wenxuan Tu, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 2 |
| 2021 | Multi-view Clustering via Deep Matrix Factorization and Partition AlignmentabstractMulti-view clustering (MVC) has been extensively studied to collect multiple source information in recent years. One typical type of MVC methods is based on matrix factorization to effectively perform dimension reduction and clustering. However, the existing approaches can be further improved with following considerations: i) The current one-layer matrix factorization framework cannot fully exploit the useful data representations. ii) Most algorithms only focus on the shared information while ignore the view-specific structure leading to suboptimal solutions. iii) The partition level information has not been utilized in existing work. To solve the above issues, we propose a novel multi-view clustering algorithm via deep matrix decomposition and partition alignment. To be specific, the partition representations of each view are obtained through deep matrix decomposition, and then are jointly utilized with the optimal partition representation for fusing multi-view information. Finally, an alternating optimization algorithm is developed to solve the optimization problem with proven convergence. The comprehensive experimental results conducted on six benchmark multi-view datasets clearly demonstrates the effectiveness of the proposed algorithm against the SOTA methods. The code address for this algorithm is https://github.com/ZCtalk/MVC-DMF-PA. Siwei Wang 0001, Jiyuan Liu 0003, Sihang Zhou 0001, Pei Zhang 0008, Xinwang Liu 0002, En Zhu, Changwang Zhang |
ACM Multimedia | 5 |
| 2021 | Improved autoencoder for unsupervised anomaly detectionabstractDeep autoencoder-based methods are the majority of deep anomaly detection. An autoencoder learning on training data is assumed to produce higher reconstruction error for the anomalous samples than the normal samples and thus can distinguish anomalies from normal data. However, this assumption does not always hold in practice, especially in unsupervised anomaly detection, where the training data is anomaly contaminated. We observe that the autoencoder generalizes so well on the training data that it can reconstruct both the normal data and the anomalous data well, leading to poor anomaly detection performance. Besides, we find that anomaly detection performance is not stable when using reconstruction error as anomaly score, which is unacceptable in the unsupervised scenario. Because there are no labels to guide on selecting a proper model. To mitigate these drawbacks for autoencoder-based anomaly detection methods, we propose an Improved AutoEncoder for unsupervised Anomaly Detection (IAEAD). Specifically, we manipulate feature space to make normal data points closer using anomaly detection-based loss as guidance. Different from previous methods, by integrating the anomaly detection-based loss and autoencoder's reconstruction loss, IAEAD can jointly optimize for anomaly detection tasks and learn representations that preserve the local data structure to avoid feature distortion. Experiments on five image data sets empirically validate the effectiveness and stability of our method. Zhen Cheng 0004, Siwei Wang 0001, Pei Zhang 0008, Siqi Wang 0001, Xinwang Liu 0002, En Zhu |
Int. J. Intell. Syst. | 3 |