VLDB 2026 Research / reviewers in the wild / expert
Weixuan Liang
dblp:274/1152
· DBLP profile ↗
27ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0002-1868-5445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Kernel Power $K$-means: Scalable and Robust Clustering with Random Fourier Features and Possibilistic MethodabstractKernel power k-means (KPKM) leverages a family of means to mitigate local minima issues in kernel k-means. However, KPKM faces two key limitations: (1) the computational burden of the full kernel matrix restricts its use on extensive data, and (2) the lack of authentic centroid-sample assignment learning reduces its noise robustness. To overcome these challenges, we propose RFF-KPKM, introducing the first approximation theory for applying random Fourier features (RFF) to KPKM. RFF-KPKM employs RFF to generate efficient, low dimensional feature maps, bypassing the need for the whole kernel matrix. Crucially, we are the first to establish strong theoretical guarantees for this combination: (1) an excess risk bound of O( k^3/n), (2) strong consistency with membership values, and (3) a (1 + ε) relative error bound achievable using the RFF of dimension poly(ε^{−1} logk). Furthermore, to improve robustness and the ability to learn multiple kernels, we propose IP-RFF-MKPKM, an improved possibilistic RFF-based multiple kernel power k-means. IP-RFF-MKPKM ensures the scalability of MKPKM via RFF and refines cluster assignments by combining the merits of the possibilistic and fuzzy membership. Experiments on large-scale datasets demonstrate the superior efficiency and clustering accuracy of the proposed methods compared to the state-of-the-art alternatives. Weixuan Liang, Xueling Zhu, Xinwang Liu 0002 |
AAAI | 2 |
| 2025 | Incremental Nyström-based Multiple Kernel ClusteringabstractExisting Multiple Kernel Clustering (MKC) algorithms commonly utilize the Nyström method to handle large-scale datasets. However, most of them employ uniform sampling for kernel matrix approximation, hence failing to accurately capture the underlying data structure, leading to large approximation errors. Additionally, they often use the same landmark points for all kernel matrix approximations, reducing kernel diversity. Moreover, in scenarios where approximate kernel matrices emerge over time, these methods require storing historical kernel information and recalculating, resulting in inefficient resource utilization. To address these issues, we propose a novel MKC algorithm, termed Incremental Nyström-based Multiple Kernel Clustering (INMKC). Specifically, leverage score sampling is utilized to reduce kernel approximation errors and enhance kernel diversity. Furthermore, we employ a consensus clustering structure that aligns with the newly emerged base kernel matrix for updates, avoiding recalculating previous kernel matrices, thus saving substantial computational resources. Additionally, we tackle the challenge of aligning incremental approximate kernels with different landmark points. Extensive experiments on the proposed INMKC demonstrate its effectiveness and efficiency compared to state-of-the-art methods. Weixuan Liang, Xinhang Wan, Jiyuan Liu 0003, Suyuan Liu, Qian Qu, Renxiang Guan, Xinwang Liu 0002 |
AAAI | 2 |
| 2025 | Large-scale Multi-view Tensor Clustering with Implicit Linear KernelsabstractMulti-view clustering is a long-standing hot topic in machine learning communities, due to its capability of integrating data information from multiple sources and modalities. By utilizing tensor Singular Value Decomposition (t-SVD) technique with the tensor rotation trick, recent advances have achieved remarkable improvements on clustering performance. However, we find this is attributed to the inadvertent use of sequential information of sorted data samples, i.e. inadvertent label use, which violates the unsupervised learning setting. On the other hand, existing large-scale approaches are mostly developed on the basis of matrix factorization or anchor techniques, thereby fail to consider the similarities among all data samples, preventing from further performance improvement. To address the above issues, we first analyze the tensor rotation trick and recommend to remove it from tensor clustering. On its basis, a novel large-scale multi-view tensor clustering method is developed by incorporating the pair-wise similarities with implicit linear kernel function. To solve the resultant optimization problem, we design an efficient algorithm of linear complexity. Moreover, extensive experiments are conducted and corresponding results well support the aforementioned finding and validate the effectiveness and efficiency of the proposed method. Jiyuan Liu 0003, Xinwang Liu 0002, Chuankun Li, Xinhang Wan, Yi Zhang 0104, Weixuan Liang, Qian Qu, Renxiang Guan, Ke Liang 0006 |
CVPR | 7 |
| 2025 | On the Adversarial Robustness of Multi-Kernel ClusteringabstractMulti-kernel clustering (MKC) has emerged as a powerful method for capturing diverse data patterns, offering robust and generalized representations of data structures. However, the increasing deployment of MKC in real-world applications raises concerns about its vulnerability to adversarial perturbations. While adversarial robustness has been extensively studied in other domains, its impact on MKC remains largely unexplored. In this paper, we address the challenge of assessing the adversarial robustness of MKC methods in a black-box setting. Specifically, we propose *AdvMKC*, a novel reinforcement-learning-based adversarial attack framework designed to inject imperceptible perturbations into data and mislead MKC methods. AdvMKC leverages proximal policy optimization with an advantage function to overcome the instability of clustering results during optimization. Additionally, it introduces a generator-clusterer framework, where a generator produces adversarial perturbations, and a clusterer approximates MKC behavior, significantly reducing computational overhead. We provide theoretical insights into the impact of adversarial perturbations on MKC and validate these findings through experiments. Evaluations across seven datasets and eleven MKC methods (seven traditional and four robust) demonstrate AdvMKC's effectiveness, robustness, and transferability. Hao Yu 0017, Weixuan Liang, Ke Liang 0006, Suyuan Liu, Meng Liu 0014, Xinwang Liu 0002 |
ICML | 2 |
| 2025 | COKE: Core Kernel for More Efficient Approximation of Kernel Weights in Multiple Kernel ClusteringabstractInspired by the well-known coreset in clustering algorithms, we introduce the definition of the core kernel for multiple kernel clustering (MKC) algorithms. The core kernel refers to running MKC algorithms on smaller-scale base kernel matrices to obtain kernel weights similar to those obtained from the original full-scale kernel matrices. Specifically, the core kernel refers to a set of kernel matrices of size $\widetilde{\mathcal{O}}(1/\varepsilon^2)$ that perform MKC algorithms on them can achieve a $(1+\varepsilon)$-approximation for the kernel weights. Subsequently, we can leverage approximated kernel weights to obtain a theoretically guaranteed large-scale extension of MKC algorithms. In this paper, we propose a core kernel construction method based on singular value decomposition and prove that it satisfies the definition of the core kernel for three mainstream MKC algorithms. Finally, we conduct experiments on several benchmark datasets to verify the correctness of theoretical results and the efficiency of the proposed method. Weixuan Liang, Xinwang Liu 0002, Ke Liang 0006, Jiyuan Liu 0003, En Zhu |
ICML | 1 |
| 2025 | Generalization Performance of Ensemble Clustering: From Theory to AlgorithmabstractEnsemble clustering has demonstrated great success in practice; however, its theoretical foundations remain underexplored. This paper examines the generalization performance of ensemble clustering, focusing on generalization error, excess risk and consistency. We derive a convergence rate of generalization error bound and excess risk bound both of $\mathcal{O}(\sqrt{\frac{\log n}{m}}+\frac{1}{\sqrt{n}})$, with $n$ and $m$ being the numbers of samples and base clusterings. Based on this, we prove that when $m$ and $n$ approach infinity and $m$ is significantly larger than log $n$, i.e., $m,n\to \infty, m\gg \log n$, ensemble clustering is consistent. Furthermore, recognizing that $n$ and $m$ are finite in practice, the generalization error cannot be reduced to zero. Thus, by assigning varying weights to finite clusterings, we minimize the error between the empirical average clusterings and their expectation. From this, we theoretically demonstrate that to achieve better clustering performance, we should minimize the deviation (bias) of base clustering from its expectation and maximize the differences (diversity) among various base clusterings. Additionally, we derive that maximizing diversity is nearly equivalent to a robust (min-max) optimization model. Finally, we instantiate our theory to develop a new ensemble clustering algorithm. Compared with SOTA methods, our approach achieves average improvements of 6.0%, 7.3%, and 6.0% on 10 datasets w.r.t. NMI, ARI, and Purity. The code is available at https://github.com/xuz2019/GPEC. Haoye Qiu, Weixuan Liang, Hui Liu 0032, Junhui Hou, Yuheng Jia |
ICML | 3 |
| 2025 | Mixture of Experts for Node ClassificationabstractNodes in the real-world graphs exhibit diverse patterns in numerous aspects, such as degree and homophily. However, most existent node predictors fail to capture a wide range of node patterns or to make predictions based on distinct node patterns, resulting in unsatisfactory classification performance. In this paper, we reveal that different node predictors are good at handling nodes with specific patterns and only apply one node predictor uniformly could lead to suboptimal result. To mitigate this gap, we propose a mixture of experts framework, MoE-NP, for node classification. Specifically, MoE-NP combines a mixture of node predictors and strategically selects models based on node patterns. Experimental results from a range of real-world datasets demonstrate significant performance improvements from MoE-NP. Yiqi Wang 0001, Weixuan Liang, Jiaxin Zhang 0030, Pan Dong, Aiping Li |
ICMR | 3 |
| 2025 | Incomplete Multi-view Deep Clustering with Data Imputation and AlignmentabstractIncomplete multi-view deep clustering is an emerging research hot-pot to incorporate data information of multiple sources or modalities when parts of them are missing. Most of existing approaches encode the available data observations into multiple view-specific latent representations and subsequently integrate them for the next clustering task. However, they ignore that the latent representations are unique to a fixed set of data samples in all views. Meanwhile, the pair-wise similarities of missing data observations are also failed to utilize in latent representation learning sufficiently, leading to unsatisfactory clustering performance. To address these issues, we propose an incomplete multi-view deep clustering method with data imputation and alignment. Assuming that each data sample corresponds to a same latent representation among all views, it projects the latent representations into feature spaces with neural networks. As a result, not only the available data observations are reconstructed, but also the missing ones can be imputed accordingly. Moreover, a linear alignment measurement of linear complexity is defined to compute the pair-wise similarities of all data observations, especially including those of the missing. By executing the above two procedures iteratively, the discriminative latent representations can be learned and used to group the data into categories with off-the-shelf clustering algorithms. In experiment, the proposed method is validated on a set of benchmark datasets and achieves state-of-the-art performances. Jiyuan Liu 0003, Xinwang Liu 0002, Xinhang Wan, Ke Liang 0006, Weixuan Liang, Sihang Zhou 0001, Huijun Wu 0001, Kehua Guo |
NeurIPS | 5 |
| 2025 | Incremental Multi-View Clustering: Exploring Stream-View Correlations to Learn Consistency and DiversityabstractMulti-view clustering (MVC) has demonstrated impressive performance due to its ability to capture both consistency and diversity information among views. However, most existing techniques assume that all views are available in advance, making them inadequate for stream-view data, such as in intelligent transportation systems and medical imaging analysis, where memory constraints or privacy concerns prevent storing all previous views. Although some methods attempt to address this issue by capturing consistency information, they often fail to effectively extract diversity information and cross-view relationships. We argue that these limitations are inherent to incremental multi-view clustering (IMVC), as the inability to retain all previous views inevitably leads to insufficient information utilization, thereby compromising performance. To address these challenges, we propose a novel algorithm, termed Incremental Multi-View Clustering with Cross-View Correlation and Diversity (CDIMVC). Unlike existing methods that only retain consistency information, CDIMVC also preserves diversity information and utilizes similarity matrices to capture cross-view relationships. To implement this method, we develop three key modules: the dynamic view correlation analysis module (DVCAM), the knowledge extraction module (KEM), and the knowledge transfer module (KTM). When a new data view arrives, DVCAM first assesses its importance and correlation with historical views. Subsequently, KEM computes its consistency and diversity information by comparing it to those in the knowledge base. Finally, KTM facilitates the effective transmission of past knowledge, preventing the loss of historical information. By integrating these modules, CDIMVC can effectively capture cross-view relationships and diversity information, facilitating efficient knowledge updating and maintenance. An alternating procedure is also designed to optimize the resulting optimization problem. Experimental results show that CDIMVC exceeds state-of-the-art methods, demonstrating its effectiveness in handling stream-view data. Weixuan Liang, Xinhang Wan, Jiyuan Liu 0003, Miaomiao Li 0001, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Prototype-Driven Multi-View Attribute-Missing Graph ClusteringabstractAttribute-missing deep graph clustering, which aims to categorize the graph nodes with partial attribute-missing samples into distinct categories in an unsupervised manner, has gained significant popularity. However, most existing researches have at least one of the following issues: 1) seldom exploit diverse clustering structural information to facilitate non-Euclidean data imputation and refine the clustering pattern and 2) ignoring the positive effect of diverse information on feature imputation and representation extraction, resulting in sub-optimal missing feature estimation and inferior clustering performance. To solve these issues, we propose a novelPrototype-drivenMulti-viewAttribute-missingGraphClustering (PMAGC) model that leverages rich structural and diverse information to assist the processes of imputing missing attributes and learning clustering-friendly features. Specifically, we design a multi-view augmentation module that extracts attribute-complete samples as node view and constructs feature and edge views using feature pre-imputation and edge masking techniques. Then, guided by clustering pseudo-labels, we promote the proximity between the prototypes of attribute-missing samples and those of attribute-complete samples within the feature space. Thus, PMAGC cleverly employs both clustering structural information and reliably attribute-complete sample data to assist feature imputation. In addition, we design a prototype-wise contrastive loss, which considers prototypes from different views within the same cluster as positive samples, while treating others as negative samples. Hence, the optimized features could more accurately guide the attribute learning process. Extensive experiments on six graph datasets with missing attributes are conducted to demonstrate the effectiveness of the proposed PMAGE. Renxiang Guan, Wenxuan Tu, Dayu Hu, Weixuan Liang, Ke Liang 0006, Yaowen Hu, Yue Liu 0008, Xinwang Liu 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Scalable Multiple Kernel Clustering: Learning Clustering Structure from ExpectationabstractIn this paper, we derive an upper bound of the difference between a kernel matrix and its expectation under a mild assumption. Specifically, we assume that the true distribution of the training data is an unknown isotropic Gaussian distribution. When the kernel function is a Gaussian kernel, and the mean of each cluster is sufficiently separated, we find that the expectation of a kernel matrix can be close to a rank-$k$ matrix, where $k$ is the cluster number. Moreover, we prove that the normalized kernel matrix of the training set deviates (w.r.t. Frobenius norm) from its expectation in the order of $\widetilde{\mathcal{O}}(1/\sqrt{d})$, where $d$ is the dimension of samples. Based on the above theoretical results, we propose a novel multiple kernel clustering framework which attempts to learn the information of the expectation kernel matrices. First, we aim to minimize the distance between each base kernel and a rank-$k$ matrix, which is a proxy of the expectation kernel. Then, we fuse these rank-$k$ matrices into a consensus rank-$k$ matrix to find the clustering structure. Using an anchor-based method, the proposed framework is flexible with the sizes of input kernel matrices and able to handle large-scale datasets. We also provide the approximation guarantee by deriving two non-asymptotic bounds for the consensus kernel and clustering indicator matrices. Finally, we conduct extensive experiments to verify the clustering performance of the proposed method and the correctness of the proposed theoretical results. Weixuan Liang, En Zhu, Shengju Yu, Xinzhong Zhu, Xinwang Liu 0002 |
ICML | 1 |
| 2024 | Towards Resource-friendly, Extensible and Stable Incomplete Multi-view ClusteringabstractIncomplete multi-view clustering (IMVC) methods typically encounter three drawbacks: (1) intense time and/or space overheads; (2) intractable hyper-parameters; (3) non-zero variance results. With these concerns in mind, we give a simple yet effective IMVC scheme, termed as ToRES. Concretely, instead of self-expression affinity, we manage to construct prototype-sample affinity for incomplete data so as to decrease the memory requirements. To eliminate hyper-parameters, besides mining complementary features among views by view-wise prototypes, we also attempt to devise cross-view prototypes to capture consensus features for jointly forming high-quality clustering representation. To avoid the variance, we successfully unify representation learning and clustering operation, and directly optimize the discrete cluster indicators from incomplete data. Then, for the resulting objective function, we provide two equivalent solutions from perspectives of feasible region partitioning and objective transformation. Many results suggest that ToRES exhibits advantages against 20 SOTA algorithms, even in scenarios with a higher ratio of incomplete data. Shengju Yu, Zhibin Dong, Siwei Wang 0001, Xinhang Wan, Yue Liu 0008, Weixuan Liang, Pei Zhang 0008, Wenxuan Tu, Xinwang Liu 0002 |
ICML | 6 |
| 2024 | A Lightweight Anchor-Based Incremental Framework for Multi-view ClusteringabstractThe rapid development of multi-media techniques boosts the emergence of multi-view data, and how to uncover its intrinsic structure and utilize it to conduct the subsequent downstream tasks is crucial in data analysis. Multi-view clustering is representative of handling multi-view data. The anchor-based method has received widespread attention for excellent performance and low time complexity. However, existing methods encounter two drawbacks, cutting down their performance, i.e., the assumption of the availability of all views and limited interaction of anchor generation among views. In some scenes, views arrive sequentially, and storing them is challenging owing to the limited space/privacy considerations, and the existing anchor-based MVC is unsuitable for this. Additionally, recent works fail to generate anchors with the guidance of other views, and it is tough to align the anchor graphs. To this end, we propose A Lightweight Anchor-Based Incremental Framework for Multi-view Clustering. Specifically, we first initialize an anchor graph with the assistance of k-means when a new view arrives. Then, the consensus one of the anchor graph is updated by the newly collected view with a permutation matrix. Our proposed method is more capable of anchor alignment because, in incremental MVC, the anchor graphs of previous views could be listed as a reference to guide the generation of anchor graphs of the coming view. Furthermore, we design a three-step iterative and convergent algorithm to address the resultant problem. Notably, the proposed algorithm shows outstanding effectiveness and time/space efficiency in extensive experiments. Qian Qu, Xinhang Wan, Weixuan Liang, Jiyuan Liu 0003, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 3 |
| 2024 | On the Consistency and Large-Scale Extension of Multiple Kernel ClusteringabstractExisting multiple kernel clustering (MKC) algorithms have two ubiquitous problems. From the theoretical perspective, most MKC algorithms lack sufficient theoretical analysis, especially the consistency of learned parameters, such as the kernel weights. From the practical perspective, the high complexity makes MKC unable to handle large-scale datasets. This paper tries to address the above two issues. We first make a consistency analysis of an influential MKC method named Simple Multiple Kernel$k$-Means (SimpleMKKM). Specifically, suppose that$\hat{\boldsymbol{\gamma }}_{n}$are the kernel weights learned by SimpleMKKM from the training samples. We also define the expected version of SimpleMKKM and denote its solution as$\boldsymbol{\gamma }^*$. We establish an upper bound of$\Vert \hat{\boldsymbol{\gamma }}_{n}-\boldsymbol{\gamma }^*\Vert _\infty$in the order of$\widetilde{\mathcal {O}}(1/\sqrt{n})$, where$n$is the sample number. Based on this result, we also derive its excess clustering risk calculated by a standard clustering loss function. For the large-scale extension, we replace the eigen decomposition of SimpleMKKM with singular value decomposition (SVD). Consequently, the complexity can be decreased to$\mathcal {O}(n)$such that SimpleMKKM can be implemented on large-scale datasets. We then deduce several theoretical results to verify the approximation ability of the proposed SVD-based method. The results of comprehensive experiments demonstrate the superiority of the proposed method. The code is publicly available athttps://github.com/weixuan-liang/SVD-based-SimpleMKKM. Weixuan Liang, Chang Tang, Xinwang Liu 0002, Yong Liu 0018, Jiyuan Liu 0003, En Zhu, Kunlun He |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Fast Continual Multi-View Clustering With Incomplete ViewsabstractMulti-view clustering (MVC) has attracted broad attention due to its capacity to exploit consistent and complementary information across views. This paper focuses on a challenging issue in MVC called the incomplete continual data problem (ICDP). Specifically, most existing algorithms assume that views are available in advance and overlook the scenarios where data observations of views are accumulated over time. Due to privacy considerations or memory limitations, previous views cannot be stored in these situations. Some works have proposed ways to handle this problem, but all of them fail to address incomplete views. Such an incomplete continual data problem (ICDP) in MVC is difficult to solve since incomplete information with continual data increases the difficulty of extracting consistent and complementary knowledge among views. We propose Fast Continual Multi-View Clustering with Incomplete Views (FCMVC-IV) to address this issue. Specifically, the method maintains a scalable consensus coefficient matrix and updates its knowledge with the incoming incomplete view rather than storing and recomputing all the data matrices. Considering that the given views are incomplete, the newly collected view might contain samples that have yet to appear; two indicator matrices and a rotation matrix are developed to match matrices with different dimensions. In addition, we design a three-step iterative algorithm to solve the resultant problem with linear complexity and proven convergence. Comprehensive experiments conducted on various datasets demonstrate the superiority of FCMVC-IV over the competing approaches. The code is publicly available at https://github.com/wanxinhang/FCMVC-IV. Xinhang Wan, Bin Xiao 0002, Xinwang Liu 0002, Jiyuan Liu 0003, Weixuan Liang, En Zhu |
IEEE Trans. Image Process. | 5 |
| 2024 | Late Fusion Multiview Clustering via Min-Max OptimizationabstractMultiview clustering (MVC) sufficiently exploits the diverse and complementary information among different views to improve the clustering performance. As a representative algorithm of MVC, the newly proposed simple multiple kernel k -means (SimpleMKKM) algorithm takes a min-max formulation and applies a gradient descent algorithm to decrease the resultant objective function. It is empirically observed that its superiority is attributed to the novel min-max formulation and the new optimization. In this article, we propose to integrate the min-max learning paradigm adopted by SimpleMKKM into late fusion MVC (LF-MVC). This leads to a tri-level max-min-max optimization problem with respect to the perturbation matrices, weight coefficient, and clustering partition matrix. To solve this intractable max-min-max optimization problem, we design an efficient two-step alternative optimization strategy. Furthermore, we analyze the generalization clustering performance of the proposed algorithm from the theoretical perspective. Comprehensive experiments have been conducted to evaluate the proposed algorithm in terms of clustering accuracy (ACC), computation time, convergence, as well as the evolution of the learned consensus clustering matrix, clustering with different numbers of samples, and analysis of the learned kernel weight. The experimental results show that the proposed algorithm is able to significantly reduce the computation time and improve the clustering ACC when compared to several state-of-the-art LF-MVC algorithms. The code of this work is publicly released at: https://xinwangliu.github.io/Under-Review. Miaomiao Li 0001, Xinwang Liu 0002, Yi Zhang 0104, Weixuan Liang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Unpaired Multi-View Graph Clustering With Cross-View Structure MatchingabstractMulti-view clustering (MVC), which effectively fuses information from multiple views for better performance, has received increasing attention. Most existing MVC methods assume that multi-view data are fully paired, which means that the mappings of all corresponding samples between views are predefined or given in advance. However, the data correspondence is often incomplete in real-world applications due to data corruption or sensor differences, referred to as the data-unpaired problem (DUP) in multi-view literature. Although several attempts have been made to address the DUP issue, they suffer from the following drawbacks: 1) most methods focus on the feature representation while ignoring the structural information of multi-view data, which is essential for clustering tasks; 2) existing methods for partially unpaired problems rely on pregiven cross-view alignment information, resulting in their inability to handle fully unpaired problems; and 3) their inevitable parameters degrade the efficiency and applicability of the models. To tackle these issues, we propose a novel parameter-free graph clustering framework termed unpaired multi-view graph clustering framework with cross-view structure matching (UPMGC-SM). Specifically, unlike the existing methods, UPMGC-SM effectively utilizes the structural information from each view to refine cross-view correspondences. Besides, our UPMGC-SM is a unified framework for both the fully and partially unpaired multi-view graph clustering. Moreover, existing graph clustering methods can adopt our UPMGC-SM to enhance their ability for unpaired scenarios. Extensive experiments demonstrate the effectiveness and generalization of our proposed framework for both paired and unpaired datasets. Yi Wen 0001, Siwei Wang 0001, Qing Liao 0001, Weixuan Liang, Ke Liang 0006, Xinhang Wan, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Auto-Weighted Multi-View Clustering for Large-Scale DataabstractMulti-view clustering has gained broad attention owing to its capacity to exploit complementary information across multiple data views. Although existing methods demonstrate delightful clustering performance, most of them are of high time complexity and cannot handle large-scale data. Matrix factorization-based models are a representative of solving this problem. However, they assume that the views share a dimension-fixed consensus coefficient matrix and view-specific base matrices, limiting their representability. Moreover, a series of large-scale algorithms that bear one or more hyperparameters are impractical in real-world applications. To address the two issues, we propose an auto-weighted multi-view clustering (AWMVC) algorithm. Specifically, AWMVC first learns coefficient matrices from corresponding base matrices of different dimensions, then fuses them to obtain an optimal consensus matrix. By mapping original features into distinctive low-dimensional spaces, we can attain more comprehensive knowledge, thus obtaining better clustering results. Moreover, we design a six-step alternative optimization algorithm proven to be convergent theoretically. Also, AWMVC shows excellent performance on various benchmark datasets compared with existing ones. The code of AWMVC is publicly available at https://github.com/wanxinhang/AAAI-2023-AWMVC. Xinhang Wan, Xinwang Liu 0002, Jiyuan Liu 0003, Siwei Wang 0001, Yi Wen 0001, Weixuan Liang, En Zhu, Zhe Liu 0001, Lu Zhou 0002 |
AAAI | 6 |
| 2023 | Consistency of Multiple Kernel ClusteringabstractConsistency plays an important role in learning theory. However, in multiple kernel clustering (MKC), the consistency of kernel weights has not been sufficiently investigated. In this work, we fill this gap with a non-asymptotic analysis on the consistency of kernel weights of a novel method termed SimpleMKKM. Under the assumptions of the eigenvalue gap, we give an infinity norm bound as $\widetilde{\mathcal{O}}(k/\sqrt{n})$, where $k$ is the number of clusters and $n$ is the number of samples. On this basis, we establish an upper bound for the excess clustering risk. Moreover, we study the difference of the kernel weights learned from $n$ samples and $r$ points sampled without replacement, and derive its upper bound as $\widetilde{\mathcal{O}}(k\cdot\sqrt{1/r-1/n})$. Based on the above results, we propose a novel strategy with Nyström method to enable SimpleMKKM to handle large-scale datasets with a theoretical learning guarantee. Finally, extensive experiments are conducted to verify the theoretical results and the effectiveness of the proposed large-scale strategy. Weixuan Liang, Xinwang Liu 0002, Yong Liu 0018, Chuan Ma 0001, Yunping Zhao, Zhe Liu 0001, En Zhu |
ICML | 1 |
| 2023 | Scalable Incomplete Multi-View Clustering with Structure AlignmentabstractThe success of existing multi-view clustering (MVC) relies on the assumption that all views are complete. However, samples are usually partially available due to data corruption or sensor malfunction, which raises the research of incomplete multi-view clustering (IMVC). Although several anchor-based IMVC methods have been proposed to process the large-scale incomplete data, they still suffer from the following drawbacks: i) Most existing approaches neglect the inter-view discrepancy and enforce cross-view representation to be consistent, which would corrupt the representation capability of the model; ii) Due to the samples disparity between different views, the learned anchor might be misaligned, which we referred as the Anchor-Unaligned Problem for Incomplete data (AUP-ID). Such the AUP-ID would cause inaccurate graph fusion and degrades clustering performance. To tackle these issues, we propose a novel incomplete anchor graph learning framework termed Scalable Incomplete Multi-View Clustering with Structure Alignment (SIMVC-SA). Specially, we construct the view-specific anchor graph to capture the complementary information from different views. In order to solve the AUP-ID, we propose a novel structure alignment module to refine the cross-view anchor correspondence. Meanwhile, the anchor graph construction and alignment are jointly optimized in our unified framework to enhance clustering quality. Through anchor graph construction instead of full graphs, the time and space complexity of the proposed SIMVC-SA is proven to be linearly correlated with the number of samples. Extensive experiments on seven incomplete benchmark datasets demonstrate the effectiveness and efficiency of our proposed method. Our code is publicly available at https://github.com/wy1019/SIMVC-SA. Yi Wen 0001, Siwei Wang 0001, Ke Liang 0006, Weixuan Liang, Xinhang Wan, Xinwang Liu 0002, Suyuan Liu, Jiyuan Liu 0003, En Zhu |
ACM Multimedia | 4 |
| 2022 | Robust Graph-Based Multi-View ClusteringabstractGraph-based multi-view clustering (G-MVC) constructs a graphical representation of each view and then fuses them to a unified graph for clustering. Though demonstrating promising clustering performance in various applications, we observe that their formulations are usually non-convex, leading to a local optimum. In this paper, we propose a novel MVC algorithm termed robust graph-based multi-view clustering (RG-MVC) to address this issue. In particular, we define a min-max formulation for robust learning and then rewrite it as a convex and differentiable objective function whose convexity and differentiability are carefully proved. Thus, we can efficiently solve the resultant problem using a reduced gradient descent algorithm, and the corresponding solution is guaranteed to be globally optimal. As a consequence, although our algorithm is free of hyper-parameters, it has shown good robustness against noisy views. Extensive experiments on benchmark datasets verify the superiority of the proposed method against the compared state-of-the-art algorithms. Our codes and appendix are available at https://github.com/wx-liang/RG-MVC. Weixuan Liang, Xinwang Liu 0002, Sihang Zhou 0001, Jiyuan Liu 0003, Siwei Wang 0001, En Zhu |
AAAI | 1 |
| 2022 | Continual Multi-view ClusteringabstractWith the increase of multimedia applications, data are often collected from multiple sensors or modalities, encouraging the rapid development of multi-view (also called multi modal) clustering technique. As a representative, late fusion multi-view clustering algorithm has attracted extensive attention due to its low computation complexity yet promising performance. However, most of them deal with the clustering problem in which all data views are available in advance, and overlook the scenarios where data observations of new views are accumulated over time. To solve this issue, we propose a continual approach on the basis of late fusion multi-view clustering framework. In specific, it only needs to maintain a consensus partition matrix and update knowledge with the incoming one of a new data view rather than keep all of them. This benefits a lot by preventing the previously learned knowledge from recomputing over and over again, saving a large amount of computation resource/time and labor force. Nevertheless, we design an alternate and convergent strategy to solve the resultant optimization problem. Also, the proposed algorithm shows excellent clustering performance and time/space efficiency in the experiment. Xinhang Wan, Jiyuan Liu 0003, Weixuan Liang, Xinwang Liu 0002, Yi Wen 0001, En Zhu |
ACM Multimedia | 3 |
| 2022 | Sample Weighted Multiple Kernel K-means via Min-Max optimizationabstractA representative multiple kernel clustering (MKC) algorithm, termed simple multiple kernel k-means (SMKKM), is recently proposed to optimally mine useful information from a set of pre-specified kernels to improve clustering performance. Different from existing min-min learning framework, it puts a novel min-max optimization manner, which attracts considerable attention in related community. Despite achieving encouraged success, we observe that SMKKM only focuses on combination coefficients among kernels and ignores the relationship among the importance of different samples. As a result, it does not sufficiently consider different contributions of each sample to clustering, and thus cannot effectively obtain the "ideal" similarity structure, leading to unsatisfying performance. To address this issue, this paper proposes a novel sample weighted multiple kernel k-means via min-max optimization (SWMKKM), which sufficiently considers the sum of relationship between one sample and the others to represent the sample weights. Such a weighting criterion helps clustering algorithm pay more attention to samples with more positive effects on clustering and avoids unreliable overestimation for samples with poor quality. Based on SMKKM, we adopt a reduced gradient algorithm with proved convergence to solve the resultant optimization problem. Comprehensive experiments on multiple benchmark datasets demonstrate that our proposed SWMKKM dramatically improves the state-of-the-art MKC algorithms, verifying the effectiveness of our proposed sample weighting criterion. Yi Zhang 0104, Weixuan Liang, Xinwang Liu 0002, Sisi Dai, Siwei Wang 0001, En Zhu |
ACM Multimedia | 2 |
| 2022 | Stability and Generalization of Kernel Clustering: from Single Kernel to Multiple KernelabstractMultiple kernel clustering (MKC) is an important research topic that has been widely studied for decades. However, current methods still face two problems: inefficient when handling out-of-sample data points and lack of theoretical study of the stability and generalization of clustering. In this paper, we propose a novel method that can efficiently compute the embedding of out-of-sample data with a solid generalization guarantee. Specifically, we approximate the eigen functions of the integral operator associated with the linear combination of base kernel functions to construct low-dimensional embeddings of out-of-sample points for efficient multiple kernel clustering. In addition, we, for the first time, theoretically study the stability of clustering algorithms and prove that the single-view version of the proposed method has uniform stability as $\mathcal{O}\left(Kn^{-3/2}\right)$ and establish an upper bound of excess risk as $\widetilde{\mathcal{O}}\left(Kn^{-3/2}+n^{-1/2}\right)$, where $K$ is the cluster number and $n$ is the number of samples. We then extend the theoretical results to multiple kernel scenarios and find that the stability of MKC depends on kernel weights. As an example, we apply our method to a novel MKC algorithm termed SimpleMKKM and derive the upper bound of its excess clustering risk, which is tighter than the current results. Extensive experimental results validate the effectiveness and efficiency of the proposed method. Weixuan Liang, Xinwang Liu 0002, Yong Liu 0018, Sihang Zhou 0001, Junjie Huang 0001, Siwei Wang 0001, Jiyuan Liu 0003, Yi Zhang 0104, En Zhu |
NeurIPS | 1 |
| 2022 | Multi-View Spectral Clustering With High-Order Optimal Neighborhood Laplacian MatrixabstractMulti-view spectral clustering can effectively reveal the intrinsic cluster structure among data by performing clustering on the learned optimal embedding across views. Though demonstrating promising performance in various applications, most of existing methods usually linearly combine a group of pre-specified first-order Laplacian matrices to construct the optimal Laplacian matrix, which may result in limited representation capability and insufficient information exploitation. Also, storing and implementing complex operations on the{$n\times n}$Laplacian matrices incurs intensive storage and computation complexity. To address these issues, this paper first proposes a multi-view spectral clustering algorithm that learns a high-order optimal neighborhood Laplacian matrix, and then extends it to the late fusion version for accurate and efficient multi-view clustering. Specifically, our proposed algorithm generates the optimal Laplacian matrix by searching the neighborhood of the linear combination of both the first-order and high-order base Laplacian matrices simultaneously. By this way, the representative capacity of the learned optimal Laplacian matrix is enhanced, which is helpful to better utilize the hidden high-order connection information among data, leading to improved clustering performance. We design an efficient algorithm with proved convergence to solve the resultant optimization problem. Extensive experimental results on nine datasets demonstrate the superiority of the proposed algorithm Weixuan Liang, Sihang Zhou 0001, Jian Xiong 0002, Xinwang Liu 0002, Siwei Wang 0001, En Zhu, Zhiping Cai, Xin Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | One-pass Multi-view Clustering for Large-scale DataabstractExisting non-negative matrix factorization based multi-view clustering algorithms compute multiple coefficient matrices respect to different data views, and learn a common consensus concurrently. The final partition is always obtained from the consensus with classical clustering techniques, such as k-means. However, the non-negativity constraint prevents from obtaining a more discriminative embedding. Meanwhile, this two-step procedure fails to unify multi-view matrix factorization with partition generation closely, resulting in unpromising performance. Therefore, we propose an one-pass multi-view clustering algorithm by removing the non-negativity constraint and jointly optimize the aforementioned two steps. In this way, the generated partition can guide multi-view matrix factorization to produce more purposive coefficient matrix which, as a feedback, improves the quality of partition. To solve the resultant optimization problem, we design an alternate strategy which is guaranteed to be convergent theoretically. Moreover, the proposed algorithm is free of parameter and of linear complexity, making it practical in applications. In addition, the proposed algorithm is compared with recent advances in literature on benchmarks, demonstrating its effectiveness, superiority and efficiency. Jiyuan Liu 0003, Xinwang Liu 0002, Yuexiang Yang, Li Liu 0002, Siqi Wang 0001, Weixuan Liang, Jiangyong Shi |
ICCV | 6 |
| 2021 | Self-Representation Subspace Clustering for Incomplete Multi-view DataabstractIncomplete multi-view clustering is an important research topic in multimedia where partial data entries of one or more views are missing. Current subspace clustering approaches mostly employ matrix factorization on the observed feature matrices to address this issue. Meanwhile, self-representation technique is left unexplored, since it explicitly relies on full data entries to construct the coefficient matrix, which is contradictory to the incomplete data setting. However, it is widely observed that self-representation subspace method enjoys a better clustering performance over the factorization based one. Therefore, we adapt it to incomplete data by jointly performing data imputation and self-representation learning. To the best of our knowledge, this is the first attempt in incomplete multi-view clustering literature. Besides, the proposed method is carefully compared with current advances in experiment with respect to different missing ratios, verifying its effectiveness. Jiyuan Liu 0003, Xinwang Liu 0002, Yi Zhang 0104, Pei Zhang 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Weixuan Liang, Siqi Wang 0001, Yuexiang Yang |
ACM Multimedia | 8 |