Quanxue Gao

dblp:63/804 · DBLP profile ↗
← Back
169ranked-venue papers
24as first author
105since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 125 · 18 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 5 first-author · 52 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Security and privacy · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Intrinsic Hierarchy for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) aims to classify unlabeled data by leveraging knowledge from labeled categories. While existing methods have achieved remarkable progress, they often treat images as flat feature sets, neglecting the intrinsic hierarchy: where key objects dominate meaning and backgrounds serve as context. For instance, in images of a dog either standing on grass or lying on a bed, the dog remains the central semantic element, whereas the background varies. Motivated by this, we propose LEArning Intrinsic Hierarchy (LEAH), a lightweight plug-and-play module designed to model hierarchical structure within images. LEAH consists of two components: a pruner that filters task-irrelevant tokens to extract key objects, and a constructor that embeds key objects and full images into hyperbolic space using adaptive entailment cones to capture compositional semantics. LEAH can be easily integrated into existing GCD frameworks with minimal modification. When applied to SimGCD, it achieves up to 13.2% accuracy improvement on fine-grained benchmarks, demonstrating its effectiveness in discovering subtle inter-class differences through hierarchical modeling.
Yu Duan 0001, Junzhi He, Zhanxuan Hu, Mengda Ji, Rong Wang 0001, Quanxue Gao
AAAI6
2026 Maximizing Schatten-p Norm Regularization Toward Balance
abstract
The Schatten-p norm, as a class of structure-inducing norms based on singular values, has been widely used to enhance model low-rankness and representation capability due to its flexibility in structural modeling and favorable mathematical properties. However, its potential in cluster distribution modeling has long been overlooked. Therefore, we explore the potential of maximizing the Schatten-p norm as a regularization strategy specifically designed to achieve balanced clustering. This work is the first to investigate its effectiveness in promoting cluster balance. To be specific, maximizing Schatten-p norm effectively guides the assignment of data points, ensuring a more balanced distribution of samples across clusters. We have conducted an in-depth theoretical analysis and validated its effectiveness through extensive clustering experiments. Experimental results demonstrate that, compared to existing methods, this regularization term significantly improves clustering quality and obtain reasonable clustering.
Fangfang Li 0005, Quanxue Gao, Yu Duan 0001, Yuzhuo Feng, Qin Li 0001
AAAI2
2026 Unified View Extraction with Low-Rankness and Smoothness Fusion for Multi-View Subspace Clustering
abstract
Tensor-based multi-view subspace clustering (MVSC) has achieved significant success by capturing high-order inter-view correlations. However, existing approaches face two principal limitations. First, most methods either exclusively emphasize the inter-view low‑rankness (R) prior while neglecting the intra-view local smoothness (S) prior, or treat R and S as two separate regularizers—complicating joint optimization. Second, conventional tensor‑based methods impose only low‑rank constraints on the representation tensor, which limits their ability to simultaneously model consistency and complementary information. To address these issues, we propose a Unified View Extraction with Low‑Rankness and Smoothness Fusion (UVELRS) method. Our framework first extracts a consistent cross‑view representation and then constructs a tensor by stacking these representations. We introduce a novel tensor total variation Schatten-p norm that simultaneously encodes both R and S priors while offering flexible singular‑value control. This unified formulation effectively captures both high-order inter-view correlations and intra-view local smoothness. Extensive experiments on real‑world datasets demonstrate UVELRS's superior performance and robustness.
Quanxue Gao, Fangfang Li 0005, Yu Yun, Ming Yang 0024
AAAI2
2026 Adversarial Fair Incomplete Multi-View Clustering
abstract
Fair incomplete multi-view clustering (FIMVC) confronts a critical yet unresolved challenge, as existing methods often fail to address the intertwined issues of data missingness and algorithmic bias simultaneously. In this paper, we propose a novel FIMVC method named Adversarial Fair Incomplete Multi-View Clustering (AFIMVC). The core of AFIMVC is a new adaptive adversarial disentanglement mechanism. This mechanism trains the feature encoder to produce representations that are invariant to sensitive attributes by adversary learning, where the adversarial intensity is dynamically controlled by the model's real-time bias. Additionally, we develop a probabilistic cross-view contrastive learning strategy to achieve semantic consistency in latent space. To handle missing data, AFIMVC employs a context-aware fusion strategy that leverages cross-sample attention to robustly synthesize a unified representation from incomplete views. Extensive experiments demonstrate that AFIMVC achieves a state-of-the-art balance between clustering accuracy and fairness, significantly outperforming existing methods.
Qianqian Wang 0001, Wei Feng 0010, Quanxue Gao
AAAI4
2026 Tensorized Label Learning via Balanced Tensor Regression
abstract
The multi-view clustering methods based on tensor regression can make full use of the potential structural information between views and achieve data-level fusion. However, existing tensor regression-based approaches for anchor graph often overlook the probabilistic nature of anchor graph, focusing solely on sample labels while ignoring the influence of anchor labels on clustering results. To overcome these limitations, we introduce Tensorized Label Learning via Balanced Tensor Regression (TLL-BTR). Our key idea is to exploit the probabilistic nature of the anchor graph by regarding the sample labels as a projection tensor that maps the anchor graph into the label space, thereby producing anchor labels. By enforcing constraints on these anchor labels, we guide the concurrent learning of sample labels and achieve co-label learning between anchors and samples. To prevent trivial solutions, we maximize the nuclear norm to promote an even distribution of samples across clusters. Extensive experiments on benchmark datasets demonstrate that TLL-BTR consistently outperforms state-of-the-art methods.
Yuzhuo Feng, Qin Li 0001, Quanxue Gao, Ming Yang 0024
AAAI4
2026 Balanced federated multi-view clustering with one-round communication
Quanxue Gao, Yu Duan 0001
Neurocomputing2
2026 Probabilistic regression-based multi-view clustering with anchor graphs
Qin Li 0001, Quanxue Gao, Yu Duan 0001, Cheng Deng 0002
Neurocomputing4
2026 Discriminative and noise-robust embedding for zero-shot learning
Cheng Deng 0002, Yu Duan 0001, Quanxue Gao
Neural Networks4
2026 Fuzzy K-means clustering without cluster centroids
Yichen Bao, Han Lu 0005, Quanxue Gao
Signal Process.4
2026 Anchor-Guided Discrete Multi-View Clustering
abstract
Multi-view clustering based on anchor graphs has attracted significant attention due to its ability to substantially reduce computational complexity, enabling the efficient processing of large-scale multimedia data. However, most existing anchor graph-based clustering methods fail to fully exploit the intrinsic properties of anchor graphs when applying regression techniques. Moreover, some approaches focus solely on sample labels while overlooking the crucial role of anchor labels in clustering. To address these limitations, we leverage the probabilistic information of the anchor graph by employing probabilistic projection to map the anchor graph into the label space, thereby obtaining anchor labels. By clustering both anchors and samples simultaneously, the anchor graph serves as a guide to induce anchor labels, which are then used to generate sample labels, facilitating anchor-guided sample clustering. Furthermore, we propose a novel regularization paradigm based on the matrix nuclear norm, ensuring that the obtained results remain discrete and that the sample distribution across clusters is balanced. Additionally, we introduce a new matrix nuclear norm optimization method based on the first-order Taylor expansion. Extensive experiments on real-world datasets demonstrate the effectiveness and robustness of our proposed method, achieving superior performance compared to state-of-the-art approaches. Our code is available at:https://github.com/harunakai/ADMC
Ran Jing, Quanxue Gao, Yu Duan 0001, Cheng Deng 0002, Ming Yang 0024
IEEE Trans. Multim.2
2025 Deep Multi-modal Graph Clustering via Graph Transformer Network
abstract
Current deep multi-modal graph clustering methods primarily rely on Graph Neural Network (GNN) to fully exploit attribute features and graph structures, including message propagation and low-dimensional feature embedding. However, these methods lack further exploration of graph structural information, such as the relationship between nodes and shortest paths. Additionally, they may not sufficiently mine complementary information among multi-modal graph data. To address these issues, we propose a novel Deep Multi-modal Graph Clustering via Graph Transformer Network method, called DMGC-GTN. This method thoroughly dissects and utilizes graph structural information, applying graph smoothing to node features and incorporating various forms of embeddings into the transformer architecture. This achieves a unified embedding of graph structure and multi-modal feature attributes, fully exploiting the complementary information within multi-modal graph data. Extensive experiments demonstrate the effectiveness of our algorithm.
Qianqian Wang 0001, Wei Feng 0010, Quanxue Gao
AAAI5
2025 Contrastive Multi-view Subspace Clustering via Tensor Transformers Autoencoder
abstract
Multi-view clustering aims to identify consistent and complementary information across multiple views to partition data into clusters, emerging as a popular unsupervised method for multi-view data analysis. However, existing methods often design view-specific encoders to extract distinct features from each view, lacking exploration of their complementarity. Additionally, current contrastive-based multi-view clustering methods may lead to erroneous negative sample pairs conflicting with the clustering objective. To address these challenges, we propose a novel Contrastive Multi-view Subspace Clustering via Tensor Transformers Autoencoder (TTAE). On the one hand, it facilitates information exchange between views by tensor transformers autoencoder, thereby enhancing complementarity. On the other hand, It learns a consistent subspace with a self-expression layer. Meanwhile, adaptive contrastive learning helps to provide more discriminative features for the self-expression learning layer, and the self-expression learning layer in turn supervises contrastive learning. Moreover, our method adaptively selects positive and negative samples for contrastive learning to mitigate the impact of inappropriate negative sample pairs. Extensive experiments on several multi-view datasets demonstrate the effectiveness and superiority of our model.
Qianqian Wang 0001, Wei Feng 0010, Zhiqiang Tao, Quanxue Gao
AAAI5
2025 Tensorized Label Learning Based Fast Fuzzy Clustering
abstract
Multi-view graph clustering methods have been widely concerned due to the ability of dealing with arbitrarily shaped datasets. However, many methods with higher time and space complexity make them challenging to deal with large-scale datasets. Besides, many fuzzy clustering methods needs additional regularization terms or hyper-parameters to obtain the membership matrix or avoid trivial solutions, which weakens the model generalization ability. Furthermore, inconsistent clustering labels can arise when there are significant discrepancies between views, making it challenging to effectively leverage the complementary information from different views. To this end, we propose Tensorized Label Learning based Fast Fuzzy Clustering (TLLFFC). Specifically, we design a novel balanced regularization term to reduce pressure of tuning regularization parameters for fuzzy clustering. The label transmission strategy with the anchor graph makes TLLFFC suitable for large-scale datasets. Moreover, incorporating the Schatten p-norm regularization on the label matrices can effectively unearth the complementary information distributed among views, thereby align the labels across views more consistently. Extensive experiments verify the superiority of TLLFFC.
Xingyu Xue, Quanxue Gao, Qianqian Wang 0001
AAAI3
2025 Deep Fair Multi-View Clustering with Attention KAN
abstract
Multi-view clustering is effective in unsupervised multi-view data analysis and has received considerable attention. However, most existing methods excessively emphasize certain attributes, resulting in unfair clustering outcomes, i.e., certain sensitive attributes dominate the clustering results. Moreover, existing methods struggle to effectively capture complex nonlinear relationships and interactions across views, limiting their ability to achieve optimal clustering performance. Therefore, in this work, we propose a novel method, Deep Fair Multi-View Clustering with Attention Kolmogorov-Arnold Network (DFMVC-AKAN), to generate fair clustering results while maintaining robust performance. DFMVC-AKAN integrates attention mechanisms into Kolmogorov-Arnold Networks (KAN) to exploit the complex nonlinear inter-view relationships. Specifically, KAN provides a nonlinear feature representation capable of efficiently approximating arbitrary multivariate continuous functions, augmented by a hybrid attention mechanism which enables the model to dynamically focus on the most relevant features. Finally, we refine the clustering assignments with a distribution alignment module to ensure fair outcomes across diverse groups while maintaining discriminative ability. Experimental results on four datasets containing sensitive attributes demonstrate that DFMVC-AKAN significantly improves fairness and clustering performance compared to state-of-the-art methods.
Qianqian Wang 0001, Boyue Wang, Quanxue Gao
CVPR4
2025 Attribute-Missing Multi-view Graph Clustering
abstract
The success of existing deep multi-view graph clustering methods is based on the assumption that node attributes are fully available across all views. However, in practical scenarios, node attributes are frequently missing due to factors such as data privacy concerns or failures in data collection devices. Although some methods have been proposed to address the issue of missing node attributes, they come with the following limitations: i) Existing methods are often not tailored specifically for clustering tasks and struggle to address missing attributes effectively. ii) They tend to ignore the relational dependencies between nodes and their neighboring nodes. This oversight results in unreliable imputations, thereby degrading clustering performance. To address the above issues, we propose an Attribute-Missing Multi-view Graph Clustering (AMMGC). Specifically, we first impute missing node attributes by leveraging neighbor-hood information through an adjacency matrix. Then, to improve the consistency, we integrate a dual structure consistency module that aligns graph structures across multiple views, reducing redundancy and retaining key information. Furthermore, we introduce a high-confidence guidance module to improve the reliability of clustering. Extensive experiment results showcase the effectiveness and superiority of our proposed method on multiple benchmark datasets.
Qianqian Wang 0001, Zhengming Ding, Quanxue Gao
CVPR4
2025 Hypergraph Clustering Network with Partial Attribute Imputation
Qianqian Wang 0001, Zhengming Ding, Wei Feng 0010, Quanxue Gao
ICCV5
2025 High-Quality Label Learning in Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) is a recently proposed open-world problem that aims to automatically classify and discover new categories based on partially labeled data. For unlabeled data, previous research commonly considers using pseudo-labels to assist in model learning. These pseudo-labels, together with the true labels of labeled data, form the learning targets for the final classifier, leading to better predictive outcomes. However, low-quality labels can inevitably hinder the learning process of the model. To address this issue, inspired by previous methods, we propose the Calibrated Generalized Category Discovery (CGCD) framework, which incorporates a projection head, a classifier head, and a calibration head. The projection head is used for representation learning. The calibration head learns high-quality labels from the robust predictions of the classifier head, and the classifier head utilizes these high-quality labels for more efficient learning. Both heads mutually enhance each other during training, ultimately leading to a superior solution. In addition, leveraging the characteristics of both the classifier head and calibration head, we designed a classifier representation distribution regularization term to further ensure consistency in their learning processes. Extensive experimental results demonstrate that the proposed CGCD framework achieves state-of-the-art performance across five general and fine-grained visual recognition datasets by leveraging high-quality label learning.
Yu Duan 0001, Junzhi He, Feiping Nie 0001, Quanxue Gao, Cheng Deng 0002
ICDM4
2025 Balanced and Discrete Projection Collaborative Clustering with Probabilistic Regression
abstract
Anchor graph-based multi-view clustering significantly reduces computational complexity, but most existing methods still exhibit the following drawbacks: 1. They neglect the probabilistic nature of the anchor graph. 2. They focus solely on the sample label matrix and fail to account for the relationship between the sample labels and the anchor labels. To address these issues, this paper proposes an balanced and discrete multi-view clustering model based on concept probabilistic regression. Specifically, according to the probabilistic characteristics of the anchor graph, we summarize the relationships between the anchor graph, anchor labels, and sample labels, and propose a probabilistic regression model that better guides the learning of sample labels by imposing constraints on the anchor labels. Moreover, unlike the widely used$\ell_{2,1}$-norm minimization methods in machine learning, we maximize the$\ell_{2,1}$-norm to ensure a balanced distribution of anchors within each clustering. Experimental results on several publicly available datasets demonstrate that the proposed model achieves superior clustering performance compared to mainstream algorithms.
Yu Duan 0001, Quanxue Gao, Xuhong Dong
ICDM3
2025 Unified K-Means Clustering with Label-Guided Manifold Learning
abstract
K-Means clustering is a classical and effective unsupervised learning method attributed to its simplicity and efficiency. However, it faces notable challenges, including sensitivity to random initial centroid selection, a limited ability to discover the intrinsic manifold structures within nonlinear datasets, and difficulty in achieving balanced clustering in practical scenarios. To overcome these weaknesses, we introduce a novel framework for K-Means that leverages manifold learning. This approach eliminates the need for centroid calculation and utilizes a cluster indicator matrix to align the manifold structures, thereby enhancing clustering accuracy. Beyond the traditional Euclidean distance, our model incorporates Gaussian kernel distance, K-nearest neighbor distance, and low-pass filtering distance to effectively manage data that is not linearly separable. Furthermore, we introduce a balanced regularizer to achieve balanced clustering results. The detailed experimental results demonstrate the efficacy of our proposed methodology.
Qianqian Wang 0001, Mengping Jiang, Zhengming Ding, Quanxue Gao
ICML4
2025 Efficient Multi-view Clustering via Reinforcement Contrastive Learning
abstract
Contrastive multi-view clustering has demonstrated remarkable potential in complex data analysis, yet existing approaches face two critical challenges: difficulty in constructing high-quality positive and negative pairs and high computational overhead due to static optimization strategies. To address these challenges, we propose an innovative efficient Multi-View Clustering framework with Reinforcement Contrastive Learning (EMVCRCL). Our key innovation is developing a reinforcement contrastive learning paradigm for dynamic clustering optimization. First, we leverage multi-view contrastive learning to obtain latent features, which are then sent to the reinforcement learning module to refine low-quality features. Specifically, it selects high-confident features to guide the positive/negative pair construction of contrastive learning. For the low-confident features, it utilizes the prior balanced distribution to adjust their assignment. Extensive experimental results showcase the effectiveness and superiority of our proposed method on multiple benchmark datasets.
Qianqian Wang 0001, Zhiqiang Tao, Quanxue Gao
IJCAI5
2025 A Simple yet Effective Hypergraph Clustering Network
abstract
Hypergraph Clustering has gained significant attention due to its capability of capturing high order structural information. Among different approaches, contrastive learning-based methods leverage self-supervised learning and data augmentation, exhibiting impressive performance. However, most of them come with the following limitations: 1) Augmentation strategies like feature dropout can potentially disrupt the intrinsic clustering structure of hypergraphs. 2) High computational demands hinder their real-world application. To address the above issues, we propose a simple yet effective Hypergraph Clustering Network framework (HCN). Specifically, HCN replaces the hypergraph convolution operation with smoothing preprocessing, which avoids high computational complexity. Besides, to retain intrinsic structure, it develops two key modules: the self-diagonal consistency module and the structure alignment mod ule. They respectively align the similarity matrix with the identity matrix and the structural affinity matrix, which ensures intra-cluster compact ness and inter-cluster separability. Extensive experiments on five benchmark datasets demonstrate HCN’s superiority over state-of-the-art methods.
Qianqian Wang 0001, Zhengming Ding, Quanxue Gao
IJCAI5
2025 Multi-view Clustering Based on Probabilistic Tensor Regression
abstract
Multi-view clustering based on anchor graph and regression is widely used to deal with high dimensional and redundant data. However, most of these methods ignore the probabilistic characteristics of anchor graph, and the effective information in different views is not fully mined. To solve these problems, we propose a multi-view clustering method based on probabilistic tensor regression (MVCPTR). Specifically, we reinterpret the regression process of the anchor graph from the perspective of probability. By modeling the anchor graph as the transition probability from samples to anchors, we construct the implicit relationship between labels of samples and anchors. In order to further mine the complementary information of multi-view data, we extend the anchor graph matrix regression to tensor regression to achieve multi-level information fusion at the representational level and decision level, and impose the Schatten p-norm constraint on the anchor label tensor and the sample label tensor to realize the bi-clustering of the anchors and samples. A large number of experiments prove the effectiveness of our proposed algorithm.
Yichen Bao, Yu Duan 0001, Jing Li 0026, Quanxue Gao
ACM Multimedia5
2025 Multi-view Collaborative Representation Learning from Noisy Labels for VHR Imagery Classification
abstract
Remote sensing image classification with noisy labels is receiving increasing attention. However, the existing methods ignore the context information of the training sample and judge whether the label is a noise label only by monitoring the loss value of a single sample, which may lead to misjudgment of the sample label. Additionally, these algorithms do not consider constructing pairs of confidence instances to obtain robust potential representations after identifying confidence instances. In this paper, a Multi-view Collaborative Representation Learning (MCRL) approach from noisy labels is proposed to improve the classification performance of very high resolution (VHR) remote sensing images. Specifically, we design a correction strategy based on spatial consistency and confidence-aware mechanisms. This strategy quantitatively measures label reliability by mining the contextual information of labelled samples within the adaptive region. Leveraging the spatial consistency principle and the confidence-aware mechanism to correct and smooth the noisy labels progressively. Moreover, we construct confidence sample pairs by establishing relationships between samples within and between views to obtain robust latent representations, which improves the model's tolerance to noisy labels. Experiments show that the MCRL can significantly reduce the impact of noisy labels on the model and is more competitive than homologous algorithms.
Guangfei Li, Quanxue Gao, Yichen Bao, Qianqian Wang 0001
ACM Multimedia2
2025 Bi-Orthogonal Non-negative Tensor tri-Factorization for Tensorized Label Learning
abstract
Recently, multimedia data analysis based on non-negative tensor factorization (NTF) has become a hot research topic, but these methods mainly focus on 2-factor factorization and cannot effectively explore the complex structures hidden in multimedia data, especially for graph multimedia data. In this paper, analysis for 3-factor NTF X = U * C * G T is provided in detail. Specifically, constrained 3-factor NTF helps provide new features to constrained 2-factor NTF. We herein study bi-orthogonal constraint due to the fact that it leads to rigorous interpretability of clustering. After that, we apply it to multimedia data label learning and produce a novel co-multi-view label learning based on bi-orthogonal 3-factor NTF. Extensive experiments show the capability of bi-orthogonal 3-factor NTF on simultaneously clustering anchors and samples of the input data matrix.
Quanxue Gao, Cheng Deng 0002
ACM Multimedia4
2025 Fast discrete multi-view collaborative clustering induced by label propagation
Quanxue Gao, Cheng Deng 0002
Expert Syst. Appl.4
2025 A transformer-based dual contrastive learning approach for zero-shot learning
Ran Jing, Fangfang Li 0005, Quanxue Gao, Cheng Deng 0002
Neurocomputing4
2025 One-step multi-view clustering via label transmission and fusion
Guangfei Li, Quanxue Gao, Cheng Deng 0002
Neurocomputing4
2025 Multi-view clustering based on feature selection and semi-non-negative anchor graph factorization
Shikun Mei, Qianqian Wang 0001, Quanxue Gao, Ming Yang 0024
Neural Networks3
2025 Manifold Based Multi-View K-Means
abstract
Although numerous clustering algorithms have been developed, many existing methods still rely on the K-means technique to identify clusters of data points. However, the performance of K-means is highly dependent on the accurate estimation of cluster centers, which is challenging to achieve optimally. Furthermore, it struggles to handle linearly non-separable data. To address these limitations, from the perspective of manifold learning, we reformulate multi-view K-means into a manifold-based multi-view clustering formulation that eliminates the need for computing centroid matrix. This reformulation ensures consistency between the manifold structure and the data labels. Building on this, we propose a novel multi-view K-means model incorporating the tensor rank constraint. Our model employs the indicator matrices from different views to construct a third-order tensor, whose rank is minimized via the tensor Schatten p-norm. This approach effectively leverages the complementary information across views. By utilizing different distance functions, our proposed model can effectively handle linearly non-separable data. Extensive experimental results on multiple databases demonstrate the superiority of our proposed model.
Quanxue Gao, Fangfang Li 0005, Qianqian Wang 0001, Xinbo Gao 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Self-Supervised Graph Embedding Clustering
abstract
Manifold learning and $K$K-means are two powerful techniques for data analysis in the field of artificial intelligence. When used for label learning, a promising strategy is to combine them directly and optimize both models simultaneously. However, a significant drawback of this approach is that it represents a naive and crude integration, requiring the optimization of all variables in both models without achieving a truly essential combination. Additionally, it introduces an extra hyperparameter and cannot ensure cluster balance. These challenges motivate us to explore whether a meaningful integration can be developed for dimensionality reduction clustering. In this paper, we propose a novel self-supervised manifold clustering framework that reformulates the two models into a unified framework, eliminating the need for additional hyperparameters while achieving dimensionality reduction clustering. Specifically, by analyzing the relationship between $K$K-means and manifold learning, we construct a meaningful low-dimensional manifold clustering model that directly produces the label matrix of the data. The label information is then used to guide the learning of the manifold structure, ensuring consistency between the manifold structure and the labels. Notably, we identify a valuable role of ${\ell _{2,p}}$ℓ2,p-norm regularization in clustering: maximizing the ${\ell _{2,p}}$ℓ2,p-norm naturally maintains class balance during clustering, and we provide a theoretical proof of this property. Extensive experimental results demonstrate the efficiency of our proposed model.
Fangfang Li 0005, Quanxue Gao, Xiaoke Ma 0001, Ming Yang 0024, Cheng Deng 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Centroid-Free K-Means With Balanced Clustering
abstract
Currently, a wide array of clustering algorithms have emerged, yet many approaches rely on K-means to detect clusters. However, K-means is highly sensitive to the selection of the initial cluster centers, which poses a significant obstacle to achieving optimal clustering results. Moreover, its capability to handle nonlinearly separable data is less than satisfactory. To overcome the limitations of traditional K-means, we draw inspiration from manifold learning to reformulate the K-means algorithm into a new clustering method based on manifold structures. This method not only eliminates the need to calculate centroids in traditional approaches, but also preserves the consistency between manifold structures and clustering labels. Furthermore, we introduce the$\ell _{2,1}$-norm to naturally maintain class balance during the clustering process. Additionally, we develop a versatile K-means variant framework that can accommodate various types of distance functions, thereby facilitating the efficient processing of nonlinearly separable data. The experimental results of several databases confirm the superiority of our proposed model.
Fan Yang 0111, Quanxue Gao
IEEE Signal Process. Lett.4
2025 Tensorized Tri-Factor Decomposition for Multi-View Clustering
abstract
Multi-view clustering leverages the complementary and compatible information among various views to achieve superior clustering outcomes. The approach of multi-view clustering through non-negative matrix factorization (NMF) has garnered extensive interest, attributed to its remarkable interpretability and clustering efficacy. Nonetheless, existing NMF-based multi-view subspace clustering methods fall short in thoroughly harnessing the complementary information across different views, potentially impairing clustering performance. To mitigate this issue, we introduce an orthogonal semi-nonnegative matrix tri-factorization model. This model excels in clustering interpretability, enabling the direct derivation of cluster labels from the clustering indicator matrix, thereby eliminating the need for post-processing. Our model employs tensor Schatten p-norm as a constraint, adeptly capturing both the complementary information and spatial structure information across views. Extensive experimental evaluations on a variety of benchmark datasets affirm the superior clustering performance of our proposed method.
Quanxue Gao, Ming Yang 0024, Qianqian Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Image Clustering With Transition Probabilities Learning
abstract
Large-scale multi-view clustering for image data has achieved impressive clustering performance and efficiency. However, most methods lack interpretability in clustering and do not fully consider the complementarity of distributions between different views. To address these problems, we introduce Multi-View Clustering with Transition Probabilities Learning (MVC-TPL). Specifically, we construct an anchor graph factorization model from the perspective of transition probabilities, while simultaneously learning transition probability matrices from samples to clusters and from anchor points to clusters, serving as soft label matrices for samples and anchor points, respectively. This model enables one-step label acquisition and provides the model with a sound probability interpretation. Moreover, since the clusters of samples and anchor points should be consistent across all views, we employ Schatten p-norm regularization on the two matrices, effectively mining the complementary information distributed among the views, thereby aligning the labels across views more consistently. Comprehensive testing on four small-scale datasets and three large-scale datasets confirms the effectiveness of this model.
Xingyu Xue, Quanxue Gao, Ming Yang 0024, Cheng Deng 0002
IEEE Trans. Image Process.3
2025 Tensorized Soft Label Learning Based on Orthogonal NMF
abstract
Recently, a strong interest has been in multiview high-dimensional data collected through cross-domain or various feature extraction mechanisms. Nonnegative matrix factorization (NMF) is an effective method for clustering these high-dimensional data with clear physical significance. However, existing multiview clustering based on NMF only measures the difference between the elements of the coefficient matrix without considering the spatial structure relationship between the elements. And they often require postprocessing to achieve clustering, making the algorithms unstable. To address this issue, we propose minimizing the Schatten p-norm of the tensor, which consists of a coefficient matrix of different views. This approach considers each element's spatial structure in the coefficient matrices, crucial for effectively capturing complementary information presented in different views. Furthermore, we apply orthogonal constraints to the cluster index matrix to make it sparse and provide a strong interpretation of the clustering. This allows us to obtain the cluster label directly without any postprocessing. To distinguish the importance of different views, we utilize adaptive weights to assign varying weights to each view. We introduce an unsupervised optimization scheme to solve and analyze the computational complexity of the model. Through comprehensive evaluations of six benchmark datasets and comparisons with several multiview clustering algorithms, we empirically demonstrate the superiority of our proposed method.
Fangfang Li 0005, Quanxue Gao, Qianqian Wang 0001, Ming Yang 0024, Cheng Deng 0002
IEEE Trans. Neural Networks Learn. Syst.2
2024 Partial Multi-View Clustering via Self-Supervised Network
abstract
Partial multi-view clustering is a challenging and practical research problem for data analysis in real-world applications, due to the potential data missing issue in different views. However, most existing methods have not fully explored the correlation information among various incomplete views. In addition, these existing clustering methods always ignore discovering discriminative features inside the data itself in this unsupervised task. To tackle these challenges, we propose Partial Multi-View Clustering via Self-Supervised \textbf{N}etwork (PVC-SSN) in this paper. Specifically, we employ contrastive learning to obtain a more discriminative and consistent subspace representation, which is guided by a self-supervised module. Self-supervised learning can exploit effective cluster information through the data itself to guide the learning process of clustering tasks. Thus, it can pull together embedding features from the same cluster and push apart these from different clusters. Extensive experiments on several benchmark datasets show that the proposed PVC-SCN method outperforms several state-of-the-art clustering methods.
Wei Feng 0010, Guoshuai Sheng, Qianqian Wang 0001, Quanxue Gao, Zhiqiang Tao, Bo Dong 0001
AAAI4
2024 Tensorized Label Learning on Anchor Graph
abstract
Graph-based multimedia data clustering has attracted much attention due to the impressive clustering performance for arbitrarily shaped multimedia data. However, existing graph-based clustering methods need post-processing to get labels for multimedia data with high computational complexity. Moreover, it is sub-optimal for label learning due to the fact that they exploit the complementary information embedded in data with different types pixel by pixel. To handle these problems, we present a novel label learning model with good interpretability for clustering. To be specific, our model decomposes anchor graph into the products of two matrices with orthogonal non-negative constraint to directly get soft label without any post-processing, which remarkably reduces the computational complexity. To well exploit the complementary information embedded in multimedia data, we introduce tensor Schatten p-norm regularization on the label tensor which is composed of soft labels of multimedia data. The solution can be obtained by iteratively optimizing four decoupled sub-problems, which can be solved more efficiently with good convergence. Experimental results on various datasets demonstrate the efficiency of our model.
Jing Li 0026, Quanxue Gao, Qianqian Wang 0001, Wei Xia 0007
AAAI2
2024 Embedded Feature Selection on Graph-Based Multi-View Clustering
abstract
Recently, anchor graph-based multi-view clustering has been proven to be highly efficient for large-scale data processing. However, most existing anchor graph-based clustering methods necessitate post-processing to obtain clustering labels and are unable to effectively utilize the information within anchor graphs. To solve these problems, we propose an Embedded Feature Selection on Graph-Based Multi-View Clustering (EFSGMC) approach to improve the clustering performance. Our method decomposes anchor graphs, taking advantage of memory efficiency, to obtain clustering labels in a single step without the need for post-processing. Furthermore, we introduce the l2,p-norm for graph-based feature selection, which selects the most relevant data for efficient graph factorization. Lastly, we employ the tensor Schatten p-norm as a tensor rank approximation function to capture the complementary information between different views, ensuring similarity between cluster assignment matrices. Experimental results on five real-world datasets demonstrate that our proposed method outperforms state-of-the-art approaches.
Guangfei Li, Haizhou Yang, Quanxue Gao, Qianqian Wang 0001
AAAI4
2024 Reconstruction Weighting Principal Component Analysis with Fusion Contrastive Learning
Qianqian Wang 0001, Wei Feng 0010, Mengping Jiang, Quanxue Gao
IJCAI6
2024 Federated Multi-View Clustering via Tensor Factorization
Wei Feng 0010, Zhenwei Wu, Qianqian Wang 0001, Bo Dong 0001, Zhiqiang Tao, Quanxue Gao
IJCAI6
2024 Efficient Federated Multi-View Clustering with Integrated Matrix Factorization and K-Means
Wei Feng 0010, Zhenwei Wu, Qianqian Wang 0001, Bo Dong 0001, Zhiqiang Tao, Quanxue Gao
IJCAI6
2024 Label Learning Method Based on Tensor Projection
abstract
Multi-view clustering method based on anchor graph has been widely concerned due to its high efficiency and effectiveness. In order to avoid post-processing, most of the existing anchor graph-based methods learn bipartite graphs with connected components. However, such methods have high requirements on parameters, and in some cases it may not be possible to obtain bipartite graphs with clear connected components. To end this, we propose a label learning method based on tensor projection (LLMTP). Specifically, we project anchor graph into the label space through an orthogonal projection matrix to obtain cluster labels directly. Considering that the spatial structure information of multi-view data may be ignored to a certain extent when projected in different views separately, we extend the matrix projection transformation to tensor projection, so that the spatial structure information between views can be fully utilized. In addition, we introduce the tensor Schatten p-norm regularization to make the clustering label matrices of different views as consistent as possible. Extensive experiments have proved the effectiveness of the proposed method.
Jing Li 0026, Quanxue Gao, Qianqian Wang 0001, Cheng Deng 0002, De-Yan Xie
KDD2
2024 Multi-View Clustering Based on Deep Non-negative Tensor Factorization
abstract
Multi-view clustering (MVC) methods based on non-negative matrix factorization (NMF) have gained popularity owing to their ability to provide interpretable clustering results. However, these NMF-based MVC methods generally process each view independently and thus ignore the potential relationship between views. Besides, they are limited in the ability to capture nonlinear data structures. To overcome these weaknesses and inspired by deep learning, we propose a multi-view clustering method based on deep non-negative tensor factorization (MVC-DNTF). With deep tensor factorization, our method can well exploit the spatial structure of the original data and is capable of extracting more deep and nonlinear features embedded in different views. To further extract the complementary information of different views, we adopt the weighted tensor Schatten p-norm regularization term. An optimization algorithm is developed to effectively solve the MVC-DNTF objective. Extensive experiments are performed to demonstrate the effectiveness and superiority of our method.
Wei Feng 0010, Dongyuan Wei, Qianqian Wang 0001, Bo Dong 0001, Quanxue Gao
ACM Multimedia5
2024 Federated Fuzzy C-means with Schatten-p Norm Minimization
abstract
Multi-view clustering has emerged as an important unsupervised method to process unlabelled multi-view data that provides a comprehensive description of an object. Existing multi-view clustering methods focus on centralized settings but ignore the fact that real-world multi-view data may be distributed across different entities. The sensitive information embedded in multi-view data hinders the cooperative training of multi-view clustering, since data of different views cannot be directly shared, leading to a great challenge to cooperatively exploit the consistent and complementary information of different views. To validate the multi-view clustering in distributed scenarios, in this paper, we propose a novel federated multi-view method named Federated Multi-View Fuzzy C-means with Schatten-p Norm Minimization (FMVFCMSP) which is based on fuzzy C-means and tensor Schatten p-norm. Specifically, we utilize the membership degrees to replace conventional hard clustering assignment in K-means, enabling improved uncertainty handling and less information loss. Moreover, we introduce a tensor Schatten p-norm-based regularizer to fully explore the inter-view complementary information and global spatial structure. We also develop a federated optimization algorithm enabling clients to collaboratively learn the clustering results. Extensive experiments on several datasets demonstrate that our proposed method exhibits superior performance in federated multi-view clustering.
Wei Feng 0010, Zhenwei Wu, Qianqian Wang 0001, Bo Dong 0001, Quanxue Gao
ACM Multimedia5
2024 DFMVC: Deep Fair Multi-view Clustering
abstract
Fair multi-view clustering aims to achieve both satisfactory clustering performance and non-discriminatory outcomes with respect to sensitive attributes. Existing fair multi-view clustering methods impose a constraint that requires the distribution of sensitive attributes to be uniform within each cluster. However, this constraint can lead to misallocation of samples with sensitive attributes. To solve this problem, we propose a novel Deep Fair Multi-View Clustering (DFMVC) method that learns a consistent and discriminative representation instructed by a fairness constraint constructed from the cluster distribution. Specifically, we incorporate contrastive constraints on semantic features from different views to obtain consistent and discriminative representations for each view. Additionally, we align the distribution of sensitive attributes with the target cluster distribution to achieve optimal fairness in clustering results. Experimental results on four datasets with sensitive attributes demonstrate that our method improves fairness and clustering performance compared with state-of-the-art multi-view clustering methods.
Qianqian Wang 0001, Zhiqiang Tao, Wei Feng 0010, Quanxue Gao
ACM Multimedia5
2024 Dual contrastive learning for multi-view clustering
Yichen Bao, Quanxue Gao, Ming Yang 0024
Neurocomputing4
2024 Anchor graph-based multiview spectral clustering
abstract
Significant advances in graph-oriented clustering methods can be attributed to their effectiveness in leveraging relationships and complex structures within multiview data. However, several limitations persist in most existing graph-based multiview clustering approaches. First, quadratic or cubic complexity is required for graph construction or eigendecomposition of the Laplacian matrix in many existing methods. Second, certain methods overlook the differences between views and employ an identical indicator matrix, which can lead to over-learning in practical scenarios. Third, existing methods often neglect spatial structure and complementary information, focusing primarily on calculating error feature-by-feature using different norms. In order to tackle these drawbacks, we propose a new multiview spectral clustering model called A nchor G raph-based M ultiview S pectral C lustering(AG-MSC). AG-MSC incorporates an adaptive weighting mechanism that assigns weights to each view, enhancing the robustness of the algorithm. Using a tensor Schatten p -norm constraint minimizes the discrepancy between indicator matrices obtained from different views, thereby preserving high-order information and spatial structure. To improve computational efficiency, we replace the full adjacency matrices of the corresponding views with anchor graphs. AG-MSC offers a distinct advantage over conventional spectral clustering by directly obtaining all sample categories without additional post-processing steps. We have validated the efficiency of our method through extensive experimental evaluations.
Zuoyuan Niu, Qianqian Wang 0001, Quanxue Gao, Ming Yang 0024
Neurocomputing4
2024 Deep contrastive representation learning for multi-modal clustering
Quanxue Gao
Neurocomputing4
2024 Deep cross-modal subspace clustering with Contrastive Neighbour Embedding
Qianqian Wang 0001, Chengquan Pei, Quanxue Gao
Neurocomputing4
2024 Learning deep representation and discriminative features for clustering of multi-layer networks
Xiaoke Ma 0001, Quan Wang 0006, Maoguo Gong, Quanxue Gao
Neural Networks5
2024 Consistent graph learning for multi-view spectral clustering
De-Yan Xie, Quanxue Gao, Yougang Zhao
Pattern Recognit.2
2024 Non-convex tensorial multi-view clustering by integrating ℓ1-based sliced-Laplacian regularization and ℓ2,p-sparsity
De-Yan Xie, Ming Yang 0024, Quanxue Gao
Pattern Recognit.3
2024 Unsupervised Cross-View Subspace Clustering via Adaptive Contrastive Learning
abstract
Cross-view subspace clustering has become a popular unsupervised method for cross-view data analysis because it can extract both the consistent and complementary features of data for different views. Nonetheless, existing methods usually ignore the discriminative features due to a lack of label supervision, which limits its further improvement in clustering performance. To address this issue, we design a novel model that leverages the self-supervision information embedded in the data itself by combining contrastive learning and self-expression learning, i.e., unsupervised cross-view subspace clustering via adaptive contrastive learning (CVCL). Specifically, CVCL employs an encoder to learn a latent subspace from the cross-view data and convert it to a consistent subspace with a self-expression layer. In this way, contrastive learning helps to provide more discriminative features for the self-expression learning layer, and the self-expression learning layer in turn supervises contrastive learning. Besides, CVCL adaptively chooses positive and negative samples for contrastive learning to reduce the noisy impact of improper negative sample pairs. Ultimately, the decoder is designed for reconstruction tasks, operating on the output of the self-expressive layer, and strives to faithfully restore the original data as much as possible, ensuring that the encoded features are potentially effective. Extensive experiments conducted across multiple cross-view datasets showcase the exceptional performance and superiority of our model.
Qianqian Wang 0001, Quanxue Gao, Chengquan Pei, Wei Feng 0010
IEEE Trans. Big Data3
2024 A Coarse-to-Fine Cell Division Approach for Hyperspectral Remote Sensing Image Classification
abstract
CNNs are widely used in remote sensing image classification because of its outstanding feature extraction ability. However, the classification performance is limited by the complexity of remote sensing scenes and the large inter-class similarity. Furthermore, the existing methods usually distinguish multiple classes of complex targets at the same time, which brings great difficulties to the classification model. To alleviate the above problems, we propose a coarse-to-fine cell division (CFCD) approach to improve HRSIs classification. The algorithm divides the limited labeled samples into two subclasses through continuous decomposition, which reduces the similarity between the ground object classes from the data level. We employ the ℓ12-norm to depict the specific distribution of the target for only two subclasses rather than multiple classes of ground objects, so that the exclusive features of targets can be selected more accurately. Moreover, we propose an optimization process of multi-level training, which not only significantly reduces the difficulty of distinguishing multi-class targets, but also improves the utilization of training samples. Experimental results show that the CFCD algorithm outperforms the state-of-the-art methods with limited training samples on three publicly available HRSIs datasets.
Guangfei Li, Quanxue Gao, Jungong Han, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Self-Supervised Edge Perceptual Learning Framework for High-Resolution Remote Sensing Images Classification
abstract
Self-supervised learning (SSL) has been successfully applied to remote sensing image classification by designing pretext tasks to extract valuable feature representations of targets. However, existing SSL methodologies overlook the edge information integral to ground objects, culminating in frequent misclassifications at target boundaries. Additionally, the scarcity of training samples often restricts the full utilization of the knowledge encapsulated in the pre-training model. To address these issues, we propose a novel self-supervised edge perception learning framework (SEPLF) to improve the classification performance of high-resolution remote sensing images (HRSI). The framework comprises self-supervised edge perception learning (SEPL) and training sample augmentation (TSA) algorithms. On the one hand, the SEPL approach leverages morphological data enhancement strategies to render the extracted invariant features more robust. It also effectively mines the potential information concealed at target edges, augmenting ground objects’s edge separability. On the other hand, the TSA algorithm not only obtains a large number of training samples but also enhances the intra-class diversity of the samples by considering different spectral features of the same category of ground objects. Experimental results validate that our proposed method outperforms state-of-the-art algorithms, particularly with limited labeled samples.
Guangfei Li, Wenbing Liu, Quanxue Gao, Qianqian Wang 0001, Jungong Han, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Efficient Multi-View -Means for Image Clustering
abstract
Nowadays, data in the real world often comes from multiple sources, but most existing multi-view ${K}$ -Means perform poorly on linearly non-separable data and require initializing the cluster centers and calculating the mean, which causes the results to be unstable and sensitive to outliers. This paper proposes an efficient multi-view ${K}$ -Means to solve the above-mentioned issues. Specifically, our model avoids the initialization and computation of clusters centroid of data. Additionally, our model use the Butterworth filters function to transform the adjacency matrix into a distance matrix, which makes the model is capable of handling linearly inseparable data and insensitive to outliers. To exploit the consistency and complementarity across multiple views, our model constructs a third tensor composed of discrete index matrices of different views and minimizes the tensor's rank by tensor Schatten ${p}$ -norm. Experiments on two artificial datasets verify the superiority of our model on linearly inseparable data, and experiments on several benchmark datasets illustrate the performance.
Han Lu 0005, Huafu Xu, Qianqian Wang 0001, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
IEEE Trans. Image Process.4
2024 Unsupervised Discriminative Feature Selection via Contrastive Graph Learning
abstract
Due to many unmarked data, there has been tremendous interest in developing unsupervised feature selection methods, among which graph-guided feature selection is one of the most representative techniques. However, the existing feature selection methods have the following limitations: (1) All of them only remove redundant features shared by all classes and neglect the class-specific properties; thus, the selected features cannot well characterize the discriminative structure of the data. (2) The existing methods only consider the relationship between the data and the corresponding neighbor points by Euclidean distance while neglecting the differences with other samples. Thus, existing methods cannot encode discriminative information well. (3) They adaptively learn the graph in the original or embedding space. Thus, the learned graph cannot characterize the data’s cluster structure. To solve these limitations, we present a novel unsupervised discriminative feature selection via contrastive graph learning, which integrates feature selection and graph learning into a uniform framework. Specifically, our model adaptively learns the affinity matrix, which helps characterize the data’s intrinsic and cluster structures in the original space and the contrastive learning. We minimize ℓ1,2-norm regularization on the projection matrix to preserve class-specific features and remove redundant features shared by all classes. Thus, the selected features encode discriminative information well and characterize the discriminative structure of the data. Generous experiments indicate that our proposed model has state-of-the-art performance.
Qianqian Wang 0001, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
IEEE Trans. Image Process.3
2024 Efficient Anchor Graph Factorization for Multi-View Clustering
abstract
Due to the excellent interpretability of non-negative matrix factorization (NMF), NMF-based multi-view clustering has attracted much attention for multi-media data analysis and processing. However, the existing clustering methods leverage NMF to cluster data matrix, resulting in high computational complexity. Moreover, they are sub-optimal to exploit the complementary information between views because they all measure the between-views error pixel by pixel. To tackle this problem, inspired by orthogonal NMF and anchor graph, we present an efficient anchor graph factorization model with orthogonal, non-negative, and tensor low-rank constraints. We use an anchor graph instead of a data matrix to get an indicator matrix without post-processing, which remarkably reduces the computational complexity. To exploit the between-views complementary information well, we introduce tensor Schatten$p$-norm regularization on the third tensor, composed of soft label matrices of views. The solution can be obtained by iteratively optimizing four decoupled sub-problems, which can be solved more efficiently with good convergence. Through experimental results on the six multi-view datasets, our approach ensures the enhancement of clustering performance while improving efficiency.
Jing Li 0026, Qianqian Wang 0001, Ming Yang 0024, Quanxue Gao, Xinbo Gao 0001
IEEE Trans. Multim.4
2024 Anchor Graph-Based Feature Selection for One-Step Multi-View Clustering
abstract
Recently, multi-view clustering methods have been widely used in handling multi-media data and have achieved impressive performances. Among the many multi-view clustering methods, anchor graph-based multi-view clustering has been proven to be highly efficient for large-scale data processing. However, most existing anchor graph-based clustering methods necessitate post-processing to obtain clustering labels and are unable to effectively utilize the information within anchor graphs. To address this issue, we draw inspiration from regression and feature selection to proposeAnchorGraph-BasedFeatureSelection forOne-stepMulti-ViewClustering (AGFS-OMVC). Our method combines embedding learning and sparse constraint to perform feature selection, allowing us to remove noisy anchor points and redundant connections in the anchor graph. This results in a clean anchor graph that can be projected into the label space, enabling us to obtain clustering labels in a single step without post-processing. Lastly, we employ the tensor Schatten$p$-norm as a tensor rank approximation function to capture the complementary information between different views, ensuring similarity between cluster assignment matrices. Experimental results on five real-world datasets demonstrate that our proposed method outperforms state-of-the-art approaches.
Qin Li 0001, Huafu Xu, Quanxue Gao, Qianqian Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.4
2024 Multi-View Subspace Clustering via Structured Multi-Pathway Network
abstract
Recently, deep multi-view clustering (MVC) has attracted increasing attention in multi-view learning owing to its promising performance. However, most existing deep multi-view methods use single-pathway neural networks to extract features of each view, which cannot explore comprehensive complementary information and multilevel features. To tackle this problem, we propose a deep structured multi-pathway network (SMpNet) for multi-view subspace clustering task in this brief. The proposed SMpNet leverages structured multi-pathway convolutional neural networks to explicitly learn the subspace representations of each view in a layer-wise way. By this means, both low-level and high-level structured features are integrated through a common connection matrix to explore the comprehensive complementary structure among multiple views. Moreover, we impose a low-rank constraint on the connection matrix to decrease the impact of noise and further highlight the consensus information of all the views. Experimental results on five public datasets show the effectiveness of the proposed SMpNet compared with several state-of-the-art deep MVC methods.
Qianqian Wang 0001, Zhiqiang Tao, Quanxue Gao, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.3
2023 Centerless Multi-View K-means Based on the Adjacency Matrix
abstract
Although K-Means clustering has been widely studied due to its simplicity, these methods still have the following fatal drawbacks. Firstly, they need to initialize the cluster centers, which causes unstable clustering performance. Secondly, they have poor performance on non-Gaussian datasets. Inspired by the affinity matrix, we propose a novel multi-view K-Means based on the adjacency matrix. It maps the affinity matrix to the distance matrix according to the principle that every sample has a small distance from the points in its neighborhood and a large distance from the points outside of the neighborhood. Moreover, this method well exploits the complementary information embedded in different views by minimizing the tensor Schatten p-norm regularize on the third-order tensor which consists of cluster assignment matrices of different views. Additionally, this method avoids initializing cluster centroids to obtain stable performance. And there is no need to compute the means of clusters so that our model is not sensitive to outliers. Experiment on a toy dataset shows the excellent performance on non-Gaussian datasets. And other experiments on several benchmark datasets demonstrate the superiority of our proposed method.
Han Lu 0005, Quanxue Gao, Qianqian Wang 0001, Ming Yang 0024, Wei Xia 0007
AAAI2
2023 Dropping Pathways Towards Deep Multi-View Graph Subspace Clustering Networks
abstract
Multi-view graph clustering aims to leverage different views to obtain consistent information and improve clustering performance by sharing the graph structure. Existing multi-view graph clustering algorithms generally adopt a single-pathway network reconstruction and consistent feature extraction, building on top of auto-encoders and graph convolutional networks (GCN). Despite their promising results, these single-pathway methods may ignore the significant complementary information between different layers and the rich multi-level context inside. On the other hand, GCN usually employs a shallow network structure (2-3 layers) due to the over-smoothing with the increase of network depth, while few multi-view graph clustering methods explore the performance of deep networks. In this work, we propose a novel Dropping Pathways strategy toward building a deep Multi-view Graph Subspace Clustering network, namely DPMGSC, to fully exploit the deep and multi-level graph network representations. The proposed method implements a multi-pathway self-expressive network to capture pairwise affinities of graph nodes among multiple views. Moreover, we empirically study the impact of a series of dropping methods on deep multi-pathway networks. Extensive experiments demonstrate the effectiveness of the proposed DPMGSC compared with its deep counterpart and state-of-the-art methods.
Qianqian Wang 0001, Zhiqiang Tao, Quanxue Gao, Wei Feng 0010
ACM Multimedia4
2023 Orthogonal Non-negative Tensor Factorization based Multi-view Clustering
abstract
Multi-view clustering (MVC) based on non-negative matrix factorization (NMF) and its variants have attracted much attention due to their advantages in clustering interpretability. However, existing NMF-based multi-view clustering methods perform NMF on each view respectively and ignore the impact of between-view. Thus, they can't well exploit the within-view spatial structure and between-view complementary information. To resolve this issue, we present orthogonal non-negative tensor factorization (Orth-NTF) and develop a novel multi-view clustering based on Orth-NTF with one-side orthogonal constraint. Our model directly performs Orth-NTF on the 3rd-order tensor which is composed of anchor graphs of views. Thus, our model directly considers the between-view relationship. Moreover, we use the tensor Schatten $p$-norm regularization as a rank approximation of the 3rd-order tensor which characterizes the cluster structure of multi-view data and exploits the between-view complementary information. In addition, we provide an optimization algorithm for the proposed method and prove mathematically that the algorithm always converges to the stationary KKT point. Extensive experiments on various benchmark datasets indicate that our proposed method is able to achieve satisfactory clustering performance.
Jing Li 0026, Quanxue Gao, Qianqian Wang 0001, Ming Yang 0024, Wei Xia 0007
NeurIPS2
2023 Multi-modal object detection via transformer network
abstract
Abstract According to the fact that single‐modal data usually contain limited information, a great deal of effort has been devoted to making use of the complementary information contained in the multi‐modal data on various patterns. Thus, this paper is concerned with an object detection method that can fully utilize multi‐modal data. First, the method introduces the transformer mechanism to realize the fusion of intra‐modal and inter‐modal features of different modal data. The aim is to take advantage of the complementarity of data between modalities, which helps to improve the performance of multi‐modal object detection. Second, a contrastive loss suitable for contrastive learning is applied. This enables the authors to effectively utilize label information. Extensive experiments are conducted on multiple object detection datasets to demonstrate the effectiveness of our proposed method.
Wenbing Liu, Quanxue Gao
IET Image Process.3
2023 Active learning based on similarity level histogram and adaptive-scale sampling for very high resolution image classification
Guangfei Li, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
Neural Networks2
2023 Joint feature selection and optimal bipartite graph learning for subspace clustering
Shikun Mei, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
Neural Networks3
2023 Enhanced tensor low-rank representation learning for multi-view clustering
De-Yan Xie, Quanxue Gao, Ming Yang 0024
Neural Networks2
2023 Low-rank discrete multi-view spectral clustering
Yu Yun, Jing Li 0026, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
Neural Networks3
2023 Contrastive self-representation learning for data clustering
Quanxue Gao, Shikun Mei, Ming Yang 0024
Neural Networks2
2023 Sparse discriminant PCA based on contrastive learning and class-specificity distribution
Quanxue Gao, Qianqian Wang 0001, Ming Yang 0024, Xinbo Gao 0001
Neural Networks2
2023 Tensorized Bipartite Graph Learning for Multi-View Clustering
abstract
Despite the impressive clustering performance and efficiency in characterizing both the relationship between the data and cluster structure, most existing graph-based multi-view clustering methods still have the following drawbacks. They suffer from the expensive time burden due to both the construction of graphs and eigen-decomposition of Laplacian matrix. Moreover, none of them simultaneously considers the similarity of inter-view and similarity of intra-view. In this article, we propose a variance-based de-correlation anchor selection strategy for bipartite construction. The selected anchors not only cover the whole classes but also characterize the intrinsic structure of data. Following that, we present a tensorized bipartite graph learning for multi-view clustering (TBGL). Specifically, TBGL exploits the similarity of inter-view by minimizing the tensor Schatten p-norm, which well exploits both the spatial structure and complementary information embedded in the bipartite graphs of views. We exploit the similarity of intra-view by using the [Formula: see text]-norm minimization regularization and connectivity constraint on each bipartite graph. So the learned graph not only well encodes discriminative information but also has the exact connected components which directly indicates the clusters of data. Moreover, we solve TBGL by an efficient algorithm which is time-economical and has good convergence. Extensive experimental results demonstrate that TBGL is superior to the state-of-the-art methods. Codes and datasets are available: https://github.com/xdweixia/TBGL-MVC.
Wei Xia 0007, Quanxue Gao, Qianqian Wang 0001, Xinbo Gao 0001, Chris Ding, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Graph Embedding Contrastive Multi-Modal Representation Learning for Clustering
abstract
Multi-modal clustering (MMC) aims to explore complementary information from diverse modalities for clustering performance facilitating. This article studies challenging problems in MMC methods based on deep neural networks. On one hand, most existing methods lack a unified objective to simultaneously learn the inter- and intra-modality consistency, resulting in a limited representation learning capacity. On the other hand, most existing processes are modeled for a finite sample set and cannot handle out-of-sample data. To handle the above two challenges, we propose a novel Graph Embedding Contrastive Multi-modal Clustering network (GECMC), which treats the representation learning and multi-modal clustering as two sides of one coin rather than two separate problems. In brief, we specifically design a contrastive loss by benefiting from pseudo-labels to explore consistency across modalities. Thus, GECMC shows an effective way to maximize the similarities of intra-cluster representations while minimizing the similarities of inter-cluster representations at both inter- and intra-modality levels. So, the clustering and representation learning interact and jointly evolve in a co-training framework. After that, we build a clustering layer parameterized with cluster centroids, showing that GECMC can learn the clustering labels with given samples and handle out-of-sample data. GECMC yields superior results than 14 competitive methods on four challenging datasets. Codes and datasets are available: https://github.com/xdweixia/GECMC.
Wei Xia 0007, Tianxiu Wang, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
IEEE Trans. Image Process.3
2023 Self-Weighted Anchor Graph Learning for Multi-View Clustering
abstract
Graph-based multi-view clustering method has attracted considerable attention in multi-media data analyse community due to its good clustering performance and efficiency in characterizing the relationship between data. But the existing graph-based clustering methods still have many shortcomings. Firstly, they have high computational complexity due to the eigenvalue decomposition. Secondly, the complementary information and spatial structure embedded in different views can affect the clustering performance. However, some existing graph-based clustering methods do not consider these two points. In this article, we use the anchor graphs of different views as input, which effectively reduces the computational complexity. And then we explicitly consider the complementary information and spatial structure between anchor graphs of different views by minimizing the tensor Schatten$p$-norm, aiming to achieve a better tensor with low-rank approximation. Finally, we learn the view-consensus anchor graph with connectivity constraints, which can directly indicate clusters by self-weighted strategy. An efficient alternating algorithm is then derived to optimize the proposed multi-view special clustering model. Furthermore, the constructed sequence was proved to converge to the stationary KKT point. Experiments show that our proposed method not only reduces the time cost, but also outperforms the most advanced methods.
Xiaochuang Shu, Quanxue Gao, Ming Yang 0024, Rong Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2023 Self-Consistent Contrastive Attributed Graph Clustering With Pseudo-Label Prompt
abstract
Attributed graph clustering, which learns node representation from node attribute and topological graph for clustering, is a fundamental and challenging task for multimedia network-structured data analysis. Recently, graph contrastive learning (GCL)-based methods have obtained impressive clustering performance on this task. Nevertheless, there still remain some limitations to be solved: 1) most existing methods fail to consider the self-consistency between latent representations and cluster structures; and 2) most methods require a post-processing operation to get clustering labels. Such a two-step learning scheme results in models that cannot handle newly generated data,i.e., out-of-sample (OOS) nodes. To address these issues in a unified framework, aSelf-consistentContrastiveAttributedGraphClustering (SCAGC) network with pseudo-label prompt is proposed in this article. In SCAGC, by clustering labels prompt information, a self-consistent contrastive loss, which aims to maximize the consistencies of intra-cluster representations while minimizing the consistencies of inter-cluster representations, is designed for representation learning. Meanwhile, a clustering module is built to directly output clustering labels by contrasting the representation of different clusters. Thus, for the OOS nodes, SCAGC can directly calculate their clustering labels. Extensive experimental results on seven benchmark datasets have shown that SCAGC consistently outperforms 16 competitive clustering methods.
Wei Xia 0007, Qianqian Wang 0001, Quanxue Gao, Ming Yang 0024, Xinbo Gao 0001
IEEE Trans. Multim.3
2023 Adversarial Multiview Clustering Networks With Adaptive Fusion
abstract
The existing deep multiview clustering (MVC) methods are mainly based on autoencoder networks, which seek common latent variables to reconstruct the original input of each view individually. However, due to the view-specific reconstruction loss, it is challenging to extract consistent latent representations over multiple views for clustering. To address this challenge, we propose adversarial MVC (AMvC) networks in this article. The proposed AMvC generates each view’s samples conditioning on the fused latent representations among different views to encourage a more consistent clustering structure. Specifically, multiview encoders are used to extract latent descriptions from all the views, and the corresponding generators are used to generate the reconstructed samples. The discriminative networks and the mean squared loss are jointly utilized for training the multiview encoders and generators to balance the distinctness and consistency of each view’s latent representation. Moreover, an adaptive fusion layer is developed to obtain a shared latent representation, on which a clustering loss and the${\ell _ {1,2}}$-norm constraint are further imposed to improve clustering performance and distinguish the latent space. Experimental results on video, image, and text datasets demonstrate that the effectiveness of our AMvC is over several state-of-the-art deep MVC methods.
Qianqian Wang 0001, Zhiqiang Tao, Wei Xia 0007, Quanxue Gao, Xiaochun Cao, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.4
2022 A multi-view clustering framework via integrating K-means and graph-cut
Han Lu 0005, Quanxue Gao, Wei Xia 0007
Neurocomputing2
2022 Contrastive deep embedded clustering
Guoshuai Sheng, Qianqian Wang 0001, Chengquan Pei, Quanxue Gao
Neurocomputing4
2022 Multi-view Spectral Clustering with Adaptive Graph Learning and Tensor Schatten p-norm
Yu Yun, Quanxue Gao
Neurocomputing5
2022 Low-rank constraint bipartite graph learning
Haizhou Yang, Quanxue Gao
Neurocomputing3
2022 Fast multiple graphs learning for multi-view clustering
Quanxue Gao
Neural Networks2
2022 Multi-view graph embedding clustering network: Joint self-supervision and block diagonal representation
Wei Xia 0007, Ming Yang 0024, Quanxue Gao, Jungong Han, Xinbo Gao 0001
Neural Networks4
2022 Attributes learning network for generalized zero-shot learning
Yu Yun, Mingzhen Hou, Quanxue Gao
Neural Networks4
2022 Enhanced nuclear norm based matrix regression for occluded face recognition
Qin Li 0001, Huihui He, Hong Lai, Tie Cai, Qianqian Wang 0001, Quanxue Gao
Pattern Recognit.6
2022 Semi-Supervised Clustering via Cannot Link Relationship for Multiview Data
abstract
Due to the diversity of data modalities, the research interest of multi-view clustering is gradually increasing, in the field of large-data analytics, particularly in clustering. However, the greater part of current multi-view clustering methods is mainly in view of unsupervised learning, which leads to unpredictable results and algorithmic instability. Besides, they ignore the diversity of graphs, which is not desirable in practical applications, because the characteristic properties of each view are different. To solve these problems, inspired by the outstanding performance of semi-supervised learning in machine learning, we propose a valid semi-supervised multi-view spectral clustering algorithm. We use the pre-set labels as prior knowledge to obtain the overall distribution of the remaining unlabeled data. Tensor minimization Schatten$p$-norm is utilized to mine the mutual information hidden in multiple views. Meanwhile, we also use the cannot-link as another semi-supervised constraint to update the graph. Our proposed algorithm is generally 5%-10% better than the comparison algorithms in view of the experimental results on five datasets, and our algorithm is relatively fast with the computational complexity of$\mathcal {O}({T({n^{2}}\log (n) + {n^{2}} + {u^{2}}l + ulc + uc\log (c))})$, where$T$denotes the number of iterations and$n$,$l$,$u$represent the number of samples, the number of labeled and unlabeled samples, respectively, which shows that our proposed method has broad application prospects.
Quanxue Gao
IEEE Trans. Circuits Syst. Video Technol.2
2022 Tensor Completion-Based Incomplete Multiview Clustering
abstract
Incomplete multiview clustering is a challenging problem in the domain of unsupervised learning. However, the existing incomplete multiview clustering methods only consider the similarity structure of intraview while neglecting the similarity structure of interview. Thus, they cannot take advantage of both the complementary information and spatial structure embedded in similarity matrices of different views. To this end, we complete the incomplete graph with missing data referring to tensor complete and present a novel and effective model to handel the incomplete multiview clustering task. To be specific, we consider the similarity of the interview graphs via the tensor Schatten p -norm-based completion technique to make use of both the complementary information and spatial structure. Meanwhile, we employ the connectivity constraint for similarity matrices of different views such that the connected components approximately represent clusters. Thus, the learned entire graph not only has the low-rank structure but also well characterizes the relationship between unmissing data. Extensive experiments show the promising performance of the proposed method comparing with several incomplete multiview approaches in the clustering tasks.
Wei Xia 0007, Quanxue Gao, Qianqian Wang 0001, Xinbo Gao 0001
IEEE Trans. Cybern.2
2022 Multiview Subspace Clustering by an Enhanced Tensor Nuclear Norm
abstract
Despite the promising preliminary results, tensor-singular value decomposition (t-SVD)-based multiview subspace is incapable of dealing with real problems, such as noise and illumination changes. The major reason is that tensor-nuclear norm minimization (TNNM) used in t-SVD regularizes each singular value equally, which does not make sense in matrix completion and coefficient matrix learning. In this case, the singular values represent different perspectives and should be treated differently. To well exploit the significant difference between singular values, we study the weighted tensor Schatten p -norm based on t-SVD and develop an efficient algorithm to solve the weighted tensor Schatten p -norm minimization (WTSNM) problem. After that, applying WTSNM to learn the coefficient matrix in multiview subspace clustering, we present a novel multiview clustering method by integrating coefficient matrix learning and spectral clustering into a unified framework. The learned coefficient matrix well exploits both the cluster structure and high-order information embedded in multiview views. The extensive experiments indicate the efficiency of our method in six metrics.
Wei Xia 0007, Quanxue Gao, Xiaochuang Shu, Jungong Han, Xinbo Gao 0001
IEEE Trans. Cybern.3
2022 View-Consistency Learning for Incomplete Multiview Clustering
abstract
In this article, we present a novel general framework for incomplete multi-view clustering by integrating graph learning and spectral clustering. In our model, a tensor low-rank constraint are introduced to learn a stable low-dimensional representation, which encodes the complementary information and takes into account the cluster structure between different views. A corresponding algorithm associated with augmented Lagrangian multipliers is established. In particular, tensor Schatten p -norm is used as a tighter approximation to the tensor rank function. Besides, both consistency and specificity are jointly exploited for subspace representation learning. Extensive experiments on benchmark datasets demonstrate that our model outperforms several baseline methods in incomplete multi-view clustering.
Ziyu Lv, Quanxue Gao, Qin Li 0001, Ming Yang 0024
IEEE Trans. Image Process.2
2022 Multiview Spectral Clustering With Bipartite Graph
abstract
Multi-view spectral clustering has become appealing due to its good performance in capturing the correlations among all views. However, on one hand, many existing methods usually require a quadratic or cubic complexity for graph construction or eigenvalue decomposition of Laplacian matrix; on the other hand, they are inefficient and unbearable burden to be applied to large scale data sets, which can be easily obtained in the era of big data. Moreover, the existing methods cannot encode the complementary information between adjacency matrices, i.e., similarity graphs of views and the low-rank spatial structure of adjacency matrix of each view. To address these limitations, we develop a novel multi-view spectral clustering model. Our model well encodes the complementary information by Schatten p -norm regularization on the third tensor whose lateral slices are composed of the adjacency matrices of the corresponding views. To further improve the computational efficiency, we leverage anchor graphs of views instead of full adjacency matrices of the corresponding views, and then present a fast model that encodes the complementary information embedded in anchor graphs of views by Schatten p -norm regularization on the tensor bipartite graph. Finally, an efficient alternating algorithm is derived to optimize our model. The constructed sequence was proved to converge to the stationary KKT point. Extensive experimental results indicate that our method has good performance.
Haizhou Yang, Quanxue Gao, Wei Xia 0007, Ming Yang 0024, Xinbo Gao 0001
IEEE Trans. Image Process.2
2022 Zero-Shot Learning Based on Quality-Verifying Adversarial Network
abstract
Recently, generative adversarial network (GAN)-based zero-shot learning methods have attracted widespread attention. However, due to the randomness of GAN generation, most existing methods cannot well guarantee to generate sufficiently reliable features and have good generalization ability. Targeting at these problems, we propose an effective Quality-Verifying Adversarial Network (QVAN) that consists of one generator and double discriminators. Adversarial learning between the former discriminator and generator is to generate visual features, which can be partitioned into pseudo-generated features and reliable-generated features. The latter discriminator is used for quality-verifying that will guide the generator to generate more reliable features that are near the real visual features. To avoid over-fitting and ensure intra-class diversity, we set the threshold for each class to distinguish pseudo-generated features and reliable-generated features. To further preserve both compactness and discriminability of the samples, we introduce the class metric constraint, which are more conducive to classification. Moreover, we introduce$\ell _{1,2}$-norm constraint to fully consider the specific distribution among different classes, thus making the generated features more discriminant. Extensive experiments on several real-world datasets show the effectiveness of the proposed approach, which demonstrate the advantage over the state-of-the-art methods.
Siyang Deng, Gang Xiang, Quanxue Gao, Wei Xia 0007, Xinbo Gao 0001
IEEE Trans. Multim.3
2022 Self-Supervised Graph Convolutional Network for Multi-View Clustering
abstract
Despite the promising preliminary results, existing graph convolutional network (GCN) based multi-view learning methods directly use the graph structure as view descriptor, which may inhibit the ability of multi-view learning for multimedia data. The major reason is that, in real multimedia applications, the graph structure may contain outliers. Moreover, they fail to take advantage of the information embedded in the inaccurate clustering labels obtained from their proposed methods, resulting in inferior clustering results. These observations motivate us to study whether there is a better alternative GCN based framework for multi-view clustering. To this end, in this paper, we propose an end-to-end self-supervised graph convolutional network for multi-view clustering (SGCMC). Specifically, SGCMC constructs a new view descriptor for graph-structured data by mapping the raw node content into the complex space via Euler transformation, which not only suppresses outliers but also reveals non-linear patterns embedded in data. Meanwhile, the proposed SGCMC uses the clustering labels to guide the learning of the latent representation and coefficient matrix, and the latter in turn is used to conduct the subsequent node clustering. By this way, clustering and representation learning are seamlessly connected, with the aim to achieve better clustering results. Extensive experiments indicate that the proposed SGCMC outperforms the state-of-the-art methods.
Wei Xia 0007, Qianqian Wang 0001, Quanxue Gao, Xinbo Gao 0001
IEEE Trans. Multim.3
2021 Deep Self-Supervised t-SNE for Multi-modal Subspace Clustering
abstract
Existing multi-modal subspace clustering methods, aiming to exploit the correlation information between different modalities, have achieved promising preliminary results. However, these methods might be incapable of handling real problems with complex heterogeneous structures between different modalities, since the large heterogeneous structure makes it difficult to directly learn a discriminative shared self-representation for multi-modal clustering. To tackle this problem, in this paper, we propose a deep Self-supervised t-SNE method (StSNE) for multi-modal subspace clustering, which learns soft label features by multi-modal encoders and utilizes the common label feature to supervise soft label feature of each modal by adversarial training and reconstruction networks. Specifically, the proposed StSNE consists of four components: 1) multi-modal convolutional encoders; 2) a self-supervised t-SNE module; 3) a self-expressive layer; 4) multi-modal convolutional decoders. Multi-modal data are fed to encoders to obtain soft label features, for which the self-supervised t-SNE module is added to make full use of the label information among different modalities. Simultaneously, the latent representations given by encoders are constrained by a self-expressive layer to capture the hierarchical information of each modal, followed by decoders reconstructing the encoded features to preserve the structure of the original data. Experimental results on several public datasets demonstrate the superior clustering performance of the proposed method over state-of-the-art methods.
Qianqian Wang 0001, Wei Xia 0007, Zhiqiang Tao, Quanxue Gao, Xiaochun Cao
ACM Multimedia4
2021 Self-representation and matrix factorization based multi-view clustering
Ying Dou, Yu Yun, Quanxue Gao
Neurocomputing3
2021 Self-supervised graph convolutional clustering by preserving latent distribution
Shiwen Kou, Wei Xia 0007, Quanxue Gao, Xinbo Gao 0001
Neurocomputing4
2021 Multi-view clustering by joint spectral embedding and spectral rotation
Zhizhen Wan, Huiling Xu, Quanxue Gao
Neurocomputing3
2021 Regression-based clustering network via combining prior information
Wei Xia 0007, Quanxue Gao, Qianqian Wang 0001, Xinbo Gao 0001
Neurocomputing2
2021 Adversarial self-supervised clustering with cluster-specificity distribution
Wei Xia 0007, Quanxue Gao, Xinbo Gao 0001
Neurocomputing3
2021 Self-representation and Class-Specificity Distribution Based Multi-View Clustering
Yu Yun, Wei Xia 0007, Quanxue Gao, Xinbo Gao 0001
Neurocomputing4
2021 Multiple graphs learning with a new weighted tensor nuclear norm
De-Yan Xie, Quanxue Gao, Siyang Deng, Xinbo Gao 0001
Neural Networks2
2021 Graph embedding clustering: Graph attention auto-encoder with cluster-specificity distribution
Huiling Xu, Wei Xia 0007, Quanxue Gao, Jungong Han, Xinbo Gao 0001
Neural Networks3
2021 Enhanced Tensor RPCA and its Application
abstract
Despite the promising results, tensor robust principal component analysis (TRPCA), which aims to recover underlying low-rank structure of clean tensor data corrupted with noise/outliers by shrinking all singular values equally, cannot well preserve the salient content of image. The major reason is that, in real applications, there is a salient difference information between all singular values of a tensor image, and the larger singular values are generally associated with some salient parts in the image. Thus, the singular values should be treated differently. Inspired by this observation, we investigate whether there is a better alternative solution when using tensor rank minimization. In this paper, we develop an enhanced TRPCA (ETRPCA) which explicitly considers the salient difference information between singular values of tensor data by the weighted tensor Schatten p-norm minimization, and then propose an efficient algorithm, which has a good convergence, to solve ETRPCA. Extensive experimental results reveal that the proposed method ETRPCA is superior to several state-of-the-art variant RPCA methods in terms of performance.
Quanxue Gao, Wei Xia 0007, De-Yan Xie, Xinbo Gao 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Cross-view classification by joint adversarial learning and class-specificity distribution
Siyang Deng, Wei Xia 0007, Quanxue Gao, Xinbo Gao 0001
Pattern Recognit.3
2021 Relation-based Discriminative Cooperation Network for Zero-Shot Classification
Yang Liu 0069, Xinbo Gao 0001, Quanxue Gao, Jungong Han, Ling Shao 0001
Pattern Recognit.3
2021 Generative Partial Multi-View Clustering With Adaptive Fusion and Cycle Consistency
abstract
Nowadays, with the rapid development of data collection sources and feature extraction methods, multi-view data are getting easy to obtain and have received increasing research attention in recent years, among which, multi-view clustering (MVC) forms a mainstream research direction and is widely used in data analysis. However, existing MVC methods mainly assume that each sample appears in all the views, without considering the incomplete view case due to data corruption, sensor failure, equipment malfunction, etc. In this study, we design and build a generative partial multi-view clustering model with adaptive fusion and cycle consistency, named as GP-MVC, to solve the incomplete multi-view problem by explicitly generating the data of missing views. The main idea of GP-MVC lies in two-fold. First, multi-view encoder networks are trained to learn common low-dimensional representations, followed by a clustering layer to capture the shared cluster structure across multiple views. Second, view-specific generative adversarial networks with multi-view cycle consistency are developed to generate the missing data of one view conditioning on the shared representation given by other views. These two steps could be promoted mutually, where the learned common representation facilitates data imputation and the generated data could further explores the view consistency. Moreover, an weighted adaptive fusion scheme is implemented to exploit the complementary information among different views. Experimental results on four benchmark datasets are provided to show the effectiveness of the proposed GP-MVC over the state-of-the-art methods.
Qianqian Wang 0001, Zhengming Ding, Zhiqiang Tao, Quanxue Gao, Yun Fu 0001
IEEE Trans. Image Process.4
2021 Adversarial Multi-Path Residual Network for Image Super-Resolution
abstract
Recently, deep convolutional neural networks have demonstrated remarkable progresses on single image super-resolution (SR) problem. However, most of them use more deeper and wider networks to improve SR performance, which is not practical in real-world applications due to large complexity, high computation cost, and low efficiency. In addition, they cannot provide high perception quality and guarantee objective quality simultaneously. To address these limitations, we in this paper propose a novel Adversarial Multipath Residual Network (AMPRN), which can largely suppress the number of network parameters and achieve a higher SR performance compared with the state-of-the-art methods. More specifically, we propose a multi-path residual block (MPRB) for multi-path residual network (MPRN) with fewer network parameters, which can extract abundant local features by fully using features from different paths generated by channel slices. These hierarchical features from all the MPRBs are then jointly aggregated by global gradual feature fusion. Following MPRN, we construct an adversarial gradient network with a gradient loss to make the gradient distribution of the generated SR images and ground truth image closer. In this way, the generated SR images of our model can provide high perception quality and objective quality. Finally, several experimental results demonstrate that our AMPRN achieves better performance in comparison with fewer parameters than the state-of-the-art methods.
Qianqian Wang 0001, Quanxue Gao, Linlu Wu, Gan Sun, Licheng Jiao
IEEE Trans. Image Process.2
2021 iCmSC: Incomplete Cross-Modal Subspace Clustering
abstract
Cross-modal clustering aims to cluster the high-similar cross-modal data into one group while separating the dissimilar data. Despite the promising cross-modal methods have developed in recent years, existing state-of-the-arts cannot effectively capture the correlations between cross-modal data when encountering with incomplete cross-modal data, which can gravely degrade the clustering performance. To well tackle the above scenario, we propose a novel incomplete cross-modal clustering method that integrates canonical correlation analysis and exclusive representation, named incomplete Cross-modal Subspace Clustering (i.e., iCmSC). To learn a consistent subspace representation among incomplete cross-modal data, we maximize the intrinsic correlations among different modalities by deep canonical correlation analysis (DCCA), while an exclusive self-expression layer is proposed after the output layers of DCCA. We exploit a ℓ1,2-norm regularization in the learned subspace to make the learned representation more discriminative, which makes samples between different clusters mutually exclusive and samples among the same cluster attractive to each other. Meanwhile, the decoding networks are employed to reconstruct the feature representation, and further preserve the structural information among the original cross-modal data. To the end, we demonstrate the effectiveness of the proposed iCmSC via extensive experiments, which can justify that iCmSC achieves consistently large improvement compared with the state-of-the-arts.
Qianqian Wang 0001, Huanhuan Lian, Gan Sun, Quanxue Gao, Licheng Jiao
IEEE Trans. Image Process.4
2021 Deep Multi-View Subspace Clustering With Unified and Discriminative Learning
abstract
Deep multi-view subspace clustering has achieved promising performance compared with other multi-view clustering. However, existing deep multi-view subspace clustering only considers the global structure for all views, and they ignore the local geometric structure among each view. In addition, they cannot learn discriminative feature on different clusters of different views, i.e., inter-cluster difference. To solve these problems, in this paper, we propose a novel Deep Multi-view Subspace Clustering with Unified and Discriminative Learning (DMSC-UDL). DMSC-UDL combines global and local structures with self-expression layer. The global and local structures help each other forward and achieve small distance between samples of the same cluster. To make samples in different clusters of different views farther, DMSC-UDL uses a discriminative constraint between different views. In this way, DMSC-UDL makes the same cluster's samples have large weights, while different clusters' samples have small weights. Thus, it can learn a better shared connection matrix for multi-view clustering. Extensive experimental results reveal that the proposed multi-view clustering method is superior to several state-of-the-art multi-view clustering methods in terms of performance.
Qianqian Wang 0001, Jiafeng Cheng, Quanxue Gao, Guoshuai Zhao 0001, Licheng Jiao
IEEE Trans. Multim.3
2020 Cross-Modal Subspace Clustering via Deep Canonical Correlation Analysis
abstract
For cross-modal subspace clustering, the key point is how to exploit the correlation information between cross-modal data. However, most hierarchical and structural correlation information among cross-modal data cannot be well exploited due to its high-dimensional non-linear property. To tackle this problem, in this paper, we propose an unsupervised framework named Cross-Modal Subspace Clustering via Deep Canonical Correlation Analysis (CMSC-DCCA), which incorporates the correlation constraint with a self-expressive layer to make full use of information among the inter-modal data and the intra-modal data. More specifically, the proposed model consists of three components: 1) deep canonical correlation analysis (Deep CCA) model; 2) self-expressive layer; 3) Deep CCA decoders. The Deep CCA model consists of convolutional encoders and correlation constraint. Convolutional encoders are used to obtain the latent representations of cross-modal data, while adding the correlation constraint for the latent representations can make full use of the information of the inter-modal data. Furthermore, self-expressive layer works on latent representations and constrain it perform self-expression properties, which makes the shared coefficient matrix could capture the hierarchical intra-modal correlations of each modality. Then Deep CCA decoders reconstruct data to ensure that the encoded features can preserve the structure of the original data. Experimental results on several real-world datasets demonstrate the proposed method outperforms the state-of-the-art methods.
Quanxue Gao, Huanhuan Lian, Qianqian Wang 0001, Gan Sun
AAAI1
2020 Tensor-SVD Based Graph Learning for Multi-View Subspace Clustering
abstract
Low-rank representation based on tensor-Singular Value Decomposition (t-SVD) has achieved impressive results for multi-view subspace clustering, but it does not well deal with noise and illumination changes embedded in multi-view data. The major reason is that all the singular values have the same contribution in tensor-nuclear norm based on t-SVD, which does not make sense in the existence of noise and illumination change. To improve the robustness and clustering performance, we study the weighted tensor-nuclear norm based on t-SVD and develop an efficient algorithm to optimize the weighted tensor-nuclear norm minimization (WTNNM) problem. We further apply the WTNNM algorithm to multi-view subspace clustering by exploiting the high order correlations embedded in different views. Extensive experimental results reveal that our WTNNM method is superior to several state-of-the-art multi-view subspace clustering methods in terms of performance.
Quanxue Gao, Zhizhen Wan, De-Yan Xie
AAAI1
2020 Multi-View Attribute Graph Convolution Networks for Clustering
abstract
Graph neural networks (GNNs) have made considerable achievements in processing graph-structured data. However, existing methods can not allocate learnable weights to different nodes in the neighborhood and lack of robustness on account of neglecting both node attributes and graph reconstruction. Moreover, most of multi-view GNNs mainly focus on the case of multiple graphs, while designing GNNs for solving graph-structured data of multi-view attributes is still under-explored. In this paper, we propose a novel Multi-View Attribute Graph Convolution Networks (MAGCN) model for the clustering task. MAGCN is designed with two-pathway encoders that map graph embedding features and learn the view-consistency information. Specifically, the first pathway develops multi-view attribute graph attention networks to reduce the noise/redundancy and learn the graph embedding features for each multi-view graph data. The second pathway develops consistent embedding encoders to capture the geometric relationship and probability distribution consistency among different views, which adaptively finds a consistent clustering embedding space for multi-view attributes. Experiments on three benchmark graph datasets show the superiority of our method compared with several state-of-the-art algorithms.
Jiafeng Cheng, Qianqian Wang 0001, Zhiqiang Tao, De-Yan Xie, Quanxue Gao
IJCAI5
2020 Fast algorithm for large-scale subspace clustering by LRR
abstract
Low‐rank representation (LRR) and its variants have been proved to be powerful tools for handling subspace clustering problems. Most of these methods involve a sub‐problem of computing the singular value decomposition of an matrix, which leads to a computation complexity of . Obviously, when n is large, it will be time consuming. To address this problem, the authors introduce a fast solution, which reformulates the large‐scale problem to an equal form with smaller size. Thus, the proposed method remarkably reduces the computation complexity by solving a small‐scale problem. Theoretical analysis proves the efficiency of the proposed model. Furthermore, we extend LRR to a general model by using Schatten p ‐norm instead of nuclear norm and present a fast algorithm to solve large‐scale problem. Experiments on MNIST and Caltech101 databse illustrate the equivalence of the proposed algorithm and the original LRR solver. Experimental results show that the proposed algorithm is remarkably faster than traditional LRR algorithm, especially in the case of large sample number.
De-Yan Xie, Feiping Nie 0001, Quanxue Gao
IET Image Process.3
2020 Discriminative comparison classifier for generalized zero-shot learning
Mingzhen Hou, Wei Xia 0007, Quanxue Gao
Neurocomputing4
2020 Double robust principal component analysis
Qianqian Wang 0001, Quanxue Gao, Gan Sun, Chris Ding
Neurocomputing2
2020 Multi-view clustering by joint manifold learning and tensor nuclear norm
De-Yan Xie, Wei Xia 0007, Qianqian Wang 0001, Quanxue Gao
Neurocomputing4
2020 On the optimal solution to maximum margin projection pursuit
De-Yan Xie, Feiping Nie 0001, Quanxue Gao
Multim. Tools Appl.3
2020 Multi-view projected clustering with graph learning
Quanxue Gao, Zhizhen Wan, Qianqian Wang 0001, Yang Liu 0069, Ling Shao 0001
Neural Networks1
2020 Label-activating framework for zero-shot learning
Yang Liu 0069, Xinbo Gao 0001, Quanxue Gao, Jungong Han, Ling Shao 0001
Neural Networks3
2020 Adaptive latent similarity learning for multi-view clustering
De-Yan Xie, Quanxue Gao, Qianqian Wang 0001, Xinbo Gao 0001
Neural Networks2
2020 Low-rank tensor constrained co-regularized multi-view spectral clustering
Huiling Xu, Wei Xia 0007, Quanxue Gao, Xinbo Gao 0001
Neural Networks4
2020 Multiview Clustering by Joint Latent Representation and Similarity Learning
abstract
Subspace learning-based multiview clustering has achieved impressive experimental results. However, the similarity matrix, which is learned by most existing methods, cannot well characterize both the intrinsic geometric structure of data and the neighbor relationship between data. To consider the fact that original data space does not well characterize the intrinsic geometric structure, we learn the latent representation of data, which is shared by different views, from the latent subspace rather than the original data space by linear transformation. Thus, the learned latent representation has a low-rank structure without solving the nuclear-norm. This reduces the computational complexity. Then, the similarity matrix is adaptively learned from the learned latent representation by manifold learning which well characterizes the local intrinsic geometric structure and neighbor relationship between data. Finally, we integrate clustering, manifold learning, and latent representation into a unified framework and develop a novel subspace learning-based multiview clustering method. Extensive experiments on benchmark datasets demonstrate the superiority of our method.
De-Yan Xie, Quanxue Gao, Jiale Han 0001, Xinbo Gao 0001
IEEE Trans. Cybern.3
2019 Deep Adversarial Multi-view Clustering Network
abstract
Multi-view clustering has attracted increasing attention in recent years by exploiting common clustering structure across multiple views. Most existing multi-view clustering algorithms use shallow and linear embedding functions to learn the common structure of multi-view data. However, these methods cannot fully utilize the non-linear property of multi-view data, which is important to reveal complex cluster structure underlying multi-view data. In this paper, we propose a novel multi-view clustering method, named Deep Adversarial Multi-view Clustering (DAMC) network, to learn the intrinsic structure embedded in multi-view data. Specifically, our model adopts deep auto-encoders to learn latent representations shared by multiple views, and meanwhile leverages adversarial training to further capture the data distribution and disentangle the latent space. Experimental results on several real-world datasets demonstrate that the proposed method outperforms the state-of art methods.
Qianqian Wang 0001, Zhiqiang Tao, Quanxue Gao, Zhaohua Yang
IJCAI4
2019 Worst-Case Discriminative Feature Selection
abstract
Feature selection plays a critical role in data mining, driven by increasing feature dimensionality in target problems. In this paper, we propose a new criterion for discriminative feature selection, worst-case discriminative feature selection (WDFS). Unlike Fisher Score and other methods based on the discriminative criteria considering the overall (or average) separation of data, WDFS adopts a new perspective called worst-case view which arguably is more suitable for classification applications. Specifically, WDFS directly maximizes the ratio of the minimum of between-class variance of all class pairs over the maximum of within-class variance, and thus it duly considers the separation of all classes. Otherwise, we take a greedy strategy by finding one feature at a time, but it is very easy to implement. Moreover, we also utilize the correlation between features to help reduce the redundancy and extend WDFS to uncorrelated WDFS (UWDFS). To evaluate the effectiveness of the proposed algorithm, we conduct classification experiments on many real data sets. In the experiment, we respectively use the original features and the score vectors of features over all class pairs to calculate the correlation coefficients, and analyze the experimental results in these two ways. Experimental results demonstrate the effectiveness of WDFS and UWDFS.
Shuangli Liao, Quanxue Gao, Feiping Nie 0001, Yang Liu 0069
IJCAI2
2019 Graph and Autoencoder Based Feature Extraction for Zero-shot Learning
abstract
Zero-shot learning (ZSL) aims to build models to recognize novel visual categories that have no associated labelled training samples. The basic framework is to transfer knowledge from seen classes to unseen classes by learning the visual-semantic embedding. However, most of approaches do not preserve the underlying sub-manifold of samples in the embedding space. In addition, whether the mapping can precisely reconstruct the original visual feature is not investigated in-depth. In order to solve these problems, we formulate a novel framework named Graph and Autoencoder Based Feature Extraction (GAFE) to seek a low-rank mapping to preserve the sub-manifold of samples. Taking the encoder-decoder paradigm, the encoder part learns a mapping from the visual feature to the semantic space, while decoder part reconstructs the original features with the learned mapping. In addition, a graph is constructed to guarantee the learned mapping can preserve the local intrinsic structure of the data. To this end, an L21 norm sparsity constraint is imposed on the mapping to identify features relevant to the target domain. Extensive experiments on five attribute datasets demonstrate the effectiveness of the proposed model.
Yang Liu 0069, De-Yan Xie, Quanxue Gao, Jungong Han, Shujian Wang, Xinbo Gao 0001
IJCAI3
2019 Hyperspectral image denoising via minimizing the partial sum of singular values and superpixel segmentation
Yang Liu 0069, Caifeng Shan, Quanxue Gao, Xinbo Gao 0001, Jungong Han, Rongmei Cui
Neurocomputing3
2019 Nuclear-norm based 2DLDA with application to face recognition
Siyang Deng, Feiping Nie 0001, Yang Liu 0069, Quanxue Gao
Neurocomputing6
2019 Adaptive robust principal component analysis
Yang Liu 0069, Xinbo Gao 0001, Quanxue Gao, Ling Shao 0001, Jungong Han
Neural Networks3
2019 Flexible unsupervised feature extraction for image classification
Yang Liu 0069, Feiping Nie 0001, Quanxue Gao, Xinbo Gao 0001, Jungong Han, Ling Shao 0001
Neural Networks3
2019 ${R}_1$ -2-DPCA and Face Recognition
abstract
2-D principal component analysis (2-DPCA) is one of the successful dimensionality reduction approaches for image classification and representation. However, 2-DPCA is not robust to outliers. To tackle this problem, we present an efficient robust method, namely R1-2-DPCA for feature extraction. R1-2-DPCA aims to seek the projection matrix such that the projected data have the maximum variance, which is measured by R1-norm. Compared with most existing robust 2-DPCA methods, our model is not only robust to outliers but also helps encode discriminant information. Accordingly, we develop a nongreedy iterative algorithm, which has not only a closed-form solution in each iteration but also a good convergence, to solve our model. Moreover, to further improve classification performance, we employ nuclear norm as the distance metric in the classification phase. Extensive experiments on several face databases illustrate that our proposed method is superior to most existing robust 2-DPCA methods.
Quanxue Gao, Sai Xu, Chris Ding, Xinbo Gao 0001, Yunsong Li 0001
IEEE Trans. Cybern.1
2018 Robust Formulation for PCA: Avoiding Mean Calculation With L2, p-norm Maximization
abstract
Most existing robust principal component analysis (PCA) involve mean estimation for extracting low-dimensional representation. However, they do not get the optimal mean for real data, which include outliers, under the different robust distances metric learning, such as L1-norm and L2,1-norm. This affects the robustness of algorithms. Motivated by the fact that the variance of data can be characterized by the variation between each pair of data, we propose a novel robust formulation for PCA. It avoids computing the mean of data in the criterion function. Our method employs L2,p-norm as the distance metric to measure the variation in the criterion function and aims to seek the projection matrix that maximizes the sum of variation between each pair of the projected data. Both theoretical analysis and experimental results demonstrate that our methods are efficient and superior to most existing robust methods for data reconstruction.
Shuangli Liao, Jin Li 0011, Yang Liu 0069, Quanxue Gao, Xinbo Gao 0001
AAAI4
2018 Euler Sparse Representation for Image Classification
abstract
Sparse representation based classification (SRC) has gained great success in image recognition. Motivated by the fact that kernel trick can capture the nonlinear similarity of features, which may help improve the separability and margin between nearby data points, we propose Euler SRC for image classification, which is essentially the SRC with Euler sparse representation. To be specific, it first maps the images into the complex space by Euler representation, which has a negligible effect for outliers and illumination, and then performs complex SRC with Euler representation. The major advantage of our method is that Euler representation is explicit with no increase of the image space dimensionality, thereby enabling this technique to be easily deployed in real applications. To solve Euler SRC, we present an efficient algorithm, which is fast and has good convergence. Extensive experimental results illustrate that Euler SRC outperforms traditional SRC and achieves better performance for image classification.
Yang Liu 0069, Quanxue Gao, Jungong Han, Shujian Wang
AAAI2
2018 Partial Multi-view Clustering via Consistent GAN
abstract
Multi-view clustering, as one of the most important methods to analyze multi-view data, has been widely used in many real-world applications. Most existing multi-view clustering methods perform well on the assumption that each sample appears in all views. Nevertheless, in real-world application, each view may well face the problem of the missing data due to noise, or malfunction. In this paper, a new consistent generative adversarial network is proposed for partial multi-view clustering. We learn a common low-dimensional representation, which can both generate the missing view data and capture a better common structure from partial multi-view data for clustering. Different from the most existing methods, we use the common representation encoded by one view to generate the missing data of the corresponding view by generative adversarial networks, then we use the encoder and clustering networks. This is intuitive and meaningful because encoding common representation and generating the missing data in our model will promote mutually. Experimental results on three different multi-view databases illustrate the superiority of the proposed method.
Qianqian Wang 0001, Zhengming Ding, Zhiqiang Tao, Quanxue Gao, Yun Fu 0001
ICDM4
2018 Zero Shot Learning via Low-rank Embedded Semantic AutoEncoder
abstract
Zero-shot learning (ZSL) has been widely researched and get successful in machine learning. Most existing ZSL methods aim to accurately recognize objects of unseen classes by learning a shared mapping from the feature space to a semantic space. However, such methods did not investigate in-depth whether the mapping can precisely reconstruct the original visual feature. Motivated by the fact that the data have low intrinsic dimensionality e.g. low-dimensional subspace. In this paper, we formulate a novel framework named Low-rank Embedded Semantic AutoEncoder (LESAE) to jointly seek a low-rank mapping to link visual features with their semantic representations. Taking the encoder-decoder paradigm, the encoder part aims to learn a low-rank mapping from the visual feature to the semantic space, while decoder part manages to reconstruct the original data with the learned mapping. In addition, a non-greedy iterative algorithm is adopted to solve our model. Extensive experiments on six benchmark datasets demonstrate its superiority over several state-of-the-art algorithms.
Yang Liu 0069, Quanxue Gao, Jin Li 0011, Jungong Han, Ling Shao 0001
IJCAI2
2018 Learning with Adaptive Neighbors for Image Clustering
abstract
Due to the importance and efficiency of learning complex structures hidden in data, graph-based methods have been widely studied and get successful in unsupervised learning. Generally, most existing graph-based clustering methods require post-processing on the original data graph to extract the clustering indicators. However, there are two drawbacks with these methods: (1) the cluster structures are not explicit in the clustering results; (2) the final clustering performance is sensitive to the construction of the original data graph. To solve these problems, in this paper, a novel learning model is proposed to learn a graph based on the given data graph such that the new obtained optimal graph is more suitable for the clustering task. We also propose an efficient algorithm to solve the model. Extensive experimental results illustrate that the proposed model outperforms other state-of-the-art clustering algorithms.
Yang Liu 0069, Quanxue Gao, Zhaohua Yang, Shujian Wang
IJCAI2
2018 Dimensionality reduction by LPP-L21
abstract
Locality preserving projection (LPP) is one of the most representative linear manifold learning methods and well exploits intrinsic structure of data. However, the performance of LPP remarkably degenerate in the presence of outliers. To alleviate this problem, the authors propose a robust LPP, namely LPP‐L21. LPP‐L21 employs L2‐norm as the distance metric in spatial dimension of data and L1‐norm as the distance metric over different data points. Moreover, the authors employ L1‐norm to construct similarity graph, this helps to improve robustness of algorithm. Accordingly, the authors present an efficient iterative algorithm to solve LPP‐L21. The authors’ proposed method not only well suppresses outliers but also retains LPP's some nice properties. Experimental results on several image data sets show its advantages.
Shujian Wang, De-Yan Xie, Quanxue Gao
IET Comput. Vis.4
2018 Nuclear-norm based semi-supervised multiple labels learning
Yang Liu 0069, Feiping Nie 0001, Quanxue Gao
Neurocomputing3
2018 Learning more distinctive representation by enhanced PCA network
Yang Liu 0069, Shuangshuang Zhao, Qianqian Wang 0001, Quanxue Gao
Neurocomputing4
2018 Euler Label Consistent K-SVD for image classification and action recognition
Yang Liu 0069, Quanxue Gao, Xinbo Gao 0001, Feiping Nie 0001, Rongmei Cui
Neurocomputing3
2018 SVM based multi-label learning with missing labels for image annotation
Yang Liu 0069, Kaiwen Wen, Quanxue Gao, Xinbo Gao 0001, Feiping Nie 0001
Pattern Recognit.3
2018 Angle 2DPCA: A New Formulation for 2DPCA
abstract
2-D principal component analysis (2DPCA), which employs squared -norm as the distance metric, has been widely used in dimensionality reduction for data representation and classification. It, however, is commonly known that squared -norm is very sensitivity to outliers. To handle this problem, we present a novel formulation for 2DPCA, namely Angle-2DPCA. It employs -norm as the distance metric and takes into consideration the relationship between reconstruction error and variance in the objective function. We present a fast iterative algorithm to solve the solution of Angle-2DPCA. Experimental results on the Extended Yale B, AR, and PIE face image databases illustrate the effectiveness of our proposed approach.
Quanxue Gao, Yang Liu 0069, Xinbo Gao 0001, Feiping Nie 0001
IEEE Trans. Cybern.1
2018 Discriminant Analysis via Joint Euler Transform and ℓ2, 1-Norm
abstract
Linear Discriminant analysis (LDA) has been widely used for face recognition. However, when identifying faces in the wild, the existence of outliers that deviate significantly from the rest of data can arbitrarily skew the desired solution. This usually deteriorates LDA's performance dramatically, thus preventing it from mass deployment in real-world applications. To handle this problem, we propose an effective distance metric learning method based LDA, namely Euler LDA-L21 (e-LDA-L21). e-LDA-L21 is carried out in two stages, in which each image is mapped into a complex space by Euler transform in the first stage and the ℓ2,1-norm is adopted as the distance metric in the second stage. This not only reveals nonlinear features but also exploits the geometric structure of data. To solve e-LDA-L21 efficiently, we propose an iterative algorithm, which is a closed-form solution at each iteration with convergence guaranteed. Finally, we extend e-LDA-L21 to Euler 2DLDA-L21 (e-2DLDA-L21) which further exploits the spatial information embedded in image pixels. Experimental results on several face databases demonstrate its superiority over the state-of-the-art algorithms.
Shuangli Liao, Quanxue Gao, Zhaohua Yang, Feiping Nie 0001, Jungong Han
IEEE Trans. Image Process.2
2018 ℓ2, p -Norm Based PCA for Image Recognition
abstract
Recently, many ℓ1-norm-based PCA approaches have been developed to improve the robustness of PCA. However, most existing approaches solve the optimal projection matrix by maximizing ℓ1-norm-based variance and do not best minimize the reconstruction error, which is the true goal of PCA. Moreover, they do not have rotational invariance. To handle these problems, we propose a generalized robust metric learning for PCA, namely, ℓ2,p-PCA, which employs ℓ2,p-norm as the distance metric for reconstruction error. The proposed method not only is robust to outliers but also retains PCA's desirable properties. For example, the solutions are the principal eigenvectors of a robust covariance matrix and the low-dimensional representation have rotational invariance. These properties are not shared by ℓ1-norm-based PCA methods. A new iteration algorithm is presented to solve ℓ2,p-PCA efficiently. Experimental results illustrate that the proposed method is more effective and robust than PCA, PCA-L1 greedy, PCA-L1 nongreedy, and HQ-PCA.
Qianqian Wang 0001, Quanxue Gao, Xinbo Gao 0001, Feiping Nie 0001
IEEE Trans. Image Process.2
2018 Robust DLPP With Nongreedy ℓ1-Norm Minimization and Maximization
abstract
Recently, discriminant locality preserving projection based on L1-norm (DLPP-L1) was developed for robust subspace learning and image classification. It obtains projection vectors by greedy strategy, i.e., all projection vectors are optimized individually through maximizing the objective function. Thus, the obtained solution does not necessarily best optimize the corresponding trace ratio optimization algorithm, which is the essential objective function for general dimensionality reduction. It results in insufficient recognition accuracy. To tackle this problem, we propose a nongreedy algorithm to solve the trace ratio formula of DLPP-L1, and analyze its convergence. Experimental results on three databases illustrate the effectiveness of our proposed algorithm.
Qianqian Wang 0001, Quanxue Gao, De-Yan Xie, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2017 Two-Dimensional PCA with F-Norm Minimization
abstract
Two-dimensional principle component analysis (2DPCA) has been widely used for face image representation and recognition. But it is sensitive to the presence of outliers. To alleviate this problem, we propose a novel robust 2DPCA, namely 2DPCA with F-norm minimization (F-2DPCA), which is intuitive and directly derived from 2DPCA. In F-2DPCA, distance in spatial dimensions (attribute dimensions) is measured in F-norm, while the summation over different data points uses 1-norm. Thus it is robust to outliers and rotational invariant as well. To solve F-2DPCA, we propose a fast iterative algorithm, which has a closed-form solution in each iteration, and prove its convergence. Experimental results on face image databases illustrate its effectiveness and advantages.
Qianqian Wang 0001, Quanxue Gao
AAAI2
2017 Angle Principal Component Analysis
abstract
Recently, many ℓ1-norm based PCA methods have been developed for dimensionality reduction, but they do not explicitly consider the reconstruction error. Moreover, they do not take into account the relationship between reconstruction error and variance of projected data. This reduces the robustness of algorithms. To handle this problem, a novel formulation for PCA, namely angle PCA, is proposed. Angle PCA employs ℓ2-norm to measure reconstruction error and variance of projected da-ta and maximizes the summation of ratio between variance and reconstruction error of each data. Angle PCA not only is robust to outliers but also retains PCA’s desirable property such as rotational invariance. To solve Angle PCA, we propose an iterative algorithm, which has closed-form solution in each iteration. Extensive experiments on several face image databases illustrate that our method is overall superior to the other robust PCA algorithms, such as PCA, PCA-L1 greedy, PCA-L1 nongreedy and HQ-PCA.
Qianqian Wang 0001, Quanxue Gao, Xinbo Gao 0001, Feiping Nie 0001
IJCAI2
2017 Trace ratio 2DLDA with L1-norm optimization
Jing Wang 0105, Qianqian Wang 0001, Quanxue Gao
Neurocomputing4
2017 F-norm distance metric based robust 2DPCA and face recognition
Quanxue Gao, De-Yan Xie
Neural Networks3
2017 Optimal mean two-dimensional principal component analysis with F-norm minimization
Qianqian Wang 0001, Quanxue Gao, Xinbo Gao 0001, Feiping Nie 0001
Pattern Recognit.2
2017 Adaptive maximum margin analysis for image recognition
Qianqian Wang 0001, Quanxue Gao, Yunsong Li 0001, Yunfang Huang, Yang Liu 0084
Pattern Recognit.3
2017 A Non-Greedy Algorithm for L1-Norm LDA
abstract
Recently, L1-norm-based discriminant subspace learning has attracted much more attention in dimensionality reduction and machine learning. However, most existing approaches solve the column vectors of the optimal projection matrix one by one with greedy strategy. Thus, the obtained optimal projection matrix does not necessarily best optimize the corresponding trace ratio objective function, which is the essential criterion function for general supervised dimensionality reduction. In this paper, we propose a non-greedy iterative algorithm to solve the trace ratio form of L1-norm-based linear discriminant analysis. We analyze the convergence of our proposed algorithm in detail. Extensive experiments on five popular image databases illustrate that our proposed algorithm can maximize the objective function value and is superior to most existing L1-LDA algorithms.
Yang Liu 0084, Quanxue Gao, Shuo Miao, Xinbo Gao 0001, Feiping Nie 0001, Yunsong Li 0001
IEEE Trans. Image Process.2
2016 Discriminant structure embedding for image recognition
Shuo Miao, Jing Wang 0105, Quanxue Gao
Neurocomputing3
2016 On the schatten norm for matrix based subspace learning and classification
Qianqian Wang 0001, Quanxue Gao, Xinbo Gao 0001, Feiping Nie 0001
Neurocomputing3
2016 Rebuttal to "Comments on 'Joint Global and Local Structure Discriminant Analysis"'
abstract
In the above paper, the authors pointed out that our motivation was flawed in our previous paper, and provided some comments. In this paper, we point out the inexact understanding and representation in the comment paper and then present a detail explanation for our previous paper.
Quanxue Gao
IEEE Trans. Inf. Forensics Secur.1
2015 Merging model-based two-dimensional principal component analysis
Quanxue Gao, Xinbo Gao 0001, De-Yan Xie
Neurocomputing2
2015 A novel semi-supervised learning for face recognition
Quanxue Gao, Yunfang Huang, Xinbo Gao 0001, Weiguo Shen
Neurocomputing1
2015 Discriminative sparsity preserving projections for image recognition
Quanxue Gao, Yunfang Huang
Pattern Recognit.1
2015 Dimensionality Reduction by Integrating Sparse Representation and Fisher Criterion and its Applications
abstract
Sparse representation shows impressive results for image classification, however, it cannot well characterize the discriminant structure of data, which is important for classification. This paper aims to seek a projection matrix such that the low-dimensional representations well characterize the discriminant structure embedded in high-dimensional data and simultaneously well fit sparse representation-based classifier (SRC). To be specific, Fisher discriminant criterion (FDC) is used to extract the discriminant structure, and sparse representation is simultaneously considered to guarantee that the projected data well satisfy the SRC. Thus, our method, called SRC-FDC, characterizes both the spatial Euclidean distribution and local reconstruction relationship, which enable SRC to achieve better performance. Extensive experiments are done on the AR, CMU-PIE, Extended Yale B face image databases, the USPS digit database, and COIL20 database, and results illustrate that the proposed method is more efficient than other feature extraction methods based on SRC.
Quanxue Gao, Qianqian Wang 0001, Yunfang Huang, Xinbo Gao 0001
IEEE Trans. Image Process.1
2014 Global-local fisher discriminant approach for face recognition
Qianqian Wang 0001, Xiaolei Hu, Quanxue Gao
Neural Comput. Appl.3
2014 Stable locality sensitive discriminant analysis for image recognition
Quanxue Gao
Neural Networks1
2013 Feature extraction using two-dimensional neighborhood margin and variation embedding
Quanxue Gao, Xiujuan Hao, Qijun Zhao, Weiguo Shen, Jingjie Ma
Comput. Vis. Image Underst.1
2013 Joint geometry and variability for image recognition
Quanxue Gao, Xiaojing Yang
Neurocomputing1
2013 Joint Global and Local Structure Discriminant Analysis
abstract
Linear discriminant analysis (LDA) only considers the global Euclidean geometrical structure of data for dimensionality reduction. However, previous works have demonstrated that the local geometrical structure is effective for dimensionality reduction. In this paper, a novel approach is proposed, namely Joint Global and Local-structure Discriminant Analysis (JGLDA), for linear dimensionality reduction. To be specific, we construct two adjacency graphs to represent the local intrinsic structure, which characterizes both the similarity and diversity of data, and integrate the local intrinsic structure into Fisher linear discriminant analysis to build a stable discriminant objective function for dimensionality reduction. Experiments on several standard image databases demonstrate the effectiveness of our algorithm.
Quanxue Gao, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2013 Two-Dimensional Maximum Local Variation Based on Image Euclidean Distance for Face Recognition
abstract
Manifold learning concerns the local manifold structure of high dimensional data, and many related algorithms are developed to improve image classification performance. None of them, however, consider both the relationships among pixels in images and the geometrical properties of various images during learning the reduced space. In this paper, we propose a linear approach, called two-dimensional maximum local variation (2DMLV), for face recognition. In 2DMLV, we encode the relationships among pixels in images using the image Euclidean distance instead of conventional Euclidean distance in estimating the variation of values of images, and then incorporate the local variation, which characterizes the diversity of images and discriminating information, into the objective function of dimensionality reduction. Extensive experiments demonstrate the effectiveness of our approach.
Quanxue Gao, Feifei Gao 0002, Hailin Zhang 0001, Xiujuan Hao, Xiaogang Wang 0001
IEEE Trans. Image Process.1
2013 Stable Orthogonal Local Discriminant Embedding for Linear Dimensionality Reduction
abstract
Manifold learning is widely used in machine learning and pattern recognition. However, manifold learning only considers the similarity of samples belonging to the same class and ignores the within-class variation of data, which will impair the generalization and stableness of the algorithms. For this purpose, we construct an adjacency graph to model the intraclass variation that characterizes the most important properties, such as diversity of patterns, and then incorporate the diversity into the discriminant objective function for linear dimensionality reduction. Finally, we introduce the orthogonal constraint for the basis vectors and propose an orthogonal algorithm called stable orthogonal local discriminate embedding. Experimental results on several standard image databases demonstrate the effectiveness of the proposed dimensionality reduction approach.
Quanxue Gao, Jingjie Ma, Xinbo Gao 0001
IEEE Trans. Image Process.1
2012 Two-dimensional margin, similarity and variation embedding
Quanxue Gao
Neurocomputing1
2012 Enhanced fisher discriminant criterion for image recognition
Quanxue Gao, Xiaojing Yang
Pattern Recognit.1
2010 Two-dimensional supervised local similarity and diversity projection
Quanxue Gao, Yiying Li, De-Yan Xie
Pattern Recognit.1
2010 Independent components extraction from image matrix
Quanxue Gao, Lei Zhang 0006, David Zhang 0001
Pattern Recognit. Lett.1
2009 Sequential row-column independent component analysis for face recognition
Quanxue Gao, Lei Zhang 0006, David Zhang 0001
Neurocomputing1
2008 Directional independent component analysis with tensor representation
abstract
Conventional independent component analysis (ICA) learns the statistical independencies of 2D variables from the training images that are unfolded to vectors. The unfolded vectors, however, make the ICA suffer from the small sample size (SSS) problem that leads to the dimensionality dilemma. This paper presents a novel directional multilinear ICA method to solve those problems by encoding the input image or high dimensional data array as a general tensor. In addition, the mode-k matrix of the tensor is re-sampled and re-arranged to form a mode-k directional image to better exploit the directional information in training. An algorithm called mode-k directional ICA is then presented for feature extraction. Compared with the conventional ICA and other subspace analysis algorithms, the proposed method can greatly alleviate the SSS problem, reduce the computational cost in the learning stage by representing the data in lower dimension, and simultaneously exploit the directional information in the high dimensional dataset. Experimental results on well-known face and palmprint databases show that the proposed method has higher recognition accuracy than many existing ICA, PCA and even supervised FLD schemes while using a low dimension of features.
Lei Zhang 0006, Quanxue Gao, David Zhang 0001
CVPR2
2007 Is two-dimensional PCA equivalent to a special case of modular PCA?
Quanxue Gao
Pattern Recognit. Lett.1
2007 Comments on "On Image Matrix Based Feature Extraction Algorithms"
abstract
A class of image-matrix-based feature extraction algorithms has been discussed earlier. The correspondence argues that 2-D principal component analysis and Fisher linear discriminant (FLD) are equivalent to block-based PCA and FLD. In this correspondence, we point out that this statement is not rigorous.
Quanxue Gao, Lei Zhang 0006, David Zhang 0001, Jian Yang 0003
IEEE Trans. Syst. Man Cybern. Part B1