Zheng Wang 0037

dblp:w/ZhengWang37 · DBLP profile ↗
← Back
69ranked-venue papers
14as first author
60since 2021 · last 2026
0000-0002-4814-1115ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 13 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 S2-Boost: Synergistic Semantic Boosting for Coarse-to-Fine Ensemble Learning
abstract
Neuroscientific evidence reveals that human visual recognition is not an instantaneous event but a hierarchical process, where the brain constructs a holistic perception by progressively integrating simple features like edges or texture into complex scenes. Ensemble learning successfully utilizes this principle, yet existing methods typically integrate models at the decision level, neglecting the rich, complementary information within the feature space itself and thus fundamentally limiting their potential. To address this, we introduce Synergistic Semantic Boosting (S2-Boosting), a framework that employs a self-supervised hierarchical semantic learning module to decompose an image into complementary, semantically meaningful parts autonomously. These parts guide a boosting procedure where a sequence of specialized learners, each focusing on a specific semantic partition, collaboratively corrects the ensemble's errors. We further present encouraging results on real-world image datasets, highlighting the intrinsic interpretability, paving the way for more robust and transparent models.
Guanxiong He, Zheng Wang 0037, Jie Wang 0164, Liaoyuan Tang, Rong Wang 0001, Feiping Nie 0001
AAAI2
2026 Towards Federated Clustering: A Client-wise Private Graph Aggregation Framework
abstract
Federated clustering addresses the critical challenge of extracting patterns from decentralized, unlabeled data. However, it is hampered by the flaw that current approaches are forced to accept a compromise between performance and privacy: transmitting embedding representations risks sensitive data leakage, while sharing only abstract cluster prototypes leads to diminished model accuracy. To resolve this dilemma, we propose Structural Privacy-Preserving Federated Graph Clustering (SPP-FGC), a novel algorithm that innovatively leverages local structural graphs as the primary medium for privacy-preserving knowledge sharing, thus moving beyond the limitations of conventional techniques. Our framework operates on a clear client-server logic; on the client-side, each participant constructs a private structural graph that captures intrinsic data relationships, which the server then securely aggregates and aligns to form a comprehensive global graph from which a unified clustering structure is derived. The framework offers two distinct modes to suit different needs. SPP-FGC is designed as an efficient one-shot method that completes its task in a single communication round, ideal for rapid analysis. For more complex, unstructured data like images, SPP-FGC+ employs an iterative process where clients and the server collaboratively refine feature representations to achieve superior downstream performance. Extensive experiments demonstrate that our framework achieves state-of-the-art performance, improving clustering accuracy by up to 10% (NMI) over federated baselines while maintaining provable privacy guarantees.
Guanxiong He, Zheng Wang 0037, Jie Wang 0164, Liaoyuan Tang, Rong Wang 0001, Feiping Nie 0001
AAAI2
2026 Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud Understanding
abstract
Self-supervised 3D point cloud understanding is crucial for scene understanding, where Masked Autoencoders (MAE) have achieved excellent performance in point cloud representation learning. However, existing MAE-style methods fail to consider spatial-semantic variations in masking strategies, and joint learning with multi-view images often overlooks view redundancy. To address these challenges, we propose an MAE framework enhanced with reliable multi-view 2D-3D Key-part alignment and Reinforced masking, named as KR-MAE. Our approach comprises three key innovations: Reinforced Masking (RM) strategically samples visible tokens based on semantic saliency to enhance reconstruction fidelity; Reliable Multi-View Selector (RVS) dynamically refines the most informative image subset by filtering occluded or low-texture views, mitigating detrimental redundancy; Reliable-view 2D-3D Key-part Aligned Transformer (KAT) establishes semantic-aligned correspondence between salient 3D point cloud parts and reliable multi-view 2D image patches, leveraging rich texture cues from 2D images to compensate for sparse geometry in point cloud. Extensive experiments on 3D classification and segmentation benchmarks demonstrate that KR-MAE achieves state-of-the-art performance, surpassing prior multi-modal methods.
Xianglong Jin, Zheng Wang 0037, Rong Wang 0001, Feiping Nie 0001
AAAI2
2026 Group-wise attentive enhancements for unsupervised feature selection
Jie Wang 0164, Yongjin Yuan, Lingyi Kong, Zheng Wang 0037, Rong Wang 0001, Feiping Nie 0001
Knowl. Based Syst.5
2026 Dual Geometry Margin Optimization for Coupled-Noisy Robust Ensemble Learning
abstract
Ensemble learning methods, such as Bagging and Boosting, are well-regarded for their ability to enhance model performance by combining diverse base learners. These approaches leverage the strengths of individual models to achieve more accurate and robust predictions. However, real-world datasets often contain noise, which can significantly impair model effectiveness. This paper focuses on two prevalent and challenging types: feature noise, which can lead to fitting instability and poor generalization, and label noise, which can lead to erroneous supervision and model overfitting. Recognizing the inherent properties of ensemble learning, particularly its focus on optimizing the decision margin to improve classification accuracy, we see an opportunity to bolster ensemble model robustness. To address both feature and label noise, we propose a novel approach called Dual Geometry Margin Boosting (DGMB). This method employs two key strategies: the Decision Plane Margin (DPM), which enhances class separation, and the Hyper-Sphere Margin (HSM), which effectively filters out potentially noisy samples during the learning process. Our experiments demonstrate the impressive ability of DGMB to resist both feature and label noise. Through rigorous testing on various noise-contaminated datasets, we show that DGMB maintains strong performance and outperforms other robust Ensemble methods.
Zheng Wang 0037, Guanxiong He, Jie Wang 0164, Runxin Zhang, Liaoyuan Tang, Rong Wang 0001, Feiping Nie 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Selective-relaxed contrastive learning for hyperspectral image classification with noisy labels
Jie Wang 0164, Zheng Wang 0037, Liaoyuan Tang, Rong Wang 0001, Feiping Nie 0001
Pattern Recognit.3
2026 Improve noise tolerance of robust feature selection via block-sparse projection learning
Jie Wang 0164, Zheng Wang 0037, Yu Guo 0006, Rong Wang 0001, Fei Wang 0008, Feiping Nie 0001
Pattern Recognit.2
2026 Max-Min Robust Unsupervised Feature Selection via Sparse Subspace
abstract
Feature selection is one of the hot issues in machine learning. It reduces storage pressure by effectively screening features and has become a very practical data preprocessing method. At present, most feature selection algorithms apply $\ell _{2,1}$ -norm on the transformation matrix to calculate the scores for all features and then select appropriate features according to these scores. But their sparsity is limited, and meaningless regularization parameters increase the cost, making it prone to falling into local optimum. To solve the above difficulties, this article proposes a novel max-min robust unsupervised feature selection via sparse subspace (MMRUFS), which considers both the reconstruction term and variance term of data, so that the model can not only fully retain the original information of data, but also make the data more dispersed. Second, $\ell _{2,0}$ -norm constraint is used on the transformation matrix to directly select the optimal feature subset, avoiding the fine-tuning of regularization parameters. To enhance the robustness, MMRUFS carefully designs mark weight vector to make the model treat normal samples and outliers differently and achieves the effect of anomaly detection. Finally, MMRUFS is solved by designing the surrogate matrix, and its convergence is strictly guaranteed, experimental results reveal that MMRUFS outperforms other feature selection algorithms on multiple real-world datasets.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Zhensheng Sun, Xuelong Li 0001
IEEE Trans. Cybern.3
2026 Explaining Neural Networks: Hierarchical Backpropagated Ensemble Learning
abstract
Deep models, characterized by complex structures and end-to-end optimization, proved effective in providing decision support based on real-world data. However, the lack of transparency in their decision-making process and the difficulty in interpreting the role of individual neurons limited their practical applicability in many critical and sensitive domains. Inspired by the parallels between neural networks and ensemble models, where performance was achieved through the collaboration of multiple weak learners, this article presents a novel perspective that reframes neural networks as hierarchical ensembles. We propose the hierarchical backpropagated ensemble (HBE) model, wherein each neuron functions both as a base learner and as part of an ensemble of preceding neurons. This framework applies ensemble learning techniques to neural networks, allowing each neuron to focus on specific subtasks while progressively constructing a network that meets global objectives. Experimental results on real-world data show that this hierarchical structure enhances the effectiveness of traditional ensemble models, and the ensemble-based explanations offer improved initialization and dynamically adjustable network structures, leading to more efficient training.
Guanxiong He, Zheng Wang 0037, Liaoyuan Tang, Runxin Zhang, Rong Wang 0001, Xuelong Li 0001, Feiping Nie 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Language Pre-training Guided Masking Representation Learning for Time Series Classification
abstract
The representation learning of time series has a wide range of downstream tasks and applications in many practical scenarios. However, due to the complexity, spatiotemporality, and continuity of sequential stream data, compared with the representation learning of structural data such as images/videos, the time series self-supervised representation learning is even more challenging. Besides, the direct application of existing contrastive learning and masked autoencoder based approaches to time series representation learning encounters inherent theoretical limitations, such as ineffective augmentation and masking strategies. To this end, we propose a Language Pre-training guided Masking Representation Learning (LPMRL) for times series classification. Specifically, we first propose a novel language pre-training guided masking encoder for adaptively sampling semantic spatiotemporal patches via natural language descriptions and improving the discriminability of latent representations. Furthermore, we present the dual-information contrastive learning mechanism to explore both local and global information by meticulously designing high-quality hard negative samples of time series data samples. As a result, we also design various experiments, such as visualization of masking position and distribution and reconstruction error to verify the reasonability of proposed language guided masking technique. Last, we evaluate the performance of proposed representation learning via classification task conducted on 106 time series datasets, which demonstrates the effectiveness of proposed method.
Liaoyuan Tang, Zheng Wang 0037, Jie Wang 0164, Guanxiong He, Zhezheng Hao, Rong Wang 0001, Feiping Nie 0001
AAAI2
2025 Point-DMAE: Point Cloud Self-supervised Learning via Density-directed Masked Autoencoders
abstract
Masked autoencoders have been extensively utilized in 3D point cloud self-supervised learning, where the fundamental approach involves masking a portion of the point cloud and subsequently reconstructing it. This process is hypothesized to enhance model learning by leveraging the inherent structure of the point cloud data. However, the information density within point clouds is inherently uneven, contrasting with the more uniform distributions found in language and 2D image data. This uneven distribution suggests that the application of random masking strategies, commonly adopted from NLP and 2D vision, may not be optimal for point cloud data, potentially leading to suboptimal learning outcomes. Based on this observation, we propose a simple yet effective Density-directed Masked Autoencoders for Point Cloud Self-supervised Learning (Point-DMAE), which learns latent semantic point cloud features using a density-directed masking strategy. Specifically, our method employs a dual-branch Transformer architecture to extract both high-level and fine-grained point features through global and local block density-directed masking, respectively. Point-DMAE demonstrates high pre-training efficiency and significantly outperforms our baseline (Point-MAE) on 3D object classification tasks within the ScanObjectNN dataset by 4.13% on OBJ-BG, 5.17% on OBJ-ONLY, and 4.17% on PB-T50-RS. Codes are available at https://github.com/jinxianglong10/Point-DMAE.
Xianglong Jin, Zheng Wang 0037, Feiping Nie 0001
CIKM2
2025 Self-Supervised Localized Topology Consistency for Noise-Robust Hyperspectral Image Classification
abstract
Label noise in hyperspectral image classification (HIC) can severely degrade model performance by leading to incorrect predictions and overfitting, especially as erroneous labels propagate and compound throughout the training process. To address this, we propose a robust learning framework called Self-Supervised Localized Topology Consistency (SSLTC), which enforces local topology consistency to enhance model resilience against noisy labels. SSLTC captures local topology via a graph-based representation, where nodes represent samples and edges encode pairwise similarities. Predictions are propagated from topologically similar nodes to central nodes, constrained by Kullback-Leibler (KL) divergence to encourage consistent predictions and reduce sensitivity to noisy labels. Additionally, a self-supervised contrastive learning strategy is used to refine spectral-spatial representations in an unsupervised manner, further improving robustness. Extensive experiments on hyperspectral benchmark datasets with varying noise levels demonstrate the superiority of SSLTC in mitigating the adverse effects of label noise compared to state-of-the-art approaches in HIC tasks.
Jie Wang 0164, Liaoyuan Tang, Guanxiong He, Zheng Wang 0037, Rong Wang 0001
ICASSP5
2025 Dynamic T-distributed stochastic neighbor graph convolutional networks for multi-modal contrastive fusion
Guoxu Li, Jie Wang 0164, Zheng Wang 0037, Jianfu Cao, Rong Wang 0001, Feiping Nie 0001
Neurocomputing4
2025 A novel linear discriminant analysis based on alternate ratio sum minimization
Chuanjie Cao, Keyi Zhou, Zheng Wang 0037, Liang Lin 0004, Feiping Nie 0001
Inf. Sci.5
2025 Toward Balance Adaptive Weighted Ensemble Clustering
abstract
Ensemble clustering, which combines the information from multiple base clusterings to obtain a better partition result, has received extensive attention due to its effectiveness and robustness. Although many algorithms have been developed in recent years that have achieved impressive results in practical applications, two challenging issues in ensemble clustering remain. First, most algorithms assume that all base clusterings have the same impact on the clustering results, assigning them the same weight. This makes the clustering performance susceptible to the influence of redundant, low-quality base clusterings. Second, co-association matrix-based algorithms often rely on additional methods, such as hierarchical agglomerative clustering, to obtain the final clustering result after constructing the weighted co-association matrix. This not only complicates optimization process but also leads to the loss of some sample-similarity information during clustering. To address this problem, we propose a novel Toward Balance Adaptive Weighted Ensemble Clustering (TBAWEC) algorithm. This method transforms the ensemble clustering problem into an optimization problem, producing the final result without requiring additional clustering algorithms. Moreover, we introduce balanced technology into ensemble clustering for the first time, significantly improving the balance of clustering results. Extensive experiments on real datasets demonstrate that the proposed algorithm outperforms the most advanced ensemble and balanced clustering algorithms simultaneously.
Runxin Zhang, Xia Wu 0001, Guanxiong He, Zheng Wang 0037, Rong Wang 0001, Feiping Nie 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Enhancing Clustering Performance With Tensorized High-Order Bipartite Graphs: A Structured Graph Learning Approach
abstract
Clustering based on structured graph learning involves acquiring a proximity matrix with an explicit clustering structure from the original one. However, the original proximity matrix often lacks some must-links compared to the groundtruth, constraining the upper bound of clustering performance. High-order proximity information can mitigate this limitation, yet traditional high-order proximity matrix-based methods are time-intensive. To tackle this, we propose the Tensorized High-order Bipartite Graphs-based structured proximity matrix learning method (THBG). Firstly, we introduce a high-order bipartite graph proximity matrix with a swift computation method, incorporating high-order information and significantly reducing computational overhead. Secondly, we apply tensor nuclear norm minimization to the tensor composed of high-order bipartite graphs, learning a low-rank tensor representation that effectively harnesses the consistency of high-order information. Concurrently, a structured bipartite graph proximity matrix with an explicit clustering structure is adaptively learned based on the low-rank tensor representation and Laplace rank constraint. Experimental results demonstrate the superiority and great potential of this method. Code available:https://anonymous.4open.science/r/THBG-D10D.
Zihua Zhao, Haonan Xin, Rong Wang 0001, Danyang Wu, Zheng Wang 0037, Feiping Nie 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Manifold-Aligned Consistency Contrastive Learning for Noise-Tolerant Hyperspectral Image Classification
abstract
Label noise in hyperspectral image classification (HIC) has posed a significant challenge, mainly due to the complexity of high-dimensional spectral-spatial data. Existing approaches that rely on sample predictions for verification and correction often amplify confirmation bias, leading to decision boundaries that overfit to noisy labels. To address these issues, we propose a Manifold-Aligned Consistency Contrastive Learning (MACCL) framework that establishes a mutually-guided consistency alignment mechanism between the spectral-spatial representation space and label space through manifold learning theory to combat label noise. Specifically, to mitigate the confirmation bias of noisy labels, we introduce a manifold-aligned consistency learning module. It leverages the manifold assumption in the representation space, modeling local neighborhood graphs and enforcing consistent prediction distributions via KL divergence. This aligns label-space classifications with the representation space’s local geometry, suppressing isolated noise through manifold continuity. Additionally, to combat representation degradation that causes decision boundaries to overfit noisy labels, we integrate a noise-tolerant contrastive representation learning module. By applying confidence-guided criteria, the module focuses on high-confidence sample pairs and regularizes gradients. This emphasizes clean pairs during training, boosting the model’s discriminative ability and preserving true semantic relationships. Through this mutual guidance, the contrastive learning refines the local geometric structure through discriminative representation learning, driving the representation space closer to the intrinsic data manifold, while the manifold alignment propagates geometric constraints to rectify label space corruptions. Finally, experiments on several benchmark datasets with varying levels of noise have validated the superiority of the proposed MACCL framework.
Jie Wang 0164, Junti Wang, Guanxiong He, Zheng Wang 0037, Rong Wang 0001, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Data Subdivision Based Dual-Weighted Robust Principal Component Analysis
abstract
Principal Component Analysis (PCA) is one of the most important unsupervised dimensionality reduction algorithms, which uses squared -norm to make it very sensitive to outliers. Those improved versions based on -norm alleviate this problem, but they have other shortcomings, such as optimization difficulties or lack of rotational invariance, etc. Besides, existing methods only vaguely divide normal samples and outliers to improve robustness, but they ignore the fact that normal samples can be more specifically divided into positive samples and hard samples, which should have different contributions to the model because positive samples are more conducive to learning the projection matrix. In this paper, we propose a novel Data Subdivision Based Dual-Weighted Robust Principal Component Analysis, namely DRPCA, which firstly designs a mark vector to distinguish normal samples and outliers, and directly removes outliers according to mark weights. Moreover, we further divide normal samples into positive samples and hard samples by self-constrained weights, and place them in relative positions, so that the weight of positive samples is larger than hard samples, which makes the projection matrix more accurate. Additionally, the optimal mean is employed to obtain a more accurate data center. To solve this problem, we carefully design an effective iterative algorithm and analyze its convergence. Experiments on real-world and RGB large-scale datasets demonstrate the superiority of our method in dimensionality reduction and anomaly detection.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Image Process.3
2025 Fuzzy Weighted Principal Component Analysis for Anomaly Detection
abstract
Principal Component Analysis (PCA) is one of the most famous unsupervised dimensionality reduction algorithms and has been widely used in many fields. However, it is very sensitive to outliers, which reduces the robustness of the algorithm. In recent years, many studies have tried to employ \(\ell_{1}\) -norm to improve the robustness of PCA, but they all lack rotation invariance or the solution is expensive. In this article, we propose a novel robust PCA, namely, Fuzzy Weighted Principal Component Analysis (FWPCA), which still uses squared \(\ell_{2}\) -norm to minimize reconstruction error and maintains rotation invariance of PCA. The biggest bright spot is that the contribution of data is restricted by fuzzy weights, so that the contribution of normal samples is much greater than noise or abnormal data, and realizes anomaly detection. Besides, a more reasonable data center can be obtained by solving the optimal mean to make projection matrix more accurate. Subsequently, an effective iterative optimization algorithm is developed to solve this problem, and its convergence is strictly proved. Extensive experimental results on face datasets and RGB anomaly detection datasets show the superiority of our proposed method.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
ACM Trans. Knowl. Discov. Data3
2025 Graph-Based Clustering: High-Order Bipartite Graph for Proximity Learning
abstract
Structured proximity matrix learning, one of the mainstream directions in clustering research, refers to learning a proximity matrix with an explicit clustering structure from the original first-order proximity matrix. Due to the complexity of the data structure, the original first-order proximity matrix always lacks some must-links compared to the groundtruth proximity matrix. It is worth noting that high-order proximity matrices can provide missed must-link information. However, the computation of high-order proximity matrices and clustering based on them are expensive. To solve the above problem, inspired by the anchor bipartite graph, we present a novel high-order bipartite graph proximity matrix and a fast method to compute it. This proposed high-order bipartite graph proximity matrix contains high-order proximity information and can significantly reduce the computational complexity of the whole clustering process. Furthermore, we introduce an efficient and simple high-order bipartite graph fusion framework that can adaptively assign weights to each order of the high-order bipartite graph matrices. Finally, under the Laplace rank constraint, a consensus structured bipartite graph proximity matrix is obtained. At the same time, an efficient solution algorithm is proposed for this model. The model's efficacy is underscored through rigorous experiments, highlighting its superior clustering performance and time efficiency. Code available:https://anonymous.4open.science/r/HBGC-F6C4.
Zihua Zhao, Danyang Wu, Rong Wang 0001, Zheng Wang 0037, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.4
2025 Reweighted-Boosting: A Gradient-Based Boosting Optimization Framework
abstract
Boosting is a well-established ensemble learning approach that aims to enhance overall performance by combining multiple weak learners with a linear combination structure. It operates on the principle of using new learners to compensate for the shortcomings of previous learners and is known for its ability to reduce computational resource requirements while mitigating the risks of overfitting. However, from the perspective of convex optimization, it becomes apparent that classical boosting methods often converge to local optima rather than global optima when minimizing the target loss due to its greedy strategy. In this article, we address the issue and propose a novel optimization framework for the boosting paradigm. Our framework focuses on refining the ensemble model by further minimizing loss function through the reallocation of base learner weights, which results in a more robust and powerful learner. We have conducted experiments on various real-world and synthetic datasets, and our findings confirm that our Reweighted-Boosting model consistently outperforms its counterparts. It also exhibits an increased classification margin for the data, making it a valuable enhancement to original boosting algorithms.
Guanxiong He, Zheng Wang 0037, Liaoyuan Tang, Weizhong Yu, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Adaptive Graph Convolutional Network for Unsupervised Generalizable Tabular Representation Learning
abstract
A challenging open problem in deep learning is the representation of tabular data. Unlike the popular domains such as image and text understanding, where the deep convolutional network is fashionable in many applications, there is still no widely used neural architecture that can effectively explore informative structure from tabular data. In addition, existing antoencoder-based nonlinear representation learning approaches that employ reconstruction loss, are incompetent to preserve discriminative information. As a step toward bridging these gaps, we propose a novel adaptive graph convolutional network (AdaGCN) for unsupervised generalizable tabular representation learning in this article. To be specific, we hypothesize that the keys to boosting the efficiency and practicality of learned representations lie in three aspects, i.e., adaptivity, unsupervised, and generalization. As a result, the adaptive graph learning module is first designed to remove the predefined rules in conventional GCN models, which can explore more local patterns on arbitrary tabular data. Moreover, our AdaGCN directly minimizes the difference between distributions of original tabular data and learned embeddings for training without any label information. Last but not least, the parametric property of AdaGCN makes the unseen data to be handled offline, which extremely expends the scope of applications. We present extensive experiments showing that AdaGCN significantly and consistently outperforms several representation learning and clustering methods on several real-world tabular datasets.
Zheng Wang 0037, Jiaxi Xie, Rong Wang 0001, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Outlier-Robust Feature Selection with ℓ2, 1-Norm Minimization and Group Row-Sparsity Induced Constraints
abstract
In the realm of high-dimensional data analysis, the existence of outliers presents a substantial hurdle to the efficacy of feature selection methods that rely on the assumption of Gaussian distribution. To tackle this issue, we propose an outlier-robust feature selection method, ORFS, which combines robust ℓ2,1-norm minimization with group row-sparsity induced constrains to achieve both robustness and discriminative prediction capabilities. Moreover, the group row-sparsity constraints subspace learning based on ℓ2,0-norm can directly select features without parameter tuning. Finally, we introduce an iterative optimization strategy to solve NP-hard problem, and extensive experiments demonstrate the efficacy of ORFS in effectively eliminating the impact of outliers and significantly improving classification performance.
Jie Wang 0164, Zheng Wang 0037, Rong Wang 0001, Feiping Nie 0001, Xuelong Li 0001
ICASSP2
2024 Multi-View Subspace Clustering With Consensus Graph Contrastive Learning
abstract
A significant challenge in multi-view clustering lies in the comprehensive extraction of consistency and complementary information from heterogeneous multi-view data. Numerous methods employ contrastive learning techniques to explore the information between views. However, the basic contrastive learning strategy does not consider cluster information when constructing sample pairs, potentially leading to the emergence of false negative pairs (FNPs). To tackle this concern, we propose a Multi-view Subspace Clustering with Consensus Graph Contrastive Learning (CGCL) model. Specifically, a self-representation layer is designed to acquire a consensus graph that elucidates the overall data distribution. Furthermore, a contrastive learning layer utilizes the cluster information embedded in the consensus graph to yield reliable sample pairs, resulting in a reduction of the detrimental FNPs and the extraction of complementary information from the various views. Extensive experiments on public datasets demonstrate the effectiveness of CGCL.
Jie Zhang 0090, Yuan Sun 0003, Yu Guo 0006, Zheng Wang 0037, Feiping Nie 0001, Fei Wang 0008
ICASSP4
2024 Perturbation Guiding Contrastive Representation Learning for Time Series Anomaly Detection
Liaoyuan Tang, Zheng Wang 0037, Guanxiong He, Rong Wang 0001, Feiping Nie 0001
IJCAI2
2024 Corrigendum to "Max-min robust principal component analysis" [Neurocomputing 521 (2023) 89-98]
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
Neurocomputing3
2024 T-distributed Stochastic Neighbor Network for unsupervised representation learning
Zheng Wang 0037, Jiaxi Xie, Feiping Nie 0001, Rong Wang 0001, Yanyan Jia, Shichang Liu
Neural Networks1
2024 Worst-Case Discriminative Feature Learning via Max-Min Ratio Analysis
abstract
We propose a novel discriminative feature learning method via Max-Min Ratio Analysis (MMRA) for exclusively dealing with the long-standing "worst-case class separation" problem. Existing technologies simply consider maximizing the minimal pairwise distance on all class pairs in the low-dimensional subspace, which is unable to separate overlapped classes entirely especially when the distribution of samples within same class is diverging. We propose a new criterion, i.e., Max-Min Ratio Analysis (MMRA) that focuses on maximizing the minimal ratio value of between-class and within-class scatter to extremely enlarge the separability on the overlapped pairwise classes. Furthermore, we develop two novel discriminative feature learning models for dimensionality reduction and metric learning based on our MMRA criterion. However, solving such a non-smooth non-convex max-min ratio problem is challenging. As an important theoretical contribution in this paper, we systematically derive an alternative iterative algorithm based on a general max-min ratio optimization framework to solve a general max-min ratio problem with rigorous proofs of convergence. More importantly, we also present another solver based on bisection search strategy to solve the SDP problem efficiently. To evaluate the effectiveness of proposed methods, we conduct extensive pattern classification and image retrieval experiments on several artificial datasets and real-world ScRNA-seq datasets, and experimental results demonstrate the effectiveness of proposed methods.
Zheng Wang 0037, Feiping Nie 0001, Canyu Zhang 0001, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Joint learning of latent subspace and structured graph for multi-view clustering
Yu Guo 0006, Zheng Wang 0037, Fei Wang 0008
Pattern Recognit.3
2024 Coordinate Descent Optimized Trace Difference Model for Joint Clustering and Feature Extraction
Fei Wang 0008, Zhongheng Li, Zheng Wang 0037, Feiping Nie 0001
Pattern Recognit.4
2024 Geometric-inspired graph-based Incomplete Multi-view Clustering
Zequn Yang, Han Zhang 0012, Yake Wei, Zheng Wang 0037, Feiping Nie 0001, Di Hu 0001
Pattern Recognit.4
2024 Sparse Trace Ratio LDA for Supervised Feature Selection
abstract
Classification is a fundamental task in the field of data mining. Unfortunately, high-dimensional data often degrade the performance of classification. To solve this problem, dimensionality reduction is usually adopted as an essential preprocessing technique, which can be divided into feature extraction and feature selection. Due to the ability to obtain category discrimination, linear discriminant analysis (LDA) is recognized as a classic feature extraction method for classification. Compared with feature extraction, feature selection has plenty of advantages in many applications. If we can integrate the discrimination of LDA and the advantages of feature selection, it is bound to play an important role in the classification of high-dimensional data. Motivated by the idea, we propose a supervised feature selection method for classification. It combines trace ratio LDA with$\ell _{2,p}$-norm regularization and imposes the orthogonal constraint on the projection matrix. The learned row-sparse projection matrix can be used to select discriminative features. Then, we present an optimization algorithm to solve the proposed method. Finally, the extensive experiments on both synthetic and real-world datasets indicate the effectiveness of the proposed method.
Feiping Nie 0001, Danyang Wu, Zheng Wang 0037, Xuelong Li 0001
IEEE Trans. Cybern.4
2024 Efficient Local Coherent Structure Learning via Self-Evolution Bipartite Graph
abstract
Dimensionality reduction (DR) targets to learn low-dimensional representations for improving discriminability of data, which is essential for many downstream machine learning tasks, such as image classification, information clustering, etc. Non-Gaussian issue as a long-standing challenge brings many obstacles to the applications of DR methods that established on Gaussian assumption. The mainstream way to address above issue is to explore the local structure of data via graph learning technique, the methods based on which however suffer from a common weakness, that is, exploring locality through pairwise points causes the optimal graph and subspace are difficult to be found, degrades the performance of downstream tasks, and also increases the computation complexity. In this article, we first propose a novel self-evolution bipartite graph (SEBG) that uses anchor points as the landmark of subclasses, and learns anchor-based rather than pairwise relationships for improving the efficiency of locality exploration. In addition, we develop an efficient local coherent structure learning (ELCS) algorithm based on SEBG, which possesses the ability of updating the edges of graph in learned subspace automatically. Finally, we also provide a multivariable iterative optimization algorithm to solve proposed problem with strict theoretical proofs. Extensive experiments have verified the superiorities of the proposed method compared to related SOTA methods in terms of performance and efficiency on several real-world benchmarks and large-scale image datasets with deep features.
Zheng Wang 0037, Qi Li 0045, Feiping Nie 0001, Rong Wang 0001, Fei Wang 0008, Xuelong Li 0001
IEEE Trans. Cybern.1
2024 Outliers Robust Unsupervised Feature Selection for Structured Sparse Subspace
abstract
Feature selection is one of the important topics of machine learning, and it has a wide range of applications in data preprocessing. At present, feature selection based on$\ell _{2,1}$-norm regularization is a relatively mature method, but it is not enough to maximize the sparsity and parameter-tuning leads to increased costs. Later scholars found that the$\ell _{2,0}$-norm constraint is more conductive to feature selection, but it is difficult to solve and lacks convergence guarantees. To address these problems, we creatively propose a novel Outliers Robust Unsupervised Feature Selection for structured sparse subspace (ORUFS), which utilizes$\ell _{2,0}$-norm constraint to learn a structured sparse subspace and avoid tuning the regularization parameter. Moreover, by adding binary weights, outliers are directly eliminated and the robustness of model is improved. More importantly, a Re-Weighted (RW) algorithm is exploited to solve our$\ell _{p}$-norm problem. For the NP-hard problem of$\ell _{2,0}$-norm constraint, we develop an effective iterative optimization algorithm with strict convergence guarantees and closed-form solution. Subsequently, we provide theoretical analysis about convergence and computational complexity. Experimental results on real-world datasets illustrate that our method is superior to the state-of-the-art methods in clustering and anomaly detection tasks.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.3
2024 An Balanced, and Scalable Graph-Based Multiview Clustering Method
abstract
In recent years, graph-based multiview clustering methods have become a research hotspot in the clustering field. However, most existing methods lack consideration of cluster balance in their results. In fact, cluster balance is crucial in many real-world scenarios. Additionally, graph-based multiview clustering methods often suffer from high time consumption and cannot handle large-scale datasets. To address these issues, this paper proposes a novel graph-based multiview clustering method. The method is built upon the bipartite graph. Specifically, it employs a label propagation mechanism to update the smaller anchor label matrix rather than the sample label matrix, significantly reducing the computational cost. The introduced balance constraint in the proposed model contributes to achieving balanced clustering results. The entire clustering model combines information from multiple views through graph fusion. The joint graph and view weight parameters in the model are obtained through task-driven self-supervised learning. Moreover, the model can directly obtain clustering results without the need for the two-stage processing typically used in general spectral clustering. Finally, extensive experiments on toy datasets and real-world datasets are conducted to validate the superiority of the proposed method in terms of clustering performance, clustering balance, and time expenditure.
Zihua Zhao, Feiping Nie 0001, Rong Wang 0001, Zheng Wang 0037, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.4
2024 Double-Structured Sparsity Guided Flexible Embedding Learning for Unsupervised Feature Selection
abstract
In this article, we propose a novel unsupervised feature selection model combined with clustering, named double-structured sparsity guided flexible embedding learning (DSFEL) for unsupervised feature selection. DSFEL includes a module for learning a block-diagonal structural sparse graph that represents the clustering structure and another module for learning a completely row-sparse projection matrix using the$\ell_{2,0}$-norm constraint to select distinctive features. Compared with the commonly used$\ell_{2,1}$-norm regularization term, the$\ell_{2,0}$-norm constraint can avoid the drawbacks of sparsity limitation and parameter tuning. The optimization of the$\ell_{2,0}$-norm constraint problem, which is a nonconvex and nonsmooth problem, is a formidable challenge, and previous optimization algorithms have only been able to provide approximate solutions. In order to address this issue, this article proposes an efficient optimization strategy that yields a closed-form solution. Eventually, through comprehensive experimentation on nine real-world datasets, it is demonstrated that the proposed method outperforms existing state-of-the-art unsupervised feature selection methods.
Yu Guo 0006, Yuan Sun 0003, Zheng Wang 0037, Feiping Nie 0001, Fei Wang 0008
IEEE Trans. Neural Networks Learn. Syst.3
2024 Semisupervised Subspace Learning With Adaptive Pairwise Graph Embedding
abstract
Graph-based semisupervised learning can explore the graph topology information behind the samples, becoming one of the most attractive research areas in machine learning in recent years. Nevertheless, existing graph-based methods also suffer from two shortcomings. On the one hand, the existing methods generate graphs in the original high-dimensional space, which are easily disturbed by noisy and redundancy features, resulting in low-quality constructed graphs that cannot accurately portray the relationships between data. On the other hand, most of the existing models are based on the Gaussian assumption, which cannot capture the local submanifold structure information of the data, thus reducing the discriminativeness of the learned low-dimensional representations. This article proposes a semisupervised subspace learning with adaptive pairwise graph embedding (APGE), which first builds a -nearest neighbor graph on the labeled data to learn local discriminant embeddings for exploring the intrinsic structure of the non-Gaussian labeled data, i.e., the submanifold structure. Then, a -nearest neighbor graph is constructed on all samples and mapped to GE learning to adaptively explore the global structure of all samples. Clustering unlabeled data and its corresponding labeled neighbors into the same submanifold, sharing the same label information, improves embedded data's discriminative ability. And the adaptive neighborhood learning method is used to learn the graph structure in the continuously optimized subspace to ensure that the optimal graph matrix and projection matrix are finally learned, which has strong robustness. Meanwhile, the rank constraint is added to the Laplacian matrix of the similarity matrix of all samples so that the connected components in the obtained similarity matrix are precisely equal to the number of classes in the sample, which makes the structure of the graph clearer and the relationship between the near-neighbor sample points more explicit. Finally, multiple experiments on several synthetic and real-world datasets show that the method performs well in exploring local structure and classification tasks.
Hebing Nie, Qi Li 0045, Zheng Wang 0037, Haifeng Zhao 0001, Feiping Nie 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Toward Robust Discriminative Projections Learning Against Adversarial Patch Attacks
abstract
As one of the most popular supervised dimensionality reduction methods, linear discriminant analysis (LDA) has been widely studied in machine learning community and applied to many scientific applications. Traditional LDA minimizes the ratio of squared norms, which is vulnerable to the adversarial examples. In recent studies, many -norm-based robust dimensionality reduction methods are proposed to improve the robustness of model. However, due to the difficulty of -norm ratio optimization and weakness on defending a large number of adversarial examples, so far, scarce works have been proposed to utilize sparsity-inducing norms for LDA objective. In this article, we propose a novel robust discriminative projections learning (rDPL) method based on the -norm trace-ratio minimization optimization algorithm. Minimizing the -norm ratio problem directly is a much more challenging problem than the traditional methods, and there is no existing optimization algorithm to solve such nonsmooth terms ratio problem. We derive a new efficient algorithm to solve this challenging problem and provide a theoretical analysis on the convergence of our algorithm. The proposed algorithm is easy to implement and converges fast in practice. Extensive experiments on both synthetic data and several real benchmark datasets show the effectiveness of the proposed method on defending the adversarial patch attack by comparison with many state-of-the-art robust dimensionality reduction methods.
Zheng Wang 0037, Feiping Nie 0001, Hua Wang 0007, Heng Huang 0001, Fei Wang 0008
IEEE Trans. Neural Networks Learn. Syst.1
2024 Robust Principal Component Analysis via Joint Reconstruction and Projection
abstract
-norm is used as distance metric. Recently, many scholars have devoted themselves to solving this difficulty. They learn the projection matrix from minimum reconstruction error or maximum projection variance as the starting point, which leads them to ignore a serious problem, that is, the original PCA learns the projection matrix by minimizing the reconstruction error and maximizing the projection variance simultaneously, but they only consider one of them, which imposes various limitations on the performance of model. To solve this problem, we propose a novel robust principal component analysis via joint reconstruction and projection, namely, RPCA-RP, which combines reconstruction error and projection variance to fully mine the potential information of data. Furthermore, we carefully design a discrete weight for model to implicitly distinguish between normal data and outliers, so as to easily remove outliers and improve the robustness of method. In addition, we also unexpectedly discovered that our method has anomaly detection capabilities. Subsequently, an effective iterative algorithm is explored to solve this problem and perform related theoretical analysis. Extensive experimental results on several real-world datasets and RGB large-scale dataset demonstrate the superiority of our method.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Joint Anchor Graph Embedding and Discrete Feature Scoring for Unsupervised Feature Selection
abstract
The success of existing unsupervised feature selection (UFS) methods heavily relies on the assumption that the intrinsic relationships among original high-dimensional (HD) data samples exist in the discriminative low-dimension (LD) subspace. However, previous UFS methods commonly construct pairwise graphs and employ$\ell_{2,1}$-norm regularization to severally preserve the local structure and calculate the score of features, which is computationally complex and easy to get stuck into local optimum, so that those approaches cannot be applied in dealing with large-scale datasets in practice. To overcome this challenge, we propose a novel UFS method, in which a novel anchor graph embedding paradigm is designed to extract the local neighborhood relationships among data samples by reducing the computational complexity of graph construction to be linear in the number of data. Moreover, to improve the optimality of selected features as well as the performance of downstream tasks, we propose a discrete feature scoring mechanism, which imposes orthogonal$\ell_{2,0}$-norm constraints on learned projections, in order to enhance the distinction of feature scores as well as reduce the probability of falling into local optimum. In addition, solving the proposed nonconvex and nonsmooth NP-hard problem is challenging, and we present an efficient optimization algorithm to address it and acquire a closed-form solution of the transformation matrix. Extensive experiments demonstrate the effectiveness and efficiency of the proposed UFS by comparison with several state-of-the-art approaches to clustering and image segmentation tasks.
Zheng Wang 0037, Dongming Wu 0001, Rong Wang 0001, Feiping Nie 0001, Fei Wang 0008
IEEE Trans. Neural Networks Learn. Syst.1
2024 Pseudo-Label Guided Structural Discriminative Subspace Learning for Unsupervised Feature Selection
abstract
In this article, we propose a new unsupervised feature selection method named pseudo-label guided structural discriminative subspace learning (PSDSL). Unlike the previous methods that perform the two stages independently, it introduces the construction of probability graph into the feature selection learning process as a unified general framework, and therefore the probability graph can be learned adaptively. Moreover, we design a pseudo-label guided learning mechanism, and combine the graph-based method and the idea of maximizing the between-class scatter matrix with the trace ratio to construct an objective function that can improve the discrimination of the selected features. Besides, the main existing strategies of selecting features are to employ -norm for feature selection, but this faces the challenges of sparsity limitations and parameter tuning. For addressing this issue, we employ the -norm constraint on the learned subspace to ensure the row sparsity of the model and make the selected feature more stable. Effective optimization strategy is given to solve such NP-hard problem with the determination of parameters and complexity analysis in theory. Ultimately, extensive experiments conducted on nine real-world datasets and three biological ScRNA-seq genes datasets verify the effectiveness of the proposed method on the data clustering downstream task.
Zheng Wang 0037, Yongjin Yuan, Rong Wang 0001, Feiping Nie 0001, Qinghua Huang, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Graph Joint Representation Clustering via Penalized Graph Contrastive Learning
abstract
Graph clustering based on graph contrastive learning (GCL) is one of the dominant paradigms in the current graph clustering research field. However, those GCL-based methods often yield false negative samples, which can distort the learned representations and limit clustering performance. In order to alleviate this issue, we propose the idea of maintaining mutual information (MI) between the representations and the inputs to mitigate the loss of semantic information of false negative samples. We demonstrate the validity of this proposal through relevant experiments. Since maximizing MI can be approximately replaced by minimizing reconstruction error, we further propose a graph clustering method based on GCL penalized by reconstruction error, in which our carefully designed reconstruction decoder, as well as reconstruction error term, improve the clustering performance. In addition, we use a pseudo-label-guided strategy to improve the GCL process and further alleviate the problem of false negative samples. Our experiment results demonstrate the superiority and great potential of our proposed graph clustering method compared with state-of-the-art algorithms.
Zihua Zhao, Rong Wang 0001, Zheng Wang 0037, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Unsupervised Feature Selection with self-Weighted and ℓ2,0-Norm Constraint
abstract
At data mining field, it is a fundamental problem to dispose of high-dimensional data. Many existing unsupervised methods select features by manifold learning or exploring spectral analysis, thus preserving the intrinsic structure of raw data. But most of them follow an assumption that all features are equally importance. To settle this problem, we draw a novel feature selection module that simultaneously performs learning of feature weights matrix, similarity graph structure and projection matrix, so that the local structure after feature weighting and subspace sparse projection is received. Finally, we solve the model based on ℓ2,0-norm directly by an iterative optimization algorithm and demonstrate the feasibility and effectiveness of our approach via extensive experiments.
Yongjin Yuan, Zheng Wang 0037, Feiping Nie 0001, Xuelong Li 0001
ICASSP2
2023 Max-Min Robust Principal Component Analysis
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
Neurocomputing3
2023 Fast spectral clustering with self-adapted bipartite graph learning
Mingjun Zhu, Yongda Cai, Zheng Wang 0037, Feiping Nie 0001
Inf. Sci.4
2023 Fuzzy C-Multiple-Means Clustering for Hyperspectral Image
abstract
Currently, unsupervised hyperspectral image (HSI) segmentation methods are mainly implemented by clustering. Nevertheless, hyperspectral data contains a large amount of noise during the acquisition process, resulting in an abnormal distribution of many pixel points. Traditional clustering algorithms suffer from inaccurate segmentation when dealing with these data. For example, FCM is sensitive to anomalies in the clustering problem of HSI, that makes the clustering accuracy degraded. To address these problems, this paper proposes a method called Fuzzy C-Multiple-Means (FCMM). The method divides data points with multiple subclusters into definedcclusters. Different from the bottom-up coalescent strategy, the proposed FCMM transforms the problem of merging multiple subclusters into an optimisation problem for the fuzzy affiliation matrix, and updates the partitioning of theqsubclusters andcclasses by an alternating iterative update method. This enhances the robustness of the algorithm and reduces the effect of outliers in the HSI datasets on the FCMM, which provides superior clustering results. Experiments on several HSI datasets validate the effectiveness of FCMM.
Mingjun Zhu, Zheng Wang 0037, Feiping Nie 0001
IEEE Geosci. Remote. Sens. Lett.4
2023 Simultaneous local clustering and unsupervised feature selection via strong space constraint
Zheng Wang 0037, Qi Li 0045, Haifeng Zhao 0001, Feiping Nie 0001
Pattern Recognit.1
2023 Sparse and Flexible Projections for Unsupervised Feature Selection
abstract
In recent decades, unsupervised feature selection methods have become increasingly popular. Nevertheless, most of the existing unsupervised feature selection methods suffer from two major problems that lead to suboptimal solutions. Many methods impose a hard linear projection constraint on original data, which is overly strict in nature and not suitable for dealing with data sampled from nonlinear manifolds. Second, most existing methods usel2,p-norm (02S and SF2SOG, which can simultaneously learn optimal flexible projections and obtain an orthogonal sparse projection to directly select discriminative features by applyingl2,0-norm constraint. Moreover, we propose to explore the local structure of flexible embedding through preserving the manifold structure of original data and adaptively constructing an optimal graph in subspace. Thirdly, the novel iterative optimization algorithms are presented to solve objective functions guaranteeing convergence theoretically. Various evaluation experiments on synthetic and real-world datasets demonstrate the effectiveness and superiority of our proposed methods.
Rong Wang 0001, Canyu Zhang 0001, Jintang Bian, Zheng Wang 0037, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.4
2023 Self-Supervised Learning for Heterogeneous Audiovisual Scene Analysis
abstract
Due to the difficulty of annotating large amounts of training data, directly learning the association of sound and its makers in natural videos is a challenging task for machines. In this paper, we present a novel audiovisual model that introduces a soft-clustering module as the audio and visual content detector, and regards the pervasive property of audiovisual concurrency as the latent supervision for inferring the correlation among detected contents. Furthermore, we discover for the first time that the complexity of data has an impact on the training efficiency and subsequent performance of audiovisual model, i.e., more complex data brings more obstacles to the model training, and degrades the performance of downstream audiovisual tasks. To address the issue of audiovisual learning, we propose a novel heterogeneous audiovisual scene analysis module that trains the model from simple to complex scene. We show that such ordered learning procedure rewards the model the merits of easy training and fast convergence. Meanwhile, our audiovisual model can also provide effective unimodal representation and cross-modal alignment performance. We further deploy the well-trained model into practical audiovisual sound localization and separation tasks. We show that our localization model significantly outperforms existing methods, based on which we show comparable performance in sound separation task by comparison to several related SOTA audiovisual learning methods without referring external visualsupervision.
Di Hu 0001, Zheng Wang 0037, Feiping Nie 0001, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Multim.2
2023 Discrete Robust Principal Component Analysis via Binary Weights Self-Learning
abstract
Principal component analysis (PCA) is a typical unsupervised dimensionality reduction algorithm, and one of its important weaknesses is that the squared$\ell _{2}$-norm cannot overcome the influence of outliers. Existing robust PCA methods based on paradigm have the following two drawbacks. First, the objective function of PCA based on the$\ell _{1}$-norm has no rotational invariance and limited robustness to outliers, and its solution mostly uses a greedy search strategy, which is expensive. Second, the robust PCA based on the$\ell _{2,1}$-norm and the$\ell _{2,p}$-norm is essential to learn probability weights for data, which only weakens the influence of outliers on the learning projection matrix and cannot be completely eliminated. Moreover, the ability to detect anomalies is also very poor. To solve these problems, we propose a novel discrete robust principal component analysis (DRPCA). Through self-learning binary weights, the influence of outliers on the projection matrix and data center estimation can be completely eliminated, and anomaly detection can be directly performed. In addition, an alternating iterative optimization algorithm is designed to solve the proposed problem and realize the automatic update of binary weights. Finally, our proposed model is successfully applied to anomaly detection applications, and experimental results demonstrate that the superiority of our proposed method compared with the state-of-the-art methods.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Local Embedding Learning via Landmark-Based Dynamic Connections
abstract
Linear discriminant analysis (LDA) is one of the most effective and popular methods to reduce the dimensionality of data with Gaussian assumption. However, LDA cannot handle non-Gaussian data because the center point is incompetent to represent the distribution of data. Some existing methods based on graph embedding focus on exploring local structures via pairwise relationships of data for addressing the non-Gaussian issue. Due to massive pairwise relationships, the computational complexity is high as well as the locally optimal solution is hard to find. To address these issues, we propose a novel and efficient local embedding learning via landmark-based dynamic connections (LDC) in which we leverage several landmarks to represent different subclusters in the same class and establish the connections between each point and landmark. Furthermore, in order to explore the relationship of landmarks pairwise more precisely, the relationship between each point and their corresponding neighbor landmarks are found in the optimal subspace, rather than the original space, which can avoid the negative influence of the noises. We also propose an efficient iterative algorithm to deal with the proposed ratio minimization problem. Extensive experiments conducted on several real-world datasets have demonstrated the advantages of the proposed method.
Feiping Nie 0001, Canyu Zhang 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Capped ℓp-norm linear discriminant analysis for robust projections learning
Zheng Wang 0037, Haojie Hu, Rong Wang 0001, Qianrong Zhang, Feiping Nie 0001, Xuelong Li 0001
Neurocomputing1
2022 Self-weighted learning framework for adaptive locality discriminant analysis
Wei Chang 0002, Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
Pattern Recognit.3
2022 Subspace Sparse Discriminative Feature Selection
abstract
In this article, we propose a novel feature selection approach via explicitly addressing the long-standing subspace sparsity issue. Leveraging$\ell _{2,1}$-norm regularization for feature selection is the major strategy in existing methods, which, however, confronts sparsity limitation and parameter-tuning trouble. To circumvent this problem, employing the$\ell _{2,0}$-norm constraint to improve the sparsity of the model has gained more attention recently whereas, optimizing the subspace sparsity constraint is still an unsolved problem, which only can acquire an approximate solution and without convergence proof. To address the above challenges, we innovatively propose a novel subspace sparsity discriminative feature selection (S2DFS) method which leverages a subspace sparsity constraint to avoid tuning parameters. In addition, the trace ratio formulated objective function extremely ensures the discriminability of selected features. Most important, an efficient iterative optimization algorithm is presented to explicitly solve the proposed problem with a closed-form solution and strict convergence proof. To the best of our knowledge, such an optimization algorithm of solving the subspace sparsity issue is first proposed in this article, and a general formulation of the optimization algorithm is provided for improving the extensibility and portability of our method. Extensive experiments conducted on several high-dimensional text and image datasets demonstrate that the proposed method outperforms related state-of-the-art methods in pattern classification and image retrieval tasks.
Feiping Nie 0001, Zheng Wang 0037, Lai Tian, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Cybern.2
2022 Adaptive Local Embedding Learning for Semi-Supervised Dimensionality Reduction
abstract
Semi-supervised learning as one of most attractive problems in machine learning research field has aroused broad attentions in recent years. In this paper, we propose a novel locality preserved dimensionality reduction framework, named Semi-supervised Adaptive Local Embedding learning (SALE), which learns a local discriminative embedding by constructing a$k_1$Nearest Neighbors ($k_1$NN) graph on labeled data, so as to explore the intrinsic structure, i.e., sub-manifolds from non-Gaussian labeled data. Then, mapping all samples into learned embedding and constructing another$k_2$NN graph on all embedded data to explore the global structure of all samples. Therefore, the unlabeled data and their corresponding labeled neighbors can be clustered into same sub-manifold, so as to improve the discriminative power of embedded data. Furthermore, we propose two semi-supervised dimensionality reduction methods with orthogonal and whitening constraints based on proposed SALE framework. An efficient alternatively iterative optimization algorithm is developed to solve the NP-hard problem in our models. Extensive experiments conducted on several synthetic and real-world data sets demonstrate the superiorities of our methods on local structure exploration and classification task.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.2
2021 Fast Local Representation Learning with Adaptive Anchor Graph
abstract
Dimension reduction is an effective technology to embed data with high dimension to lower dimension space, where Linear Discriminant Analysis (LDA), one of representative methods, only works with Gaussian distribution data. However, in order to solve non-Gaussian issue that only one cluster cannot well fit the distribution of same class, many graph-based discriminant analysis methods are proposed which capture local structure through measuring each pairwise distance. This is expense of time complexity because of the full-connections. In order to solve this issue, we propose a fast local representation learning with adaptive anchor graph to learn local structure information through similarity matrix in anchor-based graph. Notably, anchor points and similarity matrix are updated in subspace which is more precisely to capture local discriminant information. Experimental results on several synthetic and well-known datasets demonstrate the advantages of our method over the state-of-the-art methods.
Canyu Zhang 0001, Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
ICASSP3
2021 Fast local representation learning via adaptive anchor graph for image retrieval
Canyu Zhang 0001, Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
Inf. Sci.3
2021 Towards Robust Discriminative Projections Learning via Non-Greedy $\ell _{2, 1}$ℓ2, 1-Norm MinMax
abstract
Linear Discriminant Analysis (LDA) is one of the most successful supervised dimensionality reduction methods and has been widely used in many real-world applications. However,l2ℓ2-norm is employed as the distance metric in the objective of LDA, which is sensitive to outliers. Many previous works improve the robustness of LDA by usingl1ℓ1-norm distance. However, the robustness against outliers is limited and the solver ofl1ℓ1-norm is mostly based on the greedy search strategy, which is time-consuming and easy to get stuck in a local optimum. In this paper, we propose a novel robust LDA measured byl2,1ℓ2,1-norm to learn robust discriminative projections. The proposed model is challenging to solve since it needs to minimize and maximize (minmax)l2,1ℓ2,1-norm terms simultaneously. As a result, we first systematically derive an efficient iterative optimization algorithm to solve a general ratio minimization problem, and then rigorously prove its convergence. More importantly, an alternately non-greedy iterative re-weighted optimization algorithm is developed based on the preceding approach for solving proposedl2,1ℓ2,1-norm minmax problem. Besides, an optimal weighted mean mechanism is driven according to the designed objective and solver, which can be applied to other approaches for robustness improvement. Experimental results on several real-world datasets show the effectiveness of proposed method.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Zhen Wang 0004, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Local structured feature learning with dynamic maximum entropy graph
Zheng Wang 0037, Feiping Nie 0001, Rong Wang 0001, Hui Yang 0005, Xuelong Li 0001
Pattern Recognit.1
2021 Joint nonlinear feature selection and continuous values regression network
Zheng Wang 0037, Feiping Nie 0001, Canyu Zhang 0001, Rong Wang 0001, Xuelong Li 0001
Pattern Recognit. Lett.1
2020 Discriminative Feature Selection via A Structured Sparse Subspace Learning Module
abstract
In this paper, we first propose a novel Structured Sparse Subspace Learning S^3L module to address the long-standing subspace sparsity issue. Elicited by proposed module, we design a new discriminative feature selection method, named Subspace Sparsity Discriminant Feature Selection S^2DFS which enables the following new functionalities: 1) Proposed S^2DFS method directly joints trace ratio objective and structured sparse subspace constraint via L2,0-norm to learn a row-sparsity subspace, which improves the discriminability of model and overcomes the parameter-tuning trouble with comparison to the methods used L2,1-norm regularization; 2) An alternative iterative optimization algorithm based on the proposed S^3L module is presented to explicitly solve the proposed problem with a closed-form solution and strict convergence proof. To our best knowledge, such objective function and solver are first proposed in this paper, which provides a new though for the development of feature selection methods. Extensive experiments conducted on several high-dimensional datasets demonstrate the discriminability of selected features via S^2DFS with comparison to several related SOTA feature selection methods. Source matlab code: https://github.com/StevenWangNPU/L20-FS.
Zheng Wang 0037, Feiping Nie 0001, Lai Tian, Rong Wang 0001, Xuelong Li 0001
IJCAI1
2020 Multiclass discriminant analysis via adaptive weighted scheme
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001
Neurocomputing3
2020 Capped ℓp-Norm LDA for Outliers Robust Dimension Reduction
abstract
Linear discriminant analysis technique is an effective strategy to solve the long-standing issue, i.e., the “curse of dimensionality” that brings many obstacles on high-dimensional data storage and analysis. However, the projections are prone to be affected, especially when the training set contains outlier samples whose distribution deviates from the globality. In many real-world applications, the outlier samples contaminated by noisy signal or spottiness have negative effects on the classification and clustering performance. To address this issue, we propose to develop a novel capped ℓp-norm LDA model for robust dimension reduction against to outliers specifically. Proposed method integrates the capped ℓp-norm based loss into the objective, which not only suppresses the light outliers but also works well even though the training set is contaminated seriously. Furthermore, we derive an alternative iterative re-weighted optimization algorithm to minimize the proposed objective based on capped ℓp-norm with rigorous convergence proofs. Extensive experiments conducted on synthetic and real-world datasets demonstrate the robustness against to outliers of proposed method.
Zheng Wang 0037, Feiping Nie 0001, Canyu Zhang 0001, Rong Wang 0001, Xuelong Li 0001
IEEE Signal Process. Lett.1
2020 Submanifold-Preserving Discriminant Analysis With an Auto-Optimized Graph
abstract
Due to the multimodality of non-Gaussian data, traditional globality-preserved dimensionality reduction (DR) methods, such as linear discriminant analysis (LDA) and principal component analysis (PCA) are difficult to deal with. In this paper, we present a novel local DR framework via auto-optimized graph embedding to extract the intrinsic submanifold structure of multimodal data. Specifically, the proposed model seeks to learn an embedding space which can preserve the local neighborhood structure by constructing a k-nearest neighbors (kNNs) graph on data points. Different than previous works, our model employs the ℓ0-norm constraint and binary constraint on the similarity matrix to impose that there only be a k nonzero value in each row of the similarity matrix, which can ensure the k-connectivity in graph. More important, as the high-dimensional data probably contains some noises and redundant features, calculating the similarity matrix in the original space by using a kernel function is inaccurate. As a result, a mechanism of an auto-optimized graph is derived in the proposed model. Concretely, we learn the embedding space and similarity matrix simultaneously. In other words, the selection of neighbors is automatically executed in the optimal subspace rather than in the original space when the algorithm reaches convergence, which can alleviate the affect of noises and improve the robustness of the proposed model. In addition, four supervised and semi-supervised local DR methods are derived by the proposed framework which can extract the discriminative features while preserving the submanifold structure of data. Last but not least, since two variables need to be optimized simultaneously in the proposed methods, and the constraints on the similarity matrix are difficult to satisfy, which is an NP-hard problem. Consequently, an efficient iterative optimization algorithm is introduced to solve the proposed problems. Extensive experiments conducted on synthetic data and several real-world datasets have demonstrated the advantages of the proposed methods in robustness and recognition accuracy.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Xuelong Li 0001
IEEE Trans. Cybern.2
2020 Adaptive Local Linear Discriminant Analysis
abstract
Dimensionality reduction plays a significant role in high-dimensional data processing, and Linear Discriminant Analysis (LDA) is a widely used supervised dimensionality reduction approach. However, a major drawback of LDA is that it is incapable of extracting the local structure information, which is crucial for handling multimodal data. In this article, we propose a novel supervised dimensionality reduction method named Adaptive Local Linear Discriminant Analysis (ALLDA), which adaptively learns a k -nearest neighbors graph from data themselves to extract the local connectivity of data. Furthermore, the original high-dimensional data usually contains noisy and redundant features, which has a negative impact on the evaluation of neighborships and degrades the subsequent classification performance. To address this issue, our method learns the similarity matrix and updates the subspace simultaneously so that the neighborships can be evaluated in the optimal subspaces where the noises have been removed. Through the optimal graph embedding, the underlying sub-manifolds of data in intra-class can be extracted precisely. Meanwhile, an efficient iterative optimization algorithm is proposed to solve the minimization problem. Promising experimental results on synthetic and real-world datasets are provided to evaluate the effectiveness of proposed method.
Feiping Nie 0001, Zheng Wang 0037, Rong Wang 0001, Zhen Wang 0004, Xuelong Li 0001
ACM Trans. Knowl. Discov. Data2
2019 A New Formulation of Linear Discriminant Analysis for Robust Dimensionality Reduction
abstract
Dimensionality reduction is a critical technology in the domain of pattern recognition, and linear discriminant analysis (LDA) is one of the most popular supervised dimensionality reduction methods. However, whenever its distance criterion of objective function uses$L_2$-norm, it is sensitive to outliers. In this paper, we propose a new formulation of linear discriminant analysis via joint$L_{2,1}$-norm minimization on objective function to induce robustness, so as to efficiently alleviate the influence of outliers and improve the robustness of proposed method. An efficient iterative algorithm is proposed to solve the optimization problem and proved to be convergent. Extensive experiments are performed on an artificial data set, on UCI data sets, and on four face data sets, which sufficiently demonstrates the efficiency of comparing to other methods and robustness to outliers of our approach.
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.2
2018 Adaptive Neighborhood MinMax Projections
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001
Neurocomputing2
2018 Multiclass Classification and Feature Selection Based on Least Squares Regression with Large Margin
abstract
Least squares regression (LSR) is a fundamental statistical analysis technique that has been widely applied to feature learning. However, limited by its simplicity, the local structure of data is easy to neglect, and many methods have considered using orthogonal constraint for preserving more local information. Another major drawback of LSR is that the loss function between soft regression results and hard target values cannot precisely reflect the classification ability; thus, the idea of the large margin constraint is put forward. As a consequence, we pay attention to the concepts of large margin and orthogonal constraint to propose a novel algorithm, orthogonal least squares regression with large margin (OLSLM), for multiclass classification in this letter. The core task of this algorithm is to learn regression targets from data and an orthogonal transformation matrix simultaneously such that the proposed model not only ensures every data point can be correctly classified with a large margin than conventional least squares regression, but also can preserve more local data structure information in the subspace. Our efficient optimization method for solving the large margin constraint and orthogonal constraint iteratively proved to be convergent in both theory and practice. We also apply the large margin constraint in the process of generating a sparse learning model for feature selection via joint [Formula: see text]-norm minimization on both loss function and regularization terms. Experimental results validate that our method performs better than state-of-the-art methods on various real-world data sets.
Haifeng Zhao 0001, Zheng Wang 0037
Neural Comput.3
2016 Orthogonal least squares regression for feature extraction
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001
Neurocomputing2