Peng Wang 0102

dblp:95/4442-102 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-1328-5784ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 An efficient evolutionary feature selection algorithm with divide-and-conquer strategy for classification
abstract
As an important data preprocessing technique, feature selection aims to identify useful features and therefore reduce the dimensionality of the data. Recent studies have witnessed that evolutionary computation methods show significant potential in solving feature selection tasks. However, existing methods still encounter challenges due to the high computational costs, particularly when handling high-dimensional datasets. To tackle these issues, this work proposes a new evolutionary feature selection method with a divide-and-conquer strategy. The proposed method transforms a high-dimensional feature selection task into multiple low-dimensional sub-tasks. Multiple sub-populations corresponding to the multiple sub-tasks are evolved simultaneously. To further improve the quality of the generated candidate feature subsets, a set-based population update mechanism is introduced. Furthermore, diverse collaborative coevolution strategies are systematically explored and analyzed. The experiments conducted on 12 real-world classification datasets demonstrate that the proposed method achieves superior performance compared to 9 state-of-the-art feature selection methods. The results reveal that the proposed method is able to select smaller feature subsets while achieving higher classification accuracy across the majority of the used datasets with a reasonable low training time.
Ke Chen 0022, Peng Wang 0102, Jing J. Liang, Zhile Yang, Kunjie Yu
Expert Syst. Appl.3
2026 A Survey on Evolutionary Multimodal Optimization-Driven Machine Learning
abstract
Machine learning methods have obtained great success across a number of practical applications and attracted growing interest. How to improve learning performance has emerged as a prominent research area. As an important area in evolutionary computation, evolutionary multi-solution/multimodal optimization has been widely used for dealing with different learning tasks. The key attention is to find multiple optimal solutions/models using the population-based search mechanism. This survey presents a very first comprehensive review of applying evolutionary multi-solution/multimodal optimization for addressing classical machine learning tasks and the fundamental concepts related to this domain. In this context, the multi-solution characteristics of a learning task, i.e., different solutions/models with very similar or the same performance, are emphasized and analyzed. Additionally, successful applications with multi-solution characteristics are presented. This work also discusses key challenges and provides insights into potential emerging directions for the future developments of applying evolutionary multimodal optimization to machine learning.
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.1
2025 A new multi-tree Genetic Programming approach to feature construction in high-dimensional classification
Ke Chen 0022, Mingyang Dao, Ying Bi 0001, Jing J. Liang, Zhenlong Wu, Peng Wang 0102
Knowl. Based Syst.6
2024 Dimensionality Reduction for Classification Using Divide-and-Conquer Based Genetic Programming
abstract
Dimensionality reduction (DR) is to obtain meaningful low-dimensional representation concealed within high-dimensional data. Genetic programming (GP) has been used to achieve DR for classification because of the flexible representations and promising performance. However, one drawback of many current GP-based DR methods is the high computational demands, potentially limiting their utility in high-dimensional scenarios. To tackle this issue, this work develops a new GP algorithm to achieve DR for classification. A divide-and-conquer strategy based on feature clustering is proposed to split the original feature set into multiple small feature groups. Each tree in the multi-tree GP is used to learn one high-level feature from each feature group. The performance of the proposed method is examined on 10 classification datasets with varying difficulty. The results indicate that the new GP method achieves significantly better classification performance than commonly used feature construction and feature selection methods. More importantly, the proposed GP method uses fewer nodes but achieves better or comparable classification accuracy than the baseline multi-tree GP method.
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
CEC1
2023 Feature Selection Using Diversity-Based Multi-objective Binary Differential Evolution
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
Inf. Sci.1
2023 Feature clustering-Assisted feature selection with differential evolution
abstract
Modern data collection technologies may produce thousands of or even more features in a single dataset. The high dimensionality of data poses a barrier to determining discriminating features due to the curse of dimensionality. Thanks to the global search ability, many population-based feature selection approaches have been proposed. However, very few studies pay attention on that a feature selection task has multiple optimal feature subsets. To search for multiple optimal feature subsets, we propose a feature clustering-assisted feature selection method. The proposed method employs the knowledge of correlation measures to group features. And, this correlation knowledge is embedded into the encoding method and the search process. A niching-based mutation operator is also used to explore the vicinity of a target individual. The aim is to find different feature subsets with very similar or the same classification performance. In addition, a modification operator is proposed aiming to increase the population diversity to improve the feature selection performance. The experiments on 16 datasets show that the proposed algorithm outperforms other popular feature selection methods in terms of classification accuracy and feature subset size.
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
Pattern Recognit.1
2023 Multiobjective Differential Evolution for Feature Selection in Classification
abstract
Feature selection aims to reduce the number of features and improve the classification accuracy, which is an essential step in many real-world problems. Multiple feature subsets with different features selected can achieve similar or even the same objective values (e.g., maximize the classification accuracy and minimize the number of selected features). This means the optimal feature subsets of a classification problem may not be unique. However, most existing feature selection methods do not take into consideration finding multiple optimal feature subsets. In this article, a multiobjective differential evolution approach is developed to search for multiple optimal feature subsets. The contributions are three-fold. First, to provide a good starting point, an initialization method considering feature relevance is proposed. Second, a clustering method is used to divide the whole population into multiple subpopulations. In each of these subpopulations, a subarchive utilizes a developed crowding distance to ensure diversity by considering both the search space and the objective space. Finally, the nondominated solutions from all the subarchives are retained in another archive to guide the evolutionary feature selection process, together with an improved hypervolume contribution indicator. The experiments on 14 datasets of varying difficulty show that the proposed approach can evolve a better Pareto front of feature subsets compared with seven other state-of-the-art methods as well as find different feature subsets with similar or the same classification performance.
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
IEEE Trans. Cybern.1
2023 Differential Evolution With Duplication Analysis for Feature Selection in Classification
abstract
By selecting a small subset of relevant features, feature selection can reduce the dimensionality of the problem while maintaining or increasing the discriminating ability of the data. However, many existing feature selection approaches ignore the fact that there are multiple optimal solutions to a feature selection problem. Multiple feature subsets with different features selected can achieve very similar or the same classification accuracy. To search for multiple optimal feature subsets, a niching-based differential evolution (DE) method with duplication analysis is proposed. In the proposed method, the duplicated feature subsets in the population are modified by the proposed subset repairing scheme which can produce unique feature subsets. Second, the mutation operator in DE is improved, which uses both the niche and global information to produce promising feature subsets. Third, a new selection method considering the diversity among feature subsets is adopted to form a new population for the next-generation. In the experiments, the proposed method is compared with seven evolutionary feature selection algorithms and two typical feature selection methods on 18 datasets. The results show that the proposed algorithm achieves higher classification accuracy than the compared methods on most of the used datasets. Furthermore, the proposed method can find different feature subsets with very similar or the same classification accuracy.
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
IEEE Trans. Cybern.1
2023 Differential Evolution-Based Feature Selection: A Niching-Based Multiobjective Approach
abstract
Feature selection is to reduce both the dimensionality of data and the classification error rate (i.e., increase the classification accuracy) of a learning algorithm. The two objectives are often conflicting, thus a multiobjective feature selection method can obtain a set of nondominated feature subsets. Each solution in the set has a different size and a corresponding classification error rate. However, most existing feature selection algorithms have ignored that, for a given size, there can be different feature subsets with very similar or the same accuracy. This article introduces a niching-based multiobjective feature selection method that simultaneously minimizes the number of selected features and the classification error rate. The proposed method conceives to identify: 1) a set of feature subsets with good convergence and distribution and 2) multiple feature subsets choosing the same number of features with almost the same lowest classification error rate. The contributions of this article are threefold. First, a niching and global interaction mutation operator is proposed that can produce promising feature subsets. Second, a newly developed environmental selection mechanism allows equal informative feature subsets to be stored by relaxing the Pareto-dominance relationship. Finally, the proposed subset repairing mechanism can generate better feature subsets and further remove the redundant features. The proposed method is compared against seven multiobjective feature selection algorithms on 19 datasets, including both binary and multiclass classification tasks. The results show that the proposed method can evolve a rich and diverse set of nondominated solutions for different feature selection tasks, and their availability helps in understanding the relationships between features.
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
IEEE Trans. Evol. Comput.1
2021 A Grid-dominance based Multi-objective Algorithm for Feature Selection in Classification
abstract
Feature selection aims to select a small subset of relevant features while maintaining or even improving the classification performance over using all features. Feature selection can be considered as a multi-objective problem, i.e., minimizing the number of selected features and maximizing the classification accuracy (minimizing the classification error) simultaneously. Most evolutionary multi-objective algorithms encounter difficulties when handling a feature selection task due to the discrete search space, although they perform well on continuous/numeric optimization problems. This paper proposes a grid-dominance based multi-objective evolutionary algorithm to address feature selection. The aim is to explore the potential of the grid-dominance method to strengthen the selection pressure toward the optimal direction while maintaining an extensive distribution among the objective values of feature subsets. To increase the population diversity, a subset filtration mechanism is proposed. The performance of the proposed two algorithms is tested on fourteen datasets of varying difficulty. With the proposed methods, the performance metrics, hypervolume and inverted generational distance have been significantly improved compared with other commonly used multi-objective algorithms, and the population diversity has also been increased.
Peng Wang 0102, Bing Xue 0001, Mengjie Zhang 0001, Jing J. Liang
CEC1
2021 Improved Crowding Distance in Multi-objective Optimization for Feature Selection in Classification
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
EvoApplications1
2018 Multi-objective Brainstorm Optimization Algorithm for Sparse Optimization
abstract
In sparse reconstruction, a small amount of samples are used to reconstruct sparse signal. This problem can be converted into multi-objective optimization problem, which considers sparsity and measurement error two competing cost function terms. Multi-objective optimization algorithm features unique advantage for solving this nonlinear and non-derivable problem. In order to trade off the sparsity and measurement error, in this paper, a novel multi-objective brainstorm algorithm based on objective space is used to optimize these two goals. Meanwhile, the local search ability of the algorithm is enhanced by introducing iterative half-thresholding operator in the frame of L1/2 regulation. On the basis of the Pareto solution set, the knee point is selected without prior signal information. In a coarse-to-fine manner, the algorithm has a certain tolerance to noise, and we give an analysis for this. According to the results on the eighteen test functions compared with several state-of-the-art compressive sensing recover methods, the effectiveness of multiobjective brainstorm algorithm based on objective space is verified for sparse optimization problems, which achieving higher recover and smaller error and reducing the average computational time for the long signal.
Jing J. Liang, Peng Wang 0102, Caitong Yue, Kunjie Yu, Bo-Yang Qu 0001
CEC2