Mengjie Zhang 0001

dblp:z/MengjieZhang · DBLP profile ↗
← Back
19ranked-venue papers in the field
1as first author
7since 2021 · last 2025
0000-0003-4463-9538ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 10Database Systems & Data Management · 4 (1 first)Other / Interdisciplinary · 3Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2025 Improving Generalization of Genetic Programming for High-Dimensional Symbolic Regression with Shapley Value Based Feature Selection
abstract
Abstract Symbolic Regression (SR) on high-dimensional datasets often encounters significant challenges, resulting in models with poor generalization capabilities. While feature selection has the potential to enhance the generalization and learning performance in general, its application in Genetic Programming (GP) for high-dimensional SR remains a complex problem. Originating from game theory, the Shapley value is applied to additive feature attribution approaches where it distributes the difference between a model output and a baseline average across input variables. By providing an accurate assessment of each feature importance, the Shapley value offers a robust approach to select features. In this paper, we propose a novel feature selection method leveraging the Shapley value to identify and select important features in GP for high-dimensional SR. Through a series of experiments conducted on ten high-dimensional regression datasets, the results indicate that our algorithm surpasses standard GP and other GP-based feature selection methods in terms of learning and generalization performance on most datasets. Further analysis reveals that our algorithm generates more compact models, focusing on the inclusion of important features.
Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001
Data Sci. Eng.4
2024 ConvFishNet: An efficient backbone for fish classification from composited underwater images
Huishan Qu, Gaige Wang, Mengjie Zhang 0001
Inf. Sci.5
2024 Enhancing generalization in genetic programming hyper-heuristics through mini-batch sampling strategies for dynamic workflow scheduling
abstract
Genetic Programming Hyper-heuristics (GPHH) have been successfully used to evolve scheduling rules for Dynamic Workflow Scheduling (DWS) as well as other challenging combinatorial optimization problems. The method of sampling training instances has a significant impact on the generalization ability of GPHH, yet they are rarely addressed in existing research. This article aims to fill this gap by proposing a GPHH algorithm with a sampling strategy to thoroughly investigate the impact of six instance sampling strategies on algorithmic generalization, including one rotation strategy, three mini-batch strategies, and two hybrid strategies. Experiments across four scenarios with varying settings reveal that: (1) mini-batch with random sampling can outperform rotation in generalizing to unseen workflow scheduling problems under the same computational cost; (2) employing a hybrid strategy that combines rotation and mini-batch further enhances the generalization ability of GPHH; and (3) mini-batch and hybrid strategies can effectively enable heuristics trained on small-scale training instances generalizing well to large-scale unseen ones. These findings highlight the potential of mini-batch strategies in GPHH, offering improved generalization performance while maintaining diversity and suggesting promising avenues for further exploration in GPHH domains.
Yifan Yang 0002, Gang Chen 0002, Hui Ma 0001, Sven Hartmann, Mengjie Zhang 0001
Inf. Sci.5
2024 An evolutionary neural architecture search method based on performance prediction and weight inheritance
abstract
Evolutionary Neural Architecture Search (ENAS) algorithms attract great attention since they can automatically search for appropriate network architectures for a given task. However, most ENAS algorithms suffer from a prohibitive computational burden. Moreover, some of these approaches directly use performance predictors for evaluations, which may introduce inaccurate assessments and harm the evolution. To overcome these shortcomings, we propose an efficient ENAS algorithm named EPPGA. EPPGA employs a predictor to pre-select potentially high-performing offspring, enhancing the performance and accelerating the evolution. As the offspring will be further accurately evaluated, even potentially inaccurate predictions will not adversely affect the evolution. Furthermore, a weight inheritance method is suggested to accelerate the evaluation, and new genetic operations are developed to produce offspring that share a substantial proportion of beneficial genetic materials with one parent, improving the performance predictor's effectiveness and promoting weight inheritance. Finally, a new efficient backbone block structure is designed to facilitate the search for lightweight networks. The experimental results demonstrate that EPPGA is a highly competitive algorithm on three benchmarks in terms of accuracy, model size, and computational cost, reveal the superiority of the proposed block structure, and confirm the effectiveness of the proposed performance predictor and weight inheritance method.
Gonglin Yuan, Bing Xue 0001, Mengjie Zhang 0001
Inf. Sci.3
2023 Multi-objective particle swarm optimization for key quality feature selection in complex manufacturing processes
An-Da Li, Bing Xue 0001, Mengjie Zhang 0001
Inf. Sci.3
2023 Feature Selection Using Diversity-Based Multi-objective Binary Differential Evolution
Peng Wang 0102, Bing Xue 0001, Jing J. Liang, Mengjie Zhang 0001
Inf. Sci.4
2022 Using a small number of training instances in genetic programming for face image classification
Ying Bi 0001, Bing Xue 0001, Mengjie Zhang 0001
Inf. Sci.3
2020 Multi-objective feature selection using hybridization of a genetic algorithm and direct multisearch for key quality characteristic selection
An-Da Li, Bing Xue 0001, Mengjie Zhang 0001
Inf. Sci.3
2019 Self-Adaptive Particle Swarm Optimization for Large-Scale Feature Selection in Classification
abstract
Many evolutionary computation (EC) methods have been used to solve feature selection problems and they perform well on most small-scale feature selection problems. However, as the dimensionality of feature selection problems increases, the solution space increases exponentially. Meanwhile, there are more irrelevant features than relevant features in datasets, which leads to many local optima in the huge solution space. Therefore, the existing EC methods still suffer from the problem of stagnation in local optima on large-scale feature selection problems. Furthermore, large-scale feature selection problems with different datasets may have different properties. Thus, it may be of low performance to solve different large-scale feature selection problems with an existing EC method that has only one candidate solution generation strategy (CSGS). In addition, it is time-consuming to find a suitable EC method and corresponding suitable parameter values for a given large-scale feature selection problem if we want to solve it effectively and efficiently. In this article, we propose a self-adaptive particle swarm optimization (SaPSO) algorithm for feature selection, particularly for large-scale feature selection. First, an encoding scheme for the feature selection problem is employed in the SaPSO. Second, three important issues related to self-adaptive algorithms are investigated. After that, the SaPSO algorithm with a typical self-adaptive mechanism is proposed. The experimental results on 12 datasets show that the solution size obtained by the SaPSO algorithm is smaller than its EC counterparts on all datasets. The SaPSO algorithm performs better than its non-EC and EC counterparts in terms of classification accuracy not only on most training sets but also on most test sets. Furthermore, as the dimensionality of the feature selection problem increases, the advantages of SaPSO become more prominent. This highlights that the SaPSO algorithm is suitable for solving feature selection problems, particularly large-scale feature selection problems.
Yu Xue 0003, Bing Xue 0001, Mengjie Zhang 0001
ACM Trans. Knowl. Discov. Data3
2018 Pareto front feature selection based on artificial bee colony optimization
Emrah Hancer, Bing Xue 0001, Mengjie Zhang 0001, Dervis Karaboga, Bahriye Akay
Inf. Sci.3
2016 Reusing Extracted Knowledge in Genetic Programming to Solve Complex Texture Image Classification Problems
Muhammad Iqbal 0001, Bing Xue 0001, Mengjie Zhang 0001
PAKDD (2)3
2015 GraphEvol: A Graph Evolution Technique for Web Service Composition
Alexandre Sawczuk da Silva, Hui Ma 0001, Mengjie Zhang 0001
DEXA (2)3
2014 An Enhanced Genetic Algorithm for Web Service Location-Allocation
Hui Ma 0001, Mengjie Zhang 0001
DEXA (2)3
2013 Genetic Programming with Greedy Search for Web Service Composition
Hui Ma 0001, Mengjie Zhang 0001
DEXA (2)3
2013 A novel particle swarm optimisation approach to detecting continuous, thin and smooth edges in noisy images
Mahdi Setayesh, Mengjie Zhang 0001, Mark Johnston
Inf. Sci.2
2007 Automatic Data Record Detection in Web Pages
Xiaoying Gao, Le Phong Bao Vuong, Mengjie Zhang 0001
KSEM3
2006 Data Extraction from Semi-structured Web Pages by Clustering
abstract
This paper introduces an approach to the use of clustering for data extraction from semi-structured Web pages. A variant Hierarchical Agglomerative Clustering (HAC) algorithm K-neighbours-HAC is developed which uses the similarities of the data format (HTML tags) and the data content (text string values) to group similar text tokens into clusters. Using these clusters, similar text tokens are identi- fied as data fields and extracted as target information. The approach is examined and compared with a number of existing information extraction systems on two different sets of web pages and the results suggest that the new approach is effective for web information extraction and that it outperforms all of the existing approaches on these web sites.
Le Phong Bao Vuong, Xiaoying Gao, Mengjie Zhang 0001
Web Intelligence3
2003 Learning Information Extraction Patterns from Tabular Web Pages without Manual Labelling
abstract
We describe a domain independent approach to automatically constructing information extraction patterns for semistructured Web pages. The approach was tested on three corpora containing a series of tabular Web sites from different domains and achieved a success rate of at least 80%. A significant strength of the system is that it can infer extraction patterns from a single training page and does not require any manual labeling of the training page.
Xiaoying Gao, Mengjie Zhang 0001, Peter Andreae
Web Intelligence2
1999 Using Back Propagation Algorithm and Genetic Algorithm to Train and Refine Neural Networks for Object Detection
Mengjie Zhang 0001, Victor Ciesielski
DEXA1