VLDB 2026 Research / reviewers in the wild / expert
Qi Chen 0002
dblp:66/6320-2
· DBLP profile ↗
56ranked-venue papers
14as first author
40since 2021 · last 2026
0000-0001-9367-4757ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 14 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-tree Genetic Programming with Semantic Complementarity for Feature Construction in Symbolic Regression
Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
EuroGP | 2 |
| 2026 | Enhancing Generalization in Evolutionary Feature Construction for Symbolic Regression Through Vicinal Jensen Gap MinimizationabstractGenetic programming-based feature construction has achieved significant success in recent years as an automated machine learning technique to enhance learning performance. However, overfitting remains a challenge that limits its broader applicability. To improve generalization, we prove that vicinal risk, estimated through noise perturbation or mixup-based data augmentation, is bounded by the sum of empirical risk and a regularization termb–either finite difference or the vicinal Jensen gap. Leveraging this decomposition, we propose an evolutionary feature construction framework that jointly optimizes empirical risk and the vicinal Jensen gap to control overfitting. Since datasets may vary in noise levels, we develop a noise estimation strategy to dynamically adjust regularization strength. Furthermore, to mitigate manifold intrusionb–where data augmentation may generate unrealistic samples that fall outside the data manifoldb–we propose a manifold intrusion detection mechanism. Experimental results on 58 datasets demonstrate the effectiveness of Jensen gap minimization compared to other complexity measures. Comparisons with 15 machine learning algorithms further indicate that genetic programming with the proposed overfitting control strategy achieves superior performance. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2025 | Node Importance-Based Multi-Objective Genetic Programming for Enhanced Model Interpretability in Symbolic RegressionabstractInterpretability is a critical requirement in high-stakes real-world applications. Genetic Programming (GP) is one of the most interpretable machine learning algorithms currently available. While certain datasets necessitate complex GP models to capture underlying patterns, others may not require such complexity. However, GP lacks a mechanism to dynamically assess the required model complexity during evolution. As a result, GP often evolves overly complex models that are challenging to analyze, thereby reducing interpretability. This paper introduces a novel algorithm designed to evolve compact multi-objective GP models without significantly compromising regression error, thereby enhancing interpretability. The proposed method employs NSGA-II to simultaneously minimize regression error and maximize a newly introduced per-node importance metric. This metric quantifies the contribution of each node in terms of error reduction. By leveraging this metric, the algorithm discards large GP models containing many low-importance nodes, effectively avoiding unnecessarily complex solutions. Experimental results on ten regression datasets demonstrate that the proposed method consistently evolves smaller GP models with better interpretability while maintaining competitive predictive performance, particularly on high-dimensional datasets. Mohamad Rimas Mohamad Anfar, Qi Chen 0002, Mengjie Zhang 0001 |
CEC | 2 |
| 2025 | Semantics-Driven Task Similarity in Multi-Task Genetic Programming for Multi-Output Symbolic RegressionabstractMulti-Output symbolic regression predicts multiple target variables simultaneously, adding complexity overregression model with a single output, attributed to the interdependence between output variables. Effectively capturing these relationships is critical to improve prediction accuracy. Such a problem can be framed as a form of multi-task learning, where tasks share the same inputs but predict distinct output variables. This paper proposes a multi-task multi-population genetic programming algorithm that leverages the correlation between the distance matrices of the semantics of individuals in two populations to measure the similarity between the populations, thereby dynamically determining the similarity of two tasks. Furthermore, this precise similarity assessment guides a task selection strategy to determine the most appropriate task for knowledge transfer, ensuring efficient and effective transfer. Experiments on 18 real-world datasets demonstrate that the proposed method significantly enhances both training and test performance in multi-task GP, outperforming state-of-the-art methods on the majority of datasets. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2025 | A General Feature-Informed Crossover for Two-Stage Feature Selection in Symbolic RegressionabstractGenetic programming-based symbolic regression is a widely used machine learning technique, but its effectiveness can be limited as the number of input features increases. In genetic programming, two-stage feature selection has been extensively applied to enhance performance when dealing with a large number of input features. Existing two-stage feature selection methods typically require reinitializing new GP trees based on the selected features after feature selection, which disrupts the building blocks accumulated during evolution. In this paper, we propose a crossover operator that is aware of the selected features to leverage the feature selection results, thereby bypassing the need for reinitialization. This operator guides the crossover process to prioritize selected features, gradually eliminating unimportant features while preserving evolved building blocks. Experimental results validate the proposed method across three different feature-selection mechanisms on 98 datasets, demonstrating its effectiveness and broad applicability across various feature-selection strategies. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
CEC | 2 |
| 2025 | Micro-step Time-Series Regression: Insights from System Identification Using Symbolic Regression
Hengzhe Zhang, Alberto Paolo Tonda, Qi Chen 0002, Bing Xue 0001, Evelyne Lutton, Mengjie Zhang 0001 |
EuroGP | 3 |
| 2025 | Analysis of Illicit Drug Mixtures at Festivals Using Portable Near-Infrared Spectroscopy with Genetic Programming
Steven Dockter, Deepak Karunakaran, Qi Chen 0002, Yongshi Deng |
EvoApplications (2) | 3 |
| 2025 | Facial Geometric Feature Extraction for Dimensional Emotion Analysis Using Genetic Programming
Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
EvoApplications (2) | 2 |
| 2025 | RAG-SR: Retrieval-Augmented Generation for Neural Symbolic RegressionabstractSymbolic regression is a key task in machine learning, aiming to discover mathematical expressions that best describe a dataset. While deep learning has increased interest in using neural networks for symbolic regression, many existing approaches rely on pre-trained models. These models require significant computational resources and struggle with regression tasks involving unseen functions and variables. A pre-training-free paradigm is needed to better integrate with search-based symbolic regression algorithms. To address these limitations, we propose a novel framework for symbolic regression that integrates evolutionary feature construction with a neural network, without the need for pre-training. Our approach adaptively generates symbolic trees that align with the desired semantics in real-time using a language model trained via online supervised learning, providing effective building blocks for feature construction. To mitigate hallucinations from the language model, we design a retrieval-augmented generation mechanism that explicitly leverages searched symbolic expressions. Additionally, we introduce a scale-invariant data augmentation technique that further improves the robustness and generalization of the model. Experimental results demonstrate that our framework achieves state-of-the-art accuracy across 25 regression algorithms and 120 regression tasks. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
ICLR | 2 |
| 2025 | Improving Generalization of Genetic Programming for High-Dimensional Symbolic Regression with Shapley Value Based Feature SelectionabstractAbstract Symbolic Regression (SR) on high-dimensional datasets often encounters significant challenges, resulting in models with poor generalization capabilities. While feature selection has the potential to enhance the generalization and learning performance in general, its application in Genetic Programming (GP) for high-dimensional SR remains a complex problem. Originating from game theory, the Shapley value is applied to additive feature attribution approaches where it distributes the difference between a model output and a baseline average across input variables. By providing an accurate assessment of each feature importance, the Shapley value offers a robust approach to select features. In this paper, we propose a novel feature selection method leveraging the Shapley value to identify and select important features in GP for high-dimensional SR. Through a series of experiments conducted on ten high-dimensional regression datasets, the results indicate that our algorithm surpasses standard GP and other GP-based feature selection methods in terms of learning and generalization performance on most datasets. Further analysis reveals that our algorithm generates more compact models, focusing on the inclusion of important features. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
Data Sci. Eng. | 2 |
| 2025 | Semantics-guided multi-task genetic programming for multi-output regressionabstractMulti-output regression entails the simultaneous prediction of two or more output variables, presenting greater complexities than single-output regression due to the frequent interdependent relationships of these variables. Such dependencies mean that accurately predicting one variable typically requires careful analysis of its relationships with others. In this paper, multi-output regression problems are treated as multi-task problems, with a prediction of one output variable as a distinct task. A new multi-task multi-population genetic programming method is proposed to solve the problem. The method incorporates a semantics based crossover operator to identify the most informative subtree from a similar task that facilitates positive knowledge transfer. Empirical results indicate that our method significantly improves the training and testing performances of other multi-task GP methods, surpassing standard GP and GP with regressor chain on most examined regression datasets. Further analysis reveals that our proposed method can generate high-quality solutions by knowledge transfer and efficiently evolves similar GP models for analogous output variables, significantly enhancing positive knowledge transfer. • A semantics based crossover operator identifies the most informative subtree to transfer knowledge. • An origin based reservation strategy maintains structures to ensure a high-quality population. • Experiments show the proposed method improves learning efficiency and generalization of multi-task GP. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
Pattern Recognit. | 2 |
| 2024 | Feature Selection for GPSR Based on Maximal Information Coefficient and Shapley ValuesabstractFeature selection is a critical aspect of improving the interpretability of machine learning models. Genetic Programming (GP) has a built-in feature selection mechanism that explores the search space to include informative features in models. However, this built-in mechanism is insufficient for identifying important features, when dealing with high-dimensional feature spaces. To overcome this limitation, the paper introduces a novel feature importance measurement based on the Maximal Infor-mation Coefficient and Shapley Values. The proposed algorithm operates in two stages. In the first stage, it identifies the best individuals from different populations. In the second stage, the best individuals from the first stage are utilized for the calculation of the novel individual feature importance measurement. The new feature importance measurement offers valuable insights into the significance and relevance of the selected features. Regression experiments were conducted on six datasets to assess the effectiveness of the proposed method. Furthermore, comparisons were made with two other algorithms to evaluate its performance. The results indicate that the proposed approach enhances GP performance for high dimensional datasets while maintaining GP trees of similar size compared to standard GP. Mohamad Rimas, Mohamad Anfar, Qi Chen 0002, Mengjie Zhang 0001 |
CEC | 3 |
| 2024 | Genetic Programming with Multi-Task Feature Selection for Alzheimer's Disease DiagnosisabstractAlzheimer's disease (AD) has been the most common cause of dementia making cognitive score prediction and important feature identification crucial for its diagnosis. Although sparse linear regression has been used for this purpose due to its simplicity, it often selects an excessive number of features to track the disease and assumes a linear relationship between input and output, which might not always hold. To address these limitations, genetic programming-based symbolic regression (GPSR) algorithms have been proposed. GPSR can select the important features by exploring the feature space and learning a regression model without any assumption of model structure. However, the generalization ability of existing GPSR methods still needs to be improved. Considering the multiple related prediction tasks in AD studies, this work proposes a new method called linear scaled GPSR with multi-task feature selection (LSGPMTFS), to promote the prediction performance of each task by knowledge-sharing among multiple tasks. LSGPMTFS has two stages. The first stage learns a specific feature subset for each task. In the second stage, the model for each task is searched on the union of feature subsets selected from the first stage. The experimental results on authentic AD datasets demonstrate that the proposed algorithm can select a small set of important features with better learning and generalization performance compared with other GPSR methods. Shanshan Tang, Qi Chen 0002, Bing Xue 0001, Min Huang 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2024 | A New Concordance Correlation Coefficient based Fitness Function for Genetic Programming for Symbolic RegressionabstractCoefficients learning has long been challenging in genetic programming based symbolic regression (GPSR). Recent GPSR methods employ Pearson correlation coefficient for fitness assessment with post-hoc linear scaling for coefficient learning. However, this approach often leads to sub-optimal coefficient learning and inadequate consideration of nonlinear relationships between input variables and outputs. To solve those issues, this study introduces an innovative approach to integrating the Concordance Correlation Coefficient (CCC) into GPSR. Unlike Pearson correlation, CCC can effectively assess both linear and non-linear agreements between two sets of variables. Experimental results on eight regression datasets highlight the potential of CCC as a promising fitness function for GPSR without the need of a more advanced coefficient optimisation method for linear scaling. Jizhong Xu, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2024 | Improving Generalization of Evolutionary Feature Construction with Minimal Complexity Knee Points in Regression
Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
EuroGP | 2 |
| 2024 | Bias-Variance Decomposition: An Effective Tool to Improve Generalization of Genetic Programming-based Evolutionary Feature Construction for RegressionabstractEvolutionary feature construction is a technique that has been widely studied in the domain of automated machine learning. A key challenge that needs to be addressed in feature construction is its tendency to overfit the training data. Instead of the traditional approach to control overfitting by reducing model complexity, this paper proposes to control overfitting based on bias-variance decomposition. Specifically, this paper proposes reducing the variance of a model, i.e., reducing the variance of predictions when exposed to data with injected noise, to improve its generalization performance within a multi-objective optimization framework. Experiments conducted on 42 datasets demonstrate that the proposed method effectively controls overfitting and outperforms six model complexity measures for overfitting control. Moreover, further analysis reveals that controlling overfitting adhering to bias-variance decomposition outperforms several plausible variants, highlighting the importance of controlling overfitting based on solid machine learning theory. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
GECCO | 2 |
| 2024 | P-Mixup: Improving Generalization Performance of Evolutionary Feature Construction with Pessimistic Vicinal Risk Minimization
Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
PPSN (1) | 2 |
| 2024 | Multitree Genetic Programming With Feature-Based Transfer Learning for Symbolic Regression on Incomplete DataabstractData incompleteness is a serious challenge in real-world machine-learning tasks. Nevertheless, it has not received enough attention in symbolic regression (SR). Data missingness exacerbates data shortage, especially in domains with limited available data, which in turn limits the learning ability of SR algorithms. Transfer learning (TL), which aims to transfer knowledge across tasks, is a potential solution to solve this issue by making amends for the lack of knowledge. However, this approach has not been adequately investigated in SR. To fill this gap, a multitree genetic programming-based TL method is proposed in this work to transfer knowledge from complete source domains (SDs) to incomplete related target domains (TDs). The proposed method transforms the features from a complete SD to an incomplete TD. However, having many features complicates the transformation process. To mitigate this problem, we integrate a feature selection mechanism to eliminate unnecessary transformations. The method is examined on real-world and synthetic SR tasks with missing values to consider different learning scenarios. The obtained results not only show the effectiveness of the proposed method but also show its training efficiency compared with the existing TL methods. Compared to state-of-the-art methods, the proposed method reduced an average of more than 2.58% and 4% regression error on heterogeneous and homogeneous domains, respectively. Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | Evolutionary Multitasking for Multiobjective Feature Selection in ClassificationabstractEvolutionary multi-objective optimisation has shown success in feature selection. However, existing methods often address these tasks independently, disregarding their potential interconnections and shared knowledge. On the other hand, evolutionary multitasking has been utilised to address multiple related tasks simultaneously and transfer common knowledge. However, most EMTL-based feature selection methods prioritize a single task, and treat it as the main task and the other tasks as auxiliary or secondary tasks. To overcome this limitation, we propose a novel multi-objective feature selection method based on EMT in this paper. The new method introduces a novel representation that consolidates the solutions of multiple interconnected feature selection tasks into a single solution. It enables these tasks to share a common population, thereby enhancing the effectiveness and efficiency of transferring common knowledge across them. In addition, a novel searching method is devised to facilitate the evolution of the population across multiple tasks, enabling effective knowledge transfer between them. Finally, a transformation method is introduced to transfer valuable genes among the solutions for multiple tasks, thereby enhancing the overall performance of the proposed algorithm when confronted with tasks characterised by distinct features. This method effectively addresses multiple feature selection tasks simultaneously, offering a comprehensive solution to the aforementioned issue. Compared with four single-task multi-objective feature selection methods and a state-of-the-art evolutionary multitasking-based feature selection method, the proposed method demonstrates superior feature selection performance across the majority of benchmark datasets. Jiabin Lin, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2024 | Modular Multitree Genetic Programming for Evolutionary Feature Construction for RegressionabstractEvolutionary feature construction is a key technique in evolutionary machine learning, with the aim of constructing high-level features that enhance performance of a learning algorithm. In real-world applications, engineers typically construct complex features based on a combination of basic features, re-using those features as modules. However, modularity in evolutionary feature construction is still an open research topic. This paper tries to fill that gap by proposing a modular and hierarchical multitree genetic programming (GP) algorithm that allows trees to use the output values of other trees, thereby representing expressive features in a compact form. Based on this new representation, we propose a macro parent-repair strategy to reduce redundant and irrelevant features, a macro crossover operator to preserve interactive features, and an adaptive control strategy for crossover and mutation rates to dynamically balance the trade-off between exploration and exploitation. A comparison with seven bloat control methods on 98 regression datasets shows that the proposed modular representation achieves significantly better results in terms of test performance and smaller model size. Experimental results on the state-of-the-art symbolic regression benchmark demonstrate that the proposed symbolic regression method outperforms 22 existing symbolic regression and machine learning algorithms, providing empirical evidence for the superiority of the modularized evolutionary feature construction method. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2024 | A Semantic-Based Hoist Mutation Operator for Evolutionary Feature Construction in RegressionabstractIn recent years, genetic programming has achieved impressive results on evolutionary feature construction tasks. To increase search effectiveness, researchers have developed many semantic-based crossover and mutation operators to guide genetic programming searches toward the target semantics. However, semantics has not yet been explored for the hoist mutation operator, which is an operator designed for controlling the bloat effect. Although the hoist mutation operator can significantly reduce model sizes, the most informative subtree may be disrupted by the randomness in mutation. To address this issue, we develop a semantic-based hoist mutation operator in this paper to preserve the most informative subtree that has the largest cosine similarity between its semantics and the target semantics. Experimental results on 98 regression datasets from the Penn Machine Learning Benchmark show that using this operator not only significantly reduces model size, but also improves the test accuracy of features constructed by genetic programming. A comparison with seven bloat control methods shows that the proposed operator achieves the best trade-off between accuracy and model size. Moreover, an experiment on the state-of-the-art symbolic regression benchmark shows that genetic programming with the semantic-based hoist mutation operator achieves the best test accuracy and competitive model sizes compared with 22 symbolic regression and machine learning algorithms. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2024 | SR-Forest: A Genetic Programming-Based Heterogeneous Ensemble Learning MethodabstractEnsemble learning methods have been widely used in machine learning in recent years due to their high predictive performance. With the development of genetic programming-based symbolic regression methods, many papers begin to choose a popular ensemble learning method, random forests, as the baseline competitor. Instead of considering them as competitors, an alternative idea might be to consider symbolic regression as an enhancement technique for random forest. Genetic programming-based symbolic regression methods which fit a smooth function are complementary to the piecewise nature of decision trees, as the smooth variation is common in regression problems. In this article, we propose to form an ensemble model with symbolic regression-based decision trees to address this issue. Furthermore, we design a guided mutation operator to speed up the search on high-dimensional problems, a multi-fidelity evaluation strategy to reduce the computational cost and an ensemble selection mechanism to improve predictive performance. Finally, experimental results on a regression benchmark with 120 datasets show that the proposed ensemble model outperforms 25 existing symbolic regression and ensemble learning methods. Moreover, the proposed method can provide notable insights on an XGBoost hyperparameter performance prediction task, which is an important application area of ensemble learning methods. Hengzhe Zhang, Aimin Zhou, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2023 | MAP-Elites with Cosine-Similarity for Evolutionary Ensemble Learning
Hengzhe Zhang, Qi Chen 0002, Alberto Paolo Tonda, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
EuroGP | 2 |
| 2023 | AMTEA-Based Multi-task Optimisation for Multi-objective Feature Selection in Classification
Jiabin Lin, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
EvoApplications@EvoStar | 2 |
| 2023 | Relieving Genetic Programming from Coefficient Learning for Symbolic Regression via Correlation and Linear ScalingabstractThe difficulty of learning optimal coefficients in regression models using only genetic operators has long been a challenge in genetic programming for symbolic regression. As a simple but effective remedy it has been proposed to perform linear scaling of model outputs prior to a fitness evaluation. Recently, the use of a correlation coefficient-based fitness function with a post-processing linear scaling step for model alignment has been shown to outperform error-based fitness functions in generating symbolic regression models. In this study, we compare the impact of four evaluation strategies on relieving genetic programming (GP) from learning coefficients in symbolic regression and focusing on learning the more crucial model structure. The results from 12 datasets, including ten real-world tasks and two synthetic datasets, confirm that all these strategies assist GP to varying degrees in learning coefficients. Among the them, correlation fitness with one-time linear scaling as post-processing, due to be the most efficient while bringing notable benefits to the performance, is the recommended strategy to relieve GP from learning coefficients. Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
GECCO | 1 |
| 2023 | Fast and Efficient Local-Search for Genetic Programming Based Loss Function LearningabstractIn this paper, we develop upon the topic of loss function learning, an emergent meta-learning paradigm that aims to learn loss functions that significantly improve the performance of the models trained under them. Specifically, we propose a new meta-learning framework for task and model-agnostic loss function learning via a hybrid search approach. The framework first uses genetic programming to find a set of symbolic loss functions. Second, the set of learned loss functions is subsequently parameterized and optimized via unrolled differentiation. The versatility and performance of the proposed framework are empirically validated on a diverse set of supervised learning tasks. Results show that the learned loss functions bring improved convergence, sample efficiency, and inference performance on tabulated, computer vision, and natural language processing problems, using a variety of task-specific neural network architectures. Christian Raymond, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
GECCO | 2 |
| 2023 | A Double Lexicase Selection Operator for Bloat Control in Evolutionary Feature Construction for RegressionabstractEvolutionary feature construction is an important technique in the machine learning domain for enhancing learning performance. However, traditional genetic programming-based feature construction methods often suffer from bloat, which means the sizes of constructed features increase excessively without improved performance. To address this issue, this paper proposes a double-stage lexicase selection operator to control bloat while not damaging search effectiveness. This new operator contains a two-stage selection process, where the first stage selects individuals based on fitness values and the second stage selects individuals based on tree sizes. Therefore, the proposed operator can control bloat meanwhile leveraging the advantage of the lexicase selection operator. Experimental results on 98 regression datasets show that compared to the traditional bloat control method of having a depth limit, the proposed selection operator not only significantly reduces the sizes of constructed features on all datasets but also keeps a similar level of predictive performance. A comparative experiment with seven bloat control methods shows that the double lexicase selection operator achieves the best trade-off between the model performance and the model size. Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
GECCO | 2 |
| 2023 | Automatically Choosing Selection Operator Based on Semantic Information in Evolutionary Feature Construction
Hengzhe Zhang, Qi Chen 0002, Bing Xue 0001, Wolfgang Banzhaf, Mengjie Zhang 0001 |
PRICAI (2) | 2 |
| 2023 | Learning Symbolic Model-Agnostic Loss Functions via Meta-LearningabstractIn this paper, we develop upon the emerging topic of loss function learning, which aims to learn loss functions that significantly improve the performance of the models trained under them. Specifically, we propose a new meta-learning framework for learning model-agnostic loss functions via a hybrid neuro-symbolic search approach. The framework first uses evolution-based methods to search the space of primitive mathematical operations to find a set of symbolic loss functions. Second, the set of learned loss functions are subsequently parameterized and optimized via an end-to-end gradient-based training procedure. The versatility of the proposed framework is empirically validated on a diverse set of supervised learning tasks. Results show that the meta-learned loss functions discovered by the newly proposed method outperform both the cross-entropy loss and state-of-the-art loss function learning methods on a diverse range of neural network architectures and datasets. We make our code available at *retracted*. Christian Raymond, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Explainable Artificial Intelligence by Genetic Programming: A SurveyabstractExplainable artificial intelligence (XAI) has received great interest in the recent decade, due to its importance in critical application domains, such as self-driving cars, law, and healthcare. Genetic programming (GP) is a powerful evolutionary algorithm for machine learning. Compared with other standard machine learning models such as neural networks, the models evolved by GP tend to be more interpretable due to their model structure with symbolic components. However, interpretability has not been explicitly considered in GP until recently, following the surge in the popularity of XAI. This article provides a comprehensive review of the studies on GP that can potentially improve the model interpretability, both explicitly and implicitly, as a byproduct. We group the existing studies related to explainable artificial intelligence by GP into two categories. The first category considers the intrinsic interpretability, aiming to directly evolve more interpretable (and effective) models by GP. The second category focuses on post-hoc interpretability, which uses GP to explain other black-box machine learning models, or explain the models evolved by GP by simpler models such as linear models. This comprehensive survey demonstrates the strong potential of GP for improving the interpretability of machine learning models and balancing the complex tradeoff between model accuracy and interpretability. Yi Mei 0001, Qi Chen 0002, Andrew Lensen, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2022 | Multi-objective Genetic Programming with the Adaptive Weighted Splines Representation for Symbolic Regression
Christian Raymond, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
EuroGP | 2 |
| 2022 | A New Genetic Algorithm for Automated Spectral Pre-processing in Nutrient Assessment
Demelza Robinson, Qi Chen 0002, Bing Xue 0001, Daniel Killeen, Keith C. Gordon, Mengjie Zhang 0001 |
EvoApplications | 2 |
| 2022 | Genetic Programming for Instance Transfer Learning in Symbolic RegressionabstractTransfer learning has attracted more attention in the machine-learning community recently. It aims to improve the learning performance on the domain of interest with the help of the knowledge acquired from a similar domain(s). However, there is only a limited number of research on tackling transfer learning in genetic programming for symbolic regression. This article attempts to fill this gap by proposing a new instance weighting framework for transfer learning in genetic programming-based symbolic regression. In the new framework, differential evolution is employed to search for optimal weights for source-domain instances, which helps genetic programming to identify more useful source-domain instances and learn from them. Meanwhile, a density estimation method is used to provide good starting points to help the search for the optimal weights while discarding some irrelevant or less important source-domain instances before learning regression models. The experimental results show that compared with genetic programming and support vector regression that learn only from the target instances, and learning from a mixture of instances from the source and target domains without any transfer learning component, the proposed method can evolve regression models which not only achieve notably better cross-domain generalization performance in stability but also reduce the trend of overfitting effectively. Meanwhile, these models are generally much simpler than those generated by the other GP methods. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Rademacher Complexity for Enhancing the Generalization of Genetic Programming for Symbolic RegressionabstractModel complexity has a close relationship with the generalization ability and the interpretability of the learned models. Simple models are more likely to generalize well and easy to interpret. However, too much emphasis on minimizing complexity can prevent the discovery of more complex yet more accurate solutions. Genetic programming (GP) has a trend of generating overcomplex models that are difficult to interpret while not being able to generalize well. This work proposes a novel complexity measure based on the Rademacher complexity for GP for symbolic regression. The complexity of an evolved model is measured by the maximum correlation between the model and the Rademacher variables on the selected training instances. Taking minimizing the training error and the Rademacher complexity of the models as the two objectives, the proposed GP method has shown to be much superior to the standard GP on generalization performance. Compared with GP equipped with two state-of-the-art complexity measures, the proposed method still has a notable advance on generating a better front consisting of individuals with lower generalization errors and being simpler in the behavioral complexity. Further analyses reveal that compared with the state-of-the-art methods, the proposed GP method evolves models that are much closer to the target models in the model structure, and have better interpretability. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | GP with a Hybrid Tree-vector Representation for Instance Selection and Symbolic Regression on Incomplete DataabstractData incompleteness is a pervasive problem in symbolic regression, and machine learning in general. Unfortunately, most symbolic regression methods are only applicable when the given data is complete. One common approach to handling this situation is data imputation. It works by estimating missing values based on existing data. However, which existing data should be used for imputing the missing values? The answer to this question is important when dealing with incomplete data. To address this question, this work proposes a mixed tree-vector representation for genetic programming to perform instance selection and symbolic regression on incomplete data. In this representation, each individual has two components: an expression tree and a bit vector. While the tree component constructs symbolic regression models, the vector component selects the instances that are used to impute missing values by the weighted k-nearest neighbour (WKNN) imputation method. The complete imputed instances are then used to evaluate the GP-based symbolic regression model. The obtained experimental results show the applicability of the proposed method on real-world data sets with different missingness scenarios. When compared with existing methods, the proposed method not only produces more effective symbolic regression models but also achieves more efficient imputations. Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2021 | Genetic Algorithm for Feature and Latent Variable Selection for Nutrient Assessment in Horticultural ProductsabstractVibrational spectroscopy can be used for rapid determination of chemical quality markers in horticultural produce to improve quality control, optimize harvest times and maximize profits. Most commonly, spectral data are calibrated against chemical reference data (acquired using traditional, slower analytical methods) using partial least squares regression (PLSR). However, predictive performance of PLSR can be limited by the small number of instances, high dimensionality and collinearity of spectroscopic data. Here, a new genetic algorithm (GA) for PLSR feature and latent variable selection is proposed to predict concentrations of 18 important bioactive components across three New Zealand horticultural products from infrared, near-infrared and Raman spectral data sets. Models generated using the GA-enhanced PLSR method have notably better generalization and are less complex than the standard PLSR method. GA-enhanced PLSR models are produced from each spectroscopic data set individually, and from a data set that combines all three techniques. Demelza Robinson, Qi Chen 0002, Bing Xue 0001, Daniel Killeen, Sara Fraser-Miller, Keith C. Gordon, Indrawati Oey, Mengjie Zhang 0001 |
CEC | 2 |
| 2021 | Particle Swarm Optimisation for Analysing Time-Dependent Photoluminescence DataabstractNext-generation photovoltaic materials such as per-ovskites and organic photovoltaics are promising candidates for cheap, solution-processable solar cells, which have the environmental and financial advantages compared to traditional silicon-based cells. To realise commercial solar cells, the development of new materials to improve the performance is needed. Time-resolved optical spectroscopy is a powerful technique for photovoltaic materials that allows measurement of reveal the energy-dependent dynamics of photoexcited species (charges and excitons), which are responsible for the device performance. Although time-resolved spectroscopy provides contains rich information, the data analysis can be time-consuming and labour-intensive. Automated data-processing is therefore an attractive proposition to facilitate higher throughput. This paper describes a new application of evolutionary computation technique - particle swarm optimisation (PSO) - to parametrise time-resolved photoluminescence (PL) data. PSO is used to convert time- and energy-resolved photoluminescence data into decay rate distributions. From this, the excited state lifetimes can be elucidated - a key parameter for the optimisation of photovoltaic performance. The implementation of PSO in enhanced LumiML proved advantageous, yielding considerable improvements over previous techniques LumiML by two orders of magnitude. Demelza Robinson, Qi Chen 0002, Bing Xue 0001, Isabella Wagner, Paul Hume, Justin Hodgkiss, Mengjie Zhang 0001 |
CEC | 2 |
| 2021 | A new imputation method based on genetic programming and weighted KNN for symbolic regression with incomplete data
Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
Soft Comput. | 2 |
| 2021 | Multitree Genetic Programming With New Operators for Transfer Learning in Symbolic Regression With Incomplete DataabstractLack of knowledge is a common consequence of data incompleteness when learning from real-world data. To deal with such a situation, this work utilizes transfer learning (TL) to reuse knowledge from different (yet related) but complete domains. Due to its powerful feature construction ability, genetic programming (GP) is used to construct feature-based transformations that map the feature space of the source domain to that of the target domain such that their differences are reduced. Particularly, this work proposes a new multitree GP-based feature construction approach to TL in symbolic regression with missing values. It transfers knowledge related to the importance of the features and instances in the source domain to the target domain to improve the learning performance. Moreover, new genetic operators are developed to encourage minimizing the distribution discrepancy between the transformed domain and the target domain. A new probabilistic crossover is developed to make the well-constructed trees in the individuals more likely to be mated than the other trees. A new mutation operator is designed to give more probability for the poorly constructed trees to be mutated. The experimental results show that the proposed method not only achieves better performance compared with different traditional learning methods but also advances two recent TL methods on real-world data sets with various incompleteness and learning scenarios. Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2021 | Preserving Population Diversity Based on Transformed Semantics in Genetic Programming for Symbolic RegressionabstractPopulation diversity plays an important role in avoiding premature convergence in evolutionary techniques including genetic programming (GP). Obtaining an adequate level of diversity during the evolutionary process has became a concern of many previous researches in GP. This work proposes a new novelty metric for entropy-based diversity measure for GP. The new novelty metric is based on the transformed semantics of models in GP, where the semantics are the set of outputs of a model on the training data and principal component analysis is used for a transformation of the semantics. Based on the new novelty metric, a new diversity preserving framework, which incorporates a new fitness function and a new selection operator, is proposed to help GP achieve a good balance between the exploration and the exploitation, thus enhancing its learning and generalization performance. Compared with two stat-of-the-art diversity preserving methods, the new method can generalize better and reduce the overfitting trend more effectively in most cases. Further examinations on the properties of the search process confirm that the new framework notably enhances the evolvability and locality of GP. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2020 | Genetic Programming with Noise Sensitivity for Imputation Predictor Selection in Symbolic Regression with Incomplete DataabstractThis paper presents a feature selection method that incorporates a sensitivity-based single feature importance measure in a context-based feature selection approach. The single-wise importance is based on the sensitivity of the learning performance with respect to adding noise to the predictive features. Genetic programming is used as a context-based selection mechanism, where the selection of features is determined by the change in the performance of the evolved genetic programming models when the feature is injected with noise. Imputation is a key strategy to mitigate the data incompleteness problem. However, it has been rarely investigated for symbolic regression on incomplete data. In this work, an attempt to contribute to filling this gap is presented. The proposed method is applied to selecting imputation predictors (features/variables) in symbolic regression with missing values. The evaluation is performed on real-world data sets considering three performance measures: imputation accuracy, symbolic regression performance, and features' reduction ability. Compared with the benchmark methods, the experimental evaluation shows that the proposed method can achieve an enhanced imputation, improve the symbolic regression performance, and use smaller sets of selected predictors. Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2020 | Multi-Tree Genetic Programming-based Transformation for Transfer Learning in Symbolic Regression with Highly Incomplete DataabstractTransfer learning has been considered a key solution for the problem of learning when there is a lack of knowledge in some target domains. Its idea is to benefit from the learning on different (but related in some way) domains that have adequate knowledge and transfer what can improve the learning in the target domains. Although incompleteness is one of the main causes of knowledge shortage in many machine learning real-world tasks, it has received a little effort to be addressed by transfer learning. In particular, to the best of our knowledge, there is no single study to utilize transfer learning for the symbolic regression task when the underlying data are incomplete. The current work addresses this point by presenting a transfer learning method for symbolic regression on data with high ratios of missing values. A multi-tree genetic programming algorithm based feature-based transformation is proposed for transferring data from a complete source domain to a different, incomplete target domain. The experimental work has been conducted on real-world data sets considering different transfer learning scenarios each is determined based on three factors: missingness ratio, domain difference, and task similarity. In most cases, the proposed method achieved positive transductive transfer learning in both homogeneous and heterogeneous domains. Moreover, even with less significant success, the obtained results show the applicability of the proposed approach for inductive transfer learning. Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2020 | Hessian Complexity Measure for Genetic Programming-Based Imputation Predictor Selection in Symbolic Regression with Incomplete Data
Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
EuroGP | 2 |
| 2020 | Improving symbolic regression based on correlation between residuals and variablesabstractIn traditional regression analysis, a detailed examination of the residuals can provide an important way of validating the model quality. However, it has not been utilised in genetic programming based symbolic regression. This work aims to fill this gap and propose a new evaluation criterion of minimising the correlation between the residuals of regression models and the independent variables. Based on a recent association detection measure, maximal information coefficient which provides an accurate estimation of the correlation, the new evaluation measure is expected to enhance the generalisation of genetic programming by driving the evolutionary process towards models that are without unnecessary complexity and less likely learning from noise in data. The experiment results show that, compared with standard genetic programming which selects model based on the training error only and two state-of-the-art multiobjective genetic programming methods with mechanisms to prefer models with adequate structures, our new multiobjective genetic programming method minimising both the correlation between residuals and variable, and the training error has a consistently better generalisation performance and evolves simpler models. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
GECCO | 1 |
| 2020 | Multi-tree genetic programming for feature construction-based domain adaptation in symbolic regression with incomplete dataabstractNowadays, transfer learning has gained a rapid popularity in tasks with limited data available. While traditional learning limits the learning process to knowledge available in a specific (target) domain, transfer learning can use parts of knowledge extracted from learning in a different (source) domain to help learning in the target domain. This concept is of special importance when there is a lack of knowledge in the target domain. Consequently, since data incompleteness is a serious cause of knowledge shortage in real-world learning tasks, it can be typically addressed using transfer learning. One way to achieve that is feature construction-based domain adaptation. However, although it is considered as a powerful feature construction algorithm, Genetic Programming has not been fully utilized for domain adaptation. In this work, a multi-tree genetic programming method is proposed for feature construction-based domain adaptation. The main idea is to construct a transformation from the source feature space to the target feature space, which maps the source domain close to the target domain. This method is utilized for symbolic regression with missing values. The experimental work shows encouraging potential of the proposed approach when applied to real-world tasks considering different transfer learning scenarios. Baligh Al-Helali, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
GECCO | 2 |
| 2020 | Adaptive weighted splines: a new representation to genetic programming for symbolic regressionabstractGenetic Programming for Symbolic Regression is often prone to overfit the training data, resulting in poor generalization on unseen data. To address this issue, many pieces of research have been devoted to regularization via controlling the model complexity. However, due to the unstructured tree based representation of individuals the model complexity cannot be directly computed, rather approximation of the complexity must be taken. This paper proposes a new novel representation called Adaptive Weighted Splines which enables explicit control over the complexity of individuals using splines. The experimental results confirm that this new representation is significantly better than the tree-based representation at avoiding overfitting and generalizing on unseen data, demonstrating notably better and far more consistent generalization performances on all the benchmark problems. Further analysis also shows that in most cases, the new Genetic Programming method outperforms classical regression techniques such as Linear Regression, Support Vector Regression, K-Nearest Neighbour and Decision Tree Regression and performs competitively with state-of-the-art ensemble regression methods Random Forests and Gradient Boosting. Christian Raymond, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
GECCO | 2 |
| 2019 | Instance based Transfer Learning for Genetic Programming for Symbolic RegressionabstractTransfer learning aims to utilise knowledge acquired from the source domain to improve the learning performance in the target domain. It attracts increasing interests and many transfer learning approaches have been proposed. However, studies on transfer learning for genetic programming for symbolic regression are still rare, although clearly desired, due to the difficulty to evolve models with a good cross-domain generalisation ability. This work proposes a new instance weighting framework for transfer learning in genetic programming for symbolic regression. The key idea is to utilise a local weight updating scheme to identify and learn from more useful source domain instances and reduce the effort on the source domain instances, which are more different from the target domain data. The experimental results show that the proposed method notably enhances the learning capacity and the generalisation performance of genetic programming on the target domain and also outperforms some state-of-the-art regression methods. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 1 |
| 2019 | Genetic Programming with Rademacher Complexity for Symbolic RegressionabstractGenetic Programming (GP) for symbolic regression is often prone to overfitting the training data, causing poor performance on unseen data. A number of recent works in the field have been devoted to regulating this problem by investigating both the structural and functional complexity of GP individuals during the evolutionary process. This work uses the Rademacher complexity and incorporates it into the fitness function of GP, utilising it as a means of controlling the functional complexity of GP individuals. The experiment results confirm that the new GP method has a notable generalization gain compared to the standard GP and Support Vector Regression (SVR) in most of the considered problems. Further investigations also show that the new GP method generates symbolic regression models that could not only release the overfitting trend in standard GP but also are significantly smaller in size compared to their counterparts in standard GP. Christian Raymond, Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 2 |
| 2019 | Improving Generalization of Genetic Programming for Symbolic Regression With Angle-Driven Geometric Semantic OperatorsabstractGeometric semantic genetic programming (GP) has recently attracted much attention. The key innovations are inducing a unimodal fitness landscape in the semantic space and providing a theoretical framework for designing geometric semantic operators. The geometric semantic operators aim to manipulate the semantics of programs by making a bounded semantic impact and generating child programs with similar or better behavior than their parents. These properties are shown to be highly related to a notable generalization improvement in GP. However, the potential ineffectiveness and difficulties in bounding the variations in these geometric operators still limits their positive effect on generalization. This paper attempts to further explore the geometry and search space of geometric operators to gain a greater generalization improvement in GP for symbolic regression. To this end, a new angle-driven selection operator and two new angle-driven geometric search operators are proposed. The angle-awareness brings new geometric properties to these geometric operators, which are expected to provide a greater leverage for approximating the target semantics in each operation, and more importantly, be resistant to overfitting. The experiments show that compared with two state-of-the-art geometric semantic operators, our angle-driven geometric operators not only drive the evolutionary process to fit the target semantics more efficiently but also improve the generalization performance. A further comparison between the evolved models shows that the new method generally produces simpler models with a much smaller size and is more likely to evolve toward the correct structure of the target models. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2019 | Structural Risk Minimization-Driven Genetic Programming for Enhancing Generalization in Symbolic RegressionabstractGeneralization ability, which reflects the prediction ability of a learned model, is an important property in genetic programming (GP) for symbolic regression. Structural risk minimization (SRM) is a framework providing a reliable estimation of the generalization performance of prediction models. Introducing the framework into GP has the potential to drive the evolutionary process toward models with good generalization performance. However, this is tough due to the difficulty in obtaining the Vapnik-Chervonenkis (VC) dimension of nonlinear models. To address this difficulty, this paper proposes an SRM-driven GP approach, which uses an experimental method (instead of theoretical estimation) to measure the VC dimension of a mixture of linear and nonlinear regression models for the first time. The experimental method has been conducted using uniform and nonuniform settings. The results show that our method has impressive generalization gains over standard GP and GP with the 0.632 bootstrap, and that the proposed method using the nonuniform setting has further improvement than its counterpart using the uniform setting. Further analyzes reveal that the proposed method can evolve more compact models, and that the behavioral difference between these compact models and the target models is much smaller than their counterparts evolved by the other GP methods. Qi Chen 0002, Mengjie Zhang 0001, Bing Xue 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2017 | Geometric Semantic Crossover with an Angle-Aware Mating Scheme in Genetic Programming for Symbolic Regression
Qi Chen 0002, Bing Xue 0001, Yi Mei 0001, Mengjie Zhang 0001 |
EuroGP | 1 |
| 2017 | Feature Selection to Improve Generalization of Genetic Programming for High-Dimensional Symbolic RegressionabstractWhen learning from high-dimensional data for symbolic regression (SR), genetic programming (GP) typically could not generalize well. Feature selection, as a data preprocessing method, can potentially contribute not only to improving the efficiency of learning algorithms but also to enhancing the generalization ability. However, in GP for high-dimensional SR, feature selection before learning is seldom considered. In this paper, we propose a new feature selection method based on permutation to select features for high-dimensional SR using GP. A set of experiments has been conducted to investigate the performance of the proposed method on the generalization of GP for high-dimensional SR. The regression results confirm the superior performance of the proposed method over the other examined feature selection methods. Further analysis indicates that the models evolved by the proposed method are more likely to contain only the truly relevant features and have better interpretability. Qi Chen 0002, Mengjie Zhang 0001, Bing Xue 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2016 | Improving generalisation of genetic programming for high-dimensional symbolic regression with feature selectionabstractFeature selection is a desired process when learning from high-dimensional data. However, it is seldom considered in Genetic Programming (GP) for high-dimensional symbolic regression. This work aims to develop a new method, Genetic Programming with Feature Selection (GPWFS), to improve the generalisation ability of GP for symbolic regression. GPWFS is a two-stage method. The main task of the first stage is to select important/informative features from fittest individuals, and the second stage uses a set of selected features, which is a subset of original features, for regression. To investigate the learning/optimisation performance and generalisation capability of GPWFS, a set of experiments using standard GP as a baseline for comparison have been conducted on six real-world high-dimensional symbolic regression datasets. The experimental results show that GPWFS can have better performance both on the training sets and the test sets on most cases. Further analysis on the solution size, the number of distinguished features and total number of used features in the evolved models shows that using GPWFS can induce more compact models with better interpretability and lower computational costs than standard GP. Qi Chen 0002, Bing Xue 0001, Ben Niu 0002, Mengjie Zhang 0001 |
CEC | 1 |
| 2016 | Improving Generalisation of Genetic Programming for Symbolic Regression with Structural Risk MinimisationabstractGeneralisation is one of the most important performance measures for any learning algorithm, no exception to Genetic Programming (GP). A number of works have been devoted to improve the generalisation ability of GP for symbolic regression. Methods based on a reliable estimation of generalisation error of models during evolutionary process are a sensible choice to enhance the generalisation of GP. Structural risk minimisation (SRM), which is based on the VC dimension in the learning theory, provides a powerful framework for estimating the difference between the generalisation error and the empirical error. Despite its solid theoretical foundation and reliability, SRM has seldom been applied to GP. The most important reason is the difficulty in measuring the VC dimension of GP models/programs. This paper introduces SRM, which is based on an empirical method to measure the VC dimension of models, into GP to improve its generalisation performance for symbolic regression. The results of a set of experiments confirm that GP with SRM has a dramatical generalisation gain while evolving more compact/less complex models than standard GP. Further analysis also shows that in most cases, GP with SRM has better generalisation performance than GP with bias-variance decomposition, which is one of the state-of-the-art methods to control overfitting. Qi Chen 0002, Bing Xue 0001, Lin Shang 0001, Mengjie Zhang 0001 |
GECCO | 1 |
| 2016 | Proceedings in Adaptation, Learning and Optimization
Qi Chen 0002, Mengjie Zhang 0001, Bing Xue 0001 |
IES | 1 |
| 2015 | Generalisation and domain adaptation in GP with gradient descent for symbolic regressionabstractGenetic programming (GP) has been widely applied to symbolic regression problems and achieved good success. Gradient descent has also been used in GP as a complementary search to the genetic beam search to further improve symbolic regression performance. However, most existing GP approaches with gradient descent (GPGD) to symbolic regression have only been tested on the “conventional” symbolic regression problems such as benchmark function approximations and engineering practical problems with a single (training) data set only and the effectiveness on unseen data sets in the same domain and in different domains has not been fully investigated. This paper designs a series of experiment objectives to investigate the effectiveness and efficiency of GPGD with various settings for a set of symbolic regression problems applied to unseen data in the same domain and adapted to other domains. The results suggest that the existing GPGD method applying gradient descent to all evolved program trees three times at every generation can perform very well on the training set itself, but cannot generalise well on the unseen data set in the same domain and cannot be adapted to unseen data in an extended domain. Applying gradient descent to the best program in the final generation of GP can also improve the performance over the standard GP method and can generalise well on unseen data for some of the tasks in the same domain, but perform poorly on the unseen data in an extended domain. Applying gradient descent to the top 20% programs in the population can generalise reasonably well on the unseen data in not only the same domain but also in an extended domain. Qi Chen 0002, Bing Xue 0001, Mengjie Zhang 0001 |
CEC | 1 |