VLDB 2026 Research / reviewers in the wild / expert
Sara Silva
dblp:37/4200
· DBLP profile ↗
63ranked-venue papers
13as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 10 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Controlling Functional Complexity for Overfitting Reduction and Improved Interpretability in GPabstractLike other machine learning methods, Genetic Programming (GP) frequently faces the issue of overfitting when applied to supervised learning tasks. Traditional regularization techniques, though well-studied, are challenging to apply to GP due to the free-form nature of the evolved models. This work proposes a novel approach that prevents overfitting while inherently improving the interpretability of GP models. It involves a dual optimization process that minimizes loss while penalizing functional complexity using multi-objective selection mechanisms. The improved complexity measure used in this study approximates the mathematical curvature of a function in linear time. While loss minimization is common in GP, penalizing functional complexity is an additional step aimed at evolving robust and smooth functions, less prone to overfitting and potentially more interpretable. Experimental results demonstrate the effectiveness of the two variants of our method, benchmarked against standard GP and two of the seemingly best overfitting-reduction methods found in the literature. By focusing on both loss and complexity, our approach achieves state-of-the-art generalization on difficult problems and a strong feature selection that improves interpretability, making it a unified improvement of GP. Sara Silva, Inês Magessi, Leonardo Vanneschi |
IEEE Trans. Evol. Comput. | 1 |
| 2025 | Introducing Crossover in SLIM-GSGP
Gloria Pietropolli, Davide Farinati, Luca Manzoni, Mauro Castelli, Sara Silva, Leonardo Vanneschi |
EuroGP | 5 |
| 2025 | Slim_gsgp: A Python Library for Non-Bloating GSGPabstractThis paper presents slim_gsgp: an open-source Python library that provides the first ever framework for the Semantic Learning algorithm based on Inflate and deflate Mutation (SLIM-GSGP). Proposed in 2024, SLIM-GSGP is a promising non-bloating variant of Geometric Semantic Genetic Programming (GSGP). slim_gsgp includes all existing SLIM-GSGP variants, as well as traditional GSGP and standard Genetic Programming (GP), facilitating comparative analysis and benchmarking. Additionally, slim_gsgp's parallel computation and semi-modular architecture renders it not only fast but also user-friendly and easily extensible, thereby serving as a valuable resource for researchers aiming to advance this emerging and promising area of research. The source code and documentation can be accessed at https://github.com/DALabNOVA/slim. Liah Rosenfeld, Davide Farinati, Diogo Rasteiro, Gloria Pietropolli, Karina Brotto Rebuli, Sara Silva, Leonardo Vanneschi |
GECCO | 6 |
| 2025 | Jamming Attack on DSRC Communication Caused by a C-V2X Sidelink DeviceabstractAcademia and industry are constantly striving to improve the driving experience, where vehicles can now communicate, and different technologies can coexist and share the wireless medium to send and receive information. In terms of vehicle-to-vehicle (V2V) communication, vehicles can use Dedicated Short-Range Communication (DSRC) or Cellular Vehicle-to-Everything (C-V2X) sidelink technologies to communicate on the roads by sharing the 5.9 GHz band. However, coexistence between the two technologies remains challenging, as interference can occur when both technologies use co-channels to send vehicle messages in parallel, degrading the Packet Delivery Ratio (PDR) metric. In addition, malicious users can carry out various types of attacks to disrupt the vehicle connection, such as jamming attacks. In this paper, we transmitted packets using a constant bit rate (CBR). We considered Line-of-Sight (LOS) and Non-Line-of-Sight (NLOS) scenarios (i.e., in the DSRC communication) to understand how damaging the jamming attack can be. Based on our tests, the jamming attack caused by the C-V2X co-channel interference can cause an attenuation in the signal loss of up to 24 dB (i.e., using a packet size of 256 bytes) and up to 35 dB (i.e., using a packet size of 1,399 bytes), and signal-to-interference and noise ratio (SINR) by up to 26 dB (i.e., using a packet size of 256 bytes) and up to 40 dB (i.e., using a packet size of 1,399 bytes). Furthermore, we analyzed the interference effects in different data rates such as 3, 4.5, 6, 9, 12, 18, 24, and 27 Mbps. Breno Sousa, Naércio Magaia, Sara Silva, Hieu T. Nguyen 0001, Yong Liang Guan 0001 |
VTC2025-Spring | 3 |
| 2025 | Introduction to the "Best of GECCO 2023" Special IssueabstractNo abstract available. Luís Paquete, Sara Silva |
ACM Trans. Evol. Learn. Optim. | 2 |
| 2024 | Measuring Structural Complexity of GP Models for Feature Engineering over the GenerationsabstractFeature engineering is a necessary step in the machine learning pipeline. Together with other preprocessing methods, it allows the conversion of raw data into a dataset containing only the necessary features to solve the task at hand, reducing the computational complexity of inducing models and creating models that are potentially simpler, more robust, and more interpretable. We use M3GP, a wrapper-based feature engineering algorithm, to induce a set of features that are adapted in number and in shape to several classifiers with different levels of predictive power, from decision trees with depth 3 to random forests with 100 estimators and no depth limit. Intuition tells us that classifiers that are restricted in the number of features should compensate for this restriction by using features with a high degree of correlation with the target objective. By opposition, the principle behind the boosting algorithm tells us that we can create a strong classifier using a large set of weak features. This indicates that classifiers with no restrictions should prefer many but weaker features. Our results confirm this hypothesis while also revealing that M3GP induces unnecessarily complex features. We measure complexity using several structural complexity metrics found in the literature and show that, although our pipeline consistently obtains good results, the structural complexity of the induced models varies drastically across runs. Additionally, while the test performance peaks in the early stages of the evolution, the complexity of the feature engineering models continues to grow, with little to no return in test performance. This work promotes using several complexity metrics to measure model interpretability and identifies issues related to model complexity in M3GP, proposing solutions to improve the computational cost of inducing models and the complexity of the final models. João E. Batista, Adam Kotaro Pindur, Hitoshi Iba, Sara Silva |
CEC | 4 |
| 2024 | M6GP: Multiobjective Feature EngineeringabstractThe current trend in machine learning is to use powerful algorithms to induce complex predictive models that often fall under the category of “black-box models”. Thanks to this, there is also a growing interest in studying model explainabil-ity and interpretability so that human experts can understand, validate, and correct those models. With the objective of promoting the creation of inherently interpretable models, we present M6GP. This wrapper-based multi-objective automatic feature engineering algorithm combines key components of the M3GP and NSGA-II algorithms. Wrapping M6GP around another machine learning algorithm evolves a set of features optimized for this algorithm while potentially increasing its robustness. We compare our results with M3GP and M4GP, two ancestors from the same algorithm family, and verify that, by using a multi-objective approach, M6GP obtains equal or better results. In addition, by using complexity metrics on the list of objectives, the M6GP models come down to one-fifth of the size of the M3GP models, making them easier to read by comparison. João E. Batista, Nuno Miguel Rodrigues, Leonardo Vanneschi, Sara Silva |
CEC | 4 |
| 2023 | Feature Selection on Epistatic Problems Using Genetic Algorithms with Nested Classifiers
Pedro Carvalho 0002, Bruno Ribeiro 0008, Nuno M. Rodrigues, João E. Batista, Leonardo Vanneschi, Sara Silva |
EvoApplications@EvoStar | 6 |
| 2023 | An Investigation of Geometric Semantic GP with Linear ScalingabstractGeometric semantic genetic programming (GSGP) and linear scaling (LS) have both, independently, shown the ability to outperform standard genetic programming (GP) for symbolic regression. GSGP uses geometric semantic genetic operators, different from the standard ones, without altering the fitness, while LS modifies the fitness without altering the genetic operators. So far, these two methods have already been joined together in only one practical application. However, to the best of our knowledge, a methodological study on the pros and cons of integrating these two methods has never been performed. In this paper, we present a study of GSGP-LS, a system that integrates GSGP and LS. The results, obtained on five hand-tailored benchmarks and six real-life problems, indicate that GSGP-LS outperforms GSGP in the majority of the cases, confirming the expected benefit of this integration. However, for some particularly hard datasets, GSGP-LS overfits training data, being outperformed by GSGP on unseen data. Additional experiments using standard GP, with and without LS, confirm this trend also when standard crossover and mutation are employed. This contradicts the idea that LS is always beneficial for GP, warning the practitioners about its risk of overfitting in some specific cases. Giorgia Nadizar, Fraser Garrow, Berfin Sakallioglu, Lorenzo Canonne, Sara Silva, Leonardo Vanneschi |
GECCO | 5 |
| 2023 | Explainable Representations for Relation Prediction in Knowledge GraphsabstractKnowledge graphs represent real-world entities and their relations in a semantically-rich structure supported by ontologies. Exploring this data with machine learning methods often relies on knowledge graph embeddings, which produce latent representations of entities that preserve structural and local graph neighbourhood properties, but sacrifice explainability. However, in tasks such as link or relation prediction, understanding which specific features better explain a relation is crucial to support complex or critical applications. We propose SEEK, a novel approach for explainable representations to support relation prediction in knowledge graphs. It is based on identifying relevant shared semantic aspects (i.e., subgraphs) between entities and learning representations for each subgraph, producing a multi-faceted and explainable representation. We evaluate SEEK on two real-world highly complex relation prediction tasks: protein-protein interaction prediction and gene-disease association prediction. Our extensive analysis using established benchmarks demonstrates that SEEK achieves comparable or even superior performance to standard learning representation methods while identifying both sufficient and necessary explanations based on shared semantic aspects. Rita T. Sousa 0001, Sara Silva, Catia Pesquita |
KR | 2 |
| 2023 | Biomedical Knowledge Graph Embeddings with Negative Statements
Rita T. Sousa 0001, Sara Silva, Heiko Paulheim, Catia Pesquita |
ISWC | 2 |
| 2022 | Comparative study of classifier performance using automatic feature construction by M3GPabstractThe M3GP algorithm, originally designed to per-form multiclass classification with genetic programming, is also a powerful feature construction method. Here we explore its ability to evolve hyper-features that are tailored not only to the problem to be solved, but also to the learning algorithm that is used to solve it. We pair M3GP with six different machine learning algorithms and study its performance in eight classification problems from different scientific domains, with substantial variety in the number of classes, features and samples. The results show that automatic feature construction with M3GP, when compared to using the standalone classifiers without feature construction, achieves statistically significant improvements in the majority of the test cases, sometimes by a very large margin, while degrading the weighted f-measure in only one out of 48 cases. We observe the differences in the number and size of the hyper-features evolved for each case, hypothesising that the simpler the classifier, the larger the amount of problem complexity is being captured in the hyper-features. Our results also reveal that the M3GP algorithm can be improved, both in execution time and in model quality, by replacing its default classifier with support vector machines or random forest classifiers. João E. Batista, Sara Silva |
CEC | 2 |
| 2022 | SLUG: Feature Selection Using Genetic Algorithms and Genetic Programming
Nuno M. Rodrigues, João E. Batista, William G. La Cava, Leonardo Vanneschi, Sara Silva |
EuroGP | 5 |
| 2022 | Fitness landscape analysis of convolutional neural network architectures for image classificationabstractThe global structure of the hyperparameter spaces of neural networks is not well understood and it is therefore not clear which hyperparameter search algorithm will be most effective. In this paper we analyze the landscapes of convolutional neural network architecture search spaces to provide insight into appropriate search algorithms for these spaces. Using a classical fitness landscape analysis approach (fitness distance correlation) and a more recent tool (local optima networks) we study the global structure of these spaces. Our analysis on six image classification datasets reveals that the landscapes are multi-modal, but with relatively few local optima from which it is not hard to escape with a simple perturbation operator. This led us to explore the performance of iterated local search, which we found to more effectively search the training landscapes than three evolutionary algorithm variants. Evolutionary algorithms, however, outperformed iterated local search in terms of generalization on problems with larger discrepancies between the training and testing landscapes. Nuno M. Rodrigues, Katherine M. Malan, Gabriela Ochoa, Leonardo Vanneschi, Sara Silva |
Inf. Sci. | 5 |
| 2021 | μ Viz: Visualization of MicroservicesabstractMicroservice architectures have become very popular and widely adopted by the industry, because of the benefits they bring to the software development process and resulting systems, such as parallel development, modularity and scalability. However, as interfaces become more fine-grained and systems grown in size, complexity is moved from the component services to their interactions, eventually leading to intricate workflows that are hard to observe, visualize, and understand. This problem is compounded by the typically high workloads that produce intractable amounts of observation data. To deal with these challenges, operators need support from tools able to take in observation data, in particular tracing, and provide a fast and intuitive understanding of which components or workflows require attention and how are they affecting a module, service, instance, or the whole application. In this paper, we present the design of a microservice visualization application that can fill a gap that exists in leveraging tracing data, aggregating and navigating it in ways that are actionable for operators. Our application provides multiple views of the system and uses spatial and hierarchical navigation using flip zoom to simplify their exploration, while preserving context. Our application can provide a better understanding of the system than existing applications that lack navigability and do not preserve context when switching between different services, layers or views. Sara Silva, Jaime Correia, André Bento, Filipe Araújo, Raul Barbosa |
IV | 1 |
| 2020 | Improving the Detection of Burnt Areas in Remote Sensing using Hyper-features Evolved by M3GPabstractOne problem found when working with satellite images is the radiometric variations across the image and different images. Intending to improve remote sensing models for the classification of burnt areas, we set two objectives. The first is to understand the relationship between feature spaces and the predictive ability of the models, allowing us to explain the differences between learning and generalization when training and testing in different datasets. We find that training on datasets built from more than one image provides models that generalize better. These results are explained by visualizing the dispersion of values on the feature space. The second objective is to evolve hyper-features that improve the performance of different classifiers on a variety of test sets. We find the hyper-features to be beneficial, and obtain the best models with XGBoost, even if the hyper-features are optimized for a different method. João E. Batista, Sara Silva |
CEC | 2 |
| 2020 | A Study of Fitness Landscapes for NeuroevolutionabstractFitness landscapes are a useful concept to study the dynamics of meta-heuristics. In the last two decades, they have been applied with success to estimate the optimization power of several types of evolutionary algorithms, including genetic algorithms and genetic programming. However, so far they have never been used to study the performance of machine learning algorithms on unseen data, and they have never been applied to neuroevolution. This paper aims at filling both these gaps, applying for the first time fitness landscapes to neuroevolution and using them to infer useful information about the predictive ability of the method. More specifically, we use a grammar-based approach to generate convolutional neural networks, and we study the dynamics of three different mutations to evolve them. To characterize fitness landscapes, we study autocorrelation and entropic measure of ruggedness. The results show that these measures are appropriate for estimating both the optimization power and the generalization ability of the considered neuroevolution configurations. Nuno M. Rodrigues, Sara Silva, Leonardo Vanneschi |
CEC | 2 |
| 2020 | Ensemble Genetic Programming
Nuno M. Rodrigues, João E. Batista, Sara Silva |
EuroGP | 3 |
| 2020 | Is k Nearest Neighbours Regression Better Than GP?
Leonardo Vanneschi, Mauro Castelli, Luca Manzoni, Sara Silva, Leonardo Trujillo 0001 |
EuroGP | 4 |
| 2020 | Unlabeled multi-target regression with genetic programmingabstractMachine Learning (ML) has now become an important and ubiquitous tool in science and engineering, with successful applications in many real-world domains. However, there are still areas in need of improvement, and problems that are still considered difficult with off-the-shelf methods. One such problem is Multi Target Regression (MTR), where the target variable is a multidimensional tuple instead of a scalar value. In this work, we propose a more difficult variant of this problem which we call Unlabeled MTR (uMTR), where the structure of the target space is not given as part of the training data. This version of the problem lies at the intersection of MTR and clustering, an unexplored problem type. Moreover, this work proposes a solution method for uMTR, a hybrid algorithm based on Genetic Programming and RANdom SAmple Consensus (RANSAC). Using a set of benchmark problems, we are able to show that this approach can effectively solve the uMTR problem. Uriel López, Leonardo Trujillo 0001, Sara Silva, Leonardo Vanneschi, Pierrick Legrand |
GECCO | 3 |
| 2020 | The Usability Argument for Refinement Typed Genetic Programming
Alcides Fonseca, Sara Silva |
PPSN (2) | 3 |
| 2020 | Evolving knowledge graph similarity for supervised learning in complex biomedical domainsabstractBACKGROUND: In recent years, biomedical ontologies have become important for describing existing biological knowledge in the form of knowledge graphs. Data mining approaches that work with knowledge graphs have been proposed, but they are based on vector representations that do not capture the full underlying semantics. An alternative is to use machine learning approaches that explore semantic similarity. However, since ontologies can model multiple perspectives, semantic similarity computations for a given learning task need to be fine-tuned to account for this. Obtaining the best combination of semantic similarity aspects for each learning task is not trivial and typically depends on expert knowledge. RESULTS: We have developed a novel approach, evoKGsim, that applies Genetic Programming over a set of semantic similarity features, each based on a semantic aspect of the data, to obtain the best combination for a given supervised learning task. The approach was evaluated on several benchmark datasets for protein-protein interaction prediction using the Gene Ontology as the knowledge graph to support semantic similarity, and it outperformed competing strategies, including manually selected combinations of semantic aspects emulating expert knowledge. evoKGsim was also able to learn species-agnostic models with different combinations of species for training and testing, effectively addressing the limitations of predicting protein-protein interactions for species with fewer known interactions. CONCLUSIONS: evoKGsim can overcome one of the limitations in knowledge graph-based semantic similarity applications: the need to expertly select which aspects should be taken into account for a given application. Applying this methodology to protein-protein interaction prediction proved successful, paving the way to broader applications. Rita T. Sousa 0001, Sara Silva, Catia Pesquita |
BMC Bioinform. | 2 |
| 2019 | A Vectorial Approach to Genetic Programming
Irene Azzali, Leonardo Vanneschi, Sara Silva, Illya Bakurov, Mario Giacobini |
EuroGP | 3 |
| 2017 | Geometric semantic genetic programming for biomedical applications: A state of the art upgradeabstractGeometric semantic genetic programming is a hot topic in evolutionary computation and recently it has been used with success on several problems from Biology and Medicine. Given the young age of geometric semantic genetic programming, in the last few years theoretical research, aimed at improving the method, and applicative research proceeded rapidly and in parallel. As a result, the current state of the art is confused and presents some “holes”. For instance, some recent improvements of geometric semantic genetic programming have never been applied to some popular biomedical applications. The objective of this paper is to fill this gap. We consider the biomedical applications that have more frequently been used by genetic programming researchers in the last few years and we systematically test, in a consistent way, using the same parameter settings and configurations, all the most popular existing variants of geometric semantic genetic programming on all those applications. Analysing all these results, we obtain a much more homogeneous and clearer picture of the state of the art, that allows us to draw stronger conclusions. Leonardo Vanneschi, Mauro Castelli, Ivo Gonçalves, Luca Manzoni, Sara Silva |
CEC | 5 |
| 2017 | RANSAC-GP: Dealing with Outliers in Symbolic Regression with Genetic Programming
Uriel López, Leonardo Trujillo 0001, Yuliana Martínez, Pierrick Legrand, Enrique Naredo, Sara Silva |
EuroGP | 6 |
| 2017 | Genetic Programming Representations for Multi-dimensional Feature Learning in Biomedical Classification
William G. La Cava, Sara Silva, Leonardo Vanneschi, Lee Spector, Jason H. Moore |
EvoApplications (1) | 2 |
| 2017 | Unsure when to stop?: ask your semantic neighborsabstractIn iterative supervised learning algorithms it is common to reach a point in the search where no further induction seems to be possible with the available data. If the search is continued beyond this point, the risk of overfitting increases significantly. Following the recent developments in inductive semantic stochastic methods, this paper studies the feasibility of using information gathered from the semantic neighborhood to decide when to stop the search. Two semantic stopping criteria are proposed and experimentally assessed in Geometric Semantic Genetic Programming (GSGP) and in the Semantic Learning Machine (SLM) algorithm (the equivalent algorithm for neural networks). The experiments are performed on real-world high-dimensional regression datasets. The results show that the proposed semantic stopping criteria are able to detect stopping points that result in a competitive generalization for both GSGP and SLM. This approach also yields computationally efficient algorithms as it allows the evolution of neural networks in less than 3 seconds on average, and of GP trees in at most 10 seconds. The usage of the proposed semantic stopping criteria in conjunction with the computation of optimal mutation/learning steps also results in small trees and neural networks. Ivo Gonçalves, Sara Silva, Carlos M. Fonseca, Mauro Castelli |
GECCO | 2 |
| 2016 | A Machine Learning Approach for the Integration of miRNA-Target PredictionsabstractAlthough several computational methods have been developed for predicting interactions between miRNA and target genes, there are substantial differences in the achieved results. For this reason, machine learning approaches are widely used for integrating the predictions obtained from different tools. In this work we adopt a method, called M3GP, which relies on a genetic programming approach, to classify results from three tools: miRanda, TargetScan, and RNAhybrid. Such algorithm is highly parallelizable and its adoption provides great advantages while handling problems involving big datasets, since it is independent from the implementation and from the architecture on which it is executed. More precisely, we apply this technique for the classification of the achieved miRNA target predictions and we compare its results with those obtained with other classifiers. Stefano Beretta 0001, Mauro Castelli, Yuliana Martínez, Luis Muñoz, Sara Silva, Leonardo Trujillo 0001, Luciano Milanesi, Ivan Merelli |
PDP | 5 |
| 2016 | Evolving genetic programming classifiers with novelty search
Enrique Naredo, Leonardo Trujillo 0001, Pierrick Legrand, Sara Silva, Luis Muñoz |
Inf. Sci. | 4 |
| 2016 | neat Genetic Programming: Controlling bloat naturally
Leonardo Trujillo 0001, Luis Muñoz, Edgar Galván López, Sara Silva |
Inf. Sci. | 4 |
| 2015 | On the Generalization Ability of Geometric Semantic Genetic Programming
Ivo Gonçalves, Sara Silva, Carlos M. Fonseca |
EuroGP | 2 |
| 2015 | M3GP - Multiclass Classification with GP
Luis Muñoz, Sara Silva, Leonardo Trujillo 0001 |
EuroGP | 2 |
| 2015 | Geometric Semantic Genetic Programming with Local SearchabstractSince its introduction, Geometric Semantic Genetic Programming (GSGP) has aroused the interest of numerous researchers and several studies have demonstrated that GSGP is able to effectively optimize training data by means of small variation steps, that also have the effect of limiting overfitting. In order to speed up the search process, in this paper we propose a system that integrates a local search strategy into GSGP (called GSGP-LS). Furthermore, we present a hybrid approach, that combines GSGP and GSGP-LS, aimed at exploiting both the optimization speed of GSGP-LS and the ability to limit overfitting of GSGP. The experimental results we present, performed on a set of complex real-life applications, show that GSGP-LS achieves the best training fitness while converging very quickly, but severely overfits. On the other hand, GSGP converges slowly relative to the other methods, but is basically not affected by overfitting. The best overall results were achieved with the hybrid approach, allowing the search to converge quickly, while also exhibiting a noteworthy ability to limit overfitting. These results are encouraging, and suggest that future GSGP algorithms should focus on finding the correct balance between the greedy optimization of a local search strategy and the more robust geometric semantic operators. Mauro Castelli, Leonardo Trujillo 0001, Leonardo Vanneschi, Sara Silva, Emigdio Z.-Flores, Pierrick Legrand |
GECCO | 4 |
| 2014 | A Multi-dimensional Genetic Programming Approach for Multi-class Classification Problems
Vijay Ingalalli, Sara Silva, Mauro Castelli, Leonardo Vanneschi |
EuroGP | 2 |
| 2014 | ESAGP - A Semantic GP Framework Based on Alignment in the Error Space
Stefano Ruberto, Leonardo Vanneschi, Mauro Castelli, Sara Silva |
EuroGP | 4 |
| 2014 | Prediction of the Unified Parkinson's Disease Rating Scale assessment using a genetic programming system with geometric semantic genetic operators
Mauro Castelli, Leonardo Vanneschi, Sara Silva |
Expert Syst. Appl. | 3 |
| 2014 | Geometric Selective Harmony Search
Mauro Castelli, Sara Silva, Luca Manzoni, Leonardo Vanneschi |
Inf. Sci. | 2 |
| 2014 | Corrections to "Semantic Search Based Genetic Programming and the Effect of Introns Deletion"abstractThe paper above (ibid., vol. 44, no. 1, pp. 103-113, Jan. 2014), was printed with the incorrect author list as follows: M. Castelli, L. Vanneschi, S. Silva, A. Agapitos, and M. O'Neill. The correct author list is: M. Castelli, L. Vanneschi and S. Silva. Mauro Castelli, Leonardo Vanneschi, Sara Silva |
IEEE Trans. Cybern. | 3 |
| 2014 | Semantic Search-Based Genetic Programming and the Effect of Intron DeletionabstractThe concept of semantics (in the sense of input-output behavior of solutions on training data) has been the subject of a noteworthy interest in the genetic programming (GP) research community over the past few years. In this paper, we present a new GP system that uses the concept of semantics to improve search effectiveness. It maintains a distribution of different semantic behaviors and biases the search toward solutions that have similar semantics to the best solutions that have been found so far. We present experimental evidence of the fact that the new semantics-based GP system outperforms the standard GP and the well-known bacterial GP on a set of test functions, showing particularly interesting results for noncontinuous (i.e., generally harder to optimize) test functions. We also observe that the solutions generated by the proposed GP system often have a larger size than the ones returned by standard GP and bacterial GP and contain an elevated number of introns, i.e., parts of code that do not have any effect on the semantics. Nevertheless, we show that the deletion of introns during the evolution does not affect the performance of the proposed method. Mauro Castelli, Leonardo Vanneschi, Sara Silva, Alexandros Agapitos, Michael O'Neill 0001 |
IEEE Trans. Cybern. | 3 |
| 2013 | Balancing Learning and Overfitting in Genetic Programming with Interleaved Sampling of Training Data
Ivo Gonçalves, Sara Silva |
EuroGP | 2 |
| 2013 | A New Implementation of Geometric Semantic GP and Its Application to Problems in Pharmacokinetics
Leonardo Vanneschi, Mauro Castelli, Luca Manzoni, Sara Silva |
EuroGP | 4 |
| 2013 | Land Cover/Land Use Multiclass Classification Using GP with Geometric Semantic Operators
Mauro Castelli, Sara Silva, Leonardo Vanneschi, Ana I. R. Cabral, Maria J. P. de Vasconcelos, Luís Catarino, João Manuel de Brito Carreiras |
EvoApplications | 2 |
| 2013 | Prediction of Forest Aboveground Biomass: An Exercise on Avoiding Overfitting
Sara Silva, Vijay Ingalalli, Susana Vinga, João Manuel de Brito Carreiras, Joana B. Melo, Mauro Castelli, Leonardo Vanneschi, Ivo Gonçalves, José Caldas |
EvoApplications | 1 |
| 2013 | Geometric Differential Evolution for Combinatorial and Programs SpacesabstractGeometric differential evolution (GDE) is a recently introduced formal generalization of traditional differential evolution (DE) that can be used to derive specific differential evolution algorithms for both continuous and combinatorial spaces retaining the same geometric interpretation of the dynamics of the DE search across representations. In this article, we first review the theory behind the GDE algorithm, then, we use this framework to formally derive specific GDE for search spaces associated with binary strings, permutations, vectors of permutations and genetic programs. The resulting algorithms are representation-specific differential evolution algorithms searching the target spaces by acting directly on their underlying representations. We present experimental results for each of the new algorithms on a number of well-known problems comprising NK-landscapes, TSP, and Sudoku, for binary strings, permutations, and vectors of permutations. We also present results for the regression, artificial ant, parity, and multiplexer problems within the genetic programming domain. Experiments show that overall the new DE algorithms are competitive with well-tuned standard search algorithms. Alberto Moraglio, Julian Togelius, Sara Silva |
Evol. Comput. | 3 |
| 2013 | Prediction of high performance concrete strength using Genetic Programming with geometric semantic genetic operators
Mauro Castelli, Leonardo Vanneschi, Sara Silva |
Expert Syst. Appl. | 3 |
| 2012 | Random Sampling Technique for Overfitting Control in Genetic Programming
Ivo Gonçalves, Sara Silva, Joana B. Melo, João Manuel de Brito Carreiras |
EuroGP | 2 |
| 2012 | Reuse of spatial concerns based on aspectual requirements analysis patternsabstractWeb Geographic Information Systems (GIS) are systems composed by software, hardware, spatial data and computing operations, which aim to collect, model, store, share, retrieve, manipulate and display geographically referenced data. The development of online geospatial applications is currently on the rise, but this type of application often involves dealing with concerns (i.e., properties) which are inherently volatile, implying a considerable effort for system evolution. Nevertheless, geospatial concerns (e.g., temporarily blocked streets), although changeable, are reusable. However, lack of modularization in software artifacts (including system's models) can compromise reusability. In this context, the use of requirements analysis patterns, enriched with aspect-oriented modeling techniques, can support reusability and improve modularity. In this paper, we introduce requirements analysis patterns for geospatial concerns, to facilitate modularity in GIS Web applications. These patterns are generated from the domain analysis of Web GIS applications and described using a template which is supported by a comprehensive tool, enabling the completion of specific geospatial patterns. Sara Silva, João Araújo 0001, Armanda Rodrigues, Matias Urbieta, Ana Moreira 0001, Silvia E. Gordillo, Gustavo Rossi |
RCIS | 1 |
| 2011 | A Quantitative Study of Learning and Generalization in Genetic Programming
Mauro Castelli, Luca Manzoni, Sara Silva, Leonardo Vanneschi |
EuroGP | 3 |
| 2011 | An Empirical Study of Functional Complexity as an Indicator of Overfitting in Genetic Programming
Leonardo Trujillo 0001, Sara Silva, Pierrick Legrand, Leonardo Vanneschi |
EuroGP | 2 |
| 2011 | Geometric nelder-mead algorithm on the space of genetic programsabstractThe Nelder-Mead Algorithm (NMA) is a close relative of Particle Swarm Optimization (PSO) and Differential Evolution (DE). In recent work, PSO, DE and NMA have been generalized using a formal geometric framework that treats solution representations in a uniform way. These formal algorithms can be used as templates to derive rigorously specific PSO, DE and NMA for both continuous and combinatorial spaces retaining the same geometric interpretation of the search dynamics of the original algorithms across representations. In previous work, a geometric NMA has been derived for the binary string representation and permutation representation. Furthermore, PSO and DE have already been derived for the space of genetic programs. In this paper, we continue this line of research and derive formally a specific NMA for the space of genetic programs. The result is a Nelder-Mead Algorithm searching the space of genetic programs by acting directly on their tree representation. We present initial experimental results for the new algorithm. The challenge tackled in the present work compared with earlier work is that the pair NMA and genetic programs is the most complex considered so far. This combination raises a number of issues and casts light on how algorithmic features can interact with representation features to give rise to a highly peculiar search behaviour. Alberto Moraglio, Sara Silva |
GECCO | 2 |
| 2011 | Reassembling operator equalisation: a secret revealedabstractThe recent Crossover Bias theory has shown that bloat in Genetic Programming can be caused by the proliferation of small unfit individuals in the population. Inspired by this theory, Operator Equalisation is the most recent and successful bloat control method available. In this work we revisit two bloat control methods, the old Brood Recombination and the newer Dynamic Limits, hypothesizing that together they contain the two main ingredients that make Operator Equalisation so successful. We reassemble Operator Equalisation by joining these two ingredients in a hybrid method, and test it in a hard real world regression problem. The results are surprising. Operator Equalisation and the hybrid variants exhibit completely different behaviors, and an unexpected feature of Operator Equalisation is revealed, one that may be the true responsible for its success: a nearly flat length distribution target. We support this finding with additional results, and discuss its implications. Sara Silva |
GECCO | 1 |
| 2010 | A comparison of the generalization ability of different genetic programming frameworksabstractGeneralization is an important issue in machine learning. In fact, in several applications good results over training data are not as important as good results over unseen data. While this problem was deeply studied in other machine learning techniques, it has become an important issue for genetic programming only in the last few years. In this paper we compare the generalization ability of several different genetic programming frameworks, including some variants of multi-objective genetic programming and operator equalization, a recently defined bloat free genetic programming system. The test problem used is a hard regression real-life application in the field of drug discovery and development, characterized by a high number of features and where the generalization ability of the proposed solutions is a crucial issue. The results we obtained show that, at least for the considered problem, multi-optimization is effective in improving genetic programming generalization ability, outperforming all the other methods on test data. Mauro Castelli, Luca Manzoni, Sara Silva, Leonardo Vanneschi |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Geometric Differential Evolution on the Space of Genetic Programs
Alberto Moraglio, Sara Silva |
EuroGP | 2 |
| 2010 | Application of Genetic Programming Classification in an Industrial Process Resulting in Greenhouse Gas Emission Reductions
Marco Lotz, Sara Silva |
EvoApplications (2) | 2 |
| 2010 | Bloat Free Genetic Programming versus Classification Trees for Identification of Burned Areas in Satellite Imagery
Sara Silva, Maria J. P. de Vasconcelos, Joana B. Melo |
EvoApplications (1) | 1 |
| 2010 | Measuring bloat, overfitting and functional complexity in genetic programmingabstractRecent contributions clearly show that eliminating bloat in a genetic programming system does not necessarily eliminate overfitting and vice-versa. This fact seems to contradict a common agreement of many researchers known as the minimum description length principle, which states that the best model is the one that minimizes the amount of information needed to encode it. Another common agreement is that overfitting should be, in some sense, related to the functional complexity of the model. The goal of this paper is to define three measures to respectively quantify bloat, overfitting and functional complexity of solutions and show their suitability on a set of test problems including a simple bidimensional symbolic regression test function and two real-life multidimensional regression problems. The experimental results are encouraging and should pave the way to further investigation. Advantages and drawbacks of the proposed measures are discussed, and ways to improve them are suggested. In the future, these measures should be useful to study and better understand the relationship between bloat, overfitting and functional complexity of solutions. Leonardo Vanneschi, Mauro Castelli, Sara Silva |
GECCO | 3 |
| 2009 | Extending Operator Equalisation: Fitness Based Self Adaptive Length Distribution for Bloat Free GP
Sara Silva, Stephen Dignum |
EuroGP | 1 |
| 2009 | Operator equalisation, bloat and overfitting: a study on human oral bioavailability predictionabstractOperator equalisation was recently proposed as a new bloat control technique for genetic programming. By controlling the distribution of program lengths inside the population, it can bias the search towards smaller or larger programs. In this paper we propose a new implementation of operator equalisation and compare it to a previous version, using a hard real-world regression problem where bloat and overfitting are major issues. The results show that both implementations of operator equalisation are completely bloat-free, producing smaller individuals than standard genetic programming, without compromising the generalization ability. We also show that the new implementation of operator equalisation is more efficient and exhibits a more predictable and reliable behavior than the previous version. We advance some arguable ideas regarding the relationship between bloat and overfitting, and support them with our results. Sara Silva, Leonardo Vanneschi |
GECCO | 1 |
| 2008 | SWIT: an open-source web-based database management system with indexing engine integrationabstractThere are currently a considerable number of digital libraries using XML documents to describe their resources. In this paper, we describe a database solution which integrates an XML native database and a search engine to improve query performance. A web based management application is also presented which allows local and remote administration of the database. Sara Silva, Marco Fernandes, Joaquim Arnaldo Martins, Joaquim Sousa Pinto |
EATIS | 1 |
| 2005 | Comparing tree depth limits and resource-limited GPabstractIn this paper we compare two different approaches for controlling bloat in genetic programming, tree depth limits and resource-limited GP. Tree depth limits operate at the individual level, avoiding excessive code growth by imposing a maximum depth to each individual. Resource-limited GP is a new technique that operates at the population level, limiting the total amount of resources the entire population can use. We compare their dynamics and performance on three problems: symbolic regression, even parity, and artificial ant. The results suggest that resource-limited GP is superior to tree depth limits, but we question this superiority and discuss possible ways of combining the strengths of both approaches, to further improve the results Sara Silva, Ernesto Costa |
Congress on Evolutionary Computation | 1 |
| 2005 | Resource-limited genetic programming: the dynamic approachabstractResource-Limited Genetic Programming is a bloat control technique that imposes a single limit on the total amount of resources available to the entire population, where resources are tree nodes or code lines. We elaborate on this recent concept, introducing a dynamic approach to managing the amount of resources available for each generation. Initially low, this amount is increased only if it results in better population fitness. We compare the dynamic approach to the static method where a constant amount of resources is available throughout the run, and with the most traditional usage of a depth limit at the individual level. The dynamic approach does not impair performance on the Symbolic Regression of the quartic polynomial, and achieves excellent results on the Santa Fe Artificial Ant problem, obtaining the same fitness with only a small percentage of the computational effort demanded by the other techniques. Sara Silva, Ernesto Costa |
GECCO | 1 |
| 2004 | Dynamic Limits for Bloat Control: Variations on Size and Depth
Sara Silva, Ernesto Costa |
GECCO (2) | 1 |
| 2003 | Dynamic Maximum Tree Depth
Sara Silva, Jonas S. Almeida |
GECCO | 1 |