Gabriel Kronberger

dblp:82/5128 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-3012-3189ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Comparative Study of Model Selection Criteria for Symbolic Regression
abstract
Effective model selection is critical in symbolic regression (SR) to identify mathematical expressions that balance accuracy and complexity, and have low expected error on unseen data. Many modern implementations of genetic programming (GP) for SR generate a set of Pareto optimal candidate solutions, but reliable automatic selection of solutions that generalize well remains an open issue. Current literature offers various information-theoretic and Bayesian approaches, yet comprehensive comparisons of their performance across different data regimes are limited. This study presents a systematic empirical comparison of widely used selection criteria: the Akaike information criterion (AIC), the corrected AIC (AICc), the Bayesian information criterion (BIC), minimum description length (MDL), as well as Efron's bootstrap estimate for the in-sample prediction error on seven synthetic datasets with Gaussian noise. We rank candidate expressions generated by perturbing ground-truth functions to assess generalization error and selection probability of the ground-truth expression. Our findings reveal that MDL consistently identifies models with the lowest test error and the shortest length across most datasets. While no single criterion dominates all results, MDL and BIC produced the highest probability of selecting the ground-truth expressions.
Ali Soltani, Gabriel Kronberger, Fabrício Olivetti de França, Mattia Billa, Alessandro Lucantonio
GECCO2
2025 rEGGression: an Interactive and Agnostic Tool for the Exploration of Symbolic Regression Models
abstract
Regression analysis is used for prediction and to understand the effect of independent variables on dependent variables. Symbolic regression (SR) automates the search for non-linear regression models, delivering a set of hypotheses that balances accuracy with the possibility to understand the phenomena. Many SR implementations return a Pareto front allowing the choice of the best trade-off. However, this hides alternatives that are close to non-domination, limiting these choices. Equality graphs (e-graphs) allow to represent large sets of expressions compactly by efficiently handling duplicated parts occurring in multiple expressions. The e-graphs allow to efficiently store and query all solution candidates visited in one or multiple runs of different algorithms and open the possibility to analyze much larger sets of SR solution candidates. We introduce rEGGression, a tool using e-graphs to enable the exploration of a large set of symbolic expressions which provides querying, filtering, and pattern matching features creating an interactive experience to gain insights about SR models. The main highlight is its focus in the exploration of the building blocks found during the search that can help the experts to find insights about the studied phenomena. This is possible by exploiting the pattern matching capability of the e-graph data structure.
Fabrício Olivetti de França, Gabriel Kronberger
GECCO2
2025 Improving Genetic Programming for Symbolic Regression with Equality Graphs
abstract
The search for symbolic regression models with genetic programming (GP) has a tendency of revisiting expressions in their original or equivalent forms. Repeatedly evaluating equivalent expressions is inefficient, as it does not immediately lead to better solutions. However, evolutionary algorithms require diversity and should allow the accumulation of inactive building blocks that can play an important role at a later point. The equality graph is a data structure capable of compactly storing expressions and their equivalent forms allowing an efficient verification of whether an expression has been visited in any of their stored equivalent forms. We exploit the e-graph to adapt the subtree operators to reduce the chances of revisiting expressions. Our adaptation, called eggp, stores every visited expression in the e-graph, allowing us to filter out from the available selection of subtrees all the combinations that would create already visited expressions. Results show that, for small expressions, this approach improves the performance of a simple GP algorithm to compete with PySR and Operon without increasing computational cost. As a highlight, eggp was capable of reliably delivering short and at the same time accurate models for a selected set of benchmarks from SRBench and a set of real-world datasets.
Fabrício Olivetti de França, Gabriel Kronberger
GECCO2
2025 Comparative Analysis of Model Selection Criteria for Symbolic Regression Using Genetic Programming
Fitria Wulandari Ramlan, Gabriel Kronberger, Colm O'Riordan, James McDermott
IJCCI (2)2
2025 Effects of reducing redundant parameters in parameter optimization for symbolic regression using genetic programming
Gabriel Kronberger, Fabrício Olivetti de França
J. Symb. Comput.1
2024 The Inefficiency of Genetic Programming for Symbolic Regression
Gabriel Kronberger, Fabrício Olivetti de França, Harry Desmond, Deaglan J. Bartlett, Lukas Kammerer
PPSN (1)1
2024 Learning Difference Equations With Structured Grammatical Evolution for Postprandial Glycaemia Prediction
abstract
People with diabetes must carefully monitor their blood glucose levels, especially after eating. Blood glucose management requires a proper combination of food intake and insulin boluses. Glucose prediction is vital to avoid dangerous post-meal complications in treating individuals with diabetes. Although traditional methods, and also artificial neural networks, have shown high accuracy rates, sometimes they are not suitable for developing personalised treatments by physicians due to their lack of interpretability. This study proposes a novel glucose prediction method emphasising interpretability: Interpretable Sparse Identification by Grammatical Evolution. Combined with a previous clustering stage, our approach provides finite difference equations to predict postprandial glucose levels up to two hours after meals. We divide the dataset into four-hour segments and perform clustering based on blood glucose values for the two-hour window before the meal. Prediction models are trained for each cluster for the two-hour windows after meals, allowing predictions in 15-minute steps, yielding up to eight predictions at different time horizons. Prediction safety was evaluated based on Parkes Error Grid regions. Our technique produces safe predictions through explainable expressions, avoiding zones D (0.2% average) and E (0%) and reducing predictions on zone C (6.2%). In addition, our proposal has slightly better accuracy than other techniques, including sparse identification of non-linear dynamics and artificial neural networks. The results demonstrate that our proposal provides interpretable solutions without sacrificing prediction accuracy, offering a promising approach to glucose prediction in diabetes management that balances accuracy, interpretability, and computational efficiency.
Daniel Parra, David Jödicke, José Manuel Velasco, Gabriel Kronberger, J. Ignacio Hidalgo
IEEE J. Biomed. Health Informatics4
2023 Reducing Overparameterization of Symbolic Regression Models with Equality Saturation
abstract
Overparameterized models in regression analysis are often harder to interpret and can be harder to fit because of ill-conditioning. Genetic programming is prone to overparameterized models as it evolves the structure of the model without taking the location of parameters into account. One way to alleviate this is rewriting the expression and merging the redundant fitting parameters. In this paper we propose the use of equality saturation to alleviate overparameterization. We first notice that all the tested GP implementations suffer from overparameterization to different extents and then show that equality saturation together with a small set of rewriting rules is capable of reducing the number of fitting parameters to a minimum with a high probability. Compared to one of the few available alternatives, Sympy, it produces much better and consistent results. These results lead to different possible future investigations such as the simplification of expressions during the evolutionary process, and improvement of the interpretability of symbolic models.
Fabrício Olivetti de França, Gabriel Kronberger
GECCO2
2022 Comparing optimistic and pessimistic constraint evaluation in shape-constrained symbolic regression
abstract
Shape-constrained Symbolic Regression integrates prior knowledge about the function shape into the symbolic regression model. This can be used to enforce that the model has desired properties such as monotonicity, or convexity, among others. Shape-constrained Symbolic Regression can also help to create models with better extrapolation behavior and reduced sensitivity to noise. The constraint evaluation can be challenging because exact evaluation of constraints may require a search for the extrema of non-convex functions. Approximations via interval arithmetic allow to efficiently find bounds for the extrema of functions. However, interval arithmetic can lead to overly wide bounds and therefore produces a pessimistic estimation. Another possibility is to use sampling which underestimates the true range. Sampling therefore produces an optimistic estimation. In this paper we evaluate both methods and compare them on different problem instances. In particular we evaluate the sensitivity to noise and the extrapolation capabilities in combination with noise data. The results indicate that the optimistic approach works better for predicting out-of-domain points (extrapolation) and the pessimistic approach works better for high noise levels.
Christian Haider, Fabrício Olivetti de França, Gabriel Kronberger, Bogdan Burlacu
GECCO3
2022 Shape-Constrained Symbolic Regression - Improving Extrapolation with Prior Knowledge
abstract
We investigate the addition of constraints on the function image and its derivatives for the incorporation of prior knowledge in symbolic regression. The approach is called shape-constrained symbolic regression and allows us to enforce, for example, monotonicity of the function over selected inputs. The aim is to find models which conform to expected behavior and which have improved extrapolation capabilities. We demonstrate the feasibility of the idea and propose and compare two evolutionary algorithms for shape-constrained symbolic regression: (i) an extension of tree-based genetic programming which discards infeasible solutions in the selection step, and (ii) a two-population evolutionary algorithm that separates the feasible from the infeasible solutions. In both algorithms we use interval arithmetic to approximate bounds for models and their partial derivatives. The algorithms are tested on a set of 19 synthetic and four real-world regression problems. Both algorithms are able to identify models which conform to shape constraints which is not the case for the unmodified symbolic regression algorithms. However, the predictive accuracy of models with constraints is worse on the training set and the test set. Shape-constrained polynomial regression produces the best results for the test set but also significantly larger models.
Gabriel Kronberger, Fabrício Olivetti de França, Bogdan Burlacu, Christian Haider, Michael Kommenda
Evol. Comput.1
2021 Estimation of Grain-Level Residual Stresses in a Quenched Cylindrical Sample of Aluminum Alloy AA5083 Using Genetic Programming
Laura Millán, Gabriel Kronberger, J. Ignacio Hidalgo, Ricardo Fernández, Oscar Garnica, Gaspar González-Doncel
EvoApplications2
2020 Multilayer analysis of population diversity in grammatical evolution for symbolic regression
abstract
Abstract In this paper, we analyze the population diversity of grammatical evolution (GE) on multiple levels of genetic information: chromosome diversity, expression diversity, and output diversity. Thereby, we use a tree-similarity metric from tree-based GP literature to determine similarity of expression trees generated in GE. The similarity of outputs is determined via their correlation. We track the pairwise similarities for all individuals within a generation on all three levels and track the distribution of similarity values over generations. We demonstrate the analysis method using four symbolic regression problem instances and find that the visualization highlights some issues that can occur when using GE such as: large groups of individuals with highly similar outputs, a high fraction of trees with constant outputs, or short and highly similar trees in the early stages of the GE run. Especially in the early phases of GE, we see that a large subset of the population represents equivalent expressions. In early stages, rather short expressions are produced leaving large parts of the chromosome unexpressed. More complex expressions can be derived only after GE has successfully evolved well-working beginnings of chromosomes.
Gabriel Kronberger, José Manuel Colmenar, Stephan M. Winkler, J. Ignacio Hidalgo
Soft Comput.1
2019 Using Ontologies to Express Prior Knowledge for Genetic Programming
Stefan Prieschl, Dominic Girardi, Gabriel Kronberger
CD-MAKE3
2019 Online Diversity Control in Symbolic Regression via a Fast Hash-based Tree Similarity Measure
abstract
Diversity represents an important aspect of genetic programming, being directly correlated with search performance. When considered at the genotype level, diversity often requires expensive tree distance measures which have a negative impact on the algorithm's runtime performance. In this work we introduce a fast, hash-based tree distance measure to massively speed-up the calculation of population diversity during the algorithmic run. We combine this measure with the standard GA and the NSGA-II genetic algorithms to steer the search towards higher diversity. We validate the approach on a collection of benchmark problems for symbolic regression where our method consistently outperforms the standard GA as well as NSGA-II configurations with different secondary objectives.
Bogdan Burlacu, Michael Affenzeller, Gabriel Kronberger, Michael Kommenda
CEC3
2018 Predicting friction system performance with symbolic regression and genetic programming with factor variables
abstract
Friction systems are mechanical systems wherein friction is used for force transmission (e.g. mechanical braking systems or automatic gearboxes). For finding optimal and safe design parameters, engineers have to predict friction system performance. This is especially difficult in real-worlds applications, because it is affected by many parameters.
Gabriel Kronberger, Michael Kommenda, Andreas Promberger, Falk Nickel
GECCO1
2012 Knowledge Discovery through Symbolic Regression with HeuristicLab
Gabriel Kronberger, Stefan Wagner 0002, Michael Kommenda, Andreas Beham, Andreas Scheibenpflug, Michael Affenzeller
ECML/PKDD (2)1
2011 Data Mining Using Unguided Symbolic Regression on a Blast Furnace Dataset
Michael Kommenda, Gabriel Kronberger, Christoph Feilmayr, Michael Affenzeller
EvoApplications (1)2
2011 Macro-economic Time Series Modeling and Interaction Networks
Gabriel Kronberger, Stefan Fink, Michael Kommenda, Michael Affenzeller
EvoApplications (2)1
2009 On Crossover Success Rate in Genetic Programming with Offspring Selection
Gabriel Kronberger, Stephan M. Winkler, Michael Affenzeller, Stefan Wagner 0002
EuroGP1