VLDB 2026 Research / reviewers in the wild / expert
Dominik Sobania
dblp:222/3766
· DBLP profile ↗
25ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-8873-7143ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 12 first-author · 19 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ROIDS: Robust Outlier-Aware Informed Down-SamplingabstractInformed down-sampling (IDS) is known to improve performance in symbolic regression when combined with various selection strategies, especially tournament selection. However, recent work found that IDS's gains are not consistent across all problems. Our analysis reveals that IDS performance is worse for problems containing outliers. IDS systematically favors including outliers in subsets which pushes GP towards finding solutions that overfit to outliers. To address this, we introduce ROIDS (Robust Outlier-Aware Informed Down-Sampling), which excludes potential outliers from the sampling process of IDS. With ROIDS it is possible to keep the advantages of IDS without overfitting to outliers and to compete on a wide range of benchmark problems. This is also reflected in our experiments in which ROIDS shows the desired behavior on all studied benchmark problems. ROIDS consistently outperforms IDS on synthetic problems with added outliers as well as on a wide range of complex real-world problems, surpassing IDS on over 80% of the real-world benchmark problems. Moreover, compared to all studied baseline approaches, ROIDS achieves the best average rank across all tested benchmark problems. This robust behavior makes ROIDS a reliable down-sampling method for selection in symbolic regression, especially when outliers may be included in the data set. Alina Geiger, Martin Briesch, Dominik Sobania, Franz Rothlauf |
GECCO | 3 |
| 2026 | SQL3M: Token Efficient Text-to-SQL Generation
Ibrahim Ücelehan, Alina Geiger, Dominik Sobania |
SANER | 3 |
| 2026 | A Performance Analysis of Lexicase-Based and Traditional Selection Methods in GP for Symbolic RegressionabstractIn recent years, several new lexicase-based selection variants have emerged due to the success of standard lexicase selection in various application domains. For symbolic regression problems, variants that use an \(\epsilon\) -threshold or batches of training cases, among others, have led to performance improvements. Lately, especially variants that combine lexicase selection and down-sampling strategies have received a lot of attention. This article evaluates the most relevant lexicase-based selection methods as well as traditional selection methods in combination with different down-sampling strategies on a wide range of symbolic regression problems. In contrast to most work, we not only compare the methods over a given evaluation budget, but also over a given time budget as time is usually limited in practice. We find that for a given evaluation budget, \(\epsilon\) -lexicase selection in combination with a down-sampling strategy outperforms all other methods. If the given running time is very short, lexicase variants using batches of training cases perform best. Further, we find that the combination of tournament selection with informed down-sampling performs well in all studied settings. Alina Geiger, Dominik Sobania, Franz Rothlauf |
ACM Trans. Evol. Learn. Optim. | 2 |
| 2025 | Was Tournament Selection All We Ever Needed? A Critical Reflection on Lexicase Selection
Alina Geiger, Martin Briesch, Dominik Sobania, Franz Rothlauf |
EuroGP | 3 |
| 2025 | Transformer Semantic Genetic Programming for Symbolic RegressionabstractIn standard genetic programming (stdGP), solutions are varied by modifying their syntax, with uncertain effects on their semantics. Geometric-semantic genetic programming (GSGP), a popular variant of GP, effectively searches the semantic solution space using variation operations based on linear combinations, although it results in significantly larger solutions. This paper presents Transformer Semantic Genetic Programming (TSGP), a novel and flexible semantic approach that uses a generative transformer model as search operator. The transformer is trained on synthetic test problems and learns semantic similarities between solutions. Once the model is trained, it can be used to create offspring solutions with high semantic similarity also for unseen and unknown problems. Experiments on several symbolic regression problems show that TSGP generates solutions with comparable or even significantly better prediction quality than stdGP, SLIM_GSGP, DSR, and DAE-GP. Like SLIM_GSGP, TSGP is able to create new solutions that are semantically similar without creating solutions of large size. An analysis of the search dynamic reveals that the solutions generated by TSGP are semantically more similar than the solutions generated by the benchmark approaches allowing a better exploration of the semantic solution space. Philipp Anthes, Dominik Sobania, Franz Rothlauf |
GECCO | 2 |
| 2025 | ImageBreeder: Guiding Diffusion Models with Evolutionary Computation
Dominik Sobania, Martin Briesch, Franz Rothlauf |
GECCO | 1 |
| 2025 | LLM-Guided Genetic Improvement: Envisioning Semantic Aware Automated Software Evolution
Karine Even-Mendoza, Alexander E. I. Brownlee, Alina Geiger, Carol Hanna, Justyna Petke, Federica Sarro, Dominik Sobania |
ASE | 7 |
| 2025 | Large language model based mutations in genetic improvementabstractAbstract Ever since the first large language models (LLMs) have become available, both academics and practitioners have used them to aid software engineering tasks. However, little research as yet has been done in combining search-based software engineering (SBSE) and LLMs. In this paper, we evaluate the use of LLMs as mutation operators for genetic improvement (GI), an SBSE approach, to improve the GI search process. In a preliminary work, we explored the feasibility of combining the Gin Java GI toolkit with OpenAI LLMs in order to generate an edit for the tool. Here we extend this investigation involving three LLMs and three types of prompt, and five real-world software projects. We sample the edits at random, as well as using local search. We also conducted a qualitative analysis to understand why LLM-generated code edits break as part of our evaluation. Our results show that, compared with conventional statement GI edits, LLMs produce fewer unique edits, but these compile and pass tests more often, with the model finding test-passing edits 77% of the time. The and LLMs are roughly equal in finding the best run-time improvements. Simpler prompts are more successful than those providing more context and examples. The qualitative analysis reveals a wide variety of areas where LLMs typically fail to produce valid edits commonly including inconsistent formatting, generating non-Java syntax, or refusing to provide a solution. Alexander E. I. Brownlee, James Callan, Karine Even-Mendoza, Alina Geiger, Carol Hanna, Justyna Petke, Federica Sarro, Dominik Sobania |
Autom. Softw. Eng. | 8 |
| 2025 | A Comparison of Large Language Models and Genetic Programming for Program SynthesisabstractLarge language models have recently become known for their ability to generate computer programs, especially through tools, such as GitHub Copilot, a domain where genetic programming (GP) has been very successful so far. Although they require different inputs (free-text versus input/output examples) their goal is the same—program synthesis. Therefore, in this work, we compare how well GitHub Copilot and GP perform on common program synthesis benchmark problems. We study the structure and diversity of the generated programs by using well-known software metrics. We find that GitHub Copilot and GP solve a similar number of benchmark problems (85.2% versus 77.8%, respectively). We find that GitHub Copilot generated smaller and less complex programs as GP, while GP is able to find new and unique problem solving strategies. This increase in diversity of solutions comes at a cost. When analyzing the success rates for 100 runs per problem, GitHub Copilot outperforms GP on over 50% of the problems. Dominik Sobania, Justyna Petke, Martin Briesch, Franz Rothlauf |
IEEE Trans. Evol. Comput. | 1 |
| 2024 | A Comprehensive Comparison of Lexicase-Based Selection Methods for Symbolic Regression Problems
Alina Geiger, Dominik Sobania, Franz Rothlauf |
EuroGP | 2 |
| 2024 | Informed Down-Sampled Lexicase Selection: Identifying Productive Training Cases for Efficient Problem SolvingabstractGenetic Programming (GP) often uses large training sets and requires all individuals to be evaluated on all training cases during selection. Random down-sampled lexicase selection evaluates individuals on only a random subset of the training cases, allowing for more individuals to be explored with the same number of program executions. However, sampling randomly can exclude important cases from the down-sample for a number of generations, while cases that measure the same behavior (synonymous cases) may be overused. In this work, we introduce Informed Down-Sampled Lexicase Selection. This method leverages population statistics to build down-samples that contain more distinct and therefore informative training cases. Through an empirical investigation across two different GP systems (PushGP and Grammar-Guided GP), we find that informed down-sampling significantly outperforms random down-sampling on a set of contemporary program synthesis benchmark problems. Through an analysis of the created down-samples, we find that important training cases are included in the down-sample consistently across independent evolutionary runs and systems. We hypothesize that this improvement can be attributed to the ability of Informed Down-Sampled Lexicase Selection to maintain more specialist individuals over the course of evolution, while still benefiting from reduced per-evaluation costs. Ryan Bahlous-Boldi, Martin Briesch, Dominik Sobania, Alexander Lalejini, Thomas Helmuth, Franz Rothlauf, Charles Ofria, Lee Spector |
Evol. Comput. | 3 |
| 2023 | MTGP: Combining Metamorphic Testing and Genetic Programming
Dominik Sobania, Martin Briesch, Philipp Röchner, Franz Rothlauf |
EuroGP | 1 |
| 2023 | Down-Sampled Epsilon-Lexicase Selection for Real-World Symbolic Regression ProblemsabstractEpsilon-lexicase selection is a parent selection method in genetic programming that has been successfully applied to symbolic regression problems. Recently, the combination of random subsampling with lexicase selection significantly improved performance in other genetic programming domains such as program synthesis. However, the influence of subsampling on the solution quality of real-world symbolic regression problems has not yet been studied. In this paper, we propose down-sampled epsilon-lexicase selection which combines epsilon-lexicase selection with random subsampling to improve the performance in the domain of symbolic regression. Therefore, we compare down-sampled epsilon-lexicase with traditional selection methods on common real-world symbolic regression problems and analyze its influence on the properties of the population over a genetic programming run. We find that the diversity is reduced by using down-sampled epsilon-lexicase selection compared to standard epsilon-lexicase selection. This comes along with high hyperselection rates we observe for down-sampled epsilon-lexicase selection. Further, we find that down-sampled epsilon-lexicase selection outperforms the traditional selection methods on all studied problems. Overall, with down-sampled epsilon-lexicase selection we observe an improvement of the solution quality of up to 85% in comparison to standard epsilon-lexicase selection. Alina Geiger, Dominik Sobania, Franz Rothlauf |
GECCO | 2 |
| 2023 | Enhancing Genetic Improvement Mutations Using Large Language Models
Alexander E. I. Brownlee, James Callan, Karine Even-Mendoza, Alina Geiger, Carol Hanna, Justyna Petke, Federica Sarro, Dominik Sobania |
SSBSE | 8 |
| 2023 | Evaluating Explanations for Software Patches Generated by Large Language Models
Dominik Sobania, Alina Geiger, James Callan, Alexander E. I. Brownlee, Carol Hanna, Rebecca Moussa, Mar Zamorano López, Justyna Petke, Federica Sarro |
SSBSE | 1 |
| 2023 | A Comprehensive Survey on Program Synthesis With Evolutionary AlgorithmsabstractThe automatic generation of computer programs is one of the main applications with practical relevance in the field of evolutionary computation. With program synthesis techniques not only software developers could be supported in their everyday work but even users without any programming knowledge could be empowered to automate repetitive tasks and implement their own new functionality. In recent years, many novel program synthesis approaches based on evolutionary algorithms have been proposed and evaluated on common benchmark problems. Therefore, we identify and discuss in this survey the relevant evolutionary program synthesis approaches in the literature and provide an in-depth analysis of their performance. The most influential approaches we identify are stack-based, grammar-guided, as well as linear genetic programming (GP). For the stack-based approaches, we identify 37 in-scope papers, and for the grammar-guided and linear GP approaches, we identify 12 and 5 papers, respectively. Furthermore, we find that these approaches perform well on benchmark problems if there is a simple mapping from the given input to the correct output. On problems where this mapping is complex, e.g., if the problem consists of several subproblems or requires iteration/recursion for a correct solution, results tend to be worse. Consequently, for future work, we encourage researchers not only to use a program’s output for assessing the quality of a solution but also the way toward a solution (e.g., correctly solved subproblems). Dominik Sobania, Dirk Schweim, Franz Rothlauf |
IEEE Trans. Evol. Comput. | 1 |
| 2022 | Effects of the Training Set Size: A Comparison of Standard and Down-Sampled Lexicase Selection in Program SynthesisabstractFrom a practitioners perspective, the number of in-put/output examples used during the training process in program synthesis studies is too large, as in practice, these examples must be labeled by hand. Therefore, this paper analyzes the influence of different training set sizes on the performance, generalization ability, as well as the structure of the programs generated by grammar-guided genetic programming. We compare down-sampled lexicase selection with standard lexicase selection on three common problems from the general program synthesis benchmark suite. First, we find that both lexicase variants are robust against reducing the amount of training data. We find that standard lexicase has a tendency to overfit the training data on some problems. With down-sampled lexicase, in contrast, overfitting on training data is reduced and evolved programs generalize better on held-out test cases. Consequently, we suggest to use grammar-guided genetic programming with down-sampled lexicase selection in the program synthesis domain. Dirk Schweim, Dominik Sobania, Franz Rothlauf |
CEC | 2 |
| 2022 | Exploiting Knowledge from Code to Guide Program Search
Dirk Schweim, Erik Hemberg, Dominik Sobania, Una-May O'Reilly |
EuroGP | 3 |
| 2022 | Program Synthesis with Genetic Programming: The Influence of Batch Sizes
Dominik Sobania, Franz Rothlauf |
EuroGP | 1 |
| 2022 | Choose your programming copilot: a comparison of the program synthesis performance of github copilot and genetic programmingabstractGitHub Copilot, an extension for the Visual Studio Code development environment powered by the large-scale language model Codex, makes automatic program synthesis available for software developers. This model has been extensively studied in the field of deep learning, however, a comparison to genetic programming, which is also known for its performance in automatic program synthesis, has not yet been carried out. In this paper, we evaluate GitHub Copilot on standard program synthesis benchmark problems and compare the achieved results with those from the genetic programming literature. In addition, we discuss the performance of both approaches. We find that the performance of the two approaches on the benchmark problems is quite similar, however, in comparison to GitHub Copilot, the program synthesis approaches based on genetic programming are not yet mature enough to support programmers in practical software development. Genetic programming usually needs a huge amount of expensive hand-labeled training cases and takes too much time to generate solutions. Furthermore, source code generated by genetic programming approaches is often bloated and difficult to understand. For future work on program synthesis with genetic programming, we suggest researchers to focus on improving the execution time, readability, and usability. Dominik Sobania, Martin Briesch, Franz Rothlauf |
GECCO | 1 |
| 2021 | On the Generalizability of Programs Synthesized by Grammar-Guided Genetic Programming
Dominik Sobania |
EuroGP | 1 |
| 2021 | A generalizability measure for program synthesis with genetic programmingabstractThe generalizability of programs synthesized by genetic programming (GP) to unseen test cases is one of the main challenges of GP-based program synthesis. Recent work showed that increasing the amount of training data improves the generalizability of the programs synthesized by GP. However, generating training data is usually an expensive task as the output value for every training case must be calculated manually by the user. Therefore, this work suggests an approximation of the expected generalization ability of solution candidates found by GP. To obtain candidate solutions that all solve the training cases, but are structurally different, a GP run is not stopped after the first solution is found that solves all training instances but search continues for more generations. For all found candidate solutions (solving all training cases), we calculate the behavioral vector for a set of randomly generated additional inputs. The proportion of the number of different found candidate solutions generating the same behavioral vector with highest frequency compared to all other found candidate solutions with different behavior can serve as an approximation for the generalizability of the found solutions. The paper presents experimental results for a number of standard program synthesis problems confirming the high prediction accuracy. Dominik Sobania, Franz Rothlauf |
GECCO | 1 |
| 2020 | Challenges of Program Synthesis with Grammatical Evolution
Dominik Sobania, Franz Rothlauf |
EuroGP | 1 |
| 2019 | Teaching GP to program like a human software developer: using perplexity pressure to guide program synthesis approachesabstractProgram synthesis is one of the relevant applications of GP with a strong impact on new fields such as genetic improvement. In order for synthesized code to be used in real-world software, the structure of the programs created by GP must be maintainable. We can teach GP how real-world software is built by learning the relevant properties of mined human-coded software - which can be easily accessed through repository hosting services such as GitHub. So combining program synthesis and repository mining is a logical step. In this paper, we analyze if GP can write programs with properties similar to code produced by human software developers. First, we compare the structure of functions generated by different GP initialization methods to a mined corpus containing real-world software. The results show that the studied GP initialization methods produce a totally different combination of programming language elements in comparison to real-world software. Second, we propose perplexity pressure and analyze how its use changes the properties of code produced by GP. The results are very promising and show that we can guide the search to the desired program structure. Thus, we recommend using perplexity pressure as it can be easily integrated in various search-based algorithms. Dominik Sobania, Franz Rothlauf |
GECCO | 1 |
| 2018 | CovSel: A new approach for ensemble selection applied to symbolic regression problemsabstractEnsemble methods combine the predictions of a set of models to reach a better prediction quality compared to a single model's prediction. The ensemble process consists of three steps: 1) the generation phase where the models are created, 2) the selection phase where a set of possible ensembles is composed and one is selected by a selection method, 3) the fusion phase where the individual models' predictions of the selected ensemble are combined to an ensemble's estimate. This paper proposes CovSel, a selection approach for regression problems that ranks ensembles based on the coverage of adequately estimated training points and selects the ensemble with the highest coverage to be used in the fusion phase. An ensemble covers a training point if at least one of its models produces an adequate prediction for this training point. The more training points are covered this way, the higher is the ensemble's coverage. The selection of the "right" ensemble has a large impact on the final prediction. Results for two symbolic regression problems show that CovSel improves the predictions for various state-of-the-art fusion methods for ensembles composed of independently evolved GP models and also beats approaches based on single GP models. Dominik Sobania, Franz Rothlauf |
GECCO | 1 |