Thomas Gabor

dblp:166/5806 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0003-2048-8667ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 5 first-author · 19 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scalable quantum Trotterised-vs-continuous annealing for pseudo-Boolean multi-objective optimisation
Zakaria Abd El Moiz Dahi, Francisco Chicano, Gabriel Luque, Thomas Gabor
Future Gener. Comput. Syst.4
2025 Evaluating Mutation Techniques in Genetic-Algorithm-Based Quantum Circuit Synthesis
abstract
Quantum computing leverages the unique properties of qubits and quantum parallelism to solve problems intractable for classical systems, offering unparalleled computational potential. However, optimization of quantum circuits remains critical, especially for noisy intermediate-scale quantum (NISQ) devices with limited qubits and high error rates. Genetic algorithms (GAs) provide a promising approach for efficient quantum circuit synthesis by automating optimization tasks. This work examines the impact of various mutation strategies within a GA framework for quantum circuit synthesis. By analyzing how different mutations transform circuits, it identifies strategies that enhance efficiency and performance. Experiments utilized a fitness function emphasizing fidelity, while accounting for circuit depth and T-operations, to optimize circuits with four to six qubits. Our analysis revealed that, while the "swap, addition" strategy achieved the highest fidelity scores, it consistently increased circuit depth. In contrast, combining "swap, addition, delete" strategies offers a more balanced approach, delivering near-optimal results while also having the potential of reducing circuit depth.
Michael Kölle 0001, Tom Bintener, Maximilian Zorn, Gerhard Stenzel, Leo Sünkel, Thomas Gabor, Claudia Linnhoff-Popien
GECCO6
2025 Quantum Circuit Construction and Optimization through Hybrid Evolutionary Algorithms
abstract
We apply a hybrid evolutionary algorithm to minimize the depth of circuits in quantum computing. More specifically, we evaluate two different variants of the algorithm. In the first approach, we combine the evolutionary algorithm with an optimization subroutine to optimize the parameters of the rotation gates present in the quantum circuit. In the second, the algorithm solely relies on evolutionary operations (i.e., mutations and crossover). We approach the problem from two sides: (1) constructing circuits from the ground up by starting with random initializations and (2) initializing individuals with a target circuit in order to optimize it further according to the fitness function. We run experiments on random circuits with 4 and 6 qubits varying in circuit depth. Our results show that the proposed methods are able to significantly reduce the depth of circuits while still retaining a high fidelity to the target state.
Leo Sünkel, Philipp Altmann, Michael Kölle 0001, Gerhard Stenzel, Thomas Gabor, Claudia Linnhoff-Popien
GECCO5
2025 Optimizing Sensor Redundancy in Sequential Decision-Making Problems
Jonas Nüßlein, Maximilian Zorn, Fabian Ritz, Jonas Stein 0001, Gerhard Stenzel, Julian Schönberger, Thomas Gabor, Claudia Linnhoff-Popien
ICAART (1)7
2025 QMamba: Quantum Selective State Space Models for Text Generation
abstract
This book contains the proceedings of the 17th International Conference on Agents and Artificial Intelligence. This year, ICAART is held in Porto, Portugal, on February 23-25, 2025. As usual it is sponsored by the Institute for Systems and Technologies of Information, Control and Communication (INSTICC). ICAART 2025 was also organized in cooperation with other members of our AI family. We mention the ACM Special Interest Group on Artificial Intelligence, the Association for the Advancement of Artificial Intelligence, the Associação Portuguesa de Reconhecimento de Padrões, the Portuguese Association for Artificial Intelligence, the IberoAmerican Society of Artificial Intelligence and the European Society for Fuzzy Logic and Technology. The purpose of the International Conference on Agents and Artificial Intelligence is to bring together researchers, engineers and practitioners interested in the theory and applications in the areas of Agents and Artificial Intelligence, covering both applications and current (advanced) research work. On one side it focuses on Agents, Multi-Agent Systems and Software Platforms, and also Distributed Problem Solving. On the other side it focuses on Artificial Intelligence, Knowledge Representation, Planning, Learning, Scheduling, Perception. Applications are in both areas. They are using Natural Language Processing (NLP), Large Language Models (LLMs), Legal Technologies and Quantum Computing. In the last four years the research emphasis has shifted towards Explainable AI and Interpretable AI with a focus on trustworthiness, fairness, privacy, safety, security and ethical issues. A substantial amount of research work is ongoing in these knowledge areas, in an attempt to discover appropriate theories and paradigms for use in real-world applications. ICAART 2025 received 472 paper submissions from 53 countries of which 23.09% were accepted and published as full papers. A double-blind paper review was performed for each submission by at least 2 but usually 3 or more members of the International Program Committee, which is composed of established researchers and domain experts. The high quality of the ICAART 2025 program is enhanced by the keynote lecture delivered by distinguished speakers who are renowned experts in their fields: Inge Bryan (Chair of the Dutch Institute for Vulnerability Disclosure, Netherlands), Pavan Duggal (Advocate, Supreme Court of India, Chairman, International Commission on Cyber Security Law India, and Chief Executive, Artificial Intelligence Law Hub, India) and Paul Nemitz (Principal Adviser European Commission, Belgium). The conference is complemented by one workshop, two special sessions and one tutorial. They are: a Workshop on Quantum Artificial Intelligence and Optimization, chaired by Michael Kölle, a Special Session on Interpretable Artificial Intelligence Through Glass-Box Models, chaired by Mattias Wahde and a Special Session on Emotions and Affective Agents, chaired by Joaquin Taverner and Emilio Vivancos. Furthermore, a Tutorial on Self-Governing Systems will be given by Jeremy Pitt and Asimina Mertzani. All presented papers will be available at the SCITEPRESS Digital Library and will be submitted for evaluation for indexing by SCOPUS, Google Scholar, The DBLP Computer Science Bibliography, Semantic Scholar, Engineering Index and Web of Science / Conference Proceedings Citation Index. As recognition for the best contributions, several awards based on the combined marks of paper reviewing, as assessed by the Program Committee, and the quality of the presentation, as assessed by session chairs at the conference venue, are conferred at the closing session of the conference. Authors of selected papers will be invited to submit extended versions for inclusion in a forthcoming book of ICAART Selected Papers to be published by Springer, as part of the LNAI Series. Some papers will also be selected for publication of extended and revised versions in the special issue of the Springer Nature Computer Science Journal. The program for this conference required the dedicated effort of many people. Firstly, we must thank the authors, whose research efforts are herewith recorded. Next, we thank the members of the Program Committee and the auxiliary reviewers for their diligent and professional reviewing. We would also like to deeply thank the invited speakers for their invaluable contribution and for taking the time to prepare their talks. Finally, a word of appreciation for the hard work of the INSTICC team; organizing a conference of this level is a task that can only be achieved by the collaborative effort of a dedicated and highly competent team. We wish you all an exciting and inspiring conference. We hope to have contributed to the development of our research community, and we look forward to having additional research results presented at the next edition of ICAART, details of which are available at https://icaart.scitevents.org.
Gerhard Stenzel, Michael Kölle 0001, Tobias Rohe, Maximilian Balthasar Mansky, Jonas Nüßlein, Thomas Gabor
ICAART (1)6
2025 Qandle: Accelerating State Vector Simulation Using Gate-Matrix Caching and Circuit Splitting
abstract
To address the computational complexity associated with state-vector simulation for quantum circuits, we propose a combination of advanced techniques to accelerate circuit execution. Quantum gate matrix caching reduces the overhead of repeated applications of the Kronecker product when applying a gate matrix to the state vector by storing decomposed partial matrices for each gate. Circuit splitting divides the circuit into sub-circuits with fewer gates by constructing a dependency graph, enabling parallel or sequential execution on disjoint subsets of the state vector. These techniques are implemented using the PyTorch machine learning framework. We demonstrate the performance of our approach by comparing it to other PyTorch-compatible quantum state-vector simulators. Our implementation, named Qandle, is designed to seamlessly integrate with existing machine learning workflows, providing a user-friendly API and compatibility with the OpenQASM format. Qandle is an open-source project hosted on GitHub and PyPI.
Gerhard Stenzel, Sebastian Zielinski, Michael Kölle 0001, Philipp Altmann, Jonas Nüßlein, Thomas Gabor
ICAART (1)6
2025 Discriminative reward co-training
abstract
Abstract We propose discriminative reward co-training (DIRECT) as an extension to deep reinforcement learning algorithms. Building upon the concept of self-imitation learning (SIL), we introduce an imitation buffer to store beneficial trajectories generated by the policy, determined by their return. A discriminator network is trained concurrently to the policy to distinguish between trajectories generated by the current policy and beneficial trajectories generated by previous policies. The discriminator’s verdict is used to construct a reward signal for optimizing the policy. By interpolating prior experience, DIRECT is able to act as a reward surrogate, steering policy optimization toward more valuable regions of the reward landscape, thus, toward learning an optimal policy. In this article, we formally introduce the additional components, their intended purpose and parameterization, and define a unified training procedure. To reveal insights into the mechanics of the proposed architecture, we provide evaluations of the introduced hyperparameters. Further benchmark evaluations in various discrete and continuous control environments provide evidence that DIRECT is especially beneficial in environments possessing sparse rewards, hard exploration tasks, and shifting circumstances. Our results show that DIRECT outperforms state-of-the-art algorithms in those challenging scenarios by providing a surrogate reward to the policy and direct the optimization toward valuable areas.
Philipp Altmann, Fabian Ritz, Maximilian Zorn, Michael Kölle 0001, Thomy Phan, Thomas Gabor, Claudia Linnhoff-Popien
Neural Comput. Appl.6
2025 Correction: Discriminative reward co-training
Philipp Altmann, Fabian Ritz, Maximilian Zorn, Michael Kölle 0001, Thomy Phan, Thomas Gabor, Claudia Linnhoff-Popien
Neural Comput. Appl.6
2024 Quantum Federated Learning for Image Classification
Leo Sünkel, Philipp Altmann, Michael Kölle 0001, Thomas Gabor
ICAART (3)4
2024 REACT: Revealing Evolutionary Action Consequence Trajectories for Interpretable Reinforcement Learning
Philipp Altmann, Céline Davignon, Maximilian Zorn, Fabian Ritz, Claudia Linnhoff-Popien, Thomas Gabor
IJCCI6
2024 Finding Strong Lottery Ticket Networks with Genetic Algorithms
Philipp Altmann, Julian Schönberger, Maximilian Zorn, Thomas Gabor
IJCCI4
2024 Emergence in Multi-agent Systems: A Safety Perspective
Philipp Altmann, Julian Schönberger, Steffen Illium, Maximilian Zorn, Fabian Ritz, Tom Haider, Simon Burton 0001, Thomas Gabor
ISoLA (2)8
2023 Attention-Based Recurrence for Multi-Agent Reinforcement Learning under Stochastic Partial Observability
abstract
Stochastic partial observability poses a major challenge for decentralized coordination in multi-agent reinforcement learning but is largely neglected in state-of-the-art research due to a strong focus on state-based centralized training for decentralized execution (CTDE) and benchmarks that lack sufficient stochasticity like StarCraft Multi-Agent Challenge (SMAC). In this paper, we propose Attention-based Embeddings of Recurrence In multi-Agent Learning (AERIAL) to approximate value functions under stochastic partial observability. AERIAL replaces the true state with a learned representation of multi-agent recurrence, considering more accurate information about decentralized agent decisions than state-based CTDE. We then introduce MessySMAC, a modified version of SMAC with stochastic observations and higher variance in initial states, to provide a more general and configurable benchmark regarding stochastic partial observability. We evaluate AERIAL in Dec-Tiger as well as in a variety of SMAC and MessySMAC maps, and compare the results with state-based CTDE. Furthermore, we evaluate the robustness of AERIAL and state-based CTDE against various stochasticity configurations in MessySMAC.
Thomy Phan, Fabian Ritz, Philipp Altmann, Maximilian Zorn, Jonas Nüßlein, Michael Kölle 0001, Thomas Gabor, Claudia Linnhoff-Popien
ICML7
2023 The Effect of Penalty Factors of Constrained Hamiltonians on the Eigenspectrum in Quantum Annealing
abstract
Constrained optimization problems are usually translated to (naturally unconstrained) Ising formulations by introducing soft penalty terms for the previously hard constraints. In this work, we empirically demonstrate that assigning the appropriate weight to these penalty terms leads to an enlargement of the minimum spectral gap in the corresponding eigenspectrum, which also leads to a better solution quality on actual quantum annealing hardware. We apply machine learning methods to analyze the correlations of the penalty factors and the minimum spectral gap for six selected constrained optimization problems and show that regression using a neural network allows to predict the best penalty factors in our settings for various problem instances. Additionally, we observe that problem instances with a single global optimum are easier to optimize in contrast to ones with multiple global optima.
Christoph Roch, Daniel Ratke, Jonas Nüßlein, Thomas Gabor, Sebastian Feld
ACM Trans. Quantum Comput.4
2022 Empirical Analysis of Limits for Memory Distance in Recurrent Neural Networks
abstract
Common to all different kinds of recurrent neural networks (RNNs) is the intention to model relations between data points through time. When there is no immediate relationship between subsequent data points (like when the data points are generated at random, e.g.), we show that RNNs are still able to remember a few data points back into the sequence by memorizing them by heart using standard backpropagation. However, we also show that for classical RNNs, LSTM and GRU networks the distance of data points between recurrent calls that can be reproduced this way is highly limited (compared to even a loose connection between data points) and subject to various constraints imposed by the type and size of the RNN in question. This implies the existence of a hard limit (way below the information-theoretic one) for the distance between related data points within which RNNs are still able to recognize said relation.
Steffen Illium, Thore Schillman, Robert Müller 0005, Thomas Gabor, Claudia Linnhoff-Popien
ICAART (3)4
2022 Solving Large Steiner Tree Problems in Graphs for Cost-efficient Fiber-To-The-Home Network Expansion
abstract
The expansion of Fiber-To-The-Home (FTTH) networks creates high costs due to expensive excavation procedures. Optimizing the planning process and minimizing the cost of the earth excavation work therefore lead to large savings. Mathematically, the FTTH network problem can be described as a minimum Steiner Tree problem. Even though the Steiner Tree problem has already been investigated intensively in the last decades, it might be further optimized with the help of new computing paradigms and emerging approaches. This work studies upcoming technologies, such as Quantum Annealing, Simulated Annealing and nature-inspired methods like Evolutionary Algorithms or slime-mold-based optimization. Additionally, we investigate partitioning and simplifying methods. Evaluated on several real-life problem instances, we could outperform a traditional, widely-used baseline (NetworkX Approximate Solver) on most of the domains. Prior partitioning of the initial graph and the presented slime-mold-based approach were especially valuable for a cost-efficient approximation. Quantum Annealing seems promising, but was limited by the number of available qubits.
Kyrill Schmid, Daniëlle Schuman, Thomas Gabor, Markus Friedrich 0001, Marc Geitz
ICAART (3)4
2022 Case-Based Inverse Reinforcement Learning Using Temporal Coherence
Jonas Nüßlein, Steffen Illium, Robert Müller 0005, Thomas Gabor, Claudia Linnhoff-Popien
ICCBR4
2022 Capturing Dependencies Within Machine Learning via a Formal Process Model
Fabian Ritz, Thomy Phan, Andreas Sedlmeier, Philipp Altmann, Jan Wieghardt, Reiner N. Schmid, Horst Sauer, Cornel Klein, Claudia Linnhoff-Popien, Thomas Gabor
ISoLA (3)10
2022 How to Approximate any Objective Function via Quadratic Unconstrained Binary Optimization
abstract
Quadratic unconstrained binary optimization (QUBO) has become the standard format for optimization using quantum computers, i.e., for both the quantum approximate optimization algorithm (QAOA) and quantum annealing (QA). We present a toolkit of methods to transform almost arbitrary problems to QUBO by (i) approximating them as a polynomial and then (ii) translating any polynomial to QUBO. We showcase the usage of our approaches on two example problems (ratio cut and logistic regression).
Thomas Gabor, Marian Lingsch Rosenfeld, Claudia Linnhoff-Popien, Sebastian Feld
SANER1
2022 Self-Replication in Neural Networks
abstract
A key element of biological structures is self-replication. Neural networks are the prime structure used for the emergent construction of complex behavior in computers. We analyze how various network types lend themselves to self-replication. Backpropagation turns out to be the natural way to navigate the space of network weights and allows non-trivial self-replicators to arise naturally. We perform an in-depth analysis to show the self-replicators' robustness to noise. We then introduce artificial chemistry environments consisting of several neural networks and examine their emergent behavior. In extension to this work's previous version (Gabor et al., 2019), we provide an extensive analysis of the occurrence of fixpoint weight configurations within the weight space and an approximation of their respective attractor basins.
Thomas Gabor, Steffen Illium, Maximilian Zorn, Cristian Lenta, Andy Mattausch, Lenz Belzner, Claudia Linnhoff-Popien
Artif. Life1
2021 Resilient Multi-Agent Reinforcement Learning with Adversarial Value Decomposition
abstract
We focus on resilience in cooperative multi-agent systems, where agents can change their behavior due to udpates or failures of hardware and software components. Current state-of-the-art approaches to cooperative multi-agent reinforcement learning (MARL) have either focused on idealized settings without any changes or on very specialized scenarios, where the number of changing agents is fixed, e.g., in extreme cases with only one productive agent. Therefore, we propose Resilient Adversarial value Decomposition with Antagonist-Ratios (RADAR). RADAR offers a value decomposition scheme to train competing teams of varying size for improved resilience against arbitrary agent changes. We evaluate RADAR in two cooperative multi-agent domains and show that RADAR achieves better worst case performance w.r.t. arbitrary agent changes than state-of-the-art MARL.
Thomy Phan, Lenz Belzner, Thomas Gabor, Andreas Sedlmeier, Fabian Ritz, Claudia Linnhoff-Popien
AAAI3
2021 SAT-MARL: Specification Aware Training in Multi-Agent Reinforcement Learning
abstract
A characteristic of reinforcement learning is the ability to develop unforeseen strategies when solving problems. While such strategies sometimes yield superior performance, they may also result in undesired or even dangerous behavior. In industrial scenarios, a system's behavior also needs to be predictable and lie within defined ranges. To enable the agents to learn (how) to align with a given specification, this paper proposes to explicitly transfer functional and non-functional requirements into shaped rewards. Experiments are carried out on the smart factory, a multi-agent environment modeling an industrial lot-size-one production facility, with up to eight agents and different multi-agent reinforcement learning algorithms. Results indicate that compliance with functional and non-functional constraints can be achieved by the proposed approach.
Fabian Ritz, Thomy Phan, Robert Müller 0005, Thomas Gabor, Andreas Sedlmeier, Marc Zeller, Jan Wieghardt, Reiner N. Schmid, Horst Sauer, Cornel Klein, Claudia Linnhoff-Popien
ICAART (1)4
2021 VAST: Value Function Factorization with Variable Agent Sub-Teams
abstract
Value function factorization (VFF) is a popular approach to cooperative multi-agent reinforcement learning in order to learn local value functions from global rewards. However, state-of-the-art VFF is limited to a handful of agents in most domains. We hypothesize that this is due to the flat factorization scheme, where the VFF operator becomes a performance bottleneck with an increasing number of agents. Therefore, we propose VFF with variable agent sub-teams (VAST). VAST approximates a factorization for sub-teams which can be defined in an arbitrary way and vary over time, e.g., to adapt to different situations. The sub-team values are then linearly decomposed for all sub-team members. Thus, VAST can learn on a more focused and compact input representation of the original VFF operator. We evaluate VAST in three multi-agent domains and show that VAST can significantly outperform state-of-the-art VFF, when the number of agents is sufficiently large.
Thomy Phan, Fabian Ritz, Lenz Belzner, Philipp Altmann, Thomas Gabor, Claudia Linnhoff-Popien
NeurIPS5
2021 Productive fitness in diversity-aware evolutionary algorithms
abstract
Abstract In evolutionary algorithms, the notion of diversity has been adopted from biology and is used to describe the distribution of a population of solution candidates. While it has been known that maintaining a reasonable amount of diversity often benefits the overall result of the evolutionary optimization process by adjusting the exploration/exploitation trade-off, little has been known about what diversity is optimal. We introduce the notion of productive fitness based on the effect that a specific solution candidate has some generations down the evolutionary path. We derive the notion of final productive fitness, which is the ideal target fitness for any evolutionary process. Although it is inefficient to compute, we show empirically that it allows for ana posteriorianalysis of how well a given evolutionary optimization process hit the ideal exploration/exploitation trade-off, providing insight intowhydiversity-aware evolutionary optimization often performs better.
Thomas Gabor, Thomy Phan, Claudia Linnhoff-Popien
Nat. Comput.1
2020 Approximate approximation on a quantum annealer
abstract
Many problems of industrial interest are NP-complete, and quickly exhaust resources of computational devices with increasing input sizes. Quantum annealers (QA) are physical devices that aim at this class of problems by exploiting quantum mechanical properties of nature. However, they compete with efficient heuristics and probabilistic or randomised algorithms on classical machines that allow for finding approximate solutions to large NP-complete problems.
Irmi Sax, Sebastian Feld, Sebastian Zielinski, Thomas Gabor, Claudia Linnhoff-Popien, Wolfgang Mauerer
CF4
2020 Approximating Archetypal Analysis Using Quantum Annealing
Sebastian Feld, Christoph Roch, Katja Geirhos, Thomas Gabor
ESANN4
2020 Nash Equilibria in Multi-Agent Swarms
Carsten Hahn, Thomy Phan, Sebastian Feld, Christoph Roch, Fabian Ritz, Andreas Sedlmeier, Thomas Gabor, Claudia Linnhoff-Popien
ICAART (1)7
2020 Multi-agent Reinforcement Learning for Bargaining under Risk and Asymmetric Information
Kyrill Schmid, Lenz Belzner, Thomy Phan, Thomas Gabor, Claudia Linnhoff-Popien
ICAART (1)4
2020 Uncertainty-based Out-of-Distribution Classification in Deep Reinforcement Learning
abstract
Robustness to out-of-distribution (OOD) data is an important goal in building reliable machine learning systems. Especially in autonomous systems, wrong predictions for OOD inputs can cause safety critical situations. As a first step towards a solution, we consider the problem of detecting such data in a value-based deep reinforcement learning (RL) setting. Modelling this problem as a one-class classification problem, we propose a framework for uncertainty-based OOD classification: UBOOD. It is based on the effect that an agent's epistemic uncertainty is reduced for situations encountered during training (in-distribution), and thus lower than for unencountered (OOD) situations. Being agnostic towards the approach used for estimating epistemic uncertainty, combinations with different uncertainty estimation methods, e.g. approximate Bayesian inference methods or ensembling techniques are possible. We further present a first viable solution for calculating a dynamic classification threshold, based on the uncertainty distribution of the training data. Evaluation shows that the framework produces reliable classification results when combined with ensemble-based estimators, while the combination with concrete dropout-based estimators fails to reliably detect OOD situations. In summary, UBOOD presents a viable approach for OOD classification in deep RL settings by leveraging the epistemic uncertainty of the agent's value function.
Andreas Sedlmeier, Thomas Gabor, Thomy Phan, Lenz Belzner, Claudia Linnhoff-Popien
ICAART (2)2
2020 A Formal Model for Reasoning About the Ideal Fitness in Evolutionary Processes
Thomas Gabor, Claudia Linnhoff-Popien
ISoLA (2)1
2020 The scenario coevolution paradigm: adaptive quality assurance for adaptive systems
abstract
Abstract Systems are becoming increasingly more adaptive, using techniques like machine learning to enhance their behavior on their own rather than only through human developers programming them. We analyze the impact the advent of these new techniques has on the discipline of rigorous software engineering, especially on the issue of quality assurance. To this end, we provide a general description of the processes related to machine learning and embed them into a formal framework for the analysis of adaptivity, recognizing that to test an adaptive system a new approach to adaptive testing is necessary. We introduce scenario coevolution as a design pattern describing how system and test can work as antagonists in the process of software evolution. While the general pattern applies to large-scale processes (including human developers further augmenting the system), we show all techniques on a smaller-scale example of an agent navigating a simple smart factory. We point out new aspects in software engineering for adaptive systems that may be tackled naturally using scenario coevolution. This work is a substantially extended take on Gabor et al. (International symposium on leveraging applications of formal methods, Springer, pp 137–154, 2018).
Thomas Gabor, Andreas Sedlmeier, Thomy Phan, Fabian Ritz, Marie Kiermeier, Lenz Belzner, Bernhard Kempter, Cornel Klein, Horst Sauer, Reiner N. Schmid, Jan Wieghardt, Marc Zeller, Claudia Linnhoff-Popien
Int. J. Softw. Tools Technol. Transf.1
2019 Optimizing evolutionary CSG tree extraction
abstract
The extraction of 3D models represented by Constructive Solid Geometry (CSG) trees from point clouds is a common problem in reverse engineering pipelines as used by Computer Aided Design (CAD) tools. We propose three independent enhancements on state-of-the-art Genetic Algorithms (GAs) for CSG tree extraction: (1) A deterministic point cloud filtering mechanism that significantly reduces the computational effort of objective function evaluations without loss of geometric precision, (2) a graph-based partitioning scheme that divides the problem domain in smaller parts that can be solved separately and thus in parallel and (3) a 2-level improvement procedure that combines a recursive CSG tree redundancy removal technique with a local search heuristic, which significantly improves GA running times. We show in an extensive evaluation that our optimized GA-based approach provides faster running times and scales better with problem size compared to state-of-the-art GA-based approaches.
Markus Friedrich 0001, Pierre-Alain Fayolle, Thomas Gabor, Claudia Linnhoff-Popien
GECCO3
2019 Scenario co-evolution for reinforcement learning on a grid world smart factory domain
abstract
Adversarial learning has been established as a successful paradigm in reinforcement learning. We propose a hybrid adversarial learner where a reinforcement learning agent tries to solve a problem while an evolutionary algorithm tries to find problem instances that are hard to solve for the current expertise of the agent, causing the intelligent agent to co-evolve with a set of test instances or scenarios. We apply this setup, called scenario co-evolution, to a simulated smart factory problem that combines task scheduling with navigation of a grid world. We show that the so trained agent outperforms conventional reinforcement learning. We also show that the scenarios evolved this way can provide useful test cases for the evaluation of any (however trained) agent.
Thomas Gabor, Andreas Sedlmeier, Marie Kiermeier, Thomy Phan, Marcel Henrich, Monika Pichlmair, Bernhard Kempter, Cornel Klein, Horst Sauer, Reiner N. Schmid, Jan Wieghardt
GECCO1
2019 Subgoal-Based Temporal Abstraction in Monte-Carlo Tree Search
abstract
We propose an approach to general subgoal-based temporal abstraction in MCTS. Our approach approximates a set of available macro-actions locally for each state only requiring a generative model and a subgoal predicate. For that, we modify the expansion step of MCTS to automatically discover and optimize macro-actions that lead to subgoals. We empirically evaluate the effectiveness, computational efficiency and robustness of our approach w.r.t. different parameter settings in two benchmark domains and compare the results to standard MCTS without temporal abstraction.
Thomas Gabor, Jan Peter, Thomy Phan, Claudia Linnhoff-Popien
IJCAI1
2019 Adaptive Thompson Sampling Stacks for Memory Bounded Open-Loop Planning
abstract
We propose Stable Yet Memory Bounded Open-Loop (SYMBOL) planning, a general memory bounded approach to partially observable open-loop planning. SYMBOL maintains an adaptive stack of Thompson Sampling bandits, whose size is bounded by the planning horizon and can be automatically adapted according to the underlying domain without any prior domain knowledge beyond a generative model. We empirically test SYMBOL in four large POMDP benchmark problems to demonstrate its effectiveness and robustness w.r.t. the choice of hyperparameters and evaluate its adaptive memory consumption. We also compare its performance with other open-loop planning algorithms and POMCP.
Thomy Phan, Thomas Gabor, Robert Müller 0005, Christoph Roch, Claudia Linnhoff-Popien
IJCAI2
2018 Inheritance-based diversity measures for explicit convergence control in evolutionary algorithms
abstract
Diversity is an important factor in evolutionary algorithms to prevent premature convergence towards a single local optimum. In order to maintain diversity throughout the process of evolution, various means exist in literature. We analyze approaches to diversity that (a) have an explicit and quantifiable influence on fitness at the individual level and (b) require no (or very little) additional domain knowledge such as domain-specific distance functions. We also introduce the concept of genealogical diversity in a broader study. We show that employing these approaches can help evolutionary algorithms for global optimization in many cases.
Thomas Gabor, Lenz Belzner, Claudia Linnhoff-Popien
GECCO1
2018 Action Markets in Deep Multi-Agent Reinforcement Learning
Kyrill Schmid, Lenz Belzner, Thomas Gabor, Thomy Phan
ICANN (2)3
2018 The Sharer's Dilemma in Collective Adaptive Systems of Self-interested Agents
Lenz Belzner, Kyrill Schmid, Thomy Phan, Thomas Gabor, Martin Wirsing
ISoLA (3)4
2018 Adapting Quality Assurance to Adaptive Systems: The Scenario Coevolution Paradigm
Thomas Gabor, Marie Kiermeier, Andreas Sedlmeier, Bernhard Kempter, Cornel Klein, Horst Sauer, Reiner N. Schmid, Jan Wieghardt
ISoLA (3)1
2018 Mutation-Based Test Suite Evolution for Self-Organizing Systems
André Reichstaller, Thomas Gabor, Alexander Knapp
ISoLA (3)2