VLDB 2026 Research / reviewers in the wild / expert
Lenz Belzner
dblp:136/1485
· DBLP profile ↗
21ranked-venue papers
4as first author
10since 2021 · last 2025
0009-0002-4683-5460ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Evolution of Criticality in Deep Reinforcement Learning
Chidvilas Karpenahalli Ramakrishna, Adithya Mohan, Zahra Zeinaly, Lenz Belzner |
ICAART (3) | 4 |
| 2024 | Detecting Influence Structures in Multi-Agent Reinforcement LearningabstractWe consider the problem of quantifying the amount of influence one agent can exert on another in the setting of multi-agent reinforcement learning (MARL). As a step towards a unified approach to express agents’ interdependencies, we introduce the total and state influence measurement functions. Both of these are valid for all common MARL systems, such as the discounted reward setting. Additionally, we propose novel quantities, called the total impact measurement (TIM) and state impact measurement (SIM), that characterize one agent’s influence on another by the maximum impact it can have on the other agents’ expected returns and represent instances of impact measurement functions in the average reward setting. Furthermore, we provide approximation algorithms for TIM and SIM with simultaneously learning approximations of agents’ expected returns, error bounds, stability analyses under changes of the policies, and convergence guarantees. The approximation algorithm relies only on observing other agents’ actions and is, other than that, fully decentralized. Through empirical studies, we validate our approach’s effectiveness in identifying intricate influence structures in complex interactions. Our work appears to be the first study of determining influence structures in the multi-agent average reward setting with convergence guarantees. Fabian R. Pieroth, Katherine E. Fitch, Lenz Belzner |
ICML | 3 |
| 2024 | Emergent cooperation from mutual acknowledgment exchange in multi-agent reinforcement learningabstractAbstract Peer incentivization (PI) is a recent approach where all agents learn to reward or penalize each other in a distributed fashion, which often leads to emergent cooperation. Current PI mechanisms implicitly assume a flawless communication channel in order to exchange rewards. These rewards are directly incorporated into the learning process without any chance to respond with feedback. Furthermore, most PI approaches rely on global information, which limits scalability and applicability to real-world scenarios where only local information is accessible. In this paper, we propose Mutual Acknowledgment Token Exchange (MATE), a PI approach defined by a two-phase communication protocol to exchange acknowledgment tokens as incentives to shape individual rewards mutually. All agents condition their token transmissions on the locally estimated quality of their own situations based on environmental rewards and received tokens. MATE is completely decentralized and only requires local communication and information. We evaluate MATE in three social dilemma domains. Our results show that MATE is able to achieve and maintain significantly higher levels of cooperation than previous PI approaches. In addition, we evaluate the robustness of MATE in more realistic scenarios, where agents can deviate from the protocol and communication failures can occur. We also evaluate the sensitivity of MATE w.r.t. the choice of token values. Thomy Phan, Felix Sommer, Fabian Ritz, Philipp Altmann, Jonas Nüßlein, Michael Kölle 0001, Lenz Belzner, Claudia Linnhoff-Popien |
Auton. Agents Multi Agent Syst. | 7 |
| 2022 | Self-Replication in Neural NetworksabstractA key element of biological structures is self-replication. Neural networks are the prime structure used for the emergent construction of complex behavior in computers. We analyze how various network types lend themselves to self-replication. Backpropagation turns out to be the natural way to navigate the space of network weights and allows non-trivial self-replicators to arise naturally. We perform an in-depth analysis to show the self-replicators' robustness to noise. We then introduce artificial chemistry environments consisting of several neural networks and examine their emergent behavior. In extension to this work's previous version (Gabor et al., 2019), we provide an extensive analysis of the occurrence of fixpoint weight configurations within the weight space and an approximation of their respective attractor basins. Thomas Gabor, Steffen Illium, Maximilian Zorn, Cristian Lenta, Andy Mattausch, Lenz Belzner, Claudia Linnhoff-Popien |
Artif. Life | 6 |
| 2021 | Resilient Multi-Agent Reinforcement Learning with Adversarial Value DecompositionabstractWe focus on resilience in cooperative multi-agent systems, where agents can change their behavior due to udpates or failures of hardware and software components. Current state-of-the-art approaches to cooperative multi-agent reinforcement learning (MARL) have either focused on idealized settings without any changes or on very specialized scenarios, where the number of changing agents is fixed, e.g., in extreme cases with only one productive agent. Therefore, we propose Resilient Adversarial value Decomposition with Antagonist-Ratios (RADAR). RADAR offers a value decomposition scheme to train competing teams of varying size for improved resilience against arbitrary agent changes. We evaluate RADAR in two cooperative multi-agent domains and show that RADAR achieves better worst case performance w.r.t. arbitrary agent changes than state-of-the-art MARL. Thomy Phan, Lenz Belzner, Thomas Gabor, Andreas Sedlmeier, Fabian Ritz, Claudia Linnhoff-Popien |
AAAI | 2 |
| 2021 | DYME: A Dynamic Metric for Dialog Modeling Learned from Human Conversations
Florian von Unold, Monika Wintergerst, Lenz Belzner, Georg Groh |
ICONIP (5) | 3 |
| 2021 | Stochastic Market GamesabstractSome of the most relevant future applications of multi-agent systems like autonomous driving or factories as a service display mixed-motive scenarios, where agents might have conflicting goals. In these settings agents are likely to learn undesirable outcomes in terms of cooperation under independent learning, such as overly greedy behavior. Motivated from real world societies, in this work we propose to utilize market forces to provide incentives for agents to become cooperative. As demonstrated in an iterated version of the Prisoner's Dilemma, the proposed market formulation can change the dynamics of the game to consistently learn cooperative policies. Further we evaluate our approach in spatially and temporally extended settings for varying numbers of agents. We empirically find that the presence of markets can improve both the overall result and agent individual returns via their trading activities. Kyrill Schmid, Lenz Belzner, Robert Müller 0005, Johannes Tochtermann, Claudia Linnhoff-Popien |
IJCAI | 2 |
| 2021 | Distributed Emergent Agreements with Deep Reinforcement LearningabstractBuilding autonomous agents that are capable to cooperate with other machines is an essential step towards large scale application of AI systems. Especially systems comprised of multiple self-interested agents with general sum returns can profit from cooperative behavior as cooperation can help to increase the return from all agents simultaneously. A critical aspect that might undermine cooperation is given if agents cannot make credible threats or promises (called commitment problems). Inspired by this idea in this work we augment deep reinforcement learning agents with the capability to build agreements with one another, thereby enabling agents to autonomously learn at which time to cooperate with other agents. This approach, called distributed emergent agreement learning (DEAL), enables agents to commit to specific policies defined by the agreement. We evaluate DEAL with up to 16 agents, represented as Deep Q-Networks or instances of Proximal Policy Optimization in a factory domain and empirically show that agreements increase cooperation by improving both overall and agent individual returns. Kyrill Schmid, Robert Müller 0005, Lenz Belzner, Johannes Tochtermann, Claudia Linnhoff-Popien |
IJCNN | 3 |
| 2021 | VAST: Value Function Factorization with Variable Agent Sub-TeamsabstractValue function factorization (VFF) is a popular approach to cooperative multi-agent reinforcement learning in order to learn local value functions from global rewards. However, state-of-the-art VFF is limited to a handful of agents in most domains. We hypothesize that this is due to the flat factorization scheme, where the VFF operator becomes a performance bottleneck with an increasing number of agents. Therefore, we propose VFF with variable agent sub-teams (VAST). VAST approximates a factorization for sub-teams which can be defined in an arbitrary way and vary over time, e.g., to adapt to different situations. The sub-team values are then linearly decomposed for all sub-team members. Thus, VAST can learn on a more focused and compact input representation of the original VFF operator. We evaluate VAST in three multi-agent domains and show that VAST can significantly outperform state-of-the-art VFF, when the number of agents is sufficiently large. Thomy Phan, Fabian Ritz, Lenz Belzner, Philipp Altmann, Thomas Gabor, Claudia Linnhoff-Popien |
NeurIPS | 3 |
| 2021 | Synthesizing safe policies under probabilistic constraints with reinforcement learning and Bayesian model checking
Lenz Belzner, Martin Wirsing |
Sci. Comput. Program. | 1 |
| 2020 | Multi-agent Reinforcement Learning for Bargaining under Risk and Asymmetric Information
Kyrill Schmid, Lenz Belzner, Thomy Phan, Thomas Gabor, Claudia Linnhoff-Popien |
ICAART (1) | 2 |
| 2020 | Uncertainty-based Out-of-Distribution Classification in Deep Reinforcement LearningabstractRobustness to out-of-distribution (OOD) data is an important goal in building reliable machine learning systems. Especially in autonomous systems, wrong predictions for OOD inputs can cause safety critical situations. As a first step towards a solution, we consider the problem of detecting such data in a value-based deep reinforcement learning (RL) setting. Modelling this problem as a one-class classification problem, we propose a framework for uncertainty-based OOD classification: UBOOD. It is based on the effect that an agent's epistemic uncertainty is reduced for situations encountered during training (in-distribution), and thus lower than for unencountered (OOD) situations. Being agnostic towards the approach used for estimating epistemic uncertainty, combinations with different uncertainty estimation methods, e.g. approximate Bayesian inference methods or ensembling techniques are possible. We further present a first viable solution for calculating a dynamic classification threshold, based on the uncertainty distribution of the training data. Evaluation shows that the framework produces reliable classification results when combined with ensemble-based estimators, while the combination with concrete dropout-based estimators fails to reliably detect OOD situations. In summary, UBOOD presents a viable approach for OOD classification in deep RL settings by leveraging the epistemic uncertainty of the agent's value function. Andreas Sedlmeier, Thomas Gabor, Thomy Phan, Lenz Belzner, Claudia Linnhoff-Popien |
ICAART (2) | 4 |
| 2020 | The scenario coevolution paradigm: adaptive quality assurance for adaptive systemsabstractAbstract Systems are becoming increasingly more adaptive, using techniques like machine learning to enhance their behavior on their own rather than only through human developers programming them. We analyze the impact the advent of these new techniques has on the discipline of rigorous software engineering, especially on the issue of quality assurance. To this end, we provide a general description of the processes related to machine learning and embed them into a formal framework for the analysis of adaptivity, recognizing that to test an adaptive system a new approach to adaptive testing is necessary. We introduce scenario coevolution as a design pattern describing how system and test can work as antagonists in the process of software evolution. While the general pattern applies to large-scale processes (including human developers further augmenting the system), we show all techniques on a smaller-scale example of an agent navigating a simple smart factory. We point out new aspects in software engineering for adaptive systems that may be tackled naturally using scenario coevolution. This work is a substantially extended take on Gabor et al. (International symposium on leveraging applications of formal methods, Springer, pp 137–154, 2018). Thomas Gabor, Andreas Sedlmeier, Thomy Phan, Fabian Ritz, Marie Kiermeier, Lenz Belzner, Bernhard Kempter, Cornel Klein, Horst Sauer, Reiner N. Schmid, Jan Wieghardt, Marc Zeller, Claudia Linnhoff-Popien |
Int. J. Softw. Tools Technol. Transf. | 6 |
| 2019 | Memory Bounded Open-Loop Planning in Large POMDPs Using Thompson SamplingabstractState-of-the-art approaches to partially observable planning like POMCP are based on stochastic tree search. While these approaches are computationally efficient, they may still construct search trees of considerable size, which could limit the performance due to restricted memory resources. In this paper, we propose Partially Observable Stacked Thompson Sampling (POSTS), a memory bounded approach to openloop planning in large POMDPs, which optimizes a fixed size stack of Thompson Sampling bandits. We empirically evaluate POSTS in four large benchmark problems and compare its performance with different tree-based approaches. We show that POSTS achieves competitive performance compared to tree-based open-loop planning and offers a performancememory tradeoff, making it suitable for partially observable planning with highly restricted computational and memory resources. Thomy Phan, Lenz Belzner, Marie Kiermeier, Markus Friedrich 0001, Kyrill Schmid, Claudia Linnhoff-Popien |
AAAI | 2 |
| 2019 | Bayesian Surprise in Indoor EnvironmentsabstractThis paper proposes a novel method to identify unexpected structures in 2D floor plans using the concept of Bayesian Surprise. Taking into account that a person's expectation is an important aspect of the perception of space, we exploit the theory of Bayesian Surprise to robustly model expectation and thus surprise in the context of building structures. We use Isovist Analysis, which is a popular space syntax technique, to turn qualitative object attributes into quantitative environmental information. Since isovists are location-specific patterns of visibility, a sequence of isovists describes the spatial perception during a movement along multiple points in space. We then use Bayesian Surprise in a feature space consisting of these isovist readings. To demonstrate the suitability of our approach, we take "snapshots" of an agent's local environment to provide a short list of images that characterize a traversed trajectory through a 2D indoor environment. Those fingerprints represent surprising regions of a tour, characterize the traversed map and enable indoor LBS to focus more on important regions. Given this idea, we propose to use surprise as a new dimension of context in indoor location-based services (LBS). Agents of LBS, such as mobile robots or non-player characters in computer games, may use the context "surprise" to focus more on important regions of a map for a better use or understanding of the floor plan. Sebastian Feld, Andreas Sedlmeier, Markus Friedrich 0001, Jan Franz, Lenz Belzner |
SIGSPATIAL/GIS | 5 |
| 2018 | Inheritance-based diversity measures for explicit convergence control in evolutionary algorithmsabstractDiversity is an important factor in evolutionary algorithms to prevent premature convergence towards a single local optimum. In order to maintain diversity throughout the process of evolution, various means exist in literature. We analyze approaches to diversity that (a) have an explicit and quantifiable influence on fitness at the individual level and (b) require no (or very little) additional domain knowledge such as domain-specific distance functions. We also introduce the concept of genealogical diversity in a broader study. We show that employing these approaches can help evolutionary algorithms for global optimization in many cases. Thomas Gabor, Lenz Belzner, Claudia Linnhoff-Popien |
GECCO | 2 |
| 2018 | Trajectory annotation using sequences of spatial perceptionabstractIn the near future, more and more machines will perform tasks in the vicinity of human spaces or support them directly in their spatially bound activities. In order to simplify the verbal communication and the interaction between robotic units and/or humans, reliable and robust systems w.r.t. noise and processing results are needed. This work builds a foundation to address this task. By using a continuous representation of spatial perception in interiors learned from trajectory data, our approach clusters movement in dependency to its spatial context. We propose an unsupervised learning approach based on a neural autoencoding that learns semantically meaningful continuous encodings of spatio-temporal trajectory data. This learned encoding can be used to form prototypical representations. We present promising results that clear the path for future applications. Sebastian Feld, Steffen Illium, Andreas Sedlmeier, Lenz Belzner |
SIGSPATIAL/GIS | 4 |
| 2018 | Action Markets in Deep Multi-Agent Reinforcement Learning
Kyrill Schmid, Lenz Belzner, Thomas Gabor, Thomy Phan |
ICANN (2) | 2 |
| 2018 | The Sharer's Dilemma in Collective Adaptive Systems of Self-interested Agents
Lenz Belzner, Kyrill Schmid, Thomy Phan, Thomas Gabor, Martin Wirsing |
ISoLA (3) | 1 |
| 2014 | Verifiable Decisions in Autonomous Concurrent Systems
Lenz Belzner |
COORDINATION | 1 |
| 2013 | Action Programming In Rewriting Logic
Lenz Belzner |
Theory Pract. Log. Program. | 1 |