EDBT 2026 Demo / reviewers in the wild / expert
Daniel Hennes
dblp:50/4003
· DBLP profile ↗
26ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-3646-5286ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 2 first-author · 5 since 2021Systems, architecture and hardware · 6Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Planning, search and constraint satisfaction · 33% Reinforcement learning · 26% Multi-agent systems · 11% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 27 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
2.3 | 4 | 2025 | Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025 NeuPL: Neural Population Learning · ICLR 2022 Fast computation of Nash Equilibria in Imperfect Information Games · ICML 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
2.0 | 3 | 2025 | Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025 Mastering Board Games by External and Internal Planning with Language Models · ICML 2025 Interplanetary Trajectory Planning with Monte Carlo Tree Search · IJCAI 2015 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning |
1.0 | 2 | 2022 | NeuPL: Neural Population Learning · ICLR 2022 A Generalized Training Approach for Multiagent Learning · ICLR 2020 |
Machine learning › Generative modeling
generative model |
0.9 | 1 | 2025 | Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning |
0.9 | 1 | 2025 | Mastering Board Games by External and Internal Planning with Language Models · ICML 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
LLM-guided search |
0.9 | 1 | 2025 | Mastering Board Games by External and Internal Planning with Language Models · ICML 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent modeling |
0.9 | 1 | 2025 | Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent Modeling · IJCAI 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning |
0.9 | 1 | 2025 | Mastering Board Games by External and Internal Planning with Language Models · ICML 2025 |
Machine learning › Reinforcement learning
population-based learning |
0.6 | 1 | 2022 | NeuPL: Neural Population Learning · ICLR 2022 |
Knowledge, reasoning and agents › Multi-agent systems
imperfect information games |
0.4 | 1 | 2020 | Fast computation of Nash Equilibria in Imperfect Information Games · ICML 2020 |
Algorithmic game theory and mechanism design › equilibrium computation
nash equilibrium computation |
0.4 | 1 | 2020 | Fast computation of Nash Equilibria in Imperfect Information Games · ICML 2020 |
Robotics › Robot manipulation › tactile sensing › tactile perception › haptic exploration
active tactile exploration |
0.4 | 1 | 2019 | Active Multi-Contact Continuous Tactile Exploration with Gaussian Process Differential Entropy · ICRA 2019 |
Robotics › Robot manipulation › tactile sensing › tactile perception
haptic exploration |
0.4 | 1 | 2019 | Active Multi-Contact Continuous Tactile Exploration with Gaussian Process Differential Entropy · ICRA 2019 |
Robotics › Motion planning and robot control › robot control › learning control
control policy learning |
0.3 | 1 | 2018 | Learning to Control Redundant Musculoskeletal Systems with Neural Networks and SQP: Exploiting Muscle Properties · ICRA 2018 |
Robotics › Motion planning and robot control
musculoskeletal control |
0.3 | 1 | 2018 | Learning to Control Redundant Musculoskeletal Systems with Neural Networks and SQP: Exploiting Muscle Properties · ICRA 2018 |
Robotics › Motion planning and robot control
robot control |
0.3 | 1 | 2018 | Learning to Control Redundant Musculoskeletal Systems with Neural Networks and SQP: Exploiting Muscle Properties · ICRA 2018 |
Robotics › Robot navigation and mapping
localization |
0.3 | 1 | 2017 | Gaussian process estimation of odometry errors for localization and mapping · ICRA 2017 |
Robotics › Robot navigation and mapping › localization › odometry
odometry error modeling |
0.3 | 1 | 2017 | Gaussian process estimation of odometry errors for localization and mapping · ICRA 2017 |
Robotics › Robot navigation and mapping
SLAM |
0.3 | 1 | 2017 | Gaussian process estimation of odometry errors for localization and mapping · ICRA 2017 |
Robotics › Robot navigation and mapping › SLAM
visual SLAM |
0.3 | 1 | 2017 | Gaussian process estimation of odometry errors for localization and mapping · ICRA 2017 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing |
0.3 | 1 | 2025 | Mastering Board Games by External and Internal Planning with Language Models · ICML 2025 |
Robotics › Motion planning and robot control
trajectory planning |
0.2 | 1 | 2015 | Interplanetary Trajectory Planning with Monte Carlo Tree Search · IJCAI 2015 |
Knowledge, reasoning and agents › Multi-agent systems
game theory |
0.1 | 1 | 2020 | A Generalized Training Approach for Multiagent Learning · ICLR 2020 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
0.1 | 1 | 2019 | Active Multi-Contact Continuous Tactile Exploration with Gaussian Process Differential Entropy · ICRA 2019 |
Robotics › Robot manipulation
muscle co-contraction control |
0.1 | 1 | 2018 | Learning to Control Redundant Musculoskeletal Systems with Neural Networks and SQP: Exploiting Muscle Properties · ICRA 2018 |
Robotics › Robot manipulation › robot actuation
redundant actuation |
0.1 | 1 | 2018 | Learning to Control Redundant Musculoskeletal Systems with Neural Networks and SQP: Exploiting Muscle Properties · ICRA 2018 |
Robotics › Legged, aerial and field robots
space robotics |
0.1 | 1 | 2015 | Interplanetary Trajectory Planning with Monte Carlo Tree Search · IJCAI 2015 |
Methods — techniques the papers use, named apart from their topics
best response · 1.7policy gradient · 1.4policy space response oracles · 0.9internal search · 0.9in-context tree generation · 0.9external search · 0.9bargaining theory · 0.9population-based training · 0.6opponent modeling · 0.4mirror ascent · 0.4generalized training · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mastering Board Games by External and Internal Planning with Language ModelsabstractAdvancing planning and reasoning capabilities of Large Language Models (LLMs) is one of the key prerequisites towards unlocking their potential for performing reliably in complex and impactful domains. In this paper, we aim to demonstrate this across board games (Chess, Fischer Random / Chess960, Connect Four, and Hex), and we show that search-based planning can yield significant improvements in LLM game-playing strength. We introduce, compare and contrast two major approaches: In external search, the model guides Monte Carlo Tree Search (MCTS) rollouts and evaluations without calls to an external game engine, and in internal search, the model is trained to generate in-context a linearized tree of search and a resulting final choice. Both build on a language model pre-trained on relevant domain knowledge, reliably capturing the transition and value functions in the respective environments, with minimal hallucinations. We evaluate our LLM search implementations against game-specific state-of-the-art engines, showcasing substantial improvements in strength over the base model, and reaching Grandmaster-level performance in chess while operating closer to the human search budget. Our proposed approach, combining search with domain knowledge, is not specific to board games, hinting at more general future applications. John Schultz, Jakub Adámek, Matej Jusup, Marc Lanctot, Michael Kaisers, Sarah Perrin, Daniel Hennes, Jeremy Shar, Cannada A. Lewis, Anian Ruoss, Tom Zahavy, Petar Velickovic, Laurel Prince, Satinder Singh 0001, Eric Malmi, Nenad Tomasev |
ICML | 7 |
| 2025 | Combining Deep Reinforcement Learning and Search with Generative Models for Game-Theoretic Opponent ModelingabstractOpponent modeling methods typically involve two crucial steps: building a belief distribution over opponents' strategies, and exploiting this opponent model by playing a best response. However, existing approaches typically require domain-specific heurstics to come up with such a model, and algorithms for approximating best responses are hard to scale in large, imperfect information domains. In this work, we introduce a scalable and generic multiagent training regime for opponent modeling using deep game-theoretic reinforcement learning. We first propose Generative Best Respoonse (GenBR), a best response algorithm based on Monte-Carlo Tree Search (MCTS) with a learned deep generative model that samples world states during planning. This new method scales to large imperfect information domains and can be plug and play in a variety of multiagent algorithms. We use this new method under the framework of Policy Space Response Oracles (PSRO), to automate the generation of an offline opponent model via iterative game-theoretic reasoning and population-based training. We propose using solution concepts based on bargaining theory to build up an opponent mixture, which we find identifying profiles that are near the Pareto frontier. Then GenBR keeps updating an online opponent model and reacts against it during gameplay. We conduct behavioral studies where human participants negotiate with our agents in Deal-or-No-Deal, a class of bilateral bargaining games. Search with generative modeling finds stronger policies during both training time and test time, enables online Bayesian co-player prediction, and can produce agents that achieve comparable social welfare and Nash bargaining score negotiating with humans as humans trading among themselves. Zun Li 0002, Marc Lanctot, Kevin R. McKee, Luke Marris, Ian Gemp, Daniel Hennes, Paul Muller, Kate Larson, Yoram Bachrach, Michael P. Wellman |
IJCAI | 6 |
| 2022 | NeuPL: Neural Population Learning
Siqi Liu 0002, Luke Marris, Daniel Hennes, Josh Merel, Nicolas Heess, Thore Graepel |
ICLR | 3 |
| 2022 | Evolutionary Dynamics and Phi-Regret Minimization in GamesabstractRegret has been established as a foundational concept in online learning, and likewise has important applications in the analysis of learning dynamics in games. Regret quantifies the difference between a learner’s performance against a baseline in hindsight. It is well known that regret-minimizing algorithms converge to certain classes of equilibria in games; however, traditional forms of regret used in game theory predominantly consider baselines that permit deviations to deterministic actions or strategies. In this paper, we revisit our understanding of regret from the perspective of deviations over partitions of the full mixed strategy space (i.e., probability distributions over pure strategies), under the lens of the previously-established Φ-regret framework, which provides a continuum of stronger regret measures. Importantly, Φ-regret enables learning agents to consider deviations from and to mixed strategies, generalizing several existing notions of regret such as external, internal, and swap regret, and thus broadening the insights gained from regret-based analysis of learning algorithms. We prove here that the well-studied evolutionary learning algorithm of replicator dynamics (RD) seamlessly minimizes the strongest possible form of Φ-regret in generic 2 × 2 games, without any modification of the underlying algorithm itself. We subsequently conduct experiments validating our theoretical results in a suite of 144 2 × 2 games wherein RD exhibits a diverse set of behaviors. We conclude by providing empirical evidence of Φ-regret minimization by RD in some larger games, hinting at further opportunity for Φ-regret based study of such algorithms from both a theoretical and empirical perspective. Georgios Piliouras, Mark Rowland 0001, Shayegan Omidshafiei, Romuald Elie, Daniel Hennes, Jerome T. Connor, Karl Tuyls |
J. Artif. Intell. Res. | 5 |
| 2021 | Game Plan: What AI can do for Football, and What Football can do for AIabstractThe rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players’ and coordinated teams’ behaviors. The research challenges associated with predictive and prescriptive football analytics require new developments and progress at the intersection of statistical learning, game theory, and computer vision. In this paper, we provide an overarching perspective highlighting how the combination of these fields, in particular, forms a unique microcosm for AI research, while offering mutual benefits for professional teams, spectators, and broadcasters in the years to come. We illustrate that this duality makes football analytics a game changer of tremendous value, in terms of not only changing the game of football itself, but also in terms of what this domain can mean for the field of AI. We review the state-of-the-art and exemplify the types of analysis enabled by combining the aforementioned fields, including illustrative examples of counterfactual analysis using predictive models, and the combination of game-theoretic analysis of penalty kicks with statistical learning of player attributes. We conclude by highlighting envisioned downstream impacts, including possibilities for extensions to other sports (real and virtual). Karl Tuyls, Shayegan Omidshafiei, Paul Muller, Zhe Wang 0055, Jerome T. Connor, Daniel Hennes, Ian Graham, William Spearman, Tim Waskett, Dafydd Steele, Pauline Luc, Adrià Recasens, Alexandre Galashov, Gregory Thornton, Romuald Elie, Pablo Sprechmann, Pol Moreno, Kris Cao, Marta Garnelo, Praneet Dutta, Michal Valko, Nicolas Heess, Alex Bridgland, Julien Pérolat, Bart De Vylder, S. M. Ali Eslami, Mark Rowland 0001, Andrew Jaegle, Rémi Munos, Trevor Back, Razia Ahamed, Simon Bouton, Nathalie Beauguerlange, Jackson Broshear, Thore Graepel, Demis Hassabis |
J. Artif. Intell. Res. | 6 |
| 2020 | A Generalized Training Approach for Multiagent Learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland 0001, Karl Tuyls, Julien Pérolat, Siqi Liu 0002, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes 0001, Zhe Wang 0055, Guy Lever, Nicolas Heess, Thore Graepel, Rémi Munos |
ICLR | 7 |
| 2020 | Fast computation of Nash Equilibria in Imperfect Information GamesabstractWe introduce and analyze a class of algorithms, called Mirror Ascent against an Improved Opponent (MAIO), for computing Nash equilibria in two-player zero-sum games, both in normal form and in sequential form with imperfect information. These algorithms update the policy of each player with a mirror-ascent step to maximize the value of playing against an improved opponent. An improved opponent can be a best response, a greedy policy, a policy improved by policy gradient, or by any other reinforcement learning or search techniques. We establish a convergence result of the last iterate to the set of Nash equilibria and show that the speed of convergence depends on the amount of improvement offered by these improved policies. In addition, we show that under some condition, if we use a best response as improved policy, then an exponential convergence rate is achieved. Rémi Munos, Julien Pérolat, Jean-Baptiste Lespiau, Mark Rowland 0001, Bart De Vylder, Marc Lanctot, Finbarr Timbers, Daniel Hennes, Shayegan Omidshafiei, Audrunas Gruslys, Mohammad Gheshlaghi Azar, Edward Lockhart, Karl Tuyls |
ICML | 8 |
| 2019 | Active Multi-Contact Continuous Tactile Exploration with Gaussian Process Differential EntropyabstractIn the present work, we propose an active tactile exploration framework to obtain a surface model of an unknown object utilizing multiple contacts simultaneously. To incorporate these multiple contacts, the exploration strategy is based on the differential entropy of the underlying Gaussian process implicit surface model, which formalizes the exploration with multiple contacts within an information theoretic context and additionally allows for nonmyopic multi-step planning. In contrast to many previous approaches, the robot continuously slides along the surface with its end-effectors to gather the tactile stimuli, instead of touching it at discrete locations. This is realized by closely integrating the surface model into the compliant controller framework. Furthermore, we extend our recently proposed sliding based tactile exploration approach to handle non-convex objects. In the experiments, it is shown that multiple contacts simultaneously leads to a more efficient exploration of complex, non-convex objects, not only in terms of time, but also with respect to the total moved distance of all end-effectors. Finally, we demonstrate our methodology with a real PR2 robot that explores an object with both of its arms. Danny Drieß, Daniel Hennes, Marc Toussaint |
ICRA | 2 |
| 2018 | Learning to Control Redundant Musculoskeletal Systems with Neural Networks and SQP: Exploiting Muscle PropertiesabstractModeling biomechanical musculoskeletal systems reveals that the mapping from muscle stimulations to movement dynamics is highly nonlinear and complex, which makes it difficult to control those systems with classical techniques. In this work, we not only investigate whether machine learning approaches are capable of learning a controller for such systems. We are especially interested in the question if the structure of the musculoskeletal apparatus exhibits properties that are favorable for the learning task. In particular, we consider learning a control policy from target positions to muscle stimulations. To account for the high actuator redundancy of biomechanical systems, our approach uses a learned forward model represented by a neural network and sequential quadratic programming to obtain the control policy, which also enables us to alternate the co-contraction level and hence allows to change the stiffness of the system and to include optimality criteria like small muscle stimulations. Experiments on both a simulated musculoskeletal model of a human arm and a real biomimetic muscle-driven robot show that our approach is able to learn an accurate controller despite high redundancy and nonlinearity, while retaining sample efficiency. Danny Drieß, Heiko Zimmermann, Simon Wolfen, Dan Suissa, Daniel F. B. Haeufle, Daniel Hennes, Marc Toussaint, Syn Schmitt |
ICRA | 6 |
| 2017 | Gaussian process estimation of odometry errors for localization and mappingabstractSince early in robotics the performance of odometry techniques has been of constant research for mobile robots. This is due to its direct influence on localization. The pose error grows unbounded in dead-reckoning systems and its uncertainty has negative impacts in localization and mapping (i.e. SLAM). The dead-reckoning performance in terms of residuals, i.e. the difference between the expected and the real pose state, is related to the statistical error or uncertainty in probabilistic motion models. A novel approach to model odometry errors using Gaussian processes (GPs) is presented. The methodology trains a GP on the residual between the non-linear parametric motion model and the ground truth training data. The result is a GP over odometry residuals which provides an expected value and its uncertainty in order to enhance the belief with respect to the parametric model. The localization and mapping benefits from a comprehensive GP-odometry residuals model. The approach is applied to a planetary rover in an unstructured environment. We show that our approach enhances visual SLAM by efficiently computing image frames and effectively distributing keyframes. Javier Hidalgo-Carrióo, Daniel Hennes, Jakob Schwendner, Frank Kirchner |
ICRA | 2 |
| 2017 | NOctoSLAM: Fast octree surface normal mapping and registrationabstractIn this paper, we introduce a SLAM front end called NOctoSLAM. The approach adopts an octree-based map representation that implicitly enables source and reference data association for point to plane ICP registration. Additionally, the data structure is used to group map points to approximate surface normals. The multi-resolution capability of octrees, achieved by aggregating information in parent nodes, enables us to compensate for spatially unbalanced sensor data typically provided by multi-line lidar sensors. The octree-based data association is only approximate, but our empirical evaluation shows that NOctoSLAM achieves the same pose estimation accuracy as a comparable, point cloud based approach. However, NOctoSLAM can perform twice as many registration iterations per time unit. In contrast to point cloud based surface normal maps, where the map update duration depends on the current map size, we achieve a constant map update duration including surface normal recalculation. Therefore, NOctoSLAM does not require elaborate and environment dependent data filters. The results of our experiments show a mean positional error of 0.029 m and 0.019 rad, with a low standard deviation of 0.005 m and 0.006 rad, outperforming the state-of-the-art by remaining accurate while running online. Joscha-David Fossel, Karl Tuyls, Benjamin Schnieders, Daniel Claes, Daniel Hennes |
IROS | 5 |
| 2016 | Space Debris Removal: A Game Theoretic AnalysisabstractWe analyse active space debris removal efforts from a strategic, game-theoretic perspective. An active debris removal mission is a costly endeavour that has a positive effect (or risk reduction) for all satellites in the same orbital band. This leads to a dilemma: each actor (space agency, private stakeholder, etc.) has an incentive to delay its actions and wait for others to respond. The risk of the latter action is that, if everyone waits the joint outcome will be catastrophic leading to what in game theory is referred to as the ‘tragedy of the commons’. We introduce and thoroughly analyse this dilemma using simulation and empirical game theory in a two player setting. Richard Klíma, Daan Bloembergen, Rahul Savani, Karl Tuyls, Daniel Hennes, Dario Izzo |
ECAI | 5 |
| 2015 | Evolving Solutions to TSP Variants for Active Space Debris RemovalabstractThe space close to our planet is getting more and more polluted. Orbiting debris are posing an increasing threat to operational orbits and the cascading effect, known as Kessler syndrome, may result in a future where the risk of orbiting our planet at some altitudes will be unacceptable. Many argue that the debris density at the Low Earth Orbit (LEO) has already reached a level sufficient to trigger such a cascading effect. An obvious consequence is that we may soon have to actively clean space from debris. Such a space mission will involve a complex combinatorial decision as to choose which debris to remove and in what order. In this paper, we find that this part of the design of an active debris removal mission (ADR) can be mapped into increasingly complex variants to the classic Travelling Salesman Problem (TSP) and that they can be solved by the Inver-over algorithm improving the current state-of-the-art in ADR mission design. We define static and dynamic cases, according to whether we consider the debris orbits as fixed in time or subject to orbital perturbations. We are able, for the first time, to select optimally objects from debris clouds of considerable size: hundreds debris pieces considered while previous works stopped at tens. Dario Izzo, Ingmar Getzner, Daniel Hennes, Luís F. Simões |
GECCO | 3 |
| 2015 | Novelty Search for Soft Robotic Space ExplorationabstractThe use of soft robots in future space exploration is still a far-fetched idea, but an attractive one. Soft robots are inherently compliant mechanisms that are well suited for locomotion on rough terrain as often faced in extra-planetary environments. Depending on the particular application and requirements, the best shape (or body morphology) and locomotion strategy for such robots will vary substantially. Recent developments in soft robotics and evolutionary optimization showed the possibility to simultaneously evolve the morphology and locomotion strategy in simulated trials. The use of techniques such as generative encoding and neural evolution were key to these findings. In this paper, we improve further on this methodology by introducing the use of a novelty measure during the evolution process. We compare fitness search and novelty search in different gravity levels and we consistently find novelty-based search to perform as good as or better than a fitness--based search, while also delivering a greater variety of designs. We propose a combination of the two techniques using fitness-elitism in novelty search to obtain a further improvement. We then use our methodology to evolve the gait and morphology of soft robots at different gravity levels, finding a taxonomy of possible locomotion strategies that are analyzed in the context of space-exploration. Georgios Methenitis, Daniel Hennes, Dario Izzo, Arnoud Visser |
GECCO | 2 |
| 2015 | Injection, Saturation and Feedback in Meta-Heuristic InteractionsabstractMeta-heuristics have proven to be an efficient method of handling difficult global optimization tasks. A recent trend in evolutionary computation is the use of several meta-heuristics at the same time, allowing for occasional information exchange among them in hope to take advantage from the best algorithmic properties of all. Such an approach is inherently parallel and, with some restrictions, has a straight forward implementation in a heterogeneous island model. We propose a methodology for characterizing the interplay between different algorithms, and we use it to discuss their performance on real-parameter single objective optimization benchmarks. We introduce the new concepts of feedback, saturation and injection, and show how they are powerful tools to describe the interplay between different algorithms and thus to improve our understanding of the internal mechanism at work in large parallel evolutionary set-ups. Krzysztof Nowak, Dario Izzo, Daniel Hennes |
GECCO | 3 |
| 2015 | Interplanetary Trajectory Planning with Monte Carlo Tree Search
Daniel Hennes, Dario Izzo |
IJCAI | 1 |
| 2015 | Trading in markets with noisy information: an evolutionary analysisabstractWe analyse the value of information in a stock market where information can be noisy and costly, using techniques from empirical game theory. Previous work has shown that the value of information follows a J-curve, where averagely informed traders perform below market average, and only insiders prevail. Here we show that both noise and cost can change this picture, in several cases leading to opposite results where insiders perform below market average, and averagely informed traders prevail. Moreover, we investigate the effect of random explorative actions on the market dynamics, showing how these lead to a mix of traders being sustained in equilibrium. These results provide insight into the complexity of real marketplaces, and show under which conditions a broad mix of different trading strategies might be sustainable. Daan Bloembergen, Daniel Hennes, Peter McBurney, Karl Tuyls |
Connect. Sci. | 2 |
| 2015 | Preface to the special issue: Adaptive Learning Agents Part 3abstract1. An adaptive learning agent is capable of adapting its behaviour in order to react to changes in its environment and using previous experience to improve its performance with respect to some eval... Sam Devlin, Daniel Hennes, Samuel Barrett |
Connect. Sci. | 2 |
| 2015 | Evolutionary Dynamics of Multi-Agent Learning: A SurveyabstractThe interaction of multiple autonomous agents gives rise to highly dynamic and nondeterministic environments, contributing to the complexity in applications such as automated financial markets, smart grids, or robotics. Due to the sheer number of situations that may arise, it is not possible to foresee and program the optimal behaviour for all agents beforehand. Consequently, it becomes essential for the success of the system that the agents can learn their optimal behaviour and adapt to new situations or circumstances. The past two decades have seen the emergence of reinforcement learning, both in single and multi-agent settings, as a strong, robust and adaptive learning paradigm. Progress has been substantial, and a wide range of algorithms are now available. An important challenge in the domain of multi-agent learning is to gain qualitative insights into the resulting system dynamics. In the past decade, tools and methods from evolutionary game theory have been successfully employed to study multi-agent learning dynamics formally in strategic interactions. This article surveys the dynamical models that have been derived for various multi-agent reinforcement learning algorithms, making it possible to study and compare them qualitatively. Furthermore, new learning algorithms that have been introduced using these evolutionary game theoretic tools are reviewed. The evolutionary models can be used to study complex strategic interactions. Examples of such analysis are given for the domains of automated trading in stock markets and collision avoidance in multi-robot systems. The paper provides a roadmap on the progress that has been achieved in analysing the evolutionary dynamics of multi-agent learning by highlighting the main results and accomplishments. Daan Bloembergen, Karl Tuyls, Daniel Hennes, Michael Kaisers |
J. Artif. Intell. Res. | 3 |
| 2015 | Metastrategies in Large-Scale Bargaining SettingsabstractThis article presents novel methods for representing and analyzing a special class of multiagent bargaining settings that feature multiple players, large action spaces, and a relationship among players’ goals, tasks, and resources. We show how to reduce these interactions to a set of bilateral normal-form games in which the strategy space is significantly smaller than the original settings while still preserving much of their structural relationship. The method is demonstrated using the Colored Trails (CT) framework, which encompasses a broad family of games and has been used in many past studies. We define a set of heuristics (metastrategies) in multiplayer CT games that make varying assumptions about players’ strategies, such as boundedly rational play and social preferences. We show how these CT settings can be decomposed into canonical bilateral games such as the Prisoners’ Dilemma, Stag Hunt, and Ultimatum games in a way that significantly facilitates their analysis. We demonstrate the feasibility of this approach in separate CT settings involving one-shot and repeated bargaining scenarios, which are subsequently analyzed using evolutionary game-theoretic techniques. We provide a set of necessary conditions for CT games for allowing this decomposition. Our results have significance for multiagent systems researchers in mapping large multiplayer CT task settings to smaller, well-known bilateral normal-form games while preserving some of the structure of the original setting. Daniel Hennes, Steven de Jong, Karl Tuyls, Kobi Gal |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2014 | Preface to the special issue: Adaptive Learning Agents, Part 1abstract1. An Adaptive Learning Agent is capable of adapting its behaviour in order to react to changes in its environment and using previous experience to improve its performance with respect to some eval... Sam Devlin, Daniel Hennes, Enda Howley |
Connect. Sci. | 2 |
| 2014 | Preface to the special issue: Adaptive Learning Agents Part 2abstract1. An adaptive learning agent is capable of adapting its behaviour in order to react to changes in its environment and using previous experience to improve its performance with respect to some eval... Sam Devlin, Daniel Hennes, Enda Howley |
Connect. Sci. | 2 |
| 2013 | How to Win RoboCup@Work? - The Swarmlab@Work Approach Revealed
Sjriek Alers, Daniel Claes, Joscha-David Fossel, Daniel Hennes, Karl Tuyls, Gerhard Weiss 0001 |
RoboCup | 4 |
| 2012 | Evolutionary advantage of foresight in marketsabstractWe analyze the competitive advantage of price signal information for traders in simulated double auctions. Previous work has established that more information about the price development does not guarantee higher performance. In particular, traders with limited information perform below market average and are outperformed by random traders; only insiders beat the market. However, this result has only been shown in markets with a few traders and a uniform distribution over information levels. We present additional simulations of several more realistic information distributions, extending previous findings. In addition, we analyze the market dynamics with an evolutionary model of competing information levels. Results show that the highest information level will dominate if information comes for free. If information is costly, less-informed traders may prevail reflecting a more realistic distribution over information levels. Daniel Hennes, Daan Bloembergen, Michael Kaisers, Karl Tuyls, Simon Parsons |
GECCO | 1 |
| 2012 | Collision avoidance under bounded localization uncertaintyabstractWe present a multi-mobile robot collision avoidance system based on the velocity obstacle paradigm. Current positions and velocities of surrounding robots are translated to an efficient geometric representation to determine safe motions. Each robot uses on-board localization and local communication to build the velocity obstacle representation of its surroundings. Our close and error-bounded convex approximation of the localization density distribution results in collision-free paths under uncertainty. While in many algorithms the robots are approximated by circumscribed radii, we use the convex hull to minimize the overestimation in the footprint. Results show that our approach allows for safe navigation even in densely packed environments. Daniel Claes, Daniel Hennes, Karl Tuyls, Wim Meeussen |
IROS | 2 |
| 2011 | Hierarchies of octrees for efficient 3D mappingabstractThe on-chip fabrication and manipulation of microstructures are expected to be applied for single cell analysis system such as cell manipulation and measurement tools. In this paper, we previously present a methodology for fabricating and assembling microstructures inside a microfluidic channel. By the illumination of patterned UV-ray through the mask under a microscope, microstructures with arbitrary shape are made of the photo-crosslinkable resin inside microfluidic device. The microstructures are fabricated at the desired place inside microfluidic channel and manipulated by optical tweezers. Based on the technique which can manipulate multiple points simultaneously by high-speed scanning of a single laser with galvanometer mirror, a rotational microstructure made of a microgear and a rotation axis is assembled and rotated. We also report two methods of solution replacement inside microfluidic channel which reduces viscosity of solvent in order to improve manipulation performance. By adjusting the concentration of photo-crosslinkable resin and replacing solution components, the viscosity of solvent inside channel can be changed. The manipulation speed of the rotational microstructure increases when the viscosity of solvent decreases, because the viscosity resistance for the movement of microstructure is weaker inside lower viscosity solvent. We fabricate rotational microstructures inside lower viscosity solvent and evaluate the movement efficiency compared with microstructures inside former high viscosity solvent. Kai M. Wurm, Daniel Hennes, Dirk Holz, Radu Bogdan Rusu, Cyrill Stachniss, Kurt Konolige, Wolfram Burgard |
IROS | 2 |