EDBT 2026 Demo / reviewers in the wild / expert
Jilles Steeve Dibangoye
dblp:52/7118 · also Jilles Dibangoye
· DBLP profile ↗
32ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0001-8826-4438ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 9 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Reinforcement learning · 44% Multi-agent systems · 18% Planning, search and constraint satisfaction · 14% |
Topics — the 25 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
decentralized partially observable markov decision process |
1.8 | 3 | 2025 | Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach · AAAI 2025 Solving Hierarchical Information-Sharing Dec-POMDPs: An Extensive-Form Game Approach · ICML 2024 Optimally Solving Dec-POMDPs as Continuous-State MDPs · IJCAI 2013 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.4 | 3 | 2025 | Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach · AAAI 2025 Learning to Act in Decentralized Partially Observable MDPs · ICML 2018 Optimally Solving Dec-POMDPs as Continuous-State MDPs · IJCAI 2013 |
Knowledge, reasoning and agents › Multi-agent systems
partially observable stochastic games |
0.9 | 1 | 2025 | ε-Optimally Solving Two-Player Zero-Sum POSGs · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
extensive-form games |
0.8 | 1 | 2024 | Solving Hierarchical Information-Sharing Dec-POMDPs: An Extensive-Form Game Approach · ICML 2024 |
Robotics › Autonomous driving
perception |
0.7 | 1 | 2023 | LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction · ICRA 2023 |
Robotics › Robot navigation and mapping
sensor fusion |
0.7 | 1 | 2023 | LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction · ICRA 2023 |
Knowledge, reasoning and agents › Multi-agent systems
decentralized planning |
0.6 | 2 | 2020 | Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information Sharing · ICML 2020 Optimally Solving Dec-POMDPs as Continuous-State MDPs · IJCAI 2013 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
decentralized POMDP |
0.4 | 1 | 2020 | Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information Sharing · ICML 2020 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
information sharing |
0.4 | 1 | 2020 | Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information Sharing · ICML 2020 |
Robotics › Motion planning and robot control
motion planning |
0.4 | 1 | 2020 | Learning to Plan with Uncertain Topological Maps · ECCV (3) 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process |
0.4 | 2 | 2018 | rho-POMDPs have Lipschitz-Continuous epsilon-Optimal Value Functions · NeurIPS 2018 Topological Order Planner for POMDPs · IJCAI 2009 |
Robotics › Autonomous driving
driver behavior modeling |
0.3 | 1 | 2018 | Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices · ICRA 2018 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.3 | 1 | 2018 | Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices · ICRA 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint optimization
mixed-integer linear programming |
0.3 | 1 | 2018 | Learning to Act in Decentralized Partially Observable MDPs · ICML 2018 |
Robotics › Motion planning and robot control
trajectory optimization |
0.3 | 1 | 2018 | Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices · ICRA 2018 |
Machine learning › Reinforcement learning
value function approximation |
0.3 | 1 | 2018 | rho-POMDPs have Lipschitz-Continuous epsilon-Optimal Value Functions · NeurIPS 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › multi-agent planning
centralized planning |
0.3 | 1 | 2025 | Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach · AAAI 2025 |
Machine learning › Reinforcement learning › markov decision process
continuous state MDPs |
0.2 | 1 | 2015 | Exploiting Separability in Multiagent Planning with Continuous-State MDPs (Extended Abstract) · IJCAI 2015 |
Robotics › Motion planning and robot control › multi-robot control
decentralized control |
0.2 | 1 | 2015 | Structural Results for Cooperative Decentralized Control Models · IJCAI 2015 |
Robotics › Motion planning and robot control › multi-robot control
decentralized cooperative control |
0.2 | 1 | 2015 | Structural Results for Cooperative Decentralized Control Models · IJCAI 2015 |
Machine learning › Reinforcement learning
markov decision process |
0.2 | 1 | 2015 | Exploiting Separability in Multiagent Planning with Continuous-State MDPs (Extended Abstract) · IJCAI 2015 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent planning |
0.2 | 1 | 2015 | Exploiting Separability in Multiagent Planning with Continuous-State MDPs (Extended Abstract) · IJCAI 2015 |
Computer vision › 3D vision › 3d scene understanding
depth and scene understanding |
0.2 | 1 | 2023 | LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction · ICRA 2023 |
Knowledge, reasoning and agents › Multi-agent systems › decentralized planning
Dec-POMDP |
0.2 | 1 | 2013 | Optimally Solving Dec-POMDPs as Continuous-State MDPs · IJCAI 2013 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
POMDP planning |
0.1 | 1 | 2009 | Topological Order Planner for POMDPs · IJCAI 2009 |
Methods — techniques the papers use, named apart from their topics
dynamic programming · 0.9bellman principle of optimality · 0.9SARSA · 0.9hierarchical information sharing · 0.8bellman's principle of optimality · 0.8projective transform network · 0.7point cloud guidance · 0.7feature pyramid network · 0.7topological map · 0.4reinforcement learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning ApproachabstractThe centralized training for decentralized execution paradigm emerged as the state-of-the-art approach to ϵ-optimally solving decentralized partially observable Markov decision processes. However, scalability remains a significant issue. This paper presents a novel and more scalable alternative, namely the sequential-move centralized training for decentralized execution. This paradigm further pushes the applicability of the Bellman’s principle of optimality, raising three new properties. First, it allows a central planner to reason upon sufficient sequential-move statistics instead of prior simultaneous-move ones. Next, it proves that ϵ-optimal value functions are piecewise linear and convex in such sufficient sequential-move statistics. Finally, it drops the complexity of the backup operators from double exponential to polynomial at the expense of longer planning horizons. Besides, it makes it easy to use single-agent methods, e.g., SARSA algorithm enhanced with these findings, while still preserving convergence guarantees. Experiments on two- as well as many-agent domains from the literature against ϵ-optimal simultaneous-move solvers confirm the superiority of our novel approach. This paradigm opens the door for efficient planning and reinforcement learning methods for multi-agent systems. Johan Peralez, Aurélien Delage, Jacopo Castellini, Rafael F. Cunha, Jilles Steeve Dibangoye |
AAAI | 5 |
| 2025 | ε-Optimally Solving Two-Player Zero-Sum POSGs
Erwan Escudie, Matthia Sabatelli, Olivier Buffet, Jilles Steeve Dibangoye |
NeurIPS | 4 |
| 2024 | Solving Hierarchical Information-Sharing Dec-POMDPs: An Extensive-Form Game ApproachabstractA recent theory shows that a multi-player decentralized partially observable Markov decision process can be transformed into an equivalent single-player game, enabling the application of Bellman's principle of optimality to solve the single-player game by breaking it down into single-stage subgames. However, this approach entangles the decision variables of all players at each single-stage subgame, resulting in backups with a double-exponential complexity. This paper demonstrates how to disentangle these decision variables while maintaining optimality under hierarchical information sharing, a prominent management style in our society. To achieve this, we apply the principle of optimality to solve any single-stage subgame by breaking it down further into smaller subgames, enabling us to make single-player decisions at a time. Our approach reveals that extensive-form games always exist with solutions to a single-stage subgame, significantly reducing time complexity. Our experimental results show that the algorithms leveraging these findings can scale up to much larger multi-player games without compromising optimality. Johan Peralez, Aurélien Delage, Olivier Buffet, Jilles Steeve Dibangoye |
ICML | 4 |
| 2023 | Optimization of Sensor Configurations for Fault Identification in Smart BuildingsabstractIn predictive maintenance an important problem is to optimize the quantity of information to be transmitted at the control center to guarantee reliable fault detection while limiting sensor power consumption. This problem relies directly on the sensor configurations (e.g., sampling rate, coding, quantization) and the fault detection algorithm. To address this question, we introduce a codesign framework and an algorithm for joint optimization of the sensor configurations and the accuracy of the fault detection classifier. In a use case based on a dataset consisting of multiple sensor measurements and heating power levels known as the Twin House Experiment, we show that our algorithm can find efficient tradeoffs between sensor power consumption and classifier accuracy. Malcolm Egan, Jean-Marie Gorce, Jilles Steeve Dibangoye, Frédéric Le Mouël |
ICASSP | 4 |
| 2023 | LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid PredictionabstractSemantic grids can be useful representations of the scene around an autonomous system. By having information about the layout of the space around itself, a robot can leverage this type of representation for crucial tasks such as navigation or tracking. By fusing information from multiple sensors, robustness can be increased and the computational load for the task can be lowered, achieving real time performance. Our multi-scale LiDAR-Aided Perspective Transform network uses information available in point clouds to guide the projection of image features to a top-view representation, resulting in a relative improvement in the state of the art for semantic grid generation for human (+8.67%) and movable object (+49.07%) classes in the nuScenes dataset, as well as achieving results close to the state of the art for the vehicle, drivable area and walkway classes, while performing inference at 25 FPS. Manuel Diaz-Zapata, David Sierra González, Özgür Erkent, Christian Laugier, Jilles Steeve Dibangoye |
ICRA | 5 |
| 2023 | Global min-max Computation for α-Hölder Gamesabstractmin-max optimization problems recently arose in various settings. From Generative Adversarial Networks (GANs) to aerodynamic optimization through Game Theory, the assumptions on the objective function vary. Motivated by the applications to deep learning and especially GANs, most recent works assume differentiability to design local search algorithms such as Gradient Descend Ascent (GDA). In contrast, this work will only require $\alpha -$Hölder properties to tackle general game-theoretic problems with poor continuity assumptions. Focusing on the example of problems in which max and min optimization variables live in simplices, we provide a simple algorithm, based on Deterministic Optimistic Optimization (DOO), relying on an outer min-optimization using the solutions of an inner max-optimization. The algorithm is shown to converge in finite time to an -global optimum. Experimental validations are given and the time complexity of our algorithm is studied. Aurélien Delage, Olivier Buffet, Jilles Steeve Dibangoye |
ICTAI | 3 |
| 2023 | Codesigned Communication and Data Analytics for Condition-Based Maintenance in Smart BuildingsabstractWith the proliferation of cheap sensors and the ubiquity of cloud and edge computing, predictive/condition-based maintenance is expected to play an important role in smart homes and buildings. Nevertheless, a key difficulty is ensuring that sensors provide data of sufficient quality in order to reliably detect building (e.g., heating system) degradation in systems or comfort. At the same time, sensor utilization should be limited as much as possible in order to minimize power consumption, and increase the lifetimes of batteries. A solution to this problem requires careful codesign of sensor communication and data analytics. In this article, we introduce a formulation of this codesign problem, which is based on an optimization problem to jointly design how often data is collected and compression levels in order to balance the quality of fault detection with the quantity of transmitted data. To solve the optimization problem, we apply a differentiable search algorithm based on a variant of stochastic gradient descent for discrete optimization problems. We apply our codesign framework and solve the resulting optimization problem using data obtained from a building comfort experiment known as the Twin House Experiment. We also provide an extension of our algorithm to a dynamic variant of the codesign framework, where comfort levels and power consumption penalties are time varying. Numerical results show that our algorithm rapidly finds an efficient tradeoff between classifier accuracy and sensor power consumption. Malcolm Egan, Jean-Marie Gorce, Jilles Steeve Dibangoye, Frédéric Le Mouël |
IEEE Internet Things J. | 4 |
| 2022 | LAPTNet: LiDAR-Aided Perspective Transform NetworkabstractSemantic grids are a useful representation of the environment around a robot. They can be used in autonomous vehicles to concisely represent the scene around the car, capturing vital information for downstream tasks like navigation or collision assessment. Information from different sensors can be used to generate these grids. Some methods rely only on RGB images, whereas others choose to incorporate information from other sensors, such as radar or LiDAR. In this paper, we present an architecture that fuses LiDAR and camera information to generate semantic grids. By using the 3D information from a LiDAR point cloud, the LiDAR-Aided Perspective Transform Network (LAPTNet) is able to associate features in the camera plane to the bird's eye view without having to predict any depth information about the scene. Compared to state-of-the-art camera-only methods, LAPTNet achieves an improvement of up to 8.8 points (or 38.13%) over state-of-art competing approaches for the classes proposed in the NuScenes dataset validation split. Manuel Diaz-Zapata, Özgür Erkent, Christian Laugier, Jilles Steeve Dibangoye, David Sierra González |
ICARCV | 4 |
| 2022 | R-MDP: A Game Theory Approach for Fault-Tolerant Data and Service Management in Crude Oil Pipelines Monitoring Systems
Safuriyawu Ahmed, Frédéric Le Mouël, Nicolas Stouls, Jilles Steeve Dibangoye |
MobiQuitous | 4 |
| 2021 | Heuristic Search Value Iteration for Zero-Sum Stochastic GamesabstractIn sequential decision making, heuristic search algorithms allow exploiting both the initial situation and an admissible heuristic to efficiently search for an optimal solution, often for planning purposes. Such algorithms exist for problems with uncertain dynamics, partial observability, multiple criteria, or multiple collaborating agents. In this article, we look at two-player zero-sum stochastic games (zsSGs) with a discounted criterion, in a view to propose a solution tailored to the fully observable case, while solutions have been proposed for particular, though still more general, partially observable cases. This setting induces reasoning on both a lower and an upper bound of the value function, which leads us to proposing zsSG-HSVI, an algorithm based on heuristic search value iteration (HSVI), and which thus relies on generating trajectories. We demonstrate that, each player acting optimistically, and employing simple heuristic initializations, HSVI's convergence in finite time to an ∈-optimal solution is preserved. An empirical study of the resulting approach is conducted on benchmark problems of various sizes. Olivier Buffet, Jilles Steeve Dibangoye, Abdallah Saffidine, Vincent Thomas |
IEEE Trans. Games | 2 |
| 2021 | Solving Multi-Agent Routing Problems Using Deep Attention MechanismsabstractRouting delivery vehicles to serve customers in dynamic and uncertain environments like dense city centers is a challenging task that requires robustness and flexibility. Most existing approaches to routing problems produce solutions offline in the form of plans, which only apply to the situation they have been optimized for. Instead, we propose to learn a policy that provides decision rules to build the routes from online measurements of the environment state, including the customers configuration itself. Doing so, we can generalize from past experiences and quickly provide decision rules for new instances of the problem without re-optimizing any parameters of our policy. The difficulty with this approach comes from the complexity to represent this state. In this paper, we introduce a sequential multi-agent decision-making model to formalize the description and the temporal evolution of a Dynamic and Stochastic Vehicle Routing Problem. We propose a variation of Deep Neural Network using Attention Mechanisms to learn generalizable representation of the state and output online decision rules adapted to dynamic and stochastic information. Using artificially-generated data, we show promising results in these dynamic and stochastic environments, while staying competitive in deterministic ones compared to offline classical heuristics. Guillaume Bono, Jilles Steeve Dibangoye, Olivier Simonin 0001, Laëtitia Matignon, Florian Pereyron |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | Learning to Plan with Uncertain Topological Maps
Edward Beeching, Jilles Steeve Dibangoye, Olivier Simonin 0001, Christian Wolf 0001 |
ECCV (3) | 2 |
| 2020 | Leveraging Dynamic Occupancy Grids for 3D Object Detection in Point CloudsabstractTraditionally, point cloud-based 3D object detectors are trained on annotated, non-sequential samples taken from driving sequences (e.g. the KITTI dataset). However, by doing this, the developed algorithms renounce to exploit any dynamic information from the driving sequences. It is reasonable to think that this information, which is available at test time when deploying the models in the experimental vehicles, could have significant predictive potential for the object detection task. To study the advantages that this kind of information could provide, we construct a dataset of dynamic occupancy grid maps from the raw KITTI dataset and find the correspondence to each of the KITTI 3D object detection dataset samples. By training a Lidar-based state-of-the-art 3D object detector with and without the dynamic information we get insights into the predictive value of the dynamics. Our results show that having access to the environment dynamics improves by 27% the ability of the detection algorithm to predict the orientation of smaller obstacles such as pedestrians. Furthermore, the 3D and bird's eye view bounding box predictions for pedestrians in challenging cases also see a 7% improvement. Qualitatively speaking, the dynamics help with the detection of partially occluded and far-away obstacles. We illustrate this fact with numerous qualitative prediction results. David Sierra González, Anshul Paigwar, Özgür Erkent, Jilles Steeve Dibangoye, Christian Laugier |
ICARCV | 4 |
| 2020 | Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information SharingabstractOptimally solving decentralized partially observable Markov decision processes under either full or no information sharing received significant attention in recent years. However, little is known about how partial information sharing affects existing theory and algorithms. This paper addresses this question for a team of two agents, with one-sided information sharing—\ie both agents have imperfect information about the state of the world, but only one has access to what the other sees and does. From the perspective of a central planner, we show that the original problem can be reformulated into an equivalent information-state Markov decision process and solved as such. Besides, we prove that the optimal value function exhibits a specific form of uniform continuity. We also present a heuristic search algorithm utilizing this property and providing the first results for this family of problems. Jilles Steeve Dibangoye, Olivier Buffet |
ICML | 2 |
| 2020 | Deep Reinforcement Learning on a Budget: 3D Control and Reasoning Without a SupercomputerabstractAn important goal of research in Deep Reinforcement Learning in mobile robotics is to train agents capable of solving complex tasks, which require a high level of scene understanding and reasoning from an egocentric perspective. When trained from simulations, optimal environments should satisfy a currently unobtainable combination of high-fidelity photographic observations, massive amounts of different environment configurations and fast simulation speeds. In this paper we argue that research on training agents capable of complex reasoning can be simplified by decoupling from the requirement of high fidelity photographic observations. We present a suite of tasks requiring complex reasoning and exploration in continuous, partially observable 3D environments. The objective is to provide challenging scenarios and a robust baseline agent architecture that can be trained on mid-range consumer hardware in under 24h. Our scenarios combine two key advantages: (i) they are based on a simple but highly efficient 3D environment (ViZDoom) which allows high speed simulation (12000fps); (ii) the scenarios provide the user with a range of difficulty settings, in order to identify the limitations of current state of the art algorithms and network architectures. We aim to increase accessibility to the field of Deep-RL by providing baselines for challenging scenarios where new ideas can be iterated on quickly. We argue that the community should be able to address challenging problems in reasoning of mobile agents without the need for a large compute infrastructure. Code for the generation of scenarios and training of baselines is available online at the following repository1.1https://github.com/edbeeching/3d_control_deep_rl. Edward Beeching, Jilles Steeve Dibangoye, Olivier Simonin 0001, Christian Wolf 0001 |
ICPR | 2 |
| 2020 | EgoMap: Projective Mapping and Structured Egocentric Memory for Deep RL
Edward Beeching, Jilles Steeve Dibangoye, Olivier Simonin 0001, Christian Wolf 0001 |
ECML/PKDD (2) | 2 |
| 2019 | Combining Stochastic Optimization and Frontiers for Aerial Multi-Robot Exploration of 3D TerrainsabstractThis paper addresses the problem of exploring unknown terrains with a fleet of cooperating aerial vehicles. We present a novel decentralized approach which alternates gradient-free stochastic optimization and a frontier-based approach. Our method allows each robot to generate its trajectory based on the collected data and the local map built integrating the information shared by its teammates. Whenever a local optimum is reached, which corresponds to a location surrounded by already explored areas, the algorithm identifies the closest frontier to get over it and restarts the local optimization. Its low computational cost, the capability to deal with constraints and the decentralized decision-making make it particularly suitable for multi-robot applications in complex 3D environments. Simulation results show that our approach generates feasible trajectories which drive multiple robots to completely explore realistic environments. Furthermore, in terms of exploration time, our algorithm significantly outperforms a standard solution based on closest frontier points while providing similar performances compared to a computationally more expensive centralized greedy solution. Alessandro Renzaglia, Jilles Steeve Dibangoye, Vincent Le Doze, Olivier Simonin 0001 |
IROS | 2 |
| 2019 | DynFloR: A Flow Approach for Data Delivery Optimization in Multi-Robot Network PatrollingabstractDeploying fleets of mobile robots in real scenarios and environments raises several scientific challenges. One of them concerns the ability of the robots to adapt to the dynamics of their environment. We introduce DynFloR , a dynamic network flow based approach for finding optimal policies for data delivery in multi-robot network patrolling where the robots can communicate instantly and free of charge one to another when they meet, there is a periodicity of the robot meetings and the distribution of the data collected during the patrol is regular. Experiments on randomly generated synthetic examples are performed for evaluating the performance of the DynFloR method. The performed experiments empirically show that independent of the problem setting (such as number of robots, memory of the robots) the amount of data transferred to a base station per unit of time converges to an equilibrium state. The case of lost data has been also examined through various experiments, but it requires further experimentation as well as in-depth analysis. Vlad-Sebastian Ionescu, Zsuzsanna Onet-Marian, Marin-Georgian Badita, Gabriela Serban Czibula, Mihai-Ioan Popescu, Jilles Steeve Dibangoye, Olivier Simonin 0001 |
KES | 6 |
| 2019 | Learning 3D Navigation Protocols on Touch Interfaces with Cooperative Multi-agent Reinforcement Learning
Quentin Debard, Jilles Steeve Dibangoye, Stéphane Canu, Christian Wolf 0001 |
ECML/PKDD (3) | 2 |
| 2018 | Learning to Act in Decentralized Partially Observable MDPsabstractWe address a long-standing open problem of reinforcement learning in decentralized partially observable Markov decision processes. Previous attempts focussed on different forms of generalized policy iteration, which at best led to local optima. In this paper, we restrict attention to plans, which are simpler to store and update than policies. We derive, under certain conditions, the first near-optimal cooperative multi-agent reinforcement learning algorithm. To achieve significant scalability gains, we replace the greedy maximization by mixed-integer linear programming. Experiments show our approach can learn to act near-optimally in many finite domains from the literature. Jilles Steeve Dibangoye, Olivier Buffet |
ICML | 1 |
| 2018 | Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal LatticesabstractOne of the most challenging tasks in the development of path planners for intelligent vehicles is the design of the cost function that models the desired behavior of the vehicle. While this task has been traditionally accomplished by hand-tuning the model parameters, recent approaches propose to learn the model automatically from demonstrated driving data using Inverse Reinforcement Learning (IRL). To determine if the model has correctly captured the demonstrated behavior, most IRL methods require obtaining a policy by solving the forward control problem repetitively. Calculating the full policy is a costly task in continuous or large domains and thus often approximated by finding a single trajectory using traditional path-planning techniques. In this work, we propose to find such a trajectory using a conformal spatiotemporal state lattice, which offers two main advantages. First, by conforming the lattice to the environment, the search is focused only on feasible motions for the robot, saving computational power. And second, by considering time as part of the state, the trajectory is optimized with respect to the motion of the dynamic obstacles in the scene. As a consequence, the resulting trajectory can be used for the model assessment. We show how the proposed IRL framework can successfully handle highly dynamic environments by modeling the highway tactical driving task from demonstrated driving data gathered with an instrumented vehicle. David Sierra González, Özgür Erkent, Victor Romero-Cano, Jilles Steeve Dibangoye, Christian Laugier |
ICRA | 4 |
| 2018 | rho-POMDPs have Lipschitz-Continuous epsilon-Optimal Value FunctionsabstractMany state-of-the-art algorithms for solving Partially Observable Markov Decision Processes (POMDPs) rely on turning the problem into a “fully observable” problem—a belief MDP—and exploiting the piece-wise linearity and convexity (PWLC) of the optimal value function in this new state space (the belief simplex ∆). This approach has been extended to solving ρ-POMDPs—i.e., for information-oriented criteria—when the reward ρ is convex in ∆. General ρ-POMDPs can also be turned into “fully observable” problems, but with no means to exploit the PWLC property. In this paper, we focus on POMDPs and ρ-POMDPs with λ ρ -Lipschitz reward function, and demonstrate that, for finite horizons, the optimal value function is Lipschitz-continuous. Then, value function approximators are proposed for both upper- and lower-bounding the optimal value function, which are shown to provide uniformly improvable bounds. This allows proposing two algorithms derived from HSVI which are empirically evaluated on various benchmark problems. Mathieu Fehr, Olivier Buffet, Vincent Thomas, Jilles Steeve Dibangoye |
NeurIPS | 4 |
| 2018 | Cooperative Multi-agent Policy Gradient
Guillaume Bono, Jilles Steeve Dibangoye, Laëtitia Matignon, Florian Pereyron, Olivier Simonin 0001 |
ECML/PKDD (1) | 2 |
| 2017 | Open Decentralized POMDPsabstractMany real-world multi-agent applications such as rescue operations require to dynamically assemble or disassemble teams to provide a service exploiting this flexibility. While Dec-POMDPs capture real-world uncertainty in multi-agent systems, they fail to use the team flexibility. This paper introduces a new model, called the Open Dec-POMDP, allowing us to consider agents that can enter and leave the system during the execution of the process. Exploiting the team flexibility enables us to present a new best-response dynamics' algorithm that uses Monte Carlo simulations to compute joint policies. Our algorithm can dynamically adapt to the team flexibility and computes locally optimal solutions. Experiments demonstrate that Open Dec-POMDP solutions compare favorably on different fire-fighting scenarios to standard Dec-POMDPs'. Jonathan Cohen 0001, Jilles Steeve Dibangoye, Abdel-Illah Mouaddib |
ICTAI | 2 |
| 2016 | Optimally Solving Dec-POMDPs as Continuous-State MDPsabstractDecentralized partially observable Markov decision processes (Dec-POMDPs) provide a general model for decision-making under uncertainty in decentralized settings, but are difficult to solve optimally (NEXP-Complete). As a new way of solving these problems, we introduce the idea of transforming a Dec-POMDP into a continuous-state deterministic MDP with a piecewise-linear and convex value function. This approach makes use of the fact that planning can be accomplished in a centralized offline manner, while execution can still be decentralized. This new Dec-POMDP formulation, which we call an occupancy MDP, allows powerful POMDP and continuous-state MDP methods to be used for the first time. To provide scalability, we refine this approach by combining heuristic search and compact representations that exploit the structure present in multi-agent domains, without losing the ability to converge to an optimal solution. In particular, we introduce a feature-based heuristic search value iteration (FB-HSVI) algorithm that relies on feature-based compact representations, point-based updates and efficient action selection. A theoretical analysis demonstrates that FB-HSVI terminates in finite time with an optimal solution. We include an extensive empirical analysis using well-known benchmarks, thereby demonstrating that our approach provides significant scalability improvements compared to the state of the art. Jilles Steeve Dibangoye, Christopher Amato, Olivier Buffet, François Charpillet |
J. Artif. Intell. Res. | 1 |
| 2015 | Exploiting Separability in Multiagent Planning with Continuous-State MDPs (Extended Abstract)
Jilles Steeve Dibangoye, Christopher Amato, Olivier Buffet, François Charpillet |
IJCAI | 1 |
| 2015 | Structural Results for Cooperative Decentralized Control Models
Jilles Steeve Dibangoye, Olivier Buffet, Olivier Simonin 0001 |
IJCAI | 1 |
| 2015 | Distributed economic dispatch of embedded generation in smart grids
Jilles Steeve Dibangoye, Arnaud Doniec, Hicham Fakham, Frédéric Colas, Xavier Guillaud |
Eng. Appl. Artif. Intell. | 1 |
| 2014 | Error-Bounded Approximations for Infinite-Horizon Discounted Decentralized POMDPs
Jilles Steeve Dibangoye, Olivier Buffet, François Charpillet |
ECML/PKDD (1) | 1 |
| 2013 | Optimally Solving Dec-POMDPs as Continuous-State MDPs
Jilles Steeve Dibangoye, Christopher Amato, Olivier Buffet, François Charpillet |
IJCAI | 1 |
| 2012 | Scaling Up Decentralized MDPs Through Heuristic Search
Jilles Steeve Dibangoye, Christopher Amato, Arnaud Doniec |
UAI | 1 |
| 2009 | Topological Order Planner for POMDPs
Jilles Steeve Dibangoye, Guy Shani, Brahim Chaib-draa, Abdel-Illah Mouaddib |
IJCAI | 1 |