João V. Messias

dblp:22/11107 · also João Vicente Messias · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Autonomous driving · 33% Reinforcement learning · 22% Multi-agent systems · 21%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
scenario generation
0.812024
UniGen: Unified Modeling of Initial Agent States and Trajectories for Generating Autonomous Driving Scenarios · ICRA 2024
Robotics › Autonomous driving › scenario generation
traffic scenario generation
0.812024
UniGen: Unified Modeling of Initial Agent States and Trajectories for Generating Autonomous Driving Scenarios · ICRA 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent planning
0.422017
The MADP Toolbox: An Open Source Library for Planning and Learning in (Multi-)Agent Systems · J. Mach. Learn. Res. 2017
Efficient Offline Communication Policies for Factored Multiagent POMDPs · NIPS 2011
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.312017
Rapidly exploring learning trees · ICRA 2017
Machine learning › Reinforcement learning
model-based reinforcement learning
0.312017
Dynamic-Depth Context Tree Weighting · NIPS 2017
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making
0.312017
The MADP Toolbox: An Open Source Library for Planning and Learning in (Multi-)Agent Systems · J. Mach. Learn. Res. 2017
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.312017
Dynamic-Depth Context Tree Weighting · NIPS 2017
Knowledge, reasoning and agents › Multi-agent systems
partially observable stochastic games
0.312017
The MADP Toolbox: An Open Source Library for Planning and Learning in (Multi-)Agent Systems · J. Mach. Learn. Res. 2017
Robotics › Motion planning and robot control
path planning
0.312017
Rapidly exploring learning trees · ICRA 2017
Machine learning › Reinforcement learning
markov decision process
0.212013
GSMDPs for Multi-Robot Sequential Decision-Making · AAAI 2013
Knowledge, reasoning and agents › Multi-agent systems
multi-robot coordination
0.212013
GSMDPs for Multi-Robot Sequential Decision-Making · AAAI 2013
Knowledge, reasoning and agents › Multi-agent systems › decentralized planning
Dec-POMDP
0.112011
Efficient Offline Communication Policies for Factored Multiagent POMDPs · NIPS 2011
Robotics › Robot manipulation
learning from demonstration
0.112017
Rapidly exploring learning trees · ICRA 2017
Robotics › Robot manipulation › industrial robot
collaborative robot
0.012013
GSMDPs for Multi-Robot Sequential Decision-Making · AAAI 2013

Methods — techniques the papers use, named apart from their topics

global scenario embedding · 0.8autoregressive agent injection · 0.8variable-order markov model · 0.3suffix tree · 0.3reinforcement learning · 0.3maximum margin planning · 0.3decision-theoretic planning · 0.3caching · 0.3generalized semi-markov decision process · 0.2approximate solvers · 0.2
YearPublicationVenuePosition
2024 UniGen: Unified Modeling of Initial Agent States and Trajectories for Generating Autonomous Driving Scenarios
abstract
This paper introduces UniGen, a novel approach to generating new traffic scenarios for evaluating and improving autonomous driving software through simulation. Our approach models all driving scenario elements in a unified model: the position of new agents, their initial state, and their future motion trajectories. By predicting the distributions of all these variables from a shared global scenario embedding, we ensure that the final generated scenario is fully conditioned on all available context in the existing scene. Our unified modeling approach, combined with autoregressive agent injection, conditions the placement and motion trajectory of every new agent on all existing agents and their trajectories, leading to realistic scenarios with low collision rates. Our experimental results show that UniGen outperforms prior state of the art on the Waymo Open Motion Dataset.
Reza Mahjourian, Rongbing Mu, Valerii Likhosherstov, Paul Mougin, Xiukun Huang, João V. Messias, Shimon Whiteson
ICRA6
2019 Learning From Demonstration in the Wild
abstract
Learning from demonstration (LfD) is useful in settings where hand-coding behaviour or a reward function is impractical. It has succeeded in a wide range of problems but typically relies on manually generated demonstrations or specially deployed sensors and has not generally been able to leverage the copious demonstrations available in the wild: those that capture behaviours that were occurring anyway using sensors that were already deployed for another purpose, e.g., traffic camera footage capturing demonstrations of natural behaviour of vehicles, cyclists, and pedestrians. We propose video to behaviour (ViBe), a new approach to learn models of behaviour from unlabelled raw video data of a traffic scene collected from a single, monocular, initially uncalibrated camera with ordinary resolution. Our approach calibrates the camera, detects relevant objects, tracks them through time, and uses the resulting trajectories to perform LfD, yielding models of naturalistic behaviour. We apply ViBe to raw videos of a traffic intersection and show that it can learn purely from videos, without additional expert knowledge.
Feryal M. P. Behbahani, Kyriacos Shiarlis, Vitaly Kurin, Sudhanshu Kasewa, Ciprian Stirbu, Supratik Paul, Frans A. Oliehoek, João V. Messias, Shimon Whiteson
ICRA10
2017 Rapidly exploring learning trees
abstract
Inverse Reinforcement Learning (IRL) for path planning enables robots to learn cost functions for difficult tasks from demonstration, instead of hard-coding them. However, IRL methods face practical limitations that stem from the need to repeat costly planning procedures. In this paper, we propose Rapidly Exploring Learning Trees (RLT*), which learns the cost functions of Optimal Rapidly Exploring Random Trees (RRT*) from demonstration, thereby making inverse learning methods applicable to more complex tasks. Our approach extends Maximum Margin Planning to work with RRT* cost functions. Furthermore, we propose a caching scheme that greatly reduces the computational cost of this approach. Experimental results on simulated and real-robot data from a social navigation scenario show that RLT* achieves better performance at lower computational cost than existing methods. We also successfully deploy control policies learned with RLT* on a real telepresence robot.
Kyriacos Shiarlis, João V. Messias, Shimon Whiteson
ICRA2
2017 Acquiring social interaction behaviours for telepresence robots via deep learning from demonstration
abstract
As robots begin to inhabit public and social spaces, it is increasingly important to ensure that they behave in a socially appropriate way. However, manually coding social behaviours is prohibitively difficult since social norms are hard to quantify. Therefore, learning from demonstration (LfD), wherein control policies are inferred from demonstrations of correct behaviour, is a powerful tool for helping robots acquire social intelligence. In this paper, we propose a deep learning approach to learning social behaviours from demonstration. We apply this method to two challenging social tasks for a semi-autonomous telepresence robot. Our results show that our approach outperforms gradient boosting regression and performs well against a hard-coded controller. Furthermore, ablation experiments confirm that each element of our method is essential to its success.
Kyriacos Shiarlis, João V. Messias, Shimon Whiteson
IROS2
2017 Dynamic-Depth Context Tree Weighting
abstract
Reinforcement learning (RL) in partially observable settings is challenging because the agent’s observations are not Markov. Recently proposed methods can learn variable-order Markov models of the underlying process but have steep memory requirements and are sensitive to aliasing between observation histories due to sensor noise. This paper proposes dynamic-depth context tree weighting (D2-CTW), a model-learning method that addresses these limitations. D2-CTW dynamically expands a suffix tree while ensuring that the size of the model, but not its depth, remains bounded. We show that D2-CTW approximately matches the performance of state-of-the-art alternatives at stochastic time-series prediction while using at least an order of magnitude less memory. We also apply D2-CTW to model-based RL, showing that, on tasks that require memory of past observations, D2-CTW can learn without prior knowledge of a good state representation, or even the length of history upon which such a representation should depend.
João V. Messias, Shimon Whiteson
NIPS1
2017 The MADP Toolbox: An Open Source Library for Planning and Learning in (Multi-)Agent Systems
abstract
This article describes the Multiagent Decision Process (MADP) Toolbox, a software library to support planning and learning for intelligent agents and multiagent systems in uncertain environments. Key features are that it supports partially observable environments and stochastic transition models; has unified support for single- and multiagent systems; provides a large number of models for decision-theoretic decision making, including one-shot and sequential decision making under various assumptions of observability and cooperation, such as Dec-POMDPs and POSGs; provides tools and parsers to quickly prototype new problems; provides an extensive range of planning and learning algorithms for single- and multiagent systems; is released under the GNU GPL v3 license; and is written in C++ and designed to be extensible via the object-oriented paradigm.
Frans A. Oliehoek, Matthijs T. J. Spaan, Bas Terwijn, Philipp Robbel, João V. Messias
J. Mach. Learn. Res.5
2013 GSMDPs for Multi-Robot Sequential Decision-Making
abstract
Markov Decision Processes (MDPs) provide an extensive theoretical background for problems of decision-making under uncertainty. In order to maintain computational tractability, however, real-world problems are typically discretized in states and actions as well as in time. Assuming synchronous state transitions and actions at fixed rates may result in models which are not strictly Markovian, or where agents are forced to idle between actions, losing their ability to react to sudden changes in the environment. In this work, we explore the application of Generalized Semi-Markov Decision Processes (GSMDPs) to a realistic multi-robot scenario. A case study will be presented in the domain of cooperative robotics, where real-time reactivity must be preserved, and synchronous discrete-time approaches are therefore sub-optimal. This case study is tested on a team of real robots, and also in realistic simulation. By allowing asynchronous events to be modeled over continuous time, the GSMDP approach is shown to provide greater solution quality than its discrete-time counterparts, while still being approximately solvable by existing methods.
João V. Messias, Matthijs T. J. Spaan, Pedro U. Lima
AAAI1
2011 Efficient Offline Communication Policies for Factored Multiagent POMDPs
abstract
Factored Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs) form a powerful framework for multiagent planning under uncertainty, but optimal solutions require a rigid history-based policy representation. In this paper we allow inter-agent communication which turns the problem in a centralized Multiagent POMDP (MPOMDP). We map belief distributions over state factors to an agent's local actions by exploiting structure in the joint MPOMDP policy. The key point is that when sparse dependencies between the agents' decisions exist, often the belief over its local state factors is sufficient for an agent to unequivocally identify the optimal action, and communication can be avoided. We formalize these notions by casting the problem into convex optimization form, and present experimental results illustrating the savings in communication that we can obtain.
João V. Messias, Matthijs T. J. Spaan, Pedro U. Lima
NIPS1