Xiaobai Ma

dblp:217/7848 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
2since 2021 · last 2022
0000-0001-7491-3935ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 45% Graph learning · 22% Robot navigation and mapping · 11%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
0.622021
Handling Missing Data with Graph Representation Learning · NeurIPS 2020
Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships · ICRA 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.612022
Recursive Reasoning Graph for Multi-Agent Reinforcement Learning · AAAI 2022
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.612022
Recursive Reasoning Graph for Multi-Agent Reinforcement Learning · AAAI 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
recursive reasoning
0.612022
Recursive Reasoning Graph for Multi-Agent Reinforcement Learning · AAAI 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
latent state inference
0.512021
Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships · ICRA 2021
Robotics › Robot navigation and mapping › learning-based navigation
reinforcement-learning-based navigation
0.512021
Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships · ICRA 2021
Machine learning › Graph learning › heterogeneous graph learning › heterogeneous graph representation learning
bipartite graph representation learning
0.412020
Handling Missing Data with Graph Representation Learning · NeurIPS 2020
Data integration and cleaning › missing data › missing value imputation
graph-based imputation
0.412020
Handling Missing Data with Graph Representation Learning · NeurIPS 2020
Data integration and cleaning › missing data
missing value imputation
0.412020
Handling Missing Data with Graph Representation Learning · NeurIPS 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.412019
Monte Carlo Tree Search for Policy Optimization · IJCAI 2019
Machine learning › Reinforcement learning
policy optimization
0.412019
Monte Carlo Tree Search for Policy Optimization · IJCAI 2019
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal relation modeling
0.112021
Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships · ICRA 2021

Methods — techniques the papers use, named apart from their topics

graph neural network · 1.9node-level prediction · 0.9edge-level prediction · 0.9recursive reasoning model · 0.6supervised learning · 0.5deep reinforcement learning · 0.5upper confidence bound · 0.4gradient-free optimization · 0.4
YearPublicationVenuePosition
2022 Recursive Reasoning Graph for Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) provides an efficient way for simultaneously learning policies for multiple agents interacting with each other. However, in scenarios requiring complex interactions, existing algorithms can suffer from an inability to accurately anticipate the influence of self-actions on other agents. Incorporating an ability to reason about other agents' potential responses can allow an agent to formulate more effective strategies. This paper adopts a recursive reasoning model in a centralized-training-decentralized-execution framework to help learning agents better cooperate with or compete against others. The proposed algorithm, referred to as the Recursive Reasoning Graph (R2G), shows state-of-the-art performance on multiple multi-agent particle and robotics games.
Xiaobai Ma, David Isele, Jayesh K. Gupta, Kikuo Fujimura, Mykel J. Kochenderfer
AAAI1
2021 Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships
abstract
Deep reinforcement learning (DRL) provides a promising way for learning navigation in complex autonomous driving scenarios. However, identifying the subtle cues that can indicate drastically different outcomes remains an open problem with designing autonomous systems that operate in human environments. In this work, we show that explicitly inferring the latent state and encoding spatial-temporal relationships in a reinforcement learning framework can help address this difficulty. We encode prior knowledge on the latent states of other drivers through a framework that combines the reinforcement learner with a supervised learner. In addition, we model the influence passing between different vehicles through graph neural networks (GNNs). The proposed framework significantly improves performance in the context of navigating T-intersections compared with state-of-the-art baseline approaches.
Xiaobai Ma, Jiachen Li 0001, Mykel J. Kochenderfer, David Isele, Kikuo Fujimura
ICRA1
2020 Handling Missing Data with Graph Representation Learning
abstract
Machine learning with missing data has been approached in many different ways, including feature imputation where missing feature values are estimated based on observed values and label prediction where downstream labels are learned directly from incomplete data. However, existing imputation models tend to have strong prior assumptions and cannot learn from downstream tasks, while models targeting label predictions often involve heuristics and can encounter scalability issues. Here we propose GRAPE, a framework for feature imputation as well as label prediction. GRAPE tackles the missing data problem using graph representation, where the observations and features are viewed as two types of nodes in a bipartite graph, and the observed feature values as edges. Under the GRAPE framework, the feature imputation is formulated as an edge-level prediction task and the label prediction as a node-level prediction task. These tasks are then solved with Graph Neural Networks. Experimental results on nine benchmark datasets show that GRAPE yields 20% lower mean absolute error for imputation tasks and 10% lower for label prediction tasks, compared with existing state-of-the-art methods.
Jiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer, Jure Leskovec
NeurIPS2
2019 Monte Carlo Tree Search for Policy Optimization
abstract
Gradient-based methods are often used for policy optimization in deep reinforcement learning, despite being vulnerable to local optima and saddle points. Although gradient-free methods (e.g., genetic algorithms or evolution strategies) help mitigate these issues, poor initialization and local optima are still concerns in highly nonconvex spaces. This paper presents a method for policy optimization based on Monte-Carlo tree search and gradient-free optimization. Our method, called Monte-Carlo tree search for policy optimization (MCTSPO), provides a better exploration-exploitation trade-off through the use of the upper confidence bound heuristic. We demonstrate improved performance on reinforcement learning tasks with deceptive or sparse reward functions compared to popular gradient-based and deep genetic algorithm baselines.
Xiaobai Ma, Katherine Rose Driggs-Campbell, Zongzhang Zhang, Mykel J. Kochenderfer
IJCAI1
2018 Improved Robustness and Safety for Autonomous Vehicle Control with Adversarial Reinforcement Learning
abstract
To improve efficiency and reduce failures in autonomous vehicles, research has focused on developing robust and safe learning methods that take into account disturbances in the environment. Existing literature in robust reinforcement learning poses the learning problem as a two player game between the autonomous system and disturbances. This paper examines two different algorithms to solve the game, Robust Adversarial Reinforcement Learning and Neural Fictitious Self Play, and compares performance on an autonomous driving scenario. We extend the game formulation to a semi-competitive setting and demonstrate that the resulting adversary better captures meaningful disturbances that lead to better overall performance. The resulting robust policy exhibits improved driving efficiency while effectively reducing collision rates compared to baseline control policies produced by traditional reinforcement learning methods.
Xiaobai Ma, Katherine Rose Driggs-Campbell, Mykel J. Kochenderfer
Intelligent Vehicles Symposium1