EDBT 2026 Demo / reviewers in the wild / expert
Chen Tang 0001
dblp:71/7642-1
· DBLP profile ↗
17ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-7536-9983ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 11 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Reinforcement Learning for Robotics: A Survey of Real-World SuccessesabstractReinforcement learning (RL), particularly its combination with deep neural networks referred to as deep RL (DRL), has shown tremendous promise across a wide range of applications, suggesting its potential for enabling the development of sophisticated robotic behaviors. Robotics problems, however, pose fundamental difficulties for the application of RL, stemming from the complexity and cost of interacting with the physical world. These challenges notwithstanding, recent advances have enabled DRL to succeed at some real-world robotic tasks. However, state-of-the-art DRL solutions’ maturity varies significantly across robotic applications. In this talk, I will review the current progress of DRL in real-world robotic applications based on our recent survey paper (with Tang, Abbatematteo, Hu, Chandra, and Martı́n-Martı́n), with a particular focus on evaluating the real-world successes achieved with DRL in realizing several key robotic competencies, including locomotion, navigation, stationary manipulation, mobile manipulation, human-robot interaction, and multi-robot interaction. The analysis aims to identify the key factors underlying those exciting successes, reveal underexplored areas, and provide an overall characterization of the status of DRL in robotics. I will also highlight several important avenues for future work, emphasizing the need for stable and sample-efficient real-world RL paradigms, holistic approaches for discovering and integrating various competencies to tackle complex long-horizon, open-world tasks, and principled development and evaluation procedures. The talk is designed to offer insights for RL practitioners and roboticists toward harnessing RL’s power to create generally capable real-world robotic systems. Chen Tang 0001, Ben Abbatematteo, Jiaheng Hu, Rohan Chandra, Roberto Martin Martin, Peter Stone 0001 |
AAAI | 1 |
| 2025 | Residual-MPPI: Online Policy Customization for Continuous ControlabstractPolicies developed through Reinforcement Learning (RL) and Imitation Learning (IL) have shown great potential in continuous control tasks, but real-world applications often require adapting trained policies to unforeseen requirements. While fine-tuning can address such needs, it typically requires additional data and access to the original training metrics and parameters.
In contrast, an online planning algorithm, if capable of meeting the additional requirements, can eliminate the necessity for extensive training phases and customize the policy without knowledge of the original training scheme or task. In this work, we propose a generic online planning algorithm for customizing continuous-control policies at the execution time, which we call Residual-MPPI. It can customize a given prior policy on new performance metrics in few-shot and even zero-shot online settings, given access to the prior action distribution alone. Through our experiments, we demonstrate that the proposed Residual-MPPI algorithm can accomplish the few-shot/zero-shot online policy customization task effectively, including customizing the champion-level racing agent, Gran Turismo Sophy (GT Sophy) 1.0, in the challenging car racing scenario, Gran Turismo Sport (GTS) environment. Code for MuJoCo experiments is included in the supplementary and will be open-sourced upon acceptance. Demo videos are available on our website: https://sites.google.com/view/residual-mppi. Pengcheng Wang 0004, Chenran Li, Catherine Weaver, Kenta Kawamoto, Masayoshi Tomizuka, Chen Tang 0001 |
ICLR | 6 |
| 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in DrivingabstractLanguage models uncover unprecedented abilities in analyzing driving scenarios, owing to their limitless knowledge accumulated from text-based pre-training. Naturally, they should particularly excel in analyzing rule-based interactions, such as those triggered by traffic laws, which are well documented in texts. However, such interaction analysis remains underexplored due to the lack of dedicated language datasets that address it. Therefore, we propose Waymo Open Motion Dataset-Reasoning (WOMD-Reasoning), a comprehensive large-scale Q&As dataset built on WOMD focusing on describing and reasoning traffic rule-induced interactions in driving scenarios. WOMD-Reasoning also presents by far the largest multi-modal Q&A dataset, with 3 million Q&As on real-world driving scenarios, covering a wide range of driving topics from map descriptions and motion status descriptions to narratives and analyses of agents' interactions, behaviors, and intentions. To showcase the applications of WOMD-Reasoning, we design Motion-LLaVA, a motion-language model fine-tuned on WOMD-Reasoning. Quantitative and qualitative evaluations are performed on WOMD-Reasoning dataset as well as the outputs of Motion-LLaVA, supporting the data quality and wide applications of WOMD-Reasoning, in interaction predictions, traffic rule compliance plannings, etc. The dataset and its vision modal extension are available on https://waymo.com/open/download/. The codes & prompts to build it are available on https://github.com/yhli123/WOMD-Reasoning. Cunxin Fan, Chongjian Ge, Seth Z. Zhao, Chenran Li, Chenfeng Xu, Huaxiu Yao, Masayoshi Tomizuka, Bolei Zhou, Chen Tang 0001, Mingyu Ding |
ICML | 10 |
| 2024 | Optimizing Diffusion Models for Joint Trajectory Prediction and Controllable Generation
Chen Tang 0001, Lingfeng Sun, Simone Rossi 0001, Yichen Xie 0002, Chensheng Peng, Thomas Hannagan, Stefano Sabatini, Nicola Poerio, Masayoshi Tomizuka |
ECCV (29) | 2 |
| 2024 | Guided Online Distillation: Promoting Safe Reinforcement Learning by Offline DemonstrationabstractSafe Reinforcement Learning (RL) aims to find a policy that achieves high rewards while satisfying cost constraints. When learning from scratch, safe RL agents tend to be overly conservative, which impedes exploration and restrains the overall performance. In many realistic tasks, e.g. autonomous driving, large-scale expert demonstration data are available. We argue that extracting expert policy from offline data to guide online exploration is a promising solution to mitigate the conserveness issue. Large-capacity models, e.g. decision transformers (DT), have been proven to be competent in offline policy learning. However, data collected in realworld scenarios rarely contain dangerous cases (e.g., collisions), which makes it prohibitive for the policies to learn safety concepts. Besides, these bulk policy networks cannot meet the computation speed requirements at inference time on real-world tasks such as autonomous driving. To this end, we propose Guided Online Distillation (GOLD), an offline-to-online safe RL framework. GOLD distills an offline DT policy into a lightweight policy network through guided online safe RL training, which outperforms both the offline DT policy and online safe RL algorithms. Experiments in both benchmark safe RL tasks and real-world driving tasks based on the Waymo Open Motion Dataset (WOMD) [1] demonstrate that GOLD can successfully distill lightweight policies and solve decision-making problems in challenging safety-critical scenarios. Jinning Li 0002, Banghua Zhu, Jiantao Jiao, Masayoshi Tomizuka, Chen Tang 0001 |
ICRA | 6 |
| 2024 | Pre-training on Synthetic Driving Data for Trajectory PredictionabstractAccumulating substantial volumes of real-world driving data proves pivotal in the realm of trajectory forecasting for autonomous driving. Given the heavy reliance of current trajectory forecasting models on data-driven methodologies, we aim to tackle the challenge of learning general trajectory forecasting representations under limited data availability. We propose a pipeline-level solution to mitigate the issue of data scarcity in trajectory forecasting. The solution is composed of two parts: firstly, we adopt HD map augmentation and trajectory synthesis for generating driving data, and then we learn representations by pre-training on them. Specifically, we apply vector transformations to reshape the maps, and then employ a rule-based model to generate trajectories on both original and augmented scenes; thus enlarging the driving data without collecting additional real ones. To foster the learning of general representations within this augmented dataset, we comprehensively explore the different pre-training strategies, including extending the concept of a Masked AutoEncoder (MAE) for trajectory forecasting. Without bells and whistles, our proposed pipeline-level solution is general, simple, yet effective: we conduct extensive experiments to demonstrate the effectiveness of our data expansion and pre-training strategies, which outperform the baseline prediction model by large margins, e.g. 5.04%, 3.84% and 8.30% in terms of MR6, minADE6and minFDE6. The pre-training dataset and the codes for pre-training and fine-tuning are released at https://github.com/yhli123/Pretraining_on_Synthetic_Driving_Data_for_Trajectory_Prediction. Seth Z. Zhao, Chenfeng Xu, Chen Tang 0001, Chenran Li, Mingyu Ding, Masayoshi Tomizuka |
IROS | 4 |
| 2024 | Grounded Relational Inference: Domain Knowledge Driven Explainable Autonomous DrivingabstractExplainability is essential for autonomous vehicles and other robotics systems interacting with humans and other objects during operation. Humans need to understand and anticipate the actions taken by machines for trustful and safe cooperation. In this work, we aim to develop an explainable model that generates explanations consistent with both human domain knowledge and the model’s inherent causal relation. In particular, we focus on an essential building block of autonomous driving—multi-agent interaction modeling. We propose Grounded Relational Inference (GRI). It models an interactive system’s underlying dynamics by inferring an interaction graph representing the agents’ relations. We ensure a semantically meaningful interaction graph by grounding the relational latent space into semantic interactive behaviors defined with expert domain knowledge. We demonstrate that it can model interactive traffic scenarios under both simulation and real-world settings and generate semantic graphs explaining the vehicle’s behavior through their interactions. Chen Tang 0001, Nishan Srishankar, Sujitha Martin, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Residual Q-Learning: Offline and Online Policy Customization without ValueabstractImitation Learning (IL) is a widely used framework for learning imitative behavior from demonstrations. It is especially appealing for solving complex real-world tasks where handcrafting reward function is difficult, or when the goal is to mimic human expert behavior. However, the learned imitative policy can only follow the behavior in the demonstration. When applying the imitative policy, we may need to customize the policy behavior to meet different requirements coming from diverse downstream tasks. Meanwhile, we still want the customized policy to maintain its imitative nature. To this end, we formulate a new problem setting called policy customization. It defines the learning task as training a policy that inherits the characteristics of the prior policy while satisfying some additional requirements imposed by a target downstream task. We propose a novel and principled approach to interpret and determine the trade-off between the two task objectives. Specifically, we formulate the customization problem as a Markov Decision Process (MDP) with a reward function that combines 1) the inherent reward of the demonstration; and 2) the add-on reward specified by the downstream task. We propose a novel framework, Residual Q-learning, which can solve the formulated MDP by leveraging the prior policy without knowing the inherent reward or value function of the prior policy. We derive a family of residual Q-learning algorithms that can realize offline and online policy customization, and show that the proposed algorithms can effectively accomplish policy customization tasks in various environments. Demo videos and code are available on our website: https://sites.google.com/view/residualq-learning. Chenran Li, Chen Tang 0001, Haruki Nishimura, Jean Mercat, Masayoshi Tomizuka |
NeurIPS | 2 |
| 2022 | PreTraM: Self-supervised Pre-training via Connecting Trajectory and Map
Chenfeng Xu, Chen Tang 0001, Lingfeng Sun, Kurt Keutzer, Masayoshi Tomizuka, Alireza Fathi |
ECCV (39) | 3 |
| 2022 | Domain Knowledge Driven Pseudo Labels for Interpretable Goal-Conditioned Interactive Trajectory PredictionabstractMotion forecasting in highly interactive scenarios is a challenging problem in autonomous driving. In such scenarios, we need to accurately predict the joint behavior of interacting agents to ensure the safe and efficient navigation of autonomous vehicles. Recently, goal-conditioned methods have gained increasing attention due to their advantage in performance and their ability to capture the multimodality in trajec-tory distribution. In this work, we study the joint trajectory prediction problem with the goal-conditioned framework. In particular, we introduce a conditional-variational-autoencoder-based (CVAE) model to explicitly encode different interaction modes into the latent space. However, we discover that the vanilla model suffers from posterior collapse and cannot induce an informative latent space as desired. To address these issues, we propose a novel approach to avoid KL vanishing and induce an interpretable interactive latent space with pseudo labels. The proposed pseudo labels allow us to incorporate domain knowledge on interaction in a flexible manner. We motivate the proposed method using an illustrative toy example. In addition, we validate our framework on the Waymo Open Motion Dataset with both quantitative and qualitative evaluations. Lingfeng Sun, Chen Tang 0001, Yaru Niu, Enna Sachdeva, Chiho Choi, Teruhisa Misu, Masayoshi Tomizuka |
IROS | 2 |
| 2022 | Interventional Behavior Prediction: Avoiding Overly Confident Anticipation in Interactive PredictionabstractConditional behavior prediction (CBP) builds up the foundation for a coherent interactive prediction and plan-ning framework that can enable more efficient and less conser-vative maneuvers in interactive scenarios. In CBP task, we train a prediction model approximating the posterior distribution of target agents' future trajectories conditioned on the future trajectory of an assigned ego agent. However, we argue that CBP may provide overly confident anticipation on how the autonomous agent may influence the target agents' behavior. Consequently, it is risky for the planner to query a CBP model. Instead, we should treat the planned trajectory as an intervention and let the model learn the trajectory distribution under intervention. We refer to it as the interventional behavior prediction (IBP) task. Moreover, to properly evaluate an IBP model with offline datasets, we propose a Shapley-value-based metric to verify if the prediction model satisfies the inherent temporal independence of an interventional distribution. We show that the proposed metric can effectively identify a CBP model violating the temporal independence, which plays an important role when establishing IBP benchmarks. Chen Tang 0001, Masayoshi Tomizuka |
IROS | 1 |
| 2021 | Exploring Social Posterior Collapse in Variational Autoencoder for Interaction ModelingabstractMulti-agent behavior modeling and trajectory forecasting are crucial for the safe navigation of autonomous agents in interactive scenarios. Variational Autoencoder (VAE) has been widely applied in multi-agent interaction modeling to generate diverse behavior and learn a low-dimensional representation for interacting systems. However, existing literature did not formally discuss if a VAE-based model can properly encode interaction into its latent space. In this work, we argue that one of the typical formulations of VAEs in multi-agent modeling suffers from an issue we refer to as social posterior collapse, i.e., the model is prone to ignoring historical social context when predicting the future trajectory of an agent. It could cause significant prediction errors and poor generalization performance. We analyze the reason behind this under-explored phenomenon and propose several measures to tackle it. Afterward, we implement the proposed framework and experiment on real-world datasets for multi-agent trajectory prediction. In particular, we propose a novel sparse graph attention message-passing (sparse-GAMP) layer, which helps us detect social posterior collapse in our experiments. In the experiments, we verify that social posterior collapse indeed occurs. Also, the proposed measures are effective in alleviating the issue. As a result, the model attains better generalization performance when historical social context is informative for prediction. Chen Tang 0001, Masayoshi Tomizuka |
NeurIPS | 1 |
| 2020 | Application Specific System Identification for Model-Based Control in Self-Driving CarsabstractLinear Parameter Varying (LPV) models can be used to describe the vehicular lateral dynamic behavior of self-driving cars. They are particularly suitable for model-based control schemes such as model predictive control (MPC) applied to real-time trajectory tracking control, since they provide a proper trade-off between accuracy in different scenarios and reduced computation cost compared to nonlinear models. The MPC control schemes use the model for a long prediction horizon of the states, therefore prediction errors for a long time horizon should be minimized in order to increase the accuracy of the tracking. For this task, this work presents a system identification procedure for the lateral dynamics of a vehicle that combines a LPV model with a learning algorithm that has been successfully applied to other dynamic systems in the past. Simulation results show the benefits of the identified model in comparison to other well-known vehicular lateral dynamic models. Julian M. Salt Ducaju, Chen Tang 0001, Masayoshi Tomizuka, Ching-Yao Chan |
IV | 2 |
| 2020 | Disturbance-Observer-Based Tracking Controller for Neural Network Driving Policy TransferabstractThe neural network policies are widely explored in the autonomous driving field, thanks to their capability of handling complicated driving tasks. However, the practical deployment of such policies is slowed down due to their lack of robustness against modeling gap and external disturbances. In our prior work, we proposed a planner-controller architecture and applied a disturbance-observer-based (DOB) robust tracking controller to reject the disturbances and achieved zero-shot policy transfer. In this paper, we present our latest progress on improving the policy transfer performance under this framework. Concretely, we applied adaptive DOB, so as to more accurately model the inverse system dynamics and increase the cut-off frequency of the Q-filter in the DOB. A closed-loop reference path smoothing algorithm is introduced to alleviate the step disturbance input imposed by the reference trajectory re-planning. On the neural network control policy side, we applied the parallel attribute networks, a hierarchical modular policy network to dynamically handle various driving tasks. We have carried out various simulations and experiments to validate the capability of our proposed method to achieve sim-to-sim and sim-to-real policy transfer. The proposed method achieves the most outstanding performance among a series of baseline control schemes. Chen Tang 0001, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Adaptive Probabilistic Vehicle Trajectory Prediction Through Physically Feasible Bayesian Recurrent Neural NetworkabstractProbabilistic vehicle trajectory prediction is essential for robust safety of autonomous driving. Current methods for long-term trajectory prediction cannot guarantee the physical feasibility of predicted distribution. Moreover, their models cannot adapt to the driving policy of the predicted target human driver. In this work, we propose to overcome these two shortcomings by a Bayesian recurrent neural network model consisting of Bayesian-neural-network-based policy model and known physical model of the scenario. Bayesian neural network can ensemble complicated output distribution, enabling rich family of trajectory distribution. The embedded physical model ensures feasibility of the distribution. Moreover, the adopted gradient-based training method allows direct optimization for better performance in long prediction horizon. Furthermore, a particle-filter-based parameter adaptation algorithm is designed to adapt the policy Bayesian neural network to the predicted target online. Effectiveness of the proposed methods is verified with a toy example with multi-modal stochastic feedback gain and naturalistic car following data. Chen Tang 0001, Jianyu Chen 0002, Masayoshi Tomizuka |
ICRA | 1 |
| 2019 | Toward Modularization of Neural Network Autonomous Driving Policy Using Parallel Attribute NetworksabstractNeural network autonomous driving policies are widely explored. However, no matter using imitation learning or reinforcement learning, the network policies are generally hard to train, and the learned knowledge encoded in neural network policies are hard to transfer. We propose to modularize the complicated driving policies in terms of the driving attributes, and present the parallel attribute networks (PAN), which can learn to fullfill the requirements of the attributes in the driving tasks separately, and later assemble their knowledge together. Concretely, we first train a policy network that accomplish the base lane tracking attribute. The modules for the add-on attributes such as avoiding obstacles and obeying traffic rules are then trained to map the corresponding state to a satisfactory set of the vehicle action space. Finally the reference action given by the base policy is projected into the satisfactory sets so as to satisfy the requirements of all the attributes. Using the PAN, many complicated tasks that are hard to train from scratch can be easily trained; also unseen driving tasks can be solved in a zero-shot manner by assembling the pretrained attribute modules. We have validated the capability of our model on a class of autonomous driving problems with attributes of obstacle avoidance, traffic light and speed limit in simulation. Experimental results based on an obstacle avoidance task are also presented. Haonan Chang, Chen Tang 0001, Changliu Liu, Masayoshi Tomizuka |
IV | 3 |
| 2018 | Continuous Decision Making for On-road Autonomous Driving under Uncertain and Interactive EnvironmentsabstractAlthough autonomous driving techniques have achieved great improvements, challenges still exist in decision making for variety of different scenarios under uncertain and interactive environments. A good decision maker must satisfy the following requirements: (1) Be in a generic and unified form to cover as more scenarios as possible. (2) Be able to interact properly with other moving obstacles under the uncertainty of their motions. In this paper, the continuous decision making (CDM) framework is proposed to formulate different driving scenarios in a unified way, which encodes the high level decision making information into a continuous reference trajectory that can be naturally combined with a lower level trajectory planner. Within the framework, a maximum interaction defensive policy (MIDP) is proposed, which calculates the best action to interact with stochastic moving obstacles while guaranteeing safety. The method is applied to a ramp merging scenario and the stochastic behavior models of the surrounding vehicles are learned from the NGSIM dataset. Simulations are shown to visualize and analyze the results. Jianyu Chen 0002, Chen Tang 0001, Long Xin, Shengbo Eben Li, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 2 |