EDBT 2026 Demo / reviewers in the wild / expert
Ding Zhao
dblp:135/7202
· DBLP profile ↗
101ranked-venue papers
6as first author
74since 2021 · last 2026
0000-0002-9400-8446ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 3 first-author · 58 since 2021Systems, architecture and hardware · 22 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 7 since 2021Computer networks · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tailored Primitive Initialization is the Secret Key to Reinforcement LearningabstractYihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang, Ding Zhao, Zhang-Wei Hong, Chuang Gan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang 0001, Ding Zhao, Zhang-Wei Hong, Chuang Gan 0001 |
ACL (1) | 5 |
| 2026 | PHY-aware TCP BBR in Wi-Fi Networks
Yen-Chin Wang, Chunghan Lee, Ding Zhao, Seyhan Ucar, Onur Altintas, Danijela Cabric |
ICC | 3 |
| 2025 | DiffScene: Diffusion-Based Safety-Critical Scenario Generation for Autonomous VehiclesabstractThe field of Autonomous Driving (AD) has witnessed significant progress in recent years. Among the various challenges faced, the safety evaluation of autonomous vehicles (AVs) stands out as a critical concern. Traditional evaluation methods are both costly and inefficient, often requiring extensive driving mileage in order to encounter rare safety-critical scenarios, which are distributed on the long tail of the complex real-world driving landscape. In this paper, we propose a unified approach, Diffusion-Based Safety-Critical Scenario Generation (DiffScene), to generate high-quality safety-critical scenarios which are both realistic and safety-critical for efficient AV evaluation. In particular, we propose a diffusion-based generation framework, leveraging the power of approximating the distribution of low-density spaces for diffusion models. We design several adversarial optimization objectives to guide the diffusion generation under predefined adversarial budgets. These objectives, such as safety-based objective, functionality-based objective, and constraint-based objective, ensure the generation of safety-critical scenarios while adhering to specific constraints. Extensive experimentation has been conducted to validate the efficacy of our approach. Compared with 6 SOTA baselines, DiffScene generates scenarios that are (1) more safety-critical under 3 metrics, (2) more realistic under 5 distance functions, and (3) more transferable to different AV algorithms. In addition, we demonstrate that training AV algorithms with scenarios generated by DiffScene leads to significantly higher performance in terms of the safety-critical metrics compared to baselines. These findings highlight the potential of DiffScene in addressing the challenges of AV safety evaluation, paving the way for safer AV development. Chejian Xu, Aleksandr Petiushko, Ding Zhao, Bo Li 0026 |
AAAI | 3 |
| 2025 | Causal Composition Diffusion Model for Closed-loop Traffic GenerationabstractSimulation is critical for safety evaluation in autonomous driving, particularly in capturing complex interactive behaviors. However, generating realistic and controllable traffic scenarios in long-tail situations remains a significant challenge. Existing generative models suffer from the conflicting objective between user-defined controllability and realism constraints, which is amplified in safety-critical contexts. In this work, we introduce the Causal Compositional Diffusion Model (CCDiff), a structure-guided diffusion framework to address these challenges. We first formulate the learning of controllable and realistic closed-loop simulation as a constrained optimization problem. Then, CCDiff maximizes controllability while adhering to realism by automatically identifying and injecting causal structures directly into the diffusion process, providing structured guidance to enhance both realism and controllability. Through rigorous evaluations on benchmark datasets and in a closed-loop simulator, CCDiff demonstrates substantial gains over state-of-the-art approaches in generating realistic and user-preferred trajectories. Our results show CCDiff’s effectiveness in extracting and leveraging causal structures, showing improved closed-loop performance based on key metrics such as collision rate, off-road rate, FDE, and comfort. For more details, welcome to check our project website. Haohong Lin, Tung Phan, David S. Hayden, Huan Zhang 0001, Ding Zhao, Siddhartha S. Srinivasa, Eric M. Wolff, Hongge Chen |
CVPR | 6 |
| 2025 | Learning Multi-Agent Loco-Manipulation for Long-Horizon Quadrupedal PushingabstractRecently, quadrupedal locomotion has achieved significant success, but their manipulation capabilities, particularly in handling large objects, remain limited, restricting their usefulness in demanding real-world applications such as search and rescue, construction, industrial automation, and room or-ganization. This paper tackles the task of obstacle-aware, long-horizon pushing by multiple quadrupedal robots. We propose a hierarchical multi-agent reinforcement learning framework with three levels of control. The high-level controller integrates an RRT planner and a centralized adaptive policy to generate subgoals, while the mid-level controller uses a decentralized goal-conditioned policy to guide the robots toward these sub-goals. A pre-trained low-level locomotion policy executes the movement commands. We evaluate our method against several baselines in simulation, demonstrating significant improvements over baseline approaches, with 36.0% higher success rates and 24.5% reduction in completion time than the best baseline. Our framework successfully enables long-horizon, obstacle-aware manipulation tasks like Push-Cuboid and Push-Ton Gol robots in the real world. Chuye Hong, Yaru Niu, Shiqi Liu 0005, Yuxiang Yang 0007, Ding Zhao |
ICRA | 6 |
| 2025 | CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World TrajectoriesabstractSimulation is an indispensable tool in the development and testing of autonomous vehicles (AVs), offering an efficient and safe alternative to road testing. An outstanding challenge with simulation-based testing is the generation of safety-critical scenarios, which are essential to ensure that AVs can handle rare but potentially fatal situations. This paper addresses this challenge by introducing a novel framework CaDRE, to generate realistic, diverse, and controllable safetycritical scenarios. Our approach optimizes for both the quality and diversity of scenarios by employing a unique formulation and algorithm that integrates real-world scenarios, domain knowledge, and black-box optimization. We validate the effectiveness of our framework through extensive testing in three representative types of traffic scenarios. The results demonstrate superior performance in generating diverse and highquality scenarios with greater sample efficiency than existing reinforcement learning (RL) and sampling-based methods. Peide Huang, Wenhao Ding, Benjamin Stoler, Jonathan Francis, Bingqing Chen, Ding Zhao |
ICRA | 6 |
| 2025 | Agile Continuous Jumping in Discontinuous TerrainsabstractWe focus on agile, continuous, and terrain-adaptive jumping of quadrupedal robots in discontinuous terrains such as stairs and stepping stones. Unlike single-step jumping, continuous jumping requires accurately executing highly dynamic motions over long horizons, which is challenging for existing approaches. To accomplish this task, we design a hierarchical learning and control framework, which consists of a learned heightmap predictor for robust terrain perception, a reinforcement-learning-based centroidal-level motion policy for versatile and terrain-adaptive planning, and a low-level model-based leg controller for accurate motion tracking. In addition, we minimize the sim-to-real gap by accurately modeling the hardware characteristics. Our framework enables a Unitree Go1 robot to perform agile and continuous jumps on human-sized stairs and sparse stepping stones, for the first time to the best of our knowledge. In particular, the robot can cross two stair steps in each jump and completes a 3.5m long, 2.8m high, 14-step staircase in 4.5 seconds. Moreover, the same policy outperforms baselines in various other parkour tasks, such as jumping over single horizontal or vertical discontinuities. Experiment videos can be found at https://yxyang.github.io/jumping_cod/. Yuxiang Yang 0007, Guanya Shi, Changyi Lin, Xiangyun Meng, Rosario Scalise, Mateo Guaman Castro, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao, Jie Tan 0001, Byron Boots |
ICRA | 9 |
| 2025 | SatPipe: Deterministic TCP Adaptation for Highly Dynamic LEO Satellite Networks
Ding Zhao, Xinyu Zhang 0003, Myungjin Lee |
INFOCOM | 1 |
| 2025 | DTactive: A Vision-Based Tactile Sensor with Active SurfaceabstractThe development of vision-based tactile sensors has significantly enhanced robots’ perception and manipulation capabilities, especially for tasks requiring contact-rich interactions with objects. In this work, we present DTactive, a novel vision-based tactile sensor with active surfaces. DTactive inherits and modifies the tactile 3D shape reconstruction method of DTact while integrating a mechanical transmission mechanism that facilitates the mobility of its surface. Thanks to this design, the sensor is capable of simultaneously performing tactile perception and in-hand manipulation with surface movement. Leveraging the high-resolution tactile images from the sensor and the magnetic encoder data from the transmission mechanism, we propose a learning-based method to enable precise angular trajectory control during in-hand manipulation. In our experiments, we successfully achieved accurate rolling manipulation within the range of [−180°, 180°] on various objects, with the root mean square error between the desired and actual angular trajectories being less than 12° on nine trained objects and less than 19° on three novel objects. The results demonstrate the potential of DTactive for in-hand object manipulation in terms of effectiveness, robustness and precision. Jikai Xu, Changyi Lin, Ding Zhao, Huazhe Xu |
IROS | 4 |
| 2025 | QuietPaw: Learning Quadrupedal Locomotion with Versatile Noise Preference AlignmentabstractWhen operating at their full capacity, quadrupedal robots can produce loud footstep noise, which can be disruptive in human-centered environments like homes, offices, and hospitals. As a result, balancing locomotion performance with noise constraints is crucial for the successful real-world deployment of quadrupedal robots. However, achieving adaptive noise control is challenging due to (a) the trade-off between agility and noise minimization, (b) the need for generalization across diverse deployment conditions, and (c) the difficulty of effectively adjusting policies based on noise requirements. We propose QuietPaw, a framework incorporating our Conditional Noise-Constrained Policy (CNCP), a constrained learning-based algorithm that enables flexible, noise-aware locomotion by conditioning policy behavior on noise-reduction levels. We leverage value representation decomposition in the critics, disentangling state representations from condition-dependent representations and this allows a single versatile policy to generalize across noise levels without retraining while improving the Pareto trade-off between agility and noise reduction. We validate our approach in simulation and the real world, demonstrating that CNCP can effectively balance locomotion performance and noise constraints, achieving continuously adjustable noise reduction. Yuyou Zhang, Yihang Yao, Shiqi Liu 0005, Yaru Niu, Changyi Lin, Yuxiang Yang 0007, Wenhao Yu 0003, Tingnan Zhang, Jie Tan 0001, Ding Zhao |
IROS | 10 |
| 2025 | Behavior Injection: Preparing Language Models for Reinforcement LearningabstractReinforcement learning (RL) has emerged as a powerful post-training technique to incentivize the reasoning ability of large language models (LLMs). However, LLMs can respond very inconsistently to RL finetuning: some show substantial performance gains, while others plateau or even degrade. To understand this divergence, we analyze the per-step influence of the RL objective and identify two key conditions for effective post-training: (1) RL-informative rollout accuracy, and (2) strong data co-influence, which quantifies how much the training data affects performance on other samples. Guided by these insights, we propose behavior injection, a task-agnostic data augmentation scheme applied prior to RL. Behavior injection enriches the supervised finetuning (SFT) data by seeding exploratory and exploitative behaviors, effectively making the model more RL-ready. We evaluate our method across two reasoning benchmarks with multiple base models. The results demonstrate that our theoretically motivated augmentation can significantly increase the performance gain from RL over the pre-RL model. Zhepeng Cen, Yihang Yao, William Jongwon Han, Zuxin Liu, Ding Zhao |
NeurIPS | 5 |
| 2025 | Model-Based Policy Adaptation for Closed-Loop End-to-end Autonomous DrivingabstractEnd-to-end (E2E) autonomous driving models have demonstrated strong performance in open-loop evaluations but often suffer from cascading errors and poor generalization in closed-loop settings. To address this gap, we propose Model-based Policy Adaptation (MPA), a general framework that enhances the robustness and safety of pretrained E2E driving agents during deployment. MPA first generates diverse counterfactual trajectories using a geometry-consistent simulation engine, exposing the agent to scenarios beyond the original dataset. Based on this generated data, MPA trains a diffusion-based policy adapter to refine the base policy’s predictions and a multi-step Q value model to evaluate long-term outcomes. At inference time, the adapter proposes multiple trajectory candidates, and the Q value model selects the one with the highest expected utility. Experiments on the nuScenes benchmark using a photorealistic closed-loop simulator demonstrate that MPA significantly improves performance across in-domain, out-of-domain, and safety-critical scenarios. We further investigate how the scale of counterfactual data and inference-time guidance strategies affect overall effectiveness. Haohong Lin, Wenhao Ding, Ding Zhao |
NeurIPS | 5 |
| 2025 | Semantically Adversarial Scene Generation With Explicit Knowledge GuidanceabstractGenerating adversarial scenes that potentially fail autonomous driving systems provides an effective way to improve their robustness. Extending purely data-driven generative models, recent specialized models satisfy additional controllable requirements such as embedding a traffic sign in a driving scene by manipulating patternsimplicitlyat the neuron level. In this paper, we introduce a method to incorporate domain knowledgeexplicitlyin the generation process to achieveSemantically Adversarial Generation (SAG). To be consistent with the composition of driving scenes, we first categorize the knowledge into two types, the property of objects and the relationship among objects. We then propose a tree-structured variational auto-encoder (T-VAE) to learn hierarchical scene representation. By imposing semantic rules on the properties of nodes and edges into the tree structure, explicit knowledge integration enables controllable generation. To demonstrate the advantage of structural representation, we construct a synthetic example to illustrate the controllability and explainability of our method in a succinct setting. We further extend to realistic environments for autonomous vehicles, showing that our method efficiently identifies adversarial driving scenes against different state-of-the-art 3D point cloud segmentation models and satisfies the constraints specified as explicit knowledge. Wenhao Ding, Haohong Lin, Bo Li 0026, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Pixel-wise Smoothing for Certified Robustness against Camera Motion PerturbationsabstractDeep learning-based visual perception models lack robustness when faced with camera motion perturbations in practice. The current certification process for assessing robustness is costly and time-consuming due to the extensive number of image projections required for Monte Carlo sampling in the 3D camera motion space. To address these challenges, we present a novel, efficient, and practical framework for certifying the robustness of 3D-2D projective transformations against camera motion perturbations. Our approach leverages a smoothing distribution over the 2D-pixel space instead of in the 3D physical space, eliminating the need for costly camera motion sampling and significantly enhancing the efficiency of robustness certifications. With the pixel-wise smoothed classifier, we are able to fully upper bound the projection errors using a technique of uniform partitioning in camera motion space. Additionally, we extend our certification framework to a more general scenario where only a single-frame point cloud is required in the projection oracle. Through extensive experimentation, we validate the trade-off between effectiveness and efficiency enabled by our proposed method. Remarkably, our approach achieves approximately 80% certified accuracy while utilizing only 30% of the projected image frames. Hanjiang Hu, Zuxin Liu, Linyi Li 0001, Ding Zhao |
AISTATS | 5 |
| 2024 | MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosabstractMultimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility, limited size, and the absence of proper categorization, which pose significant challenges. To address these challenges and provide a comprehensive dataset for this new direction, we have meticulously curated the MMSum dataset. Our new dataset features (1) Human-validated summaries for both video and textual content, providing superior human instruction and labels for mul-timodal learning. (2) Comprehensively and meticulously arranged categorization, spanning 17 principal categories and 170 subcategories to encapsulate a diverse array of real-world scenarios. (3) Benchmark tests performed on the proposed dataset to assess various tasks and methods, including video summarization, text summarization, and multimodal summarization. To champion accessibility and collaboration, we released the MMSum dataset and the data collection tool as fully open-source resources, fostering transparency and accelerating future developments, at https://mmsum-dataset.github.io/. Jielin Qiu, William Jongwon Han, Aditesh Kumar, Karthik Mittal, Claire Jin, Zhengyuan Yang, Ding Zhao, Bo Li 0026 |
CVPR | 10 |
| 2024 | RealGen: Retrieval Augmented Generation for Controllable Traffic Scenarios
Wenhao Ding, Ding Zhao, Chaowei Xiao, Marco Pavone 0001 |
ECCV (62) | 3 |
| 2024 | Learning from Sparse Offline Datasets via Conservative Density EstimationabstractOffline reinforcement learning (RL) offers a promising direction for learning policies from pre-collected datasets without requiring further interactions with the environment. However, existing methods struggle to handle out-of-distribution (OOD) extrapolation errors, especially in sparse reward or scarce data settings. In this paper, we propose a novel training algorithm called Conservative Density Estimation (CDE), which addresses this challenge by explicitly imposing constraints on the state-action occupancy stationary distribution. CDE overcomes the limitations of existing approaches, such as the stationary distribution correction method, by addressing the support mismatch issue in marginal importance sampling. Our method achieves state-of-the-art performance on the D4RL benchmark. Notably, CDE consistently outperforms baselines in challenging tasks with sparse rewards or insufficient data, demonstrating the advantages of our approach in addressing the extrapolation error problem in offline RL. Zhepeng Cen, Zuxin Liu, Zitong Wang 0005, Yihang Yao, Henry Lam, Ding Zhao |
ICLR | 6 |
| 2024 | Meta-Evolve: Continuous Robot Evolution for One-to-many Policy TransferabstractWe investigate the problem of transferring an expert policy from a source robot to multiple different robots. To solve this problem, we propose a method named *Meta-Evolve* that uses continuous robot evolution to efficiently transfer the policy to each target robot through a set of tree-structured evolutionary robot sequences. The robot evolution tree allows the robot evolution paths to be shared, so our approach can significantly outperform naive one-to-one policy transfer. We present a heuristic approach to determine an optimized robot evolution tree. Experiments have shown that our method is able to improve the efficiency of one-to-three transfer of manipulation policy by up to 3.2$\times$ and one-to-six transfer of agile locomotion policy by 2.4$\times$ in terms of simulation cost over the baseline of launching multiple independent one-to-one policy transfers. Supplementary videos available at the project website: https://sites.google.com/view/meta-evolve. Xingyu Liu 0001, Deepak Pathak, Ding Zhao |
ICLR | 3 |
| 2024 | TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained ModelsabstractThe full potential of large pretrained models remains largely untapped in control domains like robotics. This is mainly because of the scarcity of data and the computational challenges associated with training or fine-tuning these large models for such applications. Prior work mainly emphasizes either effective \emph{pretraining} of large models for decision-making or single-task adaptation. But real-world problems will require data-efficient, \emph{continual adaptation} for new control tasks. Recognizing these constraints, we introduce TAIL (Task-specific Adapters for Imitation Learning), a framework for efficient adaptation to new control tasks. Inspired by recent advancements in parameter-efficient fine-tuning in language domains, we explore efficient fine-tuning techniques---e.g., Bottleneck Adapters, P-Tuning, and Low-Rank Adaptation (LoRA)---in TAIL to adapt large pretrained models for new tasks with limited demonstration data. Our extensive experiments comparing prevalent parameter-efficient fine-tuning techniques and adaptation baselines suggest that TAIL with LoRA can achieve the best post-adaptation performance with only 1\% of the trainable parameters of full fine-tuning, while avoiding catastrophic forgetting and preserving adaptation plasticity in continual learning settings. Zuxin Liu, Jesse Zhang, Kavosh Asadi, Yao Liu 0009, Ding Zhao, Shoham Sabach, Rasool Fakoor |
ICLR | 5 |
| 2024 | Feasibility Consistent Representation Learning for Safe Reinforcement LearningabstractIn the field of safe reinforcement learning (RL), finding a balance between satisfying safety constraints and optimizing reward performance presents a significant challenge. A key obstacle in this endeavor is the estimation of safety constraints, which is typically more difficult than estimating a reward metric due to the sparse nature of the constraint signals. To address this issue, we introduce a novel framework named Feasibility Consistent Safe Reinforcement Learning (FCSRL). This framework combines representation learning with feasibility-oriented objectives to identify and extract safety-related information from the raw state for safe RL. Leveraging self-supervised learning techniques and a more learnable safety metric, our approach enhances the policy learning and constraint estimation. Empirical evaluations across a range of vector-state and image-based tasks demonstrate that our method is capable of learning a better safety-aware embedding and achieving superior performance than previous representation learning baselines. Zhepeng Cen, Yihang Yao, Zuxin Liu, Ding Zhao |
ICML | 4 |
| 2024 | Privacy Risks in Reinforcement Learning for Household RobotsabstractThe prominence of embodied Artificial Intelligence (AI), which empowers robots to navigate, perceive, and engage within virtual environments, has attracted significant attention, owing to the remarkable advances in computer vision and large language models. Privacy emerges as a pivotal concern within the realm of embodied AI, as the robot accesses substantial personal information. However, the issue of privacy leakage in embodied AI tasks, particularly concerning reinforcement learning algorithms, has not received adequate consideration in research. This paper aims to address this gap by proposing an attack on the training process of the value-based algorithm and the gradient-based algorithm, utilizing gradient inversion to reconstruct states, actions, and supervisory signals. The choice of using gradients for the attack is motivated by the fact that commonly employed federated learning techniques solely utilize gradients computed based on private user data to optimize models, without storing or transmitting the data to public servers. Nevertheless, these gradients contain sufficient information to potentially expose private data. To validate our approach, we conducted experiments on the AI2THOR simulator and evaluated our algorithm on active perception, a prevalent task in embodied AI. The experimental results demonstrate the effectiveness of our method in successfully reconstructing all information from the data in 120 room layouts. Check our website for videos. Wenhao Ding, Ding Zhao |
ICRA | 3 |
| 2024 | Influence of Camera-LiDAR Configuration on 3D Object Detection for Autonomous DrivingabstractCameras and LiDARs are both important sensors for autonomous driving, playing critical roles in 3D object detection. Camera-LiDAR Fusion has been a prevalent solution for robust and accurate driving perception. In contrast to the vast majority of existing arts that focus on how to improve the performance of 3D target detection through cross-modal schemes, deep learning algorithms, and training tricks, we devote attention to the impact of sensor configurations on the performance of learning-based methods. To achieve this, we propose a unified information-theoretic surrogate metric for camera and LiDAR evaluation based on the proposed sensor perception model. We also design an accelerated high-quality framework for data acquisition, model training, and performance evaluation that functions with the CARLA simulator. To show the correlation between detection performance and our surrogate metrics, We conduct experiments using several camera-LiDAR placements and parameters inspired by selfdriving companies and research institutions. Extensive experimental results of representative algorithms on nuScenes dataset validate the effectiveness of our surrogate metric, demonstrating that sensor configurations significantly impact point-cloudimage fusion based detection models, which contribute up to 30% discrepancy in terms of the average precision. Hanjiang Hu, Zuxin Liu, Xiaohao Xu, Xiaonan Huang, Ding Zhao |
ICRA | 6 |
| 2024 | Generalize by Touching: Tactile Ensemble Skill Transfer for Robotic Furniture AssemblyabstractFurniture assembly remains an unsolved problem in robotic manipulation due to its long task horizon and nongeneralizable operations plan. This paper presents the Tactile Ensemble Skill Transfer (TEST) framework, a pioneering offline reinforcement learning (RL) approach that incorporates tactile feedback in the control loop. TEST’s core design is to learn a skill transition model for high-level planning, along with a set of adaptive intra-skill goal-reaching policies. Such design aims to solve the robotic furniture assembly problem in a more generalizable way, facilitating seamless chaining of skills for this long-horizon task. We first sample demonstration from a set of heuristic policies and trajectories consisting of a set of randomized sub-skill segments, enabling the acquisition of rich robot trajectories that capture skill stages, robot states, visual indicators, and crucially, tactile signals. Leveraging these trajectories, our offline RL method discerns skill termination conditions and coordinates skill transitions. Our evaluations highlight the proficiency of TEST on the in-distribution furniture assemblies, its adaptability to unseen furniture configurations, and its robustness against visual disturbances. Ablation studies further accentuate the pivotal role of two algorithmic components: the skill transition model and tactile ensemble policies. Results indicate that TEST can achieve a success rate of 90% and is over 4 times more efficient than the heuristic policy in both in-distribution and generalization settings, suggesting a scalable skill transfer approach for contact-rich manipulation. Haohong Lin, Radu Corcodel, Ding Zhao |
ICRA | 3 |
| 2024 | Reinforcement Learning in a Safety-Embedded MDP with Trajectory OptimizationabstractSafe Reinforcement Learning (RL) plays an important role in applying RL algorithms to safety-critical real-world applications, addressing the trade-off between maximizing rewards and adhering to safety constraints. This work introduces a novel approach that combines RL with trajectory optimization to manage this trade-off effectively. Our approach embeds safety constraints within the action space of a modified Markov Decision Process (MDP). The RL agent produces a sequence of actions that are transformed into safe trajectories by a trajectory optimizer, thereby effectively ensuring safety and increasing training stability. This novel approach excels in its performance on challenging Safety Gym tasks, achieving significantly higher rewards and near-zero safety violations during inference. The method’s real-world applicability is demonstrated through a safe and effective deployment in a real robot task of box-pushing around obstacles. Further insights are available from the videos and appendix on our website: https://sites.google.com/view/safemdp. Fan Yang 0144, Wenxuan Zhou 0001, Zuxin Liu, Ding Zhao, David Held |
ICRA | 4 |
| 2024 | COMPOSER: Scalable and Robust Modular Policies for Snake RobotsabstractSnake robots have showcased remarkable compliance and adaptability in their interaction with environments, mirroring the traits of their natural counterparts. While their hyper-redundant and high-dimensional characteristics add to this adaptability, they also pose great challenges to robot control. Instead of perceiving the hyper-redundancy and flexibility of snake robots as mere challenges, there lies an unexplored potential in leveraging these traits to enhance robustness and generalizability at the control policy level. We seek to develop a control policy that effectively breaks down the high dimensionality of snake robots while harnessing their redundancy. In this work, we consider the snake robot as a modular robot and formulate the control of the snake robot as a cooperative Multi-Agent Reinforcement Learning (MARL) problem. Each segment of the snake robot functions as an individual agent. Specifically, we incorporate a self-attention mechanism to enhance the cooperative behavior between agents. A high-level imagination policy is proposed to provide additional rewards to guide the low-level control policy. We validate the proposed method COMPOSER with five snake robot tasks, including goal reaching, wall climbing, shape formation, tube crossing, and block pushing. COMPOSER achieves the highest success rate across all tasks when compared to a centralized baseline and four modular policy baselines. Additionally, we show enhanced robustness against module corruption and significantly superior zero-shot generalizability in our proposed method. The videos of this work are available on our project page: https://sites.google.com/view/composer-snake/. Yuyou Zhang, Yaru Niu, Xingyu Liu 0001, Ding Zhao |
ICRA | 4 |
| 2024 | Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro J. Moreno 0001 |
INTERSPEECH | 11 |
| 2024 | LocoMan: Advancing Versatile Quadrupedal Dexterity with Lightweight Loco-ManipulatorsabstractQuadrupedal robots have emerged as versatile agents capable of locomoting and manipulating in complex environments. Traditional designs typically rely on the robot’s inherent body parts or incorporate top-mounted arms for manipulation tasks. However, these configurations may limit the robot’s operational dexterity, efficiency, and adaptability, particularly in cluttered or constrained spaces. In this work, we present LocoMan, a dexterous quadrupedal robot with a novel morphology to perform versatile manipulation in diverse constrained environments. By equipping a Unitree Go1 robot with two low-cost and lightweight modular 3-DoF loco-manipulators on its front calves, LocoMan leverages the combined mobility and functionality of the legs and grippers for complex manipulation tasks that require precise 6D positioning of the end effector in a wide workspace. To harness the loco-manipulation capabilities of LocoMan, we introduce a unified control framework that extends the whole-body controller (WBC) to integrate the dynamics of loco-manipulators. Through experiments, we validate that the proposed whole-body controller can accurately and stably follow desired 6D trajectories of the end effector and torso, which, when combined with the large workspace from our design, facilitates a diverse set of challenging dexterous loco-manipulation tasks in confined spaces, such as opening doors, plugging into sockets, picking objects in narrow and low-lying spaces, and bimanual manipulation. Changyi Lin, Xingyu Liu 0001, Yuxiang Yang 0007, Yaru Niu, Wenhao Yu 0003, Tingnan Zhang, Jie Tan 0001, Byron Boots, Ding Zhao |
IROS | 9 |
| 2024 | Embodied Executable Policy Learning with Language-based Scene SummarizationabstractJielin Qiu, Mengdi Xu, William Han, Seungwhan Moon, Ding Zhao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jielin Qiu, Mengdi Xu, William Jongwon Han, Seungwhan Moon, Ding Zhao |
NAACL-HLT | 5 |
| 2024 | BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement LearningabstractOffline model-based reinforcement learning (MBRL) enhances data efficiency by utilizing pre-collected datasets to learn models and policies, especially in scenarios where exploration is costly or infeasible. Nevertheless, its performance often suffers from the objective mismatch between model and policy learning, resulting in inferior performance despite accurate model predictions. This paper first identifies the primary source of this mismatch comes from the underlying confounders present in offline data for MBRL. Subsequently, we introduce **B**ilin**E**ar **CAUS**al r**E**presentation (BECAUSE), an algorithm to capture causal representation for both states and actions to reduce the influence of the distribution shift, thus mitigating the objective mismatch problem. Comprehensive evaluations on 18 tasks that vary in data quality and environment context demonstrate the superior performance of BECAUSE over existing offline RL algorithms. We show the generalizability and robustness of BECAUSE under fewer samples or larger numbers of confounders. Additionally, we offer theoretical analysis of BECAUSE to prove its error bound and sample efficiency when integrating causal representation into offline MBRL. See more details in our project page: [https://sites.google.com/view/be-cause](https://sites.google.com/view/be-cause). Haohong Lin, Wenhao Ding, Laixi Shi, Ding Zhao |
NeurIPS | 7 |
| 2024 | OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement LearningabstractOffline safe reinforcement learning (RL) aims to train a policy that satisfies con- straints using a pre-collected dataset. Most current methods struggle with the mismatch between imperfect demonstrations and the desired safe and rewarding performance. In this paper, we mitigate this issue from a data-centric perspective and introduce OASIS (cOnditionAl diStributIon Shaping), a new paradigm in offline safe RL designed to overcome these critical limitations. OASIS utilizes a conditional diffusion model to synthesize offline datasets, thus shaping the data dis- tribution toward a beneficial target domain. Our approach makes compliance with safety constraints through effective data utilization and regularization techniques to benefit offline safe RL training. Comprehensive evaluations on public benchmarks and varying datasets showcase OASIS’s superiority in benefiting offline safe RL agents to achieve high-reward behavior while satisfying the safety constraints, out- performing established baselines. Furthermore, OASIS exhibits high data efficiency and robustness, making it suitable for real-world applications, particularly in tasks where safety is imperative and high-quality demonstrations are scarce. More details are available at the website https://sites.google.com/view/saferl-oasis/home. Yihang Yao, Zhepeng Cen, Wenhao Ding, Haohong Lin, Shiqi Liu 0005, Tingnan Zhang, Wenhao Yu 0003, Ding Zhao |
NeurIPS | 8 |
| 2024 | Functional optimal transport: regularized map estimation and domain adaptation for functional dataabstractWe introduce a formulation of regularized optimal transport problem for distributions on function spaces, where the stochastic map between functional domains can be approximated in terms of an (infinite-dimensional) Hilbert-Schmidt operator mapping a Hilbert space of functions to another. For numerous machine learning applications, data can be naturally viewed as samples drawn from spaces of functions, such as curves and surfaces, in high dimensions. Optimal transport for functional data analysis provides a useful framework of treatment for such domains. Since probability measures in infinite dimensional spaces generally lack absolute continuity (i.e., with respect to non-degenerate Gaussian measures), the Monge map in the standard optimal transport theory for finite dimensional spaces typically does not exist in the functional settings arising in such machine learning applications. This necessitates a suitable notion of approximation for the best pushforward measure to be obtained via a transport map. Indeed, our approach to the transportation problem in functional spaces is by a suitable regularization technique --- we restrict the class of transport maps to be a Hilbert-Schmidt space of operators.Within this regularization framework, we develop an efficient algorithm for finding the stochastic transport map between functional domains and provide theoretical guarantees on the existence, uniqueness, and consistency of our estimate for the Hilbert-Schmidt space of compact linear operators. We validate our method on synthetic datasets and examine the functional properties of the transport map. Experiments on real-world datasets of robot arm trajectories further demonstrate the effectiveness of our method on applications in domain adaptation. Aritra Guha, Dat Do, Mengdi Xu, XuanLong Nguyen, Ding Zhao |
J. Mach. Learn. Res. | 6 |
| 2023 | Group Distributionally Robust Reinforcement Learning with Hierarchical Latent VariablesabstractOne key challenge for multi-task Reinforcement learning (RL) in practice is the absence of task specifications. Robust RL has been applied to deal with task ambiguity but may result in over-conservative policies. To balance the worst-case (robustness) and average performance, we propose Group Distributionally Robust Markov Decision Process (GDR-MDP), a flexible hierarchical MDP formulation that encodes task groups via a latent mixture model. GDR-MDP identifies the optimal policy that maximizes the expected return under the worst-possible qualified belief over task groups within an ambiguity set. We rigorously show that GDR-MDP’s hierarchical structure improves distributional robustness by adding regularization to the worst possible outcomes. We then develop deep RL algorithms for GDR-MDP for both value-based and policy-based RL methods. Extensive experiments on Box2D control tasks, MuJoCo benchmarks, and Google football platforms show that our algorithms outperform classic robust training algorithms across diverse environments in terms of robustness under belief uncertainties. Demos are available on our project page (https://sites.google.com/view/gdr-rl/home). Mengdi Xu, Peide Huang, Yaru Niu, Visak Kumar, Jielin Qiu, Kuan-Hui Lee, Xuewei Qi, Henry Lam, Bo Li 0026, Ding Zhao |
AISTATS | 11 |
| 2023 | Structured Two-Stage True-Time-Delay Array Code book Design for Multi - User Data CommunicationabstractWideband millimeter-wave and terahertz (THz) systems can facilitate simultaneous data communication with multiple spatially separated users. It is desirable to orthogonalize users across sub-bands by deploying frequency-dependent beams with a sub-band-specific spatial response. True-Time-Delay (TTD) antenna arrays are a promising wide band architecture to implement sub-band-specific dispersion of beams across space using a single radio frequency (RF) chain. This paper proposes a structured design of analog TTD codebooks to generate beams that exhibit quantized sub-band-to-angle mapping. We introduce a structured Staircase TTD codebook and analyze the frequency-spatial behaviour of the resulting beam patterns. We develop the closed-form two-stage design of the proposed codebook to achieve the desired sub-band-specific beams and evaluate their performance in multi-user communication networks. Aditya Wadaskar, Ding Zhao, Ibrahim Pehlivan, Danijela Cabric |
GLOBECOM | 2 |
| 2023 | Sharing Low Rank Conformer Weights for Tiny Always-On Ambient Speech Recognition ModelsabstractContinued improvements in machine learning techniques offer exciting new opportunities through the use of larger models and larger training datasets. However, there is a growing need to offer these new capabilities on-board low-powered devices such as smart-phones, wearables and other embedded environments where only low memory is available. Towards this, we consider methods to reduce the model size of Conformer-based speech recognition models which typically require models with greater than 100M parameters down to just 5M parameters while minimizing impact on model quality. Such a model allows us to achieve always-on ambient speech recognition on edge devices with low-memory neural processors. We propose model weight reuse at different levels within our model architecture: (i) repeating full conformer block layers, (ii) sharing specific conformer modules across layers, (iii) sharing sub-components per conformer module, and (iv) sharing decomposed sub-component weights after low-rank decomposition. By sharing weights at different levels of our model, we can retain the full model in-memory while increasing the number of virtual trans-formations applied to the input. Through a series of ablation studies and evaluations, we find that with weight sharing and a low-rank architecture, we can achieve a WER of 2.84 and 2.94 for Librispeech dev-clean and test-clean respectively with a 5M parameter model. Steven M. Hernandez, Ding Zhao, Shaojin Ding, Antoine Bruguier, Rohit Prabhavalkar, Tara N. Sainath, Yanzhang He, Ian McGraw |
ICASSP | 2 |
| 2023 | Cardiac Disease Diagnosis on Imbalanced Electrocardiography Data Through Optimal Transport AugmentationabstractIn this paper, we focus on a new method of data augmentation to solve the data imbalance problem within imbalanced ECG datasets to improve the robustness and accuracy of heart disease detection. By using Optimal Transport, we augment the ECG disease data from normal ECG beats to balance the data among different categories. We build a Multi-Feature Transformer (MF-Transformer) as our classification model, where different features are extracted from both time and frequency domains to diagnose various heart conditions. Our results demonstrate 1) the classification models’ ability to make competitive predictions on five ECG categories; 2) improvements in accuracy and robustness reflecting the effectiveness of our data augmentation method. Jielin Qiu, Mengdi Xu, Peide Huang, Michael A. Rosenberg, Douglas Weber, Emerson Liu, Ding Zhao |
ICASSP | 8 |
| 2023 | Multi-Output RNN-T Joint Networks for Multi-Task Learning of ASR and Auxiliary TasksabstractWe propose a multi-output joint network architecture for RNN-T transducer, for multi-task modeling of ASR and auxiliary tasks that rely on ASR outputs. Each output of the joint network predicts tar-get labels with disjoint vocabularies for each task, while sharing the same audio features by the encoder and language model features by the prediction network. Each task is trained with an RNN-T loss that marginalizes over all possible paths, and we allow multiple tasks to share the blank logit so that they are synchronized. We demonstrate our method on two auxiliary tasks, namely capitalization and pause prediction, and discuss different considerations for modeling and inference procedures. For capitalization, we successfully distill capitalization labels from a standalone text normalization model, and achieve competitive Uppercase Error Rate (UER) while offering streaming capability and improved inference efficiency. In addition, our model has similar capitalization accuracy compared to a mixed-case ASR model, but obtains improved WERs if integrated with external language models. For pause prediction, we achieve the same performance as the previous two-step approach while providing a simpler training recipe without affecting ASR accuracy. Ding Zhao, Shaojin Ding, Hao Zhang 0010, Shuo-Yiin Chang, David Rybach, Tara N. Sainath, Yanzhang He, Ian McGraw, Shankar Kumar |
ICASSP | 2 |
| 2023 | On the Robustness of Safe Reinforcement Learning under Observational Perturbations
Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Jie Tan 0001, Bo Li 0026, Ding Zhao |
ICLR | 7 |
| 2023 | Hyper-Decision Transformer for Efficient Online Policy Adaptation
Mengdi Xu, Yikang Shen, Ding Zhao, Chuang Gan 0001 |
ICLR | 5 |
| 2023 | Bayesian Reparameterization of Reward-Conditioned Reinforcement Learning with Energy-based ModelsabstractRecently, reward-conditioned reinforcement learning (RCRL) has gained popularity due to its simplicity, flexibility, and off-policy nature. However, we will show that current RCRL approaches are fundamentally limited and fail to address two critical challenges of RCRL -- improving generalization on high reward-to-go (RTG) inputs, and avoiding out-of-distribution (OOD) RTG queries during testing time. To address these challenges when training vanilla RCRL architectures, we propose Bayesian Reparameterized RCRL (BR-RCRL), a novel set of inductive biases for RCRL inspired by Bayes' theorem. BR-RCRL removes a core obstacle preventing vanilla RCRL from generalizing on high RTG inputs -- a tendency that the model treats different RTG inputs as independent values, which we term ``RTG Independence". BR-RCRL also allows us to design an accompanying adaptive inference method, which maximizes total returns while avoiding OOD queries that yield unpredictable behaviors in vanilla RCRL methods. We show that BR-RCRL achieves state-of-the-art performance on the Gym-Mujoco and Atari offline RL benchmarks, improving upon vanilla RCRL by up to 11%. Wenhao Ding, Tong Che, Ding Zhao, Marco Pavone 0001 |
ICML | 3 |
| 2023 | Towards Robust and Safe Reinforcement Learning with Benign Off-policy DataabstractPrevious work demonstrates that the optimal safe reinforcement learning policy in a noise-free environment is vulnerable and could be unsafe under observational attacks. While adversarial training effectively improves robustness and safety, collecting samples by attacking the behavior agent online could be expensive or prohibitively dangerous in many applications. We propose the robuSt vAriational ofF-policy lEaRning (SAFER) approach, which only requires benign training data without attacking the agent. SAFER obtains an optimal non-parametric variational policy distribution via convex optimization and then uses it to improve the parameterized policy robustly via supervised learning. The two-stage policy optimization facilitates robust training, and extensive experiments on multiple robot platforms show the efficiency of SAFER in learning a robust and safe policy: achieving the same reward with much fewer constraint violations during training than on-policy baselines. Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Yihang Yao, Hanjiang Hu, Ding Zhao |
ICML | 7 |
| 2023 | Constrained Decision Transformer for Offline Safe Reinforcement LearningabstractSafe reinforcement learning (RL) trains a constraint satisfaction policy by interacting with the environment. We aim to tackle a more challenging problem: learning a safe policy from an offline dataset. We study the offline safe RL problem from a novel multi-objective optimization perspective and propose the $\epsilon$-reducible concept to characterize problem difficulties. The inherent trade-offs between safety and task performance inspire us to propose the constrained decision transformer (CDT) approach, which can dynamically adjust the trade-offs during deployment. Extensive experiments show the advantages of the proposed method in learning an adaptive, safe, robust, and high-reward policy. CDT outperforms its variants and strong offline safe RL baselines by a large margin with the same hyperparameters across all tasks, while keeping the zero-shot adaptation capability to different constraint thresholds, making our approach more suitable for real-world RL under constraints. Zuxin Liu, Zijian Guo 0002, Yihang Yao, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao |
ICML | 7 |
| 2023 | Interpolation for Robust Learning: Data Augmentation on Wasserstein GeodesicsabstractWe propose to study and promote the robustness of a model as per its performance on a continuous geodesic interpolation of subpopulations, e.g., a class of samples in a classification problem. Specifically, (1) we augment the data by finding the worst-case Wasserstein barycenter on the geodesic connecting subpopulation distributions. (2) we regularize the model for smoother performance on the continuous geodesic path connecting subpopulation distributions. (3) Additionally, we provide a theoretical guarantee of robustness improvement and investigate how the geodesic location and the sample size contribute, respectively. Experimental validations of the proposed strategy on four datasets including CIFAR-100 and ImageNet, establish the efficacy of our method, e.g., our method improves the baselines’ certifiable robustness on CIFAR10 upto 7.7%, with 16.8% on empirical robustness on CIFAR-100. Our work provides a new perspective of model robustness through the lens of Wasserstein geodesic-based interpolation with a practical off-the-shelf strategy that can be combined with existing robust training methods. Jielin Qiu, Aritra Guha, XuanLong Nguyen, Bo Li 0026, Ding Zhao |
ICML | 7 |
| 2023 | Learning to View: Decision Transformers for Active Object DetectionabstractActive perception describes a broad class of techniques that couple planning and perception systems to move the robot in a way to give the robot more information about the environment. In most robotic systems, perception is typically independent of motion planning. For example, traditional object detection is passive: it operates only on the images it receives. However, we have a chance to improve the results if we allow planning to consume detection signals and move the robot to collect views that maximize the quality of the results. In this paper, we use reinforcement learning (RL) methods to control the robot in order to obtain images that maximize the detection quality. Specifically, we propose using a Decision Transformer with online fine-tuning, which first optimizes the policy with a pre-collected expert dataset and then improves the learned policy by exploring better solutions in the environment. We evaluate the performance of proposed method on an interactive dataset collected from an indoor scenario simulator. Experimental results demonstrate that our method outperforms all baselines, including expert policy and pure offline RL methods. We also provide exhaustive analyses of the reward distribution and observation space. Wenhao Ding, Nathalie Majcherczyk, Mohit Deshpande, Xuewei Qi, Ding Zhao, Rajasimman Madhivanan, Arnie Sen |
ICRA | 5 |
| 2023 | SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark Under Multiple EnvironmentsabstractDifferent environments pose a great challenge to the outdoor robust visual perception for long-term autonomous driving, and the generalization of learning-based algorithms on different environments is still an open problem. Although monocular depth prediction has been well studied recently, few works focus on the robustness of learning-based depth prediction across different environments, e.g. changing illumination and seasons, owing to the lack of such a multi-environment real-world dataset and benchmark. To this end, the cross-season monocular depth prediction dataset and benchmark, SeasonDepth, is introduced to benchmark the depth estimation performance under different environments. We investigate several state-of-the-art representative open-source supervised and self-supervised depth prediction methods using newly-formulated metrics. Through extensive experimental evaluation on the proposed dataset and cross-dataset evaluation with current autonomous driving datasets, the performance and robustness against the influence of multiple environments are analyzed qualitatively and quantitatively. We show that long-term monocular depth prediction is still challenging and believe our work can boost further research on the long-term robustness and generalization for outdoor visual perception. The dataset is available on https://seasondepth.github.io. Hanjiang Hu, Baoquan Yang, Zhijian Qiao, Shiqi Liu 0005, Zuxin Liu, Wenhao Ding, Ding Zhao, Hesheng Wang 0001 |
IROS | 8 |
| 2023 | GOATS: Goal Sampling Adaptation for Scooping with Curriculum Reinforcement LearningabstractIn this work, we first formulate the problem of robotic water scooping using goal-conditioned reinforcement learning. This task is particularly challenging due to the complex dynamics of fluid and the need to achieve multi-modal goals. The policy is required to successfully reach both position goals and water amount goals, which leads to a large convoluted goal state space. To overcome these challenges, we introduce Goal Sampling Adaptation for Scooping (GOATS), a curriculum reinforcement learning method that can learn an effective and generalizable policy for robot scooping tasks. Specifically, we use a goal-factorized reward formulation and interpolate position goal distributions and amount goal distributions to create curriculum throughout the learning process. As a result, our proposed method can outperform the baselines in simulation and achieves 5.46% and 8.71% amount errors on bowl scooping and bucket scooping tasks, respectively, under 1000 variations of initial water states in the tank and a large goal state space. Besides being effective in simulation environments, our method can efficiently adapt to noisy real-robot water-scooping scenarios with diverse physical configurations and unseen settings, demonstrating superior efficacy and generalizability. The videos of this work are available on our project page: https://sites.google.com/view/goatscooping. Yaru Niu, Shiyu Jin, Zeqing Zhang, Ding Zhao, Liangjun Zhang |
IROS | 5 |
| 2023 | Seeing is not Believing: Robust Reinforcement Learning against Spurious CorrelationabstractRobustness has been extensively studied in reinforcement learning (RL) to handle various forms of uncertainty such as random perturbations, rare events, and malicious attacks. In this work, we consider one critical type of robustness against spurious correlation, where different portions of the state do not have correlations induced by unobserved confounders. These spurious correlations are ubiquitous in real-world tasks, for instance, a self-driving car usually observes heavy traffic in the daytime and light traffic at night due to unobservable human activity. A model that learns such useless or even harmful correlation could catastrophically fail when the confounder in the test case deviates from the training one. Although motivated, enabling robustness against spurious correlation poses significant challenges since the uncertainty set, shaped by the unobserved confounder and causal structure, is difficult to characterize and identify. Existing robust algorithms that assume simple and unstructured uncertainty sets are therefore inadequate to address this challenge. To solve this issue, we propose Robust State-Confounded Markov Decision Processes (RSC-MDPs) and theoretically demonstrate its superiority in avoiding learning spurious correlations compared with other robust RL counterparts. We also design an empirical algorithm to learn the robust optimal policy for RSC-MDPs, which outperforms all baselines in eight realistic self-driving and manipulation tasks. Wenhao Ding, Laixi Shi, Yuejie Chi, Ding Zhao |
NeurIPS | 4 |
| 2023 | Learning Shared Safety Constraints from Multi-task DemonstrationsabstractRegardless of the particular task we want to perform in an environment, there are often shared safety constraints we want our agents to respect. For example, regardless of whether it is making a sandwich or clearing the table, a kitchen robot should not break a plate. Manually specifying such a constraint can be both time-consuming and error-prone. We show how to learn constraints from expert demonstrations of safe task completion by extending inverse reinforcement learning (IRL) techniques to the space of constraints. Intuitively, we learn constraints that forbid highly rewarding behavior that the expert could have taken but chose not to. Unfortunately, the constraint learning problem is rather ill-posed and typically leads to overly conservative constraints that forbid all behavior that the expert did not take. We counter this by leveraging diverse demonstrations that naturally occur in multi-task setting to learn a tighter set of constraints. We validate our method with simulation experiments on high-dimensional continuous control tasks. Konwoo Kim, Gokul Swamy 0001, Zuxin Liu, Ding Zhao, Sanjiban Choudhury, Steven Z. Wu |
NeurIPS | 4 |
| 2023 | Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement LearningabstractSafe reinforcement learning (RL) focuses on training reward-maximizing agents subject to pre-defined safety constraints. Yet, learning versatile safe policies that can adapt to varying safety constraint requirements during deployment without retraining remains a largely unexplored and challenging area. In this work, we formulate the versatile safe RL problem and consider two primary requirements: training efficiency and zero-shot adaptation capability. To address them, we introduce the Conditioned Constrained Policy Optimization (CCPO) framework, consisting of two key modules: (1) Versatile Value Estimation (VVE) for approximating value functions under unseen threshold conditions, and (2) Conditioned Variational Inference (CVI) for encoding arbitrary constraint thresholds during policy optimization. Our extensive experiments demonstrate that CCPO outperforms the baselines in terms of safety and task performance while preserving zero-shot adaptation capabilities to different constraint thresholds data-efficiently. This makes our approach suitable for real-world dynamic applications. Yihang Yao, Zuxin Liu, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao |
NeurIPS | 7 |
| 2023 | LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream VideosabstractLivestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the sessions, making them valuable materials. However, Livestream tutorial videos are usually hours long, recorded, and uploaded to the Internet directly after the live sessions, making it hard for other people to catch up quickly. An outline will be a beneficial solution, which requires the video to be temporally segmented according to topics. In this work, we introduced a large Livestream video dataset named MultiLive, and formulated the temporal segmentation of the long Livestream videos (TSLLV) task. We propose LiveSeg, an unsupervised Livestream video temporal Segmentation solution, which takes advantage of multimodal features from different domains. Our method achieved a 16.8% F1-score performance improvement compared with the state-of-the-art method. Jielin Qiu, Franck Dernoncourt, Trung Bui, Ding Zhao, Hailin Jin |
WACV | 5 |
| 2023 | A Survey on Safety-Critical Driving Scenario Generation - A Methodological PerspectiveabstractAutonomous driving systems have witnessed significant development during the past years thanks to the advance in machine learning-enabled sensing and decision-making algorithms. One critical challenge for their massive deployment in the real world is their safety evaluation. Most existing driving systems are still trained and evaluated on naturalistic scenarios collected from daily life or heuristically-generated adversarial ones. However, the large population of cars, in general, leads to an extremely low collision rate, indicating that safety-critical scenarios are rare in the collected real-world data. Thus, methods to artificially generate scenarios become crucial to measure the risk and reduce the cost. In this survey, we focus on the algorithms of safety-critical scenario generation in autonomous driving. We first provide a comprehensive taxonomy of existing algorithms by dividing them into three categories: data-driven generation, adversarial generation, and knowledge-based generation. Then, we discuss useful tools for scenario generation, including simulation platforms and packages. Finally, we extend our discussion to five main challenges of current works– fidelity, efficiency, diversity, transferability, controllability– and research opportunities lighted up by these challenges. Wenhao Ding, Chejian Xu, Mansur Arief, Haohong Lin, Bo Li 0026, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Semantic Traffic Law Adaptive Decision-Making for Self-Driving VehiclesabstractFacts proved that obeying traffic laws keeps the promise to promote the safety of self-driving vehicles. Current self-driving vehicles usually have fixed algorithms during autonomous driving, however the traffic laws may differ or change in different regions or times, e.g., tidal lanes. It raises a crucial requirement to make self-driving vehicles adapt to the newly received traffic laws. The challenges are that traffic laws are usually semantic and manually designed, but the original algorithms may not always contain the pre-designed interface to adapt to emerging laws. To this end, this work proposes a traffic law adaptive decision-making platform, which uses the linear temporal logic (LTL) formula to consistently describe the semantic traffic laws. Then, an LTL-based reinforcement learning framework is designed to estimate the probability of illegal behavior under different traffic laws. Finally, a law-specific backup policy is designed to maintain the performance threshold by monitoring the probability of illegal behavior. This work takes three typical scenarios where the traffic laws differ for instance to prove the effectiveness of the proposed approach, i.e., law amendment presented by the government, law difference between different regions, and temporary traffic control. The results show that the proposed method can help the original decision-making algorithms adapt to the traffic laws well without pre-defined interfaces. This method provides a way to administer on-road driving self-driving vehicles. Hong Wang 0014, Zhong Cao 0003, Wenhao Yu 0006, Chengxiang Zhao, Ding Zhao, Diange Yang, Jun Li 0082 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Investigating the Impact of Multi-LiDAR Placement on Object Detection for Autonomous DrivingabstractThe past few years have witnessed an increasing interest in improving the perception performance of LiDARs on au-tonomous vehicles. While most of the existing works focus on developing new deep learning algorithms or model ar-chitectures, we study the problem from the physical design perspective, i.e., how different placements of multiple Li-DARs influence the learning-based perception. To this end, we introduce an easy-to-compute information-theoretic sur-rogate metric to quantitatively and fast evaluate LiDAR placement for 3D detection of different types of objects. We also present a new data collection, detection model training and evaluation framework in the realistic CARLA simula-tor to evaluate disparate multi-LiDAR configurations. Using several prevalent placements inspired by the designs of self-driving companies, we show the correlation between our surrogate metric and object detection performance of different representative algorithms on KITTI through exten-sive experiments, validating the effectiveness of our LiDAR placement evaluation approach. Our results show that sen-sor placement is non-negligible in 3D point cloud-based ob-ject detection, which will contribute to 5% ~ 10% performance discrepancy in terms of average precision in chal-lenging 3D object detection settings. We believe that this is one of the first studies to quantitatively investigate the influence of LiDAR placement on perception performance. Hanjiang Hu, Zuxin Liu, Sharad Chitlangia, Akhil Agnihotri, Ding Zhao |
CVPR | 5 |
| 2022 | COPA: Certifying Robust Policies for Offline Reinforcement Learning against Poisoning Attacks
Fan Wu 0011, Linyi Li 0001, Huan Zhang 0001, Bhavya Kailkhura, Krishnaram Kenthapadi, Ding Zhao, Bo Li 0026 |
ICLR | 6 |
| 2022 | CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing
Fan Wu 0011, Linyi Li 0001, Zijian Huang 0002, Yevgeniy Vorobeychik, Ding Zhao, Bo Li 0026 |
ICLR | 5 |
| 2022 | Constrained Variational Policy Optimization for Safe Reinforcement LearningabstractSafe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications. Previous primal-dual style approaches suffer from instability issues and lack optimality guarantees. This paper overcomes the issues from the perspective of probabilistic inference. We introduce a novel Expectation-Maximization approach to naturally incorporate constraints during the policy learning: 1) a provable optimal non-parametric variational distribution could be computed in closed form after a convex optimization (E-step); 2) the policy parameter is improved within the trust region based on the optimal variational distribution (M-step). The proposed algorithm decomposes the safe RL problem into a convex optimization phase and a supervised learning phase, which yields a more stable training performance. A wide range of experiments on continuous robotic tasks shows that the proposed method achieves significantly better constraint satisfaction performance and better sample efficiency than baselines. The code is available at https://github.com/liuzuxin/cvpo-safe-rl. Zuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Steven Z. Wu, Bo Li 0026, Ding Zhao |
ICML | 7 |
| 2022 | Prompting Decision Transformer for Few-Shot Policy GeneralizationabstractHuman can leverage prior experience and learn novel tasks from a handful of demonstrations. In contrast to offline meta-reinforcement learning, which aims to achieve quick adaptation through better algorithm design, we investigate the effect of architecture inductive bias on the few-shot learning capability. We propose a Prompt-based Decision Transformer (Prompt-DT), which leverages the sequential modeling ability of the Transformer architecture and the prompt framework to achieve few-shot adaptation in offline RL. We design the trajectory prompt, which contains segments of the few-shot demonstrations, and encodes task-specific information to guide policy generation. Our experiments in five MuJoCo control benchmarks show that Prompt-DT is a strong few-shot learner without any extra finetuning on unseen target tasks. Prompt-DT outperforms its variants and strong meta offline RL baselines by a large margin with a trajectory prompt containing only a few timesteps. Prompt-DT is also robust to prompt length changes and can generalize to out-of-distribution (OOD) environments. Project page: \href{https://mxu34.github.io/PromptDT/}{https://mxu34.github.io/PromptDT/}. Mengdi Xu, Yikang Shen, Ding Zhao, Josh Tenenbaum, Chuang Gan 0001 |
ICML | 5 |
| 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial TrainingabstractRobust Reinforcement Learning (RL) focuses on improving performances under model errors or adversarial attacks, which facilitates the real-life deployment of RL agents. Robust Adversarial Reinforcement Learning (RARL) is one of the most popular frameworks for robust RL. However, most of the existing literature models RARL as a zero-sum simultaneous game with Nash equilibrium as the solution concept, which could overlook the sequential nature of RL deployments, produce overly conservative agents, and induce training instability. In this paper, we introduce a novel hierarchical formulation of robust RL -- a general-sum Stackelberg game model called RRL-Stack -- to formalize the sequential nature and provide extra flexibility for robust training. We develop the Stackelberg Policy Gradient algorithm to solve RRL-Stack, leveraging the Stackelberg learning dynamics by considering the adversary's response. Our method generates challenging yet solvable adversarial environments which benefit RL agents' robust learning. Our algorithm demonstrates better training stability and robustness against different testing conditions in the single-agent robotics control and multi-agent highway merging tasks. Peide Huang, Mengdi Xu, Fei Fang 0001, Ding Zhao |
IJCAI | 4 |
| 2022 | A Unified Cascaded Encoder ASR Model for Dynamic Model SizesabstractIn this paper, we propose a dynamic cascaded encoder Automatic Speech Recognition (ASR) model, which unifies models for different deployment scenarios. Moreover, the model can significantly reduce model size and power consumption without loss of quality. Namely, with the dynamic cascaded encoder model, we explore three techniques to maximally boost the performance of each model size: 1) Use separate decoders for each sub-model while sharing the encoders; 2) Use funnel-pooling to improve the encoder efficiency; 3) Balance the size of causal and non-causal encoders to improve quality and fit deployment constraints. Overall, the proposed large-medium model has 30% smaller size and reduces power consumption by 33%, compared to the baseline cascaded encoder model. The triple-size model that unifies the large, medium, and small models achieves 37% total size reduction with minimal quality loss, while substantially reducing the engineering efforts of having separate models. Shaojin Ding, Ding Zhao, Tara N. Sainath, Yanzhang He, Robert David 0002, Rami Botros, Xin Wang 0116, Rina Panigrahy, Qiao Liang 0001, Dongseong Hwang, Ian McGraw, Rohit Prabhavalkar, Trevor Strohman |
INTERSPEECH | 3 |
| 2022 | Federated Pruning: Improving Neural Network Efficiency with Federated LearningabstractAutomatic Speech Recognition models require large amount of speech data for training, and the collection of such data often leads to privacy concerns.Federated learning has been widely used and is considered to be an effective decentralized technique by collaboratively learning a shared prediction model while keeping the data local on different clients devices.However, the limited computation and communication resources on clients devices present practical difficulties for large models.To overcome such challenges, we propose Federated Pruning to train a reduced model under the federated setting, while maintaining similar performance compared to the full model.Moreover, the vast amount of clients data can also be leveraged to improve the pruning results compared to centralized training.We explore different pruning schemes and provide empirical evidence of the effectiveness of our methods. Rongmei Lin, Yonghui Xiao, Tien-Ju Yang, Ding Zhao, Li Xiong 0001, Giovanni Motta, Françoise Beaufays |
INTERSPEECH | 4 |
| 2022 | Improve Speech Enhancement using Perception-High-Related Time-Frequency Loss
Ding Zhao, Yuehai Wang |
INTERSPEECH | 1 |
| 2022 | Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event SamplingabstractEvaluating rare but high-stakes events is one of the main challenges in obtaining reliable reinforcement learning policies, especially in large or infinite state/action spaces where limited scalability dictates a prohibitively large number of testing iterations. On the other hand, a biased or inaccurate policy evaluation in a safety-critical system could potentially cause unexpected catastrophic failures during deployment. This paper proposes the Accelerated Policy Evaluation (APE) method, which simultaneously uncovers rare events and estimates the rare event probability in Markov decision processes. The APE method treats the environment nature as an adversarial agent and learns towards, through adaptive importance sampling, the zero-variance sampling distribution for the policy evaluation. Moreover, APE is scalable to large discrete or continuous spaces by incorporating function approximators. We investigate the convergence property of APE in the tabular setting. Our empirical studies show that APE can estimate the rare event probability with a smaller bias while only using orders of magnitude fewer samples than baselines in multi-agent and single-agent environments. Mengdi Xu, Peide Huang, Fengpei Li, Xuewei Qi, Kentaro Oguchi 0001, Henry Lam, Ding Zhao |
IROS | 9 |
| 2022 | Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningabstractAs a pivotal component to attaining generalizable solutions in human intelligence, reasoning provides great potential for reinforcement learning (RL) agents' generalization towards varied goals by summarizing part-to-whole arguments and discovering cause-and-effect relations. However, how to discover and represent causalities remains a huge gap that hinders the development of causal RL. In this paper, we augment Goal-Conditioned RL (GCRL) with Causal Graph (CG), a structure built upon the relation between objects and events. We novelly formulate the GCRL problem into variational likelihood maximization with CG as latent variables. To optimize the derived objective, we propose a framework with theoretical performance guarantees that alternates between two steps: using interventional data to estimate the posterior of CG; using CG to learn generalizable models and interpretable policies. Due to the lack of public benchmarks that verify generalization capability under reasoning, we design nine tasks and then empirically show the effectiveness of the proposed method against five baselines on these tasks. Further theoretical analysis shows that our performance improvement is attributed to the virtuous cycle of causal discovery, transition modeling, and policy training, which aligns with the experimental evidence in extensive ablation studies. Wenhao Ding, Haohong Lin, Bo Li 0026, Ding Zhao |
NeurIPS | 4 |
| 2022 | Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationabstractCurriculum Reinforcement Learning (CRL) aims to create a sequence of tasks, starting from easy ones and gradually learning towards difficult tasks. In this work, we focus on the idea of framing CRL as interpolations between a source (auxiliary) and a target task distribution. Although existing studies have shown the great potential of this idea, it remains unclear how to formally quantify and generate the movement between task distributions. Inspired by the insights from gradual domain adaptation in semi-supervised learning, we create a natural curriculum by breaking down the potentially large task distributional shift in CRL into smaller shifts. We propose GRADIENT which formulates CRL as an optimal transport problem with a tailored distance metric between tasks. Specifically, we generate a sequence of task distributions as a geodesic interpolation between the source and target distributions, which are actually the Wasserstein barycenter. Different from many existing methods, our algorithm considers a task-dependent contextual distance metric and is capable of handling nonparametric distributions in both continuous and discrete context settings. In addition, we theoretically show that GRADIENT enables smooth transfer between subsequent stages in the curriculum under certain conditions. We conduct extensive experiments in locomotion and manipulation tasks and show that our proposed GRADIENT achieves higher performance than baselines in terms of learning efficiency and asymptotic performance. Peide Huang, Mengdi Xu, Laixi Shi, Fei Fang 0001, Ding Zhao |
NeurIPS | 6 |
| 2022 | SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous VehiclesabstractAs shown by recent studies, machine intelligence-enabled systems are vulnerable to test cases resulting from either adversarial manipulation or natural distribution shifts. This has raised great concerns about deploying machine learning algorithms for real-world applications, especially in safety-critical domains such as autonomous driving (AD). On the other hand, traditional AD testing on naturalistic scenarios requires hundreds of millions of driving miles due to the high dimensionality and rareness of the safety-critical scenarios in the real world. As a result, several approaches for autonomous driving evaluation have been explored, which are usually, however, based on different simulation platforms, types of safety-critical scenarios, scenario generation algorithms, and driving route variations. Thus, despite a large amount of effort in autonomous driving testing, it is still challenging to compare and understand the effectiveness and efficiency of different testing scenario generation algorithms and testing mechanisms under similar conditions. In this paper, we aim to provide the first unified platform SafeBench to integrate different types of safety-critical testing scenarios, scenario generation algorithms, and other variations such as driving routes and environments. In particular, we consider 8 safety-critical testing scenarios following National Highway Traffic Safety Administration (NHTSA) and develop 4 scenario generation algorithms considering 10 variations for each scenario. Meanwhile, we implement 4 deep reinforcement learning-based AD algorithms with 4 types of input (e.g., bird’s-eye view, camera) to perform fair comparisons on SafeBench. We find our generated testing scenarios are indeed more challenging and observe the trade-off between the performance of AD agents under benign and safety-critical testing scenarios. We believe our unified platform SafeBench for large-scale and effective autonomous driving testing will motivate the development of new testing scenario generation and safe AD algorithms. SafeBench is available at https://safebench.github.io. Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Yihan He, Hanjiang Hu, Ding Zhao, Bo Li 0026 |
NeurIPS | 8 |
| 2022 | Joint Transmit and Reflective Beamforming for IRS-Assisted Integrated Sensing and CommunicationabstractThis paper studies an intelligent reflecting surface (IRS)-assisted integrated sensing and communication (ISAC) system, in which one IRS with a uniform linear array (ULA) is deployed to not only assist the wireless communication from a multi-antenna base station (BS) to a single-antenna communication user (CU), but also create virtual line-of-sight (LoS) links for sensing potential targets at areas with LoS links blocked. We consider that the BS transmits combined information and sensing signals for ISAC. Under this setup, we jointly optimize the transmit information and sensing beamforming at the BS and the reflective beamforming at the IRS, to maximize the IRS’s minimum beampattern gain towards the desired sensing angles, subject to the minimum signal-to-noise ratio (SNR) requirement at the CU and the maximum transmit power constraint at the BS. Although the formulated SNR-constrained beampattern gain maximization problem is non-convex and difficult to solve, we present an efficient algorithm to obtain a high-quality solution by using the techniques of alternating optimization and semi-definite relaxation (SDR). Numerical results show that the proposed joint beamforming design achieves improved sensing performance while ensuring the communication requirement as compared to benchmarks without such joint optimization. It is also shown that the use of dedicated sensing beams is beneficial in enhancing the performance for IRS-assisted ISAC. Xianxin Song, Ding Zhao, Haocheng Hua, Tony Xiao Han, Xun Yang 0009, Jie Xu 0002 |
WCNC | 2 |
| 2022 | Guest Editorial Special Issue on Artificial Intelligence for Autonomous Unmanned System ApplicationsabstractThis special issue of the IEEE TRANSACTIONS ON AUTOMATION SCIENCE AND ENGINEERING (T-ASE) focuses on how the state-of-the-art achievements and applications in the general area of artificial intelligence in automation for autonomous unmanned systems applications. As Guest Editors, we are very pleased to present the selected 16 articles, whose topics are specifically related to artificial intelligence real-time object detection, recognition, localization, control optimization, motion planning, formation control, adaptive control, and autonomous decision-making. Hongbo Gao 0001, Ming Liu 0001, Fei Chen 0007, Xiaoxiang Na, Ding Zhao, Linghe Kong, Keqiang Li 0002, Chun-Yi Su |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2022 | Understanding V2V Driving Scenarios Through Traffic PrimitivesabstractUnderstanding driver interaction behavioral semantics has potential benefits to autonomous car’s decision-making design. This article presents a framework of analyzing various encountering behaviors through decomposing driving encounter sequential data into small building blocks, called traffic primitives, using a Bayesian nonparametric learning (BNPL) approach. This framework offers a flexible way to gain semantic insights into complex driving encounters without any prerequisite knowledge of interaction behavior categories. Its effectiveness is then validated using 976 naturalistic driving encounters from which more than 4000 traffic primitives were learned with the BNPL approach. After that, a dynamic time warping method integrated with$k$-means clustering is then developed to cluster all these extracted traffic primitives into groups. Experimental results identify 20 kinds of traffic primitives capable of representing the essential components of driving encounters in our database. Based on the results, we conclude that the proposed primitive-based analysis could prove useful for autonomous vehicle applications. Wenshuo Wang 0001, Weiyang Zhang, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Deep Probabilistic Accelerated Evaluation: A Robust Certifiable Rare-Event Simulation Methodology for Black-Box Safety-Critical SystemsabstractEvaluating the reliability of intelligent physical systems against rare safety-critical events poses a huge testing burden for real-world applications. Simulation provides a useful platform to evaluate the extremal risks of these systems before their deployments. Importance Sampling (IS), while proven to be powerful for rare-event simulation, faces challenges in handling these learning-based systems due to their black-box nature that fundamentally undermines its efficiency guarantee, which can lead to under-estimation without diagnostically detected. We propose a framework called Deep Probabilistic Accelerated Evaluation (Deep-PrAE) to design statistically guaranteed IS, by converting black-box samplers that are versatile but could lack guarantees, into one with what we call a relaxed efficiency certificate that allows accurate estimation of bounds on the safety-critical event probability. We present the theory of Deep-PrAE that combines the dominating point concept with rare-event set learning via deep neural network classifiers, and demonstrate its effectiveness in numerical examples including the safety-testing of an intelligent driving algorithm. Mansur Arief, Guru Koushik Senthil Kumar, Yuanlu Bai, Shengyi He, Wenhao Ding, Henry Lam, Ding Zhao |
AISTATS | 8 |
| 2021 | Dynamic Sparsity Neural Networks for Automatic Speech RecognitionabstractIn automatic speech recognition (ASR), model pruning is a widely adopted technique that reduces model size and latency to deploy neural network models on edge devices with resource constraints. However, multiple models with different sparsity levels usually need to be separately trained and deployed to heterogeneous target hardware with different resource specifications and for applications that have various latency requirements. In this paper, we present Dynamic Sparsity Neural Networks (DSNN) that, once trained, can instantly switch to any predefined sparsity configuration at run-time. We demonstrate the effectiveness and flexibility of DSNN using experiments on internal production datasets with Google Voice Search data, and show that the performance of a DSNN model is on par with that of individually trained single sparsity networks. Our trained DSNN model, therefore, can greatly ease the training process and simplify deployment in diverse scenarios with resource constraints. Zhaofeng Wu, Ding Zhao, Qiao Liang 0001, Anmol Gulati, Ruoming Pang |
ICASSP | 2 |
| 2021 | Context-Aware Safe Reinforcement Learning for Non-Stationary EnvironmentsabstractSafety is a critical concern when deploying reinforcement learning agents for realistic tasks. Recently, safe reinforcement learning algorithms have been developed to optimize the agent’s performance while avoiding violations of safety constraints. However, few studies have addressed the nonstationary disturbances in the environments, which may cause catastrophic outcomes. In this paper, we propose the context-aware safe reinforcement learning (CASRL) method, a metal-earning framework to realize safe adaptation in non-stationary environments. We use a probabilistic latent variable model to achieve fast inference of the posterior environment transition distribution given the context data. Safety constraints are then evaluated with uncertainty-aware trajectory sampling. Prior safety constraints are formulated with domain knowledge to improve safety during exploration. The algorithm is evaluated in realistic safety-critical environments with non-stationary disturbances. Results show that the proposed algorithm significantly outperforms existing baselines in terms of safety and robustness. Baiming Chen, Zuxin Liu, Mengdi Xu, Wenhao Ding, Liang Li 0004, Ding Zhao |
ICRA | 7 |
| 2021 | Personalized Keyphrase Detection Using Speaker and Environment InformationabstractIn this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The system is implemented with an end-to-end trained automatic speech recognition (ASR) model and a text-independent speaker verification model. To address the challenge of detecting these keyphrases under various noisy conditions, a speaker separation model is added to the feature frontend of the speaker verification model, and an adaptive noise cancellation (ANC) algorithm is included to exploit cross-microphone noise coherence. Our experiments show that the text-independent speaker verification model largely reduces the false triggering rate of the keyphrase detection, while the speaker separation model and adaptive noise cancellation largely reduce false rejections. Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Ding Zhao, Yiteng Huang, Arun Narayanan, Ian McGraw |
Interspeech | 5 |
| 2021 | Delay-aware model-based reinforcement learning for continuous control
Baiming Chen, Mengdi Xu, Liang Li 0004, Ding Zhao |
Neurocomputing | 4 |
| 2021 | Highway Exiting Planner for Automated Vehicles Using Reinforcement LearningabstractExiting from highways in crowded dynamic traffic is an important path planning task for autonomous vehicles (AVs). This task can be challenging because of the uncertain motion of surrounding vehicles and limited sensing/observing window. Conventional path planning methods usually compute a mandatory lane change (MLC) command, but the lane change behavior (e.g., vehicle speed and gap acceptance) should also adapt to traffic conditions and the urgency for exiting. In this paper, we propose a reinforcement learning-enhanced highway-exit planner. The learning-based strategy learns from past failures and adjusts the vehicle motion when the AV fails to exit. The reinforcement learning is based on the Monte Carlo tree search (MCTS) approach. The proposed learning-enhanced highway-exit planner is tested 6000 times in stochastic simulations. The results indicate that the proposed planner achieves a higher probability of successful highway exiting than a benchmark MLC planner. Zhong Cao 0003, Diange Yang, Shaobing Xu, Huei Peng, Boqi Li 0001, Shuo Feng 0002, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2021 | How to Evaluate Proving Grounds for Self-Driving? A Quantitative ApproachabstractProving ground has been a critical component in testing and validation for Connected and Automated Vehicles (CAV). Although quite a few world-class testing facilities have been under construction over the years, the evaluation of proving grounds themselves as testing approaches has rarely been studied. In this paper, we present the first attempt to systematically evaluate CAV proving grounds and contribute to a generative sample-based approach to assessing the representation of traffic scenarios in proving grounds. Leveraging typical use cases extracted from naturalistic driving events, we establish a strong link between proving ground testing results of CAVs and their anticipated public street performance. We present benchmark results of our approach on three world-class CAV testing facilities: Mcity, Almono (Uber ATG), and Kcity. We successfully show the overall evaluation of these proving grounds in terms of their capability to accommodate real-world traffic scenarios. We believe that when the effectiveness of a testing ground itself is validated, the testing results would grant more confidence for CAV public deployment. Rui Chen 0030, Mansur Arief, Weiyang Zhang, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | A Streaming On-Device End-To-End Model Surpassing Server-Side Conventional Model Quality and LatencyabstractThus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e., the time the hypothesis is finalized after the user stops speaking. In this paper, we develop a first-pass Recurrent Neural Network Transducer (RNN-T) model and a second-pass Listen, Attend, Spell (LAS) rescorer that surpasses a conventional model in both quality and latency. On the quality side, we incorporate a large number of utterances across varied domains [1] to increase acoustic diversity and the vocabulary seen by the model. We also train with accented English speech to make the model more robust to different pronunciations. In addition, given the increased amount of training data, we explore a varied learning rate schedule. On the latency front, we explore using the end-of-sentence decision emitted by the RNN-T model to close the microphone, and also introduce various optimizations to improve the speed of LAS rescoring. Overall, we find that RNN-T+LAS offers a better WER and latency tradeoff compared to a conventional model. For example, for the same latency, RNN-T+LAS obtains a 8% relative improvement in WER, while being more than 400-times smaller in model size. Tara N. Sainath, Yanzhang He, Bo Li 0028, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo-Yiin Chang, Wei Li 0133, Raziel Alvarez, Chung-Cheng Chiu, Alexander Gruenstein, Anjuli Kannan, Qiao Liang 0001, Ian McGraw, Cal Peyser, Rohit Prabhavalkar, Golan Pundak, David Rybach, Yuan Shangguan, Yash Sheth, Trevor Strohman, Mirkó Visontai, Yu Zhang 0033, Ding Zhao |
ICASSP | 28 |
| 2020 | CMTS: A Conditional Multiple Trajectory Synthesizer for Generating Safety-Critical Driving ScenariosabstractNaturalistic driving trajectory generation is crucial for the development of autonomous driving algorithms. However, most of the data is collected in collision-free scenarios leading to the sparsity of the safety-critical cases. When considering safety, testing algorithms in near-miss scenarios that rarely show up in off-the-shelf datasets and are costly to accumulate is a vital part of the evaluation. As a remedy, we propose a safety-critical data synthesizing framework based on variational Bayesian methods and term it as Conditional Multiple Trajectory Synthesizer (CMTS). We extend a generative model to connect safe and collision driving data by representing their distribution in the latent space and use conditional probability to adapt to different maps. Sampling from the mixed distribution enables us to synthesize the safety-critical data not shown in the safe or collision datasets. Experimental results demonstrate that the generated dataset covers many different realistic scenarios, especially the near-misses. We conclude that the use of data generated by CMTS can improve the accuracy of trajectory predictions and autonomous vehicle safety. Wenhao Ding, Mengdi Xu, Ding Zhao |
ICRA | 3 |
| 2020 | Learning to Collide: An Adaptive Safety-Critical Scenarios Generating MethodabstractLong-tail and rare event problems become crucial when autonomous driving algorithms are applied in the real world. For the purpose of evaluating systems in challenging settings, we propose a generative framework to create safety-critical scenarios for evaluating specific task algorithms. We first represent the traffic scenarios with a series of autoregressive building blocks and generate diverse scenarios by sampling from the joint distribution of these blocks. We then train the generative model as an agent (or a generator) to search the risky scenario parameters for a given driving algorithm. We treat the driving algorithm as an environment that returns high reward to the agent when a risky scenario is generated. The whole process is optimized by the policy gradient reinforcement learning method. Through the experiments conducted on several scenarios in the simulation, we demonstrate that the proposed framework generates safety-critical scenarios more efficiently than grid search or human design methods. Another advantage of this method is its adaptiveness to the routes and parameters. Wenhao Ding, Baiming Chen, Minjun Xu, Ding Zhao |
IROS | 4 |
| 2020 | MAPPER: Multi-Agent Path Planning with Evolutionary Reinforcement Learning in Mixed Dynamic EnvironmentsabstractMulti-agent navigation in dynamic environments is of great industrial value when deploying a large scale fleet of robot to real-world applications. This paper proposes a decentralized partially observable multi-agent path planning with evolutionary reinforcement learning (MAPPER) method to learn an effective local planning policy in mixed dynamic environments. Reinforcement learning-based methods usually suffer performance degradation on long-horizon tasks with goal-conditioned sparse rewards, so we decompose the long-range navigation task into many easier sub-tasks under the guidance of a global planner, which increases agents' performance in large environments. Moreover, most existing multi-agent planning approaches assume either perfect information of the surrounding environment or homogeneity of nearby dynamic agents, which may not hold in practice. Our approach models dynamic obstacles' behavior with an image-based representation and trains a policy in mixed dynamic environments without homogeneity assumption. To ensure multi-agent training stability and performance, we propose an evolutionary training approach that can be easily scaled to large and complex environments. Experiments show that MAPPER is able to achieve higher success rates and more stable performance when exposed to a large number of non-cooperative dynamic obstacles compared with traditional reaction-based planner LRA* and the state-of-the-art learning-based method. Zuxin Liu, Baiming Chen, Hongyi Zhou, Guru Koushik Senthil Kumar, Martial Hebert, Ding Zhao |
IROS | 6 |
| 2020 | Multi-Vehicle Interaction Scenarios Generation with Interpretable Traffic Primitives and Gaussian Process RegressionabstractGenerating multi-vehicle interaction scenarios can benefit motion planning and decision making of autonomous vehicles when on-road data is insufficient. This paper presents an efficient approach to generate varied multi-vehicle interaction scenarios that can both adapt to different road geometries and inherit the key interaction patterns in real-world driving. Towards this end, the available multi-vehicle interaction scenarios are temporally segmented into several interpretable fundamental building blocks, called traffic primitives, via the Bayesian nonparametric learning. Then, the changepoints of traffic primitives are transformed into the desired road to generate collision-free interaction trajectories through a sampling-based path planning algorithm. The Gaussian process regression is finally introduced to control the variance and smoothness of the generated multi-vehicle interaction trajectories. Experiments with simulation results of three multi-vehicle trajectories at different road conditions are carried out. The experimental results demonstrate that our proposed method can generate a bunch of human-like multi-vehicle interaction trajectories that can fit different road conditions remaining the key interaction patterns of agents in the provided scenarios, which is import to the development of autonomous vehicles. Weiyang Zhang, Wenshuo Wang 0001, Ding Zhao |
IV | 4 |
| 2020 | Task-Agnostic Online Reinforcement Learning with an Infinite Mixture of Gaussian ProcessesabstractContinuously learning to solve unseen tasks with limited experience has been extensively pursued in meta-learning and continual learning, but with restricted assumptions such as accessible task distributions, independently and identically distributed tasks, and clear task delineations. However, real-world physical tasks frequently violate these assumptions, resulting in performance degradation. This paper proposes a continual online model-based reinforcement learning approach that does not require pre-training to solve task-agnostic problems with unknown task boundaries. We maintain a mixture of experts to handle nonstationarity, and represent each different type of dynamics with a Gaussian Process to efficiently leverage collected data and expressively model uncertainty. We propose a transition prior to account for the temporal dependencies in streaming data and update the mixture online via sequential variational inference. Our approach reliably handles the task distribution shift by generating new models for never-before-seen dynamics and reusing old models for previously seen dynamics. In experiments, our approach outperforms alternative methods in non-stationary tasks, including classic control with changing dynamics and decision making in different driving scenarios. Mengdi Xu, Wenhao Ding, Zuxin Liu, Baiming Chen, Ding Zhao |
NeurIPS | 6 |
| 2019 | Streaming End-to-end Speech Recognition for Mobile DevicesabstractEnd-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition. E2E models, however, present numerous challenges: In order to be truly useful, such models must decode speech utterances in a streaming fashion, in real time; they must be robust to the long tail of use cases; they must be able to leverage user-specific context (e.g., contact lists); and above all, they must be extremely accurate. In this work, we describe our efforts at building an E2E speech recog-nizer using a recurrent neural network transducer. In experimental evaluations, we find that the proposed approach can outperform a conventional CTC-based model in terms of both latency and accuracy in a number of evaluation categories. Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Ruoming Pang, Qiao Liang 0001, Deepti Bhatia, Yuan Shangguan, Bo Li 0028, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo-Yiin Chang, Kanishka Rao, Alexander Gruenstein |
ICASSP | 6 |
| 2019 | A Multi-Vehicle Trajectories Generator to Simulate Vehicle-to-Vehicle Encountering ScenariosabstractGenerating multi-vehicle trajectories from existing limited data can provide rich resources for autonomous vehicle development and testing. This paper introduces a multi-vehicle trajectory generator (MTG) that can encode multi-vehicle interaction scenarios (called driving encounters) into an interpretable representation from which new driving encounter scenarios are generated by sampling. The MTG consists of a bi-directional encoder and a multi-branch decoder. A new disentanglement metric is then developed for model analyses and comparisons in terms of model robustness and the independence of the latent codes. Comparison of our proposed MTG with β-VAE and InfoGAN demonstrates that the MTG has stronger capability to purposely generate rational vehicle-to-vehicle encounters through operating the disentangled latent codes. Thus the MTG could provide more data for engineers and researchers to develop testing and evaluation scenarios for autonomous vehicles. Wenhao Ding, Wenshuo Wang 0001, Ding Zhao |
ICRA | 3 |
| 2019 | Where Should We Place LiDARs on the Autonomous Vehicle? - An Optimal Design ApproachabstractAutonomous vehicle manufacturers recognize that LiDAR provides accurate 3D views and precise distance measures under highly uncertain driving conditions. Its practical implementation, however, remains costly. This paper investigates the optimal LiDAR configuration problem to achieve utility maximization. We use the perception area and non-detectable subspace to construct the design procedure as solving a min-max optimization problem and propose a bio-inspired measure - volume to surface area ratio (VSR) - as an easy-to-evaluate cost function representing the notion of the size of the non-detectable subspaces of a given configuration. We then adopt a cuboid-based approach to show that the proposed VSR-based measure is a well-suited proxy for object detection rate. It is found that the Artificial Bee Colony evolutionary algorithm yields a tractable cost function computation. Our experiments highlight the effectiveness of our proposed VSR measure in terms of cost-effectiveness configuration as well as providing insightful analyses that can improve the design of AV systems. Zuxin Liu, Mansur Arief, Ding Zhao |
ICRA | 3 |
| 2019 | Shallow-Fusion End-to-End Contextual Biasing
Ding Zhao, Tara N. Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li 0028, Ruoming Pang |
INTERSPEECH | 1 |
| 2019 | Destriping Algorithms Based on Statistics and Spatial Filtering for Visible-to-Thermal Infrared Pushbroom Hyperspectral ImageryabstractFollowing calibration, hyperspectral images remain affected by spatial dimension nonuniformity, i.e., stripe noise, due to stray light interference, slit contamination, and instrument instability. The full spectrum airborne hyperspectral imager (FAHI) is a Chinese next-generation pushbroom sensor with a spectral range covering the visible near-infrared, shortwave-infrared (SWIR), and thermal infrared regions with spectral sampling intervals of 2.34, 3, and 32 nm, respectively. However, the residual stripe noise remains in FAHI images after relative radiometric correction based on the laboratory calibration, especially in low signal-to-noise ratio bands. To solve this problem, a new technique combining image statistics and spatial filtering algorithms has been developed for FAHI image correction. In this method, image statistics are obtained to calculate the gain and offset of each pixel for image nonuniformity correction. Then, a spatial filter removes the residual stripes. This paper presents the principles of this method along with details of validation experiments and results. To validate the effectiveness of the proposed method, comparison with two destriping methods and quantitative analyses are carried out. Moreover, the method is applied to an SWIR hyperspectral image from the TianGong-1 spacecraft, yielding good results. The experimental results suggest that the proposed method is convenient and practical for improving the relative radiometric accuracy of airborne/spaceborne hyperspectral images. Jianxin Jia, Yueming Wang 0002, Liyin Yuan, Ding Zhao, Xiaoqiong Zhuang, Rong Shu, Jianyu Wang 0017 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Improving Localization Accuracy in Connected Vehicle Networks Using Rao-Blackwellized Particle Filters: Theory, Simulations, and ExperimentsabstractA crucial function for automated vehicle technologies is accurate localization. Lane-level accuracy is not readily available from low-cost global navigation satellite system (GNSS) receivers because of factors such as multipath error and atmospheric bias. Approaches such as differential GNSS can improve localization accuracy, but usually require investment in expensive base stations. Connected vehicle technologies provide an alternative approach in improving the localization accuracy. It will be shown in this paper that localization accuracy can be enhanced using crude GNSS measurements from a group of connected vehicles, by matching their locations to a digital map. A Rao-Blackwellized particle filter is used to jointly estimate the common biases of the pseudo-ranges and the vehicle positions. Multipath biases, which introduce receiver-specific (non-common) error, are mitigated by a multi-hypothesis detection-rejection approach. The temporal correlation of the estimations is exploited through the prediction-update process. The proposed approach is compared with existing methods using both simulations and experimental results. It was found that the proposed algorithm can eliminate the common biases and reduce the localization error to below 1 m under open sky conditions. Macheng Shen, Jing Sun 0003, Huei Peng, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | Driving Style Analysis Using Primitive Driving Patterns With Bayesian Nonparametric ApproachesabstractDriving style analysis plays a pivotal role in intelligent vehicle design. This paper presents a novel framework for driving style analysis based on primitive driving patterns. To this end, a Bayesian nonparametric approach based on a hidden semi-Markov model (HSMM) is introduced to extract the primitive driving patterns from muti-dimensional time-series driving data without prior knowledge of these driving patterns. For the Bayesian nonparametric approach, a hierarchical Dirichlet process (HDP) is applied to learn the unknown smooth dynamical modes in the HSMM, called primitive driving patterns. Two other types of Bayesian nonparametric approaches (HDP-HMM and sticky HDP-HMM) are developed as comparatives in order to show the advantages of the HDP-HSMM. The naturalistic car-following data of 18 drivers are collected from the University of Michigan Safety Pilot Model Deployment database. For each driver, 75 primitive driving patterns are semantically predefined according to their physical and psychological perception thresholds. The individual driving styles are then semantically analyzed based on the distribution over primitive driving patterns, and the similarity of driving styles among drivers is then evaluated. Experimental results demonstrate that the utilization of driving primitive pattern provides a semantically interpretable way to analyze driver's behavior and driving style. Wenshuo Wang 0001, Junqiang Xi, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Cluster Naturalistic Driving Encounters Using Deep Unsupervised LearningabstractLearning knowledge from driving encounters could help self-driving cars make appropriate decisions when driving in complex settings with nearby vehicles engaged. This paper develops an unsupervised classifier to group naturalistic driving encounters into distinguishable clusters by combining an auto-encoder with k-means clustering (AE-kMC). The effectiveness of AE-kMC was validated using the data of 10,000 naturalistic driving encounters which were collected by the University of Michigan, Ann Arbor in the past five years. We compare our developed method with the k-means clustering methods and experimental results demonstrate that the AE-kMC method outperforms the original k-means clustering method. Wenshuo Wang 0001, Zhaobin Mo, Ding Zhao |
Intelligent Vehicles Symposium | 4 |
| 2018 | Deep Context: End-to-end Contextual Speech RecognitionabstractIn automatic speech recognition (ASR) what a user says depends on the particular context she is in. Typically, this context is represented as a set of word n-grams. In this work, we present a novel, all-neural, end-to-end (E2E) ASR system that utilizes such context. Our approach, which we refer to as Contextual Listen, Attend and Spell (CLAS) jointly-optimizes the ASR components along with embeddings of the context n-grams. During inference, the CLAS system can be presented with context phrases which might contain-of-vocabulary (OOV) terms not seen during training. We compare our proposed system to a more traditional contextualization approach, which performs shallow-fusion between independently trained LAS and contextual n-gram models during beam search. Across a number of tasks, we find that the proposed CLAS system outperforms the baseline method by as much as 68% relative WER, indicating the advantage of joint optimization over individually trained components. Golan Pundak, Tara N. Sainath, Rohit Prabhavalkar, Anjuli Kannan, Ding Zhao |
SLT | 5 |
| 2018 | Accelerated Evaluation of Automated Vehicles Using Piecewise Mixture ModelsabstractThe process to certify highly automated vehicles has not yet been defined by any country in the world. Currently, companies test automated vehicles on public roads, which is time-consuming and inefficient. We proposed the accelerated evaluation concept, which uses a modified statistics of the surrounding vehicles and the importance sampling theory to reduce the evaluation time by several orders of magnitude, while ensuring the evaluation results are statistically accurate. In this paper, we further improve the accelerated evaluation concept by using piecewise mixture distribution models, instead of single parametric distribution models. We developed and applied this idea to forward collision control system reacting to vehicles making cutin lane changes. The behavior of the cutin vehicles was modeled based on more than 403,581 lane changes collected by the University of Michigan Safety Pilot Model Deployment Program. Simulation results confirm that the accuracy and efficiency of the piecewise mixture distribution method outperformed single parametric distribution methods in accuracy and efficiency, and accelerated the evaluation process by almost four orders of magnitude. Henry Lam, David J. LeBlanc, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2018 | The Impact of Road Configuration in V2V-Based Cooperative Localization: Mathematical Analysis and Real-World EvaluationabstractCooperative map matching (CMM) uses the global navigation satellite system (GNSS) position information of a group of vehicles to improve the standalone localization accuracy. While increasing accuracy is expected by increasing the number of participating vehicles, fundamental questions on how the vehicle membership within CMM affects the performance of the CMM results need to be addressed to provide guidelines for design and optimization of the vehicle network. This paper presents a theoretical study that establishes a framework for quantitative evaluation of the impact of the road constraints on the CMM accuracy. More specifically, a closed-form expression of the CMM error in terms of the road constraints and GNSS error is derived based on a simple CMM rule. The asymptotic decay of the CMM error as the number of vehicles increases is established and justified through numerical simulations. Moreover, it is proved that the CMM error can be minimized if the directions of the roads on which the connected vehicles travel obey a uniform distribution. Finally, the localization accuracy of CMM is evaluated based on the real traffic data. The contributions of this paper include establishing a theoretical foundation for CMM as well as providing insight and motivation for applications of CMM. Macheng Shen, Jing Sun 0003, Ding Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Accelerated Evaluation of Automated Vehicles in Car-Following ManeuversabstractThe safety of automated vehicles (AVs) must be assured before their release and deployment. The current approach to evaluation relies primarily on 1) testing AVs on public roads or 2) track testing with scenarios defined in a test matrix. These two methods have completely opposing drawbacks: the former, while offering realistic scenarios, takes too much time to execute and the latter, though it can be completed in a short amount of time, has no clear correlation to safety benefits in the real world. To avoid the aforementioned problems, we propose accelerated evaluation, focusing on the car-following scenario. The stochastic human-controlled vehicle (HV) motions are modeled based on 1.3 million miles of naturalistic driving data collected by the University of Michigan Safety Pilot Model Deployment Program. The statistics of the HV behaviors are then modified to generate more intense interactions between HVs and AVs to accelerate the evaluation procedure. The importance sampling theory was used to ensure that the safety benefits of AVs are accurately assessed under accelerated tests. Crash, injury and conflict rates for a simulated AV are simulated to demonstrate the proposed approach. Results show that test duration is reduced by a factor of 300 to 100 000 compared with the non-accelerated (naturalistic) evaluation. In other words, the proposed techniques have great potential for accelerating the AV evaluation process. Ding Zhao, Xianan Huang, Huei Peng, Henry Lam, David J. LeBlanc |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Evaluation of automated vehicles in the frontal cut-in scenario - An enhanced approach using piecewise mixture modelsabstractEvaluation and testing are critical for the development of Automated Vehicles (AVs). Currently, companies test AVs on public roads, which is very time-consuming and inefficient. We proposed the Accelerated Evaluation concept which uses a modified statistics of the surrounding vehicles and the Importance Sampling theory to reduce the evaluation time by several orders of magnitude, while ensuring the final evaluation results are accurate. In this paper, we further extend this idea by using Piecewise Mixture Distribution models instead of Single Distribution models. We demonstrate this idea to evaluate vehicle safety in lane change scenarios. The behavior of the cut-in vehicles was modeled based on more than 400,000 naturalistic driving lane changes collected by the University of Michigan Safety Pilot Model Deployment Program. Simulation results confirm that the accuracy and efficiency of the Piecewise Mixture Distribution method are better than the single distribution. Ding Zhao, Henry Lam, David J. LeBlanc, Huei Peng |
ICRA | 2 |
| 2017 | Evaluation of automated vehicles encountering pedestrians at unsignalized crossingsabstractInteractions between vehicles and pedestrians have always been a major problem in traffic safety. Experienced human drivers are able to analyze the environment and choose driving strategies that will help them avoid crashes. What is not yet clear, however, is how automated vehicles will interact with pedestrians. This paper proposes a new method for evaluating the safety and feasibility of the driving strategy of automated vehicles when encountering unsignalized crossings. MobilEye® sensors installed on buses in Ann Arbor, Michigan, collected data on 2,973 valid crossing events. A stochastic interaction model was then created using a multivariate Gaussian mixture model. This model allowed us to simulate the movements of pedestrians reacting to an oncoming vehicle when approaching unsignalized crossings, and to evaluate the passing strategies of automated vehicles. A simulation was then conducted to demonstrate the evaluation procedure. Baiming Chen, Ding Zhao, Huei Peng |
Intelligent Vehicles Symposium | 2 |
| 2017 | Towards secure and safe appified automated vehiclesabstractThe advancement in Autonomous Vehicles (AVs) has created an enormous market for the development of self-driving functionalities, raising the question of how it will transform the traditional vehicle development process. One adventurous proposal is to open the AV platform to third-party developers, so that AV functionalities can be developed in a crowd-sourcing way, which could provide tangible benefits to both automakers and end users. Some pioneering companies in the automotive industry have made the move to open the platform so that developers are allowed to test their code on the road. Such openness, however, brings serious security and safety issues by allowing untrusted code to run on the vehicle. In this paper, we introduce the concept of an Appified AV platform that opens the development framework to third-party developers. To further address the safety challenges, we propose an enhanced appified AV design schema called AVGUARD, which focuses primarily on mitigating the threats brought about by untrusted code, leveraging theory in the vehicle evaluation field, and conducting program analysis techniques in the cyber security area. Our study provides guidelines and suggested practice for the future design of open AV platforms. Yunhan Jia, Ding Zhao, Qi Alfred Chen, Z. Morley Mao |
Intelligent Vehicles Symposium | 2 |
| 2017 | Analysis of unprotected intersection left-turn conflicts based on naturalistic driving dataabstractAnalyzing and reconstructing driving scenarios is crucial for testing and evaluating highly automated vehicles (HAVs). This research analyzed left-turn / straight-driving conflicts at unprotected intersections by extracting actual vehicle motion data from a naturalistic driving database collected by the University of Michigan. Nearly 7,000 left turn across path - opposite direction (LTAP/OD) events involving heavy trucks and light vehicles were extracted and used to build a stochastic model of such LTAP/OD scenario, which is among the top priority light-vehicle pre-crash scenarios identified by National Highway Traffic Safety Administration (NHTSA). Statistical analysis showed that vehicle type is a significant factor, whereas the change of season seems to have limited influence on the statistical nature of the conflict. The results can be used to build testing environments for HAVs to simulate the LTAP/OD crash cases in a stochastic manner. Xinpeng Wang 0002, Ding Zhao, Huei Peng, David J. LeBlanc |
Intelligent Vehicles Symposium | 2 |
| 2017 | Evaluation of a semi-autonomous lane departure correction system using naturalistic driving dataabstractEvaluating the effectiveness and benefits of driver assistance systems is essential for improving the system performance. In this paper, we propose an efficient evaluation method for a semi-autonomous lane departure correction system. To achieve this, we apply a bounded Gaussian mixture model to describe drivers' stochastic lane departure behavior learned from naturalistic driving data, which can regenerate departure behaviors to evaluate the lane departure correction system. In the stochastic lane departure model, we conduct a dimension reduction to reduce the computation cost. Finally, to show the advantages of our proposed evaluation approach, we compare steering systems with and without lane departure assistance based on the stochastic lane departure model. The simulation results show that the proposed method can effectively evaluate the lane departure correction system. Ding Zhao, Wenshuo Wang 0001, David J. LeBlanc |
Intelligent Vehicles Symposium | 1 |
| 2017 | The Impact of Road Configuration on V2V-Based Cooperative LocalizationabstractCooperative localization with map matching has been shown to reduce Global Navigation Satellite System (GNSS) localization error from several meters to sub-meter level by fusing the GNSS measurements of four vehicles in our previous work. While further error reduction is expected to be achievable by increasing the number of vehicles, the quantitative relationship between the estimation error and the number of connected vehicles has neither been systematically investigated nor analytically proved. In this work, a theoretical study is presented that analytically proves the correlation between the localization error and the number of connected vehicles in two cases of practical interest. More specifically, it is shown that, under the assumption of small non-common error, the expected square error of the GNSS common error correction is inversely proportional to the number of vehicles, if the road directions obey a uniform distribution, or inversely proportional to logarithm of the number of vehicles, if the road directions obey a Bernoulli distribution. Numerical simulations are conducted to justify these analytic results. Moreover, the simulation results show that the aforementioned error decrement rates hold even when the assumption of small non-common error is violated. Macheng Shen, Ding Zhao, Jing Sun 0003 |
VTC Spring | 2 |
| 2017 | Empirical Study of DSRC Performance Based on Safety Pilot Model Deployment DataabstractDedicated short range communication (DSRC) was designed to provide reliable wireless communication for intelligent transportation system applications. Sharing information among cars and between cars and the infrastructure, pedestrians, or “the cloud” has great potential to improve safety, mobility, and fuel economy. DSRC is being considered by the U.S. Department of Transportation to be required for ground vehicles. In the past, their performance has been assessed thoroughly in the laboratories and limited field testing, but not on a large fleet. In this paper, we present the analysis of DSRC performance using data from the world's largest connected vehicle test program-Safety Pilot Model Deployment lead by the University of Michigan. We first investigate their maximum and effective range, and then study the effect of environmental factors, such as trees/foliage, weather, buildings, vehicle travel direction, and road elevation. The results can be used to guide future DSRC equipment placement and installation, and can be used to develop DSRC communication models for numerical simulations. Xianan Huang, Ding Zhao, Huei Peng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Accelerated Evaluation of Automated Vehicles Safety in Lane-Change Scenarios Based on Importance Sampling TechniquesabstractAutomated vehicles (AVs) must be thoroughly evaluated before their release and deployment. A widely used evaluation approach is the Naturalistic-Field Operational Test (N-FOT), which tests prototype vehicles directly on the public roads. Due to the low exposure to safety-critical scenarios, N-FOTs are time consuming and expensive to conduct. In this paper, we propose an accelerated evaluation approach for AVs. The results can be used to generate motions of the other primary vehicles to accelerate the verification of AVs in simulations and controlled experiments. Frontal collision due to unsafe cut-ins is the target crash type of this paper. Human-controlled vehicles making unsafe lane changes are modeled as the primary disturbance to AVs based on data collected by the University of Michigan Safety Pilot Model Deployment Program. The cut-in scenarios are generated based on skewed statistics of collected human driver behaviors, which generate risky testing scenarios while preserving the statistical information so that the safety benefits of AVs in nonaccelerated cases can be accurately estimated. The cross-entropy method is used to recursively search for the optimal skewing parameters. The frequencies of the occurrences of conflicts, crashes, and injuries are estimated for a modeled AV, and the achieved accelerated rate is around 2000 to 20 000. In other words, in the accelerated simulations, driving for 1000 miles will expose the AV with challenging scenarios that will take about 2 to 20 million miles of real-world driving to encounter. This technique thus has the potential to greatly reduce the development and validation time for AVs. Ding Zhao, Henry Lam, Huei Peng, Shan Bao, David J. LeBlanc, Kazutoshi Nobukawa, Christopher S. Pan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Gap Acceptance During Lane Changes by Large-Truck Drivers - An Image-Based AnalysisabstractThis paper presents an analysis of rearward gap acceptance characteristics of drivers of large trucks in highway lane change scenarios. The range between the vehicles was inferred from camera images using the estimated lane width obtained from the lane tracking camera as the reference. Six-hundred lane change events were acquired from a large-scale naturalistic driving data set. The kinematic variables from the image-based gap analysis were filtered by the weighted linear least squares in order to extrapolate them at the lane change time. In addition, the time-to-collision and required deceleration were computed, and potential safety threshold values are provided. The resulting range and range rate distributions showed directional discrepancies, i.e., in left lane changes, large trucks are often slower than other vehicles in the target lane, whereas they are usually faster in right lane changes. Video observations have confirmed that major motivations for changing lanes are different depending on the direction of move, i.e., moving to the left (faster) lane occurs due to a slower vehicle ahead or a merging vehicle on the right-hand side, whereas right lane changes are frequently made to return to the original lane after passing. Kazutoshi Nobukawa, Shan Bao, David J. LeBlanc, Ding Zhao, Huei Peng, Christopher S. Pan |
IEEE Trans. Intell. Transp. Syst. | 4 |