VLDB 2026 Research / reviewers in the wild / expert
Changliu Liu
dblp:166/3563
· DBLP profile ↗
46ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0002-3767-5517ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 3 first-author · 29 since 2021Systems, architecture and hardware · 17 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Synthesis of Formally Verified Neural Value Function for Hamilton-Jacobi Reachability Analysis (Abstract Reprint)abstractHamilton-Jacobi (HJ) reachability analysis provides a formal method for guaranteeing safety in constrained control problems. It synthesizes a value function to represent a long-term safe set called feasible region. Early synthesis methods based on state space discretization cannot scale to high-dimensional problems, while recent methods that use neural networks to approximate value functions result in unverifiable feasible regions. To achieve both scalability and verifiability, we propose a framework for synthesizing verified neural value functions for HJ reachability analysis. Our framework consists of three stages: pre-training, adversarial training, and verification-guided training. We design three techniques to address three challenges to improve scalability respectively: boundary-guided backtracking (BGB) to improve counterexample search efficiency, entering state regularization (ESR) to enlarge feasible region, and activation pattern alignment (APA) to accelerate neural network verification. We also provide a neural safety certificate synthesis and verification benchmark called Cersyve-9, which includes nine commonly used safe control tasks and supplements existing neural network verification benchmarks. Our framework successfully synthesizes verified neural value functions on all tasks, and our proposed three techniques exhibit superior scalability and efficiency compared with existing methods. Hanjiang Hu, Tianhao Wei, Shengbo Eben Li, Changliu Liu |
AAAI | 5 |
| 2026 | ε-retraining reinforcement learning algorithmsabstractAbstract We present $$\varepsilon $$ , a general exploration strategy for reinforcement learning (RL) that encourages adherence to behavioral preferences while preserving the convergence guarantees of the underlying RL algorithm. $$\varepsilon $$ maintains a dynamic collection of retrain areas—regions of the state space where the agent previously violated a specified preference—and mixes the standard uniform restart distribution with states from these areas, according to a decaying parameter $$\varepsilon $$ . This mixed retraining thus focuses on enforcing the desired behaviors in the collected areas. We develop the theory for both policy and value-based methods, showing that: (i) in policy-based settings, our method retains monotonic improvement bounds; and (ii) in value-based settings, $$\varepsilon $$ preserves convergence properties without additional assumptions. The approach is simple to integrate into existing RL algorithms and improves sample efficiency and behavioral adherence in the locomotion, power systems, and navigation tasks tested. These results establish $$\varepsilon $$ as a lightweight, theoretically grounded mechanism for incorporating behavioral preferences into RL. Luca Marzari, Changliu Liu, Priya L. Donti, Enrico Marchesini |
Auton. Agents Multi Agent Syst. | 2 |
| 2025 | ModelVerification.jl: A Comprehensive Toolbox for Formally Verifying Deep Neural NetworksabstractAbstract Deep Neural Networks (DNN) are crucial in approximating nonlinear functions across diverse applications, ranging from image classification to control. Verifying specific input-output properties can be a highly challenging task due to the lack of a single, self-contained framework that allows a complete range of various model architecture and input-output properties. To this end, we present ( https://github.com/intelligent-control-lab/ModelVerification.jl ), the first comprehensive, cutting-edge toolbox that contains a suite of state-of-the-art methods for verifying different types of DNNs and input-output specifications. This versatile toolbox is designed to empower developers and machine learning practitioners with robust tools for verifying and ensuring the trustworthiness of their DNN models. Tianhao Wei, Hanjiang Hu, Luca Marzari, Kai S. Yun, Peizhi Niu, Xusheng Luo, Changliu Liu |
CAV (2) | 7 |
| 2025 | Generating Physically Stable and Buildable Brick Structures from Text
Ava Pun, Kangle Deng, Ruixuan Liu, Deva Ramanan, Changliu Liu, Jun-Yan Zhu |
ICCV | 5 |
| 2025 | ThinkBot: Embodied Instruction Following with Thought Chain ReasoningabstractEmbodied Instruction Following (EIF) requires agents to complete human instruction by interacting objects in complicated surrounding environments. Conventional methods directly consider the sparse human instruction to generate action plans for agents, which usually fail to achieve human goals because of the instruction incoherence in action descriptions. On the contrary, we propose ThinkBot that reasons the thought chain in human instruction to recover the missing action descriptions, so that the agent can successfully complete human goals by following the coherent instruction. Specifically, we first design an instruction completer based on large language models to recover the missing actions with interacted objects between consecutive human instruction, where the perceived surrounding environments and the completed sub-goals are considered for instruction completion. Based on the partially observed scene semantic maps, we present an object localizer to infer the position of interacted objects and the related Bayesian uncertainty for close-loop planning. Extensive experiments in the simulated environment show that our ThinkBot outperforms the state-of-the-art EIF methods by a sizable margin in both success rate and execution efficiency. Project page: https://guanxinglu.github.io/thinkbot/. Guanxing Lu, Ziwei Wang 0010, Changliu Liu, Jiwen Lu, Yansong Tang |
ICLR | 3 |
| 2025 | HOVER: Versatile Neural Whole-Body Controller for Humanoid RobotsabstractHumanoid whole-body control requires adapting to diverse tasks such as navigation, loco-manipulation, and tabletop manipulation, each demanding a different mode of control. For example, navigation relies on root velocity or position tracking, while tabletop manipulation prioritizes upper-body joint angle tracking. Existing approaches typically train individual policies tailored to a specific command space, limiting their transferability across modes. We present the key insight that full-body kinematic motion imitation can serve as a common abstraction for all these tasks and provide general-purpose motor skills for learning multiple modes of whole-body control. Building on this, we propose HOVER (Humanoid Versatile Controller), a multi-mode policy distillation framework that consolidates diverse control modes into a unified policy. HOVER enables seamless transitions between control modes while preserving the distinct advantages of each, offering a robust and scalable solution for humanoid control across a wide range of modes. By eliminating the need for policy retraining for each control mode, our approach improves efficiency and flexibility for future humanoid applications. Tairan He, Toru Lin, Zhengyi Luo 0002, Zhenjia Xu, Zhenyu Jiang 0002, Jan Kautz, Changliu Liu, Guanya Shi, Xiaolong Wang 0004, Linxi Fan, Yuke Zhu |
ICRA | 8 |
| 2025 | Robots that Learn to Safely Influence via Prediction-Informed Reach-Avoid Dynamic Games
Ravi Pandya, Changliu Liu, Andrea Bajcsy |
ICRA | 2 |
| 2025 | Safe Control of Quadruped in Varying Dynamics via Safety Index AdaptationabstractVarying dynamics pose a fundamental difficulty when deploying safe control laws in the real world. Safety Index Synthesis (SIS) deeply relies on the system dynamics and once the dynamics change, the previously synthesized safety index becomes invalid. In this work, we show the real-time efficacy of Safety Index Adaptation (SIA) in varying dynamics. SIA enables real-time adaptation to the changing dynamics so that the adapted safe control law can still guarantee 1) forward invariance within a safe region and 2) finite time convergence to that safe region. This work employs SIA on a packagecarrying quadruped robot, where the payload weight changes in real-time. SIA updates the safety index when the dynamics change, e.g., a change in payload weight, so that the quadruped can avoid obstacles while achieving its performance objectives. Numerical study provides theoretical guarantees for SIA and a series of hardware experiments demonstrate the effectiveness of SIA in real-world deployment in avoiding obstacles under varying dynamics. Kai S. Yun, Rui Chen 0030, Chase Dunaway, John M. Dolan, Changliu Liu |
ICRA | 5 |
| 2025 | Improving Policy Optimization via ε-Retrain
Luca Marzari, Priya L. Donti, Changliu Liu, Enrico Marchesini |
AAMAS | 3 |
| 2025 | Time-Optimal Trajectory Generation with Multi-level Continuous Kinodynamics ConstraintsabstractTime-optimal trajectory generation (TOTG) is critical in robotics applications to minimize travel time and increase robot task efficiency. To ensure the trajectory is feasible and executable by the robot, it is important to constrain the trajectory kinodynamics subject to the robot actuator limits. A typical actuator has multiple limits, 1) peak limit, and 2) multi-level continuous limits with different operation time windows. The peak limit bounds the instantaneous kinodynamics (IKD), whereas the continuous limits bound the system continuous kinodynamics (CKD). Existing works only constrain IKD, usually by the actuator peak limit, to achieve time optimality. However, a joint capable of operating at its peak limit momentarily will overheat and damage robot life if the motion continues. Alternatively, users can constrain the IKD with a reduced peak limit to avoid violating continuous limits. However, the reduced peak limit would inevitably sacrifice task efficiency. To address the challenge, this paper studies TOTG with both IKD and CKD, and proposes TOTG-C. It formulates the TOTG as a nonlinear programming (NLP). In particular, it proposes a novel formulation to encode the multi-level CKD constraints efficiently. To the best of our knowledge, TOTG-C is the first work that explicitly considers multi-level CKD constraints. We demonstrate the effectiveness and robustness of the proposed TOTG-C both in simulation and real robot experiments. Ruixuan Liu, Changliu Liu, Jessica Leu |
IROS | 2 |
| 2025 | Eye-In-Finger: Smart Fingers for Delicate Assembly and Disassembly of LEGOabstractManipulation and insertion of small and tight-toleranced objects in robotic assembly remain a critical challenge for vision-based robotics systems due to the required precision and cluttered environment. Conventional global or wrist-mounted cameras often suffer from occlusions when either assembling or disassembling from an existing structure. To address the challenge, this paper introduces "Eye-In-Finger", a novel tool design approach that enhances robotic manipulation by embedding low-cost, high-resolution perception directly at the tool tip. We validate our approach using LEGO assembly and disassembly tasks, which require the robot to manipulate in a cluttered environment and achieve sub-millimeter accuracy and robust error correction due to the tight tolerances. Experimental results demonstrate that our proposed system enables real-time, fine corrections to alignment error, increasing the tolerance of calibration error from 0.4mm to up to 2.0mm for the LEGO manipulation robot. Zhenran Tang, Ruixuan Liu, Changliu Liu |
IROS | 3 |
| 2025 | Passing-Order Decision for Three-to-Two Lane Merging of Connected and Autonomous Vehicles
Cheng-Pei Chien, Ben-Hau Chia, Ching-Yun Chang, Shang-Chien Lin, Iris Hui-Ru Jiang, Changliu Liu, Chung-Wei Lin |
RTCSA | 6 |
| 2025 | Scalable Synthesis of Formally Verified Neural Value Function for Hamilton-Jacobi Reachability AnalysisabstractHamilton-Jacobi (HJ) reachability analysis provides a formal method for guaranteeing safety in constrained control problems. It synthesizes a value function to represent a long-term safe set called feasible region. Early synthesis methods based on state space discretization cannot scale to high-dimensional problems, while recent methods that use neural networks to approximate value functions result in unverifiable feasible regions. To achieve both scalability and verifiability, we propose a framework for synthesizing verified neural value functions for HJ reachability analysis. Our framework consists of three stages: pre-training, adversarial training, and verification-guided training. We design three techniques to address three challenges to improve scalability respectively: boundary-guided backtracking (BGB) to improve counterexample search efficiency, entering state regularization (ESR) to enlarge feasible region, and activation pattern alignment (APA) to accelerate neural network verification. We also provide a neural safety certificate synthesis and verification benchmark called Cersyve-9, which includes nine commonly used safe control tasks and supplements existing neural network verification benchmarks. Our framework successfully synthesizes verified neural value functions on all tasks, and our proposed three techniques exhibit superior scalability and efficiency compared with existing methods. Hanjiang Hu, Tianhao Wei, Shengbo Eben Li, Changliu Liu |
J. Artif. Intell. Res. | 5 |
| 2025 | Implicit Safe Set Algorithm for Provably Safe Reinforcement LearningabstractDeep reinforcement learning (DRL) has demonstrated remarkable performance in many continuous control tasks. However, a significant obstacle to the real-world application of DRL is the lack of safety guarantees. Although DRL agents can satisfy system safety in expectation through reward shaping, designing agents to consistently meet hard constraints (e.g., safety specifications) at every time step remains a formidable challenge. In contrast, existing work in the field of safe control provides guarantees on persistent satisfaction of hard safety constraints. However, these methods require explicit analytical system dynamics models to synthesize safe control, which are typically inaccessible in DRL settings. In this paper, we present a model-free safe control algorithm, the implicit safe set algorithm, for synthesizing safeguards for DRL agents that ensure provable safety throughout training. The proposed algorithm synthesizes a safety index (barrier certificate) and a subsequent safe control law solely by querying a black-box dynamic function (e.g., a digital twin simulator). Moreover, we theoretically prove that the implicit safe set algorithm guarantees finite time convergence to the safe set and forward invariance for both continuous-time and discrete-time systems. We validate the proposed algorithm on the state-of-the-art Safety Gym benchmark, where it achieves zero safety violations while gaining 95% ± 9% cumulative reward compared to state-of-the-art safe DRL methods. Furthermore, the resulting algorithm scales well to high-dimensional systems with parallel computing. Weiye Zhao, Feihan Li, Tairan He, Changliu Liu |
J. Artif. Intell. Res. | 4 |
| 2025 | Certifying Robustness of Learning-Based Keypoint Detection and Pose Estimation MethodsabstractThis work addresses the certification of the local robustness of vision-based two-stage 6D object pose estimation. The two-stage method for object pose estimation achieves superior accuracy over the single-stage approach by first employing deep neural network-driven keypoint regression and then applying a Perspective-n-Point (PnP) technique. Despite advancements, the certification of these methods’ robustness, especially in safety-critical scenarios, remains scarce. This research aims to fill this gap with a focus on their local robustness on the system level—the capacity to maintain robust estimations amidst semantic input perturbations. The core idea is to transform the certification of local robustness into a process of neural network verification for classification tasks. The challenge is to develop model, input, and output specifications that align with off-the-shelf verification tools. To facilitate verification, we modify the keypoint detection model by substituting non-linear operations with those more amenable to the verification processes. Instead of merely injecting random noise into images, as is common, we employ a convex hull representation of images as input specifications to more accurately depict semantic perturbations. Furthermore, by conducting a sensitivity analysis, we propagate the robustness criteria from pose estimation to keypoint accuracy, and then formulating an optimal error threshold allocation problem that allows for the setting of a maximally permissible keypoint deviation thresholds. Viewing each pixel as an individual class, these thresholds result in linear, classification-akin output specifications. Under certain conditions, we demonstrate that the main components of our certification framework are both sound and complete, and validate its effects through extensive evaluations on realistic perturbations. To our knowledge, this is the first study to certify the robustness of large-scale, keypoint-based pose estimation given images in real-world scenarios. Xusheng Luo, Tianhao Wei, Ziwei Wang 0010, Luis Mattei-Mendez, Taylor Loper, Joshua Neighbor, Casidhe Hutchison, Changliu Liu |
ACM Trans. Cyber Phys. Syst. | 9 |
| 2025 | Learn Zero-Constraint-Violation Safe Policy in Model-Free Constrained Reinforcement LearningabstractWe focus on learning the zero-constraint-violation safe policy in model-free reinforcement learning (RL). Existing model-free RL studies mostly use the posterior penalty to penalize dangerous actions, which means they must experience the danger to learn from the danger. Therefore, they cannot learn a zero-violation safe policy even after convergence. To handle this problem, we leverage the safety-oriented energy functions to learn zero-constraint-violation safe policies and propose the safe set actor-critic (SSAC) algorithm. The energy function is designed to increase rapidly for potentially dangerous actions, locating the safe set on the action space. Therefore, we can identify the dangerous actions prior to taking them and achieve zero-constraint violation. Our major contributions are twofold. First, we use the data-driven methods to learn the energy function, which releases the requirement of known dynamics. Second, we formulate a constrained RL problem to solve the zero-violation policies. We prove that our Lagrangian-based constrained RL solutions converge to the constrained optimal zero-violation policies theoretically. The proposed algorithm is evaluated on the complex simulation environments and a hardware-in-loop (HIL) experiment with a real autonomous vehicle controller. Experimental results suggest that the converged policies in all environments achieve zero-constraint violation and comparable performance with model-based baseline. Haitong Ma, Changliu Liu, Shengbo Eben Li, Sifa Zheng, Jianyu Chen 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Simultaneous Task Allocation and Planning for Multirobots Under Hierarchical Temporal Logic SpecificationsabstractResearch in robotic planning with temporal logic specifications, such as linear temporal logic (LTL), has relied on single formulas. However, as task complexity increases, LTL formulas become lengthy, making them difficult to interpret and generate, and straining the computational capacities of planners. To address this, we introduce a hierarchical structure for a widely used specification type—LTL on finite traces (LTL$_{f}$). The resulting language, termed H-LTL$_{f}$, is defined with both its syntax and semantics. We further prove that H-LTL$_{f}$is more expressive than its standard “flat” counterparts. Moreover, we conducted a user study that compared the standard LTL$_{f}$with our hierarchical version and found that users could more easily comprehend complex tasks using the hierarchical structure. We develop a search-based approach to synthesize plans for multirobot systems, achieving simultaneous task allocation and planning. This method approximates the search space by loosely interconnected subspaces, each corresponding to an LTL$_{f}$specification. The search primarily focuses on a single subspace, transitioning to another under conditions determined by the decomposition of automata. We develop multiple heuristics to significantly expedite the search. Our theoretical analysis, conducted under mild assumptions, addresses completeness and optimality. Compared to existing methods used in various simulators for service tasks, our approach improves planning times while maintaining comparable solution quality. Xusheng Luo, Changliu Liu |
IEEE Trans. Robotics | 2 |
| 2024 | ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
Guanxing Lu, Ziwei Wang 0010, Changliu Liu, Jiwen Lu, Yansong Tang |
ECCV (35) | 4 |
| 2024 | Absolute Policy Optimization: Enhancing Lower Probability Bound of Performance with High ConfidenceabstractIn recent years, trust region on-policy reinforcement learning has achieved impressive results in addressing complex control tasks and gaming scenarios. However, contemporary state-of-the-art algorithms within this category primarily emphasize improvement in expected performance, lacking the ability to control over the worst-case performance outcomes. To address this limitation, we introduce a novel objective function, optimizing which leads to guaranteed monotonic improvement in the lower probability bound of performance with high confidence. Building upon this groundbreaking theoretical advancement, we further introduce a practical solution called Absolute Policy Optimization (APO). Our experiments demonstrate the effectiveness of our approach across challenging continuous control benchmark tasks and extend its applicability to mastering Atari games. Our findings reveal that APO as well as its efficient variation Proximal Absolute Policy Optimization (PAPO) significantly outperforms state-of-the-art policy gradient algorithms, resulting in substantial improvements in worst-case performance, as well as expected performance. Weiye Zhao, Feihan Li, Yifan Sun 0011, Rui Chen 0030, Tianhao Wei, Changliu Liu |
ICML | 6 |
| 2024 | Towards Proactive Safe Human-Robot Collaborations via Data-Efficient Conditional Behavior PredictionabstractWe focus on the problem of how we can enable a robot to collaborate seamlessly with a human partner, specifically in scenarios where preexisting data is sparse. Much prior work in human-robot collaboration uses observational models of humans (i.e. models that treat the robot purely as an observer) to choose the robot’s behavior, but such models do not account for the influence the robot has on the human’s actions, which may lead to inefficient interactions. We instead formulate the problem of optimally choosing a collaborative robot’s behavior based on a conditional model of the human that depends on the robot’s future behavior. First, we propose a novel model-based formulation of conditional behavior prediction that allows the robot to infer the human’s intentions based on its future plan in data-sparse environments. We then show how to utilize a conditional model for proactive goal selection and safe trajectory generation around human collaborators. Finally, we use our proposed proactive controller in a collaborative task with real users to show that it can improve users’ interactions with a robot collaborator quantitatively and qualitatively. Ravi Pandya, Yorie Nakahira, Changliu Liu |
ICRA | 4 |
| 2024 | Multi-Agent Strategy Explanations for Human-Robot CollaborationabstractAs robots are deployed in human spaces, it is important that they are able to coordinate their actions with the people around them. Part of such coordination involves ensuring that people have a good understanding of how a robot will act in the environment. This can be achieved through explanations of the robot’s policy. Much prior work in explainable AI and RL focuses on generating explanations for single-agent policies, but little has been explored in generating explanations for collaborative policies. In this work, we investigate how to generate multi-agent strategy explanations for human-robot collaboration. We formulate the problem using a generic multi-agent planner, show how to generate visual explanations through strategy-conditioned landmark states and generate textual explanations by giving the landmarks to an LLM. Through a user study, we find that when presented with explanations from our proposed framework, users are able to better explore the full space of strategies and collaborate more efficiently with new robot partners. Ravi Pandya, Michelle Zhao, Changliu Liu, Reid G. Simmons, Henny Admoni |
ICRA | 3 |
| 2024 | Optimizing Multi-Touch Textile and Tactile Skin Sensing Through Circuit Parameter EstimationabstractTactile and textile skin technologies have become increasingly important for enhancing human-robot interaction and allowing robots to adapt to different environments. Despite notable advancements, there are ongoing challenges in skin signal processing, particularly in achieving both accuracy and speed in dynamic touch sensing. This paper introduces a new framework that poses the touch sensing problem as an estimation problem of resistive sensory arrays. Utilizing a Regularized Least Squares objective function—which estimates the resistance distribution of the skin—we enhance the touch sensing accuracy and mitigate the ghosting effects, where false or misleading touches may be registered. Furthermore, our study presents a streamlined skin design that simplifies manufacturing processes without sacrificing performance. Experimental outcomes substantiate the effectiveness of our method, showing 26.9% improvement in multi-touch force-sensing accuracy for the tactile skin. Bo Ying Su, Chengtao Wen, Changliu Liu |
ICRA | 4 |
| 2024 | Learning Human-to-Humanoid Real-Time Whole-Body TeleoperationabstractWe present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of human movements for humanoid robots, we propose a scalable "sim-to-data" process to filter and pick feasible motions using a privileged motion imitator. Afterwards, we train a robust real-time humanoid motion imitator in simulation using these refined motions and transfer it to the real humanoid robot in a zero-shot manner. We successfully achieve teleoperation of dynamic whole-body motions in real-world scenarios, including walking, back jumping, kicking, turning, waving, pushing, boxing, etc. To the best of our knowledge, this is the first demonstration to achieve learning-based real-time whole-body humanoid teleoperation. Tairan He, Zhengyi Luo 0002, Kris Makoto Kitani, Changliu Liu, Guanya Shi |
IROS | 6 |
| 2024 | NN4SysBench: Characterizing Neural Network Verification for Computer SystemsabstractWe present NN4SysBench, a benchmark suite for neural network verification that is composed of applications from the domain of computer systems. We call these neural networks for computer systems or NN4Sys. NN4Sys is booming: there are many proposals for using neural networks in computer systems—for example, databases, OSes, and networked systems—many of which are safety critical. Neural network verification is a technique to formally verify whether neural networks satisfy safety properties. We however observe that NN4Sys has some unique characteristics that today’s verification tools overlook and have limited support. Therefore, this benchmark suite aims at bridging the gap between NN4Sys and the verification by using impactful NN4Sys applications as benchmarks to illustrate computer systems’ unique challenges. We also build a compatible version of NN4SysBench, so that today’s verifiers can also work on these benchmarks with approximately the same verification difficulties. The code is available at https://github.com/lydialin1212/NN4Sys_Benchmark. Shuyi Lin 0001, Tianhao Wei, Kaidi Xu, Gagandeep Singh 0001, Changliu Liu |
NeurIPS | 7 |
| 2023 | AutoCost: Evolving Intrinsic Cost for Zero-Violation Reinforcement LearningabstractSafety is a critical hurdle that limits the application of deep reinforcement learning to real-world control tasks. To this end, constrained reinforcement learning leverages cost functions to improve safety in constrained Markov decision process. However, constrained methods fail to achieve zero violation even when the cost limit is zero. This paper analyzes the reason for such failure, which suggests that a proper cost function plays an important role in constrained RL. Inspired by the analysis, we propose AutoCost, a simple yet effective framework that automatically searches for cost functions that help constrained RL to achieve zero-violation performance. We validate the proposed method and the searched cost function on the safety benchmark Safety Gym. We compare the performance of augmented agents that use our cost function to provide additive intrinsic costs to a Lagrangian-based policy learner and a constrained-optimization policy learner with baseline agents that use the same policy learners but with only extrinsic costs. Results show that the converged policies with intrinsic costs in all environments achieve zero constraint violation and comparable performance with baselines. Tairan He, Weiye Zhao, Changliu Liu |
AAAI | 3 |
| 2023 | Learning from Physical Human Feedback: An Object-Centric One-Shot Adaptation MethodabstractFor robots to be effectively deployed in novel environments and tasks, they must be able to understand the feedback expressed by humans during intervention. This can either correct undesirable behavior or indicate additional preferences. Existing methods either require repeated episodes of interactions or assume prior known reward features, which is data-inefficient and can hardly transfer to new tasks. We relax these assumptions by describing human tasks in terms of object-centric sub-tasks and interpreting physical interventions in relation to specific objects. Our method, Object Preference Adaptation (OPA), is composed of two key stages: 1) pre-training a base policy to produce a wide variety of behaviors, and 2) online-updating according to human feedback. The key to our fast, yet simple adaptation is that general interaction dynamics between agents and objects are fixed, and only object-specific preferences are updated. Our adaptation occurs online, requires only one human intervention (one-shot), and produces new behaviors never seen during training. Trained on cheap synthetic data instead of expensive human demonstrations, our policy correctly adapts to human perturbations on realistic tasks on a physical 7DOF robot. Videos, code, and supplementary material: https://alvinosaur.github.io/AboutMe/projects/opa. Alvin Shek, Bo Ying Su, Rui Chen 0030, Changliu Liu |
ICRA | 4 |
| 2023 | State-wise Safe Reinforcement Learning: A SurveyabstractDespite the tremendous success of Reinforcement Learning (RL) algorithms in simulation environments, applying RL to real-world applications still faces many challenges. A major concern is safety, in another word, constraint satisfaction. State-wise constraints are one of the most common constraints in real-world applications and one of the most challenging constraints in Safe RL. Enforcing state-wise constraints is necessary and essential to many challenging tasks such as autonomous driving, robot manipulation. This paper provides a comprehensive review of existing approaches that address state-wise constraints in RL. Under the framework of State-wise Constrained Markov Decision Process (SCMDP), we will discuss the connections, differences, and trade-offs of existing approaches in terms of (i) safety guarantee and scalability, (ii) safety and reward performance, and (iii) safety after convergence and during training. We also summarize limitations of current methods and discuss potential future directions. Weiye Zhao, Tairan He, Rui Chen 0030, Tianhao Wei, Changliu Liu |
IJCAI | 5 |
| 2023 | Space-Time Conflict Spheres for Constrained Multi-Agent Motion PlanningabstractMulti-agent motion planning (MAMP) is a critical challenge in applications such as connected autonomous vehicles and multi-robot systems. In this paper, we propose a space-time conflict resolution approach for MAMP. We formulate the problem using a novel, flexible sphere-based discretization for trajectories. Our approach leverages a depth-first conflict search strategy to provide the scalability of decoupled approaches while maintaining the computational guarantees of coupled approaches. We compose procedures for evading discretization error and adhering to kinematic constraints in generated solutions. Theoretically, we prove the continuous-time feasibility and formulation-space completeness of our algorithm. Experimentally, we demonstrate that our algorithm matches the performance of the current state of the art with respect to both runtime and solution quality, while expanding upon the abilities of current work through accommodation for both static and dynamic obstacles. We evaluate our algorithm in various unsignalized traffic intersection scenarios using CARLA, an open-source vehicle simulator. Results show significant success rate improvement in spatially constrained settings, involving both connected and non-connected vehicles. Furthermore, we maintain a reasonable suboptimality ratio that scales well among increasingly complex scenarios. Anirudh Chari, Rui Chen 0030, Changliu Liu |
IV | 3 |
| 2023 | Consensus-Based Fault-Tolerant Platooning for Connected and Autonomous VehiclesabstractPlatooning is a representative application of connected and autonomous vehicles. The information exchanged between connected functions and the precise control of autonomous functions provide great safety and traffic capacity. In this paper, we develop an advanced consensus-based approach for platooning. By applying consensus-based fault detection and adaptive gains to controllers, we can detect faulty position and speed information from vehicles and reinstate the normal behavior of the platooning. Experimental results demonstrate that the developed approach outperforms the state-of-the-art approaches and achieves small steady state errors and small settling times under scenarios with faults. Tzu-Yen Tseng, Ding-Jiun Huang, Jia-You Lin, Po-Jui Chang, Chung-Wei Lin, Changliu Liu |
IV | 6 |
| 2023 | First three years of the international verification of neural networks competition (VNN-COMP)abstractAbstract This paper presents a summary and meta-analysis of the first three iterations of the annual International Verification of Neural Networks Competition (VNN-COMP), held in 2020, 2021, and 2022. In the VNN-COMP, participants submit software tools that analyze whether given neural networks satisfy specifications describing their input-output behavior. These neural networks and specifications cover a variety of problem classes and tasks, corresponding to safety and robustness properties in image classification, neural control, reinforcement learning, and autonomous systems. We summarize the key processes, rules, and results, present trends observed over the last three years, and provide an outlook into possible future developments. Christopher Brix, Mark Niklas Müller, Stanley Bak, Taylor T. Johnson, Changliu Liu |
Int. J. Softw. Tools Technol. Transf. | 5 |
| 2023 | BioSLAM: A Bioinspired Lifelong Memory System for General Place RecognitionabstractWe present BioSLAM, a lifelong (lifelong simultaneous localization and mapping) SLAM framework for learning various new appearances incrementally and maintaining accurate place recognition for previously visited areas. Unlike humans, artificial neural networks suffer from catastrophic forgetting and may forget the previously visited areas when trained with new arrivals. For humans, researchers discover that there exists a memory replay mechanism in the brain to keep the neuron active for previous events. Inspired by this discovery, BioSLAM designs a gated generative replay to control the robot's learning behavior based on the feedback rewards. Specifically, BioSLAM provides a novel dual-memory mechanism for the maintenance of: 1) a dynamic memory to efficiently learn new observations; and 2) a static memory to balance new–old knowledge. When the agent is encountered with different appearances under new domains, the complete processing pipeline can help to incrementally update the place recognition ability, robust to the increasing complexity of long-term place recognition. We demonstrate BioSLAM in three incremental SLAM scenarios as follows. 1) A 120 km city-scale trajectories with LiDAR-based inputs. 2) A multivisited 4.5 km campus-scale trajectories with LiDAR-vision inputs. 3) An official Oxford dataset with 10 km visual inputs under different environmental conditions. We show that BioSLAM can incrementally update the agent's place recognition ability and outperform the state-of-the-art incremental approach, generative replay, by 24% in terms of place recognition accuracy. To the best of our knowledge, BioSLAM is the first memory-enhanced lifelong SLAM system to help incremental place recognition in long-term navigation tasks. Peng Yin 0001, Abulikemu Abuduweili, Changliu Liu, Sebastian A. Scherer |
IEEE Trans. Robotics | 5 |
| 2022 | A Composable Framework for Policy Design, Learning, and Transfer Toward Safe and Efficient Industrial InsertionabstractDelicate industrial insertion tasks (e.g., PC board assembly) remain challenging for industrial robots. The chal-lenges include low error tolerance, delicacy of the components, and large task variations with respect to the components to be inserted. To deliver a feasible robotic solution for these insertion tasks, we also need to account for hardware limits of existing robotic systems and minimize the integration effort. This paper proposes a composable framework for efficient integration of a safe insertion policy on existing robotic platforms to accomplish these insertion tasks. The policy has an interpretable modularized design and can be learned efficiently on hardware and transferred to new tasks easily. In particular, the policy includes a safe insertion agent as a baseline policy for insertion, an optimal configurable Cartesian tracker as an interface to robot hardware, a probabilistic inference module to handle component variety and insertion errors, and a safe learning module to optimize the parameters in the aforementioned modules to achieve the best performance on designated hard-ware. The experiment results on a URIO robot show that the proposed framework achieves safety (for the delicacy of components), accuracy (for low tolerance), robustness (against perception error and component defection), adaptability and transferability (for task variations), as well as task efficiency during execution plus data and time efficiency during learning. Rui Chen 0030, Tianhao Wei, Changliu Liu |
IROS | 4 |
| 2022 | Safe and Efficient Exploration of Human Models During Human-Robot InteractionabstractMany collaborative human-robot tasks require the robot to stay safe and work efficiently around humans. Since the robot can only stay safe with respect to its own model of the human, we want the robot to learn a good model of the human in order to act both safely and efficiently. This paper studies methods that enable a robot to safely explore the space of a human-robot system to improve the robot's model of the human, which will consequently allow the robot to access a larger state space and better work with the human. In particular, we introduce active exploration under the framework of energy-function based safe control, investigate the effect of different active exploration strategies, and finally analyze the effect of safe active exploration on both analytical and neural network human models. Ravi Pandya, Changliu Liu |
IROS | 2 |
| 2021 | Distributed Motion Coordination Using Convex Feasible Set Based Model Predictive ControlabstractThe implementation of optimization-based motion coordination approaches in real world multi-agent systems remains challenging due to their high computational complexity and potential deadlocks. This paper presents a distributed model predictive control (MPC) approach based on convex feasible set (CFS) algorithm for multi-vehicle motion coordination in autonomous driving. By using CFS to convexify the collision avoidance constraints, collision-free trajectories can be computed in real time. We analyze the potential deadlocks and show that a deadlock can be resolved by changing vehicles’ desired speeds. The MPC structure ensures that our algorithm is robust to low-level tracking errors. The proposed distributed method has been tested in multiple challenging multi-vehicle environments, including unstructured road, intersection, crossing, platoon formation, merging, and overtaking scenarios. The numerical results and comparison with other approaches (including a centralized MPC and reciprocal velocity obstacles) show that the proposed method is computationally efficient and robust, and avoids deadlocks. Changliu Liu |
ICRA | 2 |
| 2021 | Deadlock Analysis and Resolution for Multi-robot Systems
Jaskaran Grover, Changliu Liu, Katia P. Sycara |
WAFR | 2 |
| 2020 | A Dynamic Programming Approach to Optimal Lane Merging of Connected and Autonomous VehiclesabstractLane merging is one of the major sources causing traffic congestion and delay. With the help of vehicle-to-vehicle or vehicle-to-infrastructure communication and autonomous driving technology, there are opportunities to alleviate congestion and delay resulting from lane merging. In this paper, we first summarize modern features and requirements for lane merging, along with the advance of vehicular technology. We then formulate and propose a dynamic programming algorithm to find the optimal solution for a two-lane merging scenario. It schedules the passing order for vehicles while minimizing the time needed for all vehicles to go through the merging point (equivalent to the time that the last vehicle goes through the merging point). We further extend the problem to a consecutive lane-merging scenario. We show the difficulty to apply the original dynamic programming to the consecutive lane-merging scenario and propose an improved version to solve it. Experimental results show that our dynamic programming algorithm can efficiently minimize the time needed for all vehicles to go through the merging point and reduce the average delay of all vehicles, compared with some greedy methods. Shang-Chien Lin, Hsiang Hsu, Chung-Wei Lin, Iris Hui-Ru Jiang, Changliu Liu |
IV | 6 |
| 2019 | Simulating Emergent Properties of Human Driving Behavior Using Multi-Agent Reward Augmented Imitation LearningabstractRecent developments in multi-agent imitation learning have shown promising results for modeling the behavior of human drivers. However, it is challenging to capture emergent traffic behaviors that are observed in real-world datasets. Such behaviors arise due to the many local interactions between agents that are not commonly accounted for in imitation learning. This paper proposes Reward Augmented Imitation Learning (RAIL), which integrates reward augmentation into the multi-agent imitation learning framework and allows the designer to specify prior knowledge in a principled fashion. We prove that convergence guarantees for the imitation learning process are preserved under the application of reward augmentation. This method is validated in a driving scenario, where an entire traffic scene is controlled by driving policies learned using our proposed algorithm. Further, we demonstrate improved performance in comparison to traditional imitation learning algorithms both in terms of the local actions of a single agent and the behavior of emergent properties in complex, multi-agent settings. Raunak P. Bhattacharyya, Derek J. Phillips, Changliu Liu, Jayesh K. Gupta, Katherine Rose Driggs-Campbell, Mykel J. Kochenderfer |
ICRA | 3 |
| 2019 | AGen: Adaptable Generative Prediction Networks for Autonomous DrivingabstractIn highly interactive driving scenarios, accurate prediction of other road participants is critical for safe and efficient navigation of autonomous cars. Prediction is challenging due to the difficulty in modeling various driving behavior, or learning such a model. The model should be interactive and reflect individual differences. Imitation learning methods, such as parameter sharing generative adversarial imitation learning (PS-GAIL), are able to learn interactive models. However, the learned models average out individual differences. When used to predict trajectories of individual vehicles, these models are biased. This paper introduces an adaptable generative prediction framework (AGen), which performs online adaptation of the offline learned models to recover individual differences for better prediction. In particular, we combine the recursive least square parameter adaptation algorithm (RLS-PAA) with the offline learned model from PS-GAIL. RLS-PAA has analytical solutions and is able to adapt the model for every single vehicle efficiently online. The proposed method is able to reduce the root mean squared prediction error in a 2.5 s time window by 60%, compared with PS-GAIL. Wenwen Si, Tianhao Wei, Changliu Liu |
IV | 3 |
| 2019 | Toward Modularization of Neural Network Autonomous Driving Policy Using Parallel Attribute NetworksabstractNeural network autonomous driving policies are widely explored. However, no matter using imitation learning or reinforcement learning, the network policies are generally hard to train, and the learned knowledge encoded in neural network policies are hard to transfer. We propose to modularize the complicated driving policies in terms of the driving attributes, and present the parallel attribute networks (PAN), which can learn to fullfill the requirements of the attributes in the driving tasks separately, and later assemble their knowledge together. Concretely, we first train a policy network that accomplish the base lane tracking attribute. The modules for the add-on attributes such as avoiding obstacles and obeying traffic rules are then trained to map the corresponding state to a satisfactory set of the vehicle action space. Finally the reference action given by the base policy is projected into the satisfactory sets so as to satisfy the requirements of all the attributes. Using the PAN, many complicated tasks that are hard to train from scratch can be easily trained; also unseen driving tasks can be solved in a zero-shot manner by assembling the pretrained attribute modules. We have validated the capability of our model on a class of autonomous driving problems with attributes of obstacle avoidance, traffic light and speed limit in simulation. Experimental results based on an obstacle avoidance task are also presented. Haonan Chang, Chen Tang 0001, Changliu Liu, Masayoshi Tomizuka |
IV | 4 |
| 2019 | Graph-Based Modeling, Scheduling, and Verification for Intersection Management of Intelligent VehiclesabstractIntersection management is one of the most representative applications of intelligent vehicles with connected and autonomous functions. The connectivity provides environmental information that a single vehicle cannot sense, and the autonomy supports precise vehicular control that a human driver cannot achieve. Intersection management solves the fundamental conflict resolution problem for vehicles—two vehicles should not appear at the same location at the same time, and, if they intend to do that, an order should be decided to optimize certain objectives such as the traffic throughput or smoothness. In this paper, we first propose a graph-based model for intersection management. The model is general and applicable to different granularities of intersections and other conflicting scenarios. We then derive formal verification approaches which can guarantee deadlock-freeness. Based on the graph-based model and the verification approaches, we develop a centralized cycle removal algorithm for the graph-based model to schedule vehicles to go through the intersection safely (without collisions) and efficiently without deadlocks. Experimental results demonstrate the expressiveness of the proposed model and the effectiveness and efficiency of the proposed algorithm. Hsiang Hsu, Shang-Chien Lin, Chung-Wei Lin, Iris Hui-Ru Jiang, Changliu Liu |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2018 | Fast Robot Motion Planning with Collision Avoidance and Temporal OptimizationabstractConsidering the growing demand of real-time motion planning in robot applications, this paper proposes a fast robot motion planner (FRMP) to plan collision-free and time-optimal trajectories, which applies the convex feasible set algorithm (CFS) to solve both the trajectory planning problem and the temporal optimization problem. The performance of CFS in trajectory planning is compared to the sequential quadratic programming (SQP) in simulation, which shows a significant decrease in iteration numbers and computation time to converge a solution. The effectiveness of temporal optimization is shown on the operational time reduction in the experiment on FANUC LR Mate 200iD/7L. Hsien-Chung Lin, Changliu Liu, Masayoshi Tomizuka |
ICARCV | 2 |
| 2017 | Boundary layer heuristic for search-based nonholonomic path planning in maze-like environmentsabstractAutomatic valet parking is widely viewed as a milestone towards fully autonomous driving. One of the key problems is nonholonomic path planning in maze-like environments (e.g. parking lots). To balance efficiency and passenger comfort, the planner needs to minimize the length of the path as well as the number of gear shifts. Lattice A* search is widely adopted for optimal path planning. However, existing heuristics do not evaluate the nonholonomic dynamic constraint and the collision avoidance constraint simultaneously, which may mislead the search. To efficiently search the environment, the boundary layer heuristic is proposed which puts large cost in the area that the vehicle must shift gear to escape. Such area is called the boundary layer. A simple and efficient geometric method to compute the boundary layer is proposed. The admissibility and consistency of the additive combination of the boundary layer heuristic and existing heuristics are proved in the paper. The simulation results verify that the introduction of the boundary layer heuristic improves the search performance by reducing the computation time by 56.1%. Changliu Liu, Yizhou Wang 0003, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 1 |
| 2017 | Speed profile planning in dynamic environments via temporal optimizationabstractTo generate safe and efficient trajectories for an automated vehicle in dynamic environments, a layered approach is usually considered, which separates path planning and speed profile planning. This paper is focused on speed profile planning for a given path that is represented by a set of waypoints. The speed profile will be generated using temporal optimization which optimizes the time stamps for all waypoints along the given path. The formulation of the problem under urban driving scenarios is discussed. To speed up the computation, the non-convex temporal optimization is approximated by a set of quadratic programs which are solved iteratively using the slack convex feasible set (SCFS) algorithm. The simulations in various urban driving scenarios validate the effectiveness of the method. Changliu Liu, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 1 |
| 2017 | Spatially-partitioned environmental representation and planning architecture for on-road autonomous drivingabstractConventional layered planning architecture temporally partitions the spatiotemporal motion planning by the path and speed, which is not suitable for lane change and overtaking scenarios with moving obstacles. In this paper, we propose to spatially partition the motion planning by longitudinal and lateral motions along the rough reference path in the Frenét Frame, which makes it possible to create linearized safety constraints for each layer in a variety of on-road driving scenarios. A generic environmental representation methodology is proposed with three topological elements and corresponding longitudinal constraints to compose all driving scenarios mentioned in this paper according to the overlap between the potential path of the autonomous vehicle and predicted path of other road users. Planners combining A* search and quadratic programming (QP) are designed to plan both rough long-term longitudinal motions and short-term trajectories to exploit the advantages of both search-based and optimization-based methods. Limits of vehicle kinematics and dynamics are considered in the planners to handle extreme cases. Simulation results show that the proposed framework can plan collision-free motions with high driving quality under complicated scenarios and emergency situations. Jianyu Chen 0002, Ching-Yao Chan, Changliu Liu, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 4 |
| 2016 | Algorithmic safety measures for intelligent industrial co-robotsabstractIn factories of the future, humans and robots are expected to be co-workers and co-inhabitants in the flexible production lines. It is important to ensure that humans and robots do not harm each other. This paper is concerned with functional issues to ensure safe and efficient interactions among human workers and the next generation intelligent industrial co-robots. The robot motion planning and control problem in a human involved environment is posed as a constrained optimal control problem. A modularized parallel controller structure is proposed to solve the problem online, which includes a baseline controller that ensures efficiency, and a safety controller that addresses real time safety by making a safe set invariant. Capsules are used to represent the complicated geometry of humans and robots. The design considerations of each module are discussed. Simulation studies which reproduce realistic scenarios are performed on a planar robot arm and a 6 DoF robot arm. The simulation results confirm the effectiveness of the method. Changliu Liu, Masayoshi Tomizuka |
ICRA | 1 |
| 2016 | Robotic manipulation of deformable objects by tangent space mapping and non-rigid registrationabstractRecent works of non-rigid registration have shown promising applications on tasks of deformable manipulation. Those approaches use thin plate spline-robust point matching (TPS-RPM) algorithm to regress a transformation function, which could generate a corresponding manipulation trajectory given a new pose/shape of the object. However, this method regards the object as a bunch of discrete and independent points. Structural information, such as shape and length, is lost during the transformation. This limitation makes the object's final shape to differ from training to test, and can sometimes cause damage to the object because of excessive stretching. To deal with these problems, this paper introduces a tangent space mapping (TSM) algorithm, which maps the deformable object in the tangent space instead of the Cartesian space to maintain structural information. The new algorithm is shown to be robust to the changes in the object's pose/shape, and the object's final shape is similar to that of training. It is also guaranteed not to overstretch the object during manipulation. A series of rope manipulation tests are performed to validate the effectiveness of the proposed algorithm. Te Tang, Changliu Liu, Masayoshi Tomizuka |
IROS | 2 |