VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0013
dblp:10/4661-13
· DBLP profile ↗
44ranked-venue papers
3as first author
28since 2021 · last 2024
0000-0002-7511-2870ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 26 since 2021Systems, architecture and hardware · 28 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Theory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose RefinementabstractObject pose refinement is essential for robust object pose estimation. Previous work has made significant progress to-wards instance-level object pose refinement. Yet, category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrep-ancies between the target object and the shape prior. To address these challenges, we introduce a novel architecture for category-level object pose refinement. Our approach in-tegrates an HS-Iayer and learnable affine transformations, which aims to enhance the extraction and alignment of Geometric information. Additionally, we introduce a cross-cloud transformation mechanism that efficiently merges di-verse data sources. Finally, we push the limits of our model by incorporating the shape prior information for translation and size error prediction. We conducted extensive ex-periments to demonstrate the effectiveness of the proposed framework. Through extensive quantitative experiments, we demonstrate significant improvement over the baseline method by a large margin across all metrics.11Project page: https://lynne-zheng-linfang.github.io/georef.github.io Linfang Zheng, Tze Ho Elden Tse, Chen Wang 0123, Yinghan Sun, Hua Chen 0007, Ales Leonardis, Wei Zhang 0013, Hyung Jin Chang |
CVPR | 7 |
| 2024 | Data-Driven Latent Space Representation for Robust Bipedal Locomotion LearningabstractThis paper presents a novel framework for learning robust bipedal walking by combining a data-driven state representation with a Reinforcement Learning (RL) based locomotion policy. The framework utilizes an autoencoder to learn a low-dimensional latent space that captures the complex dynamics of bipedal locomotion from existing locomotion data. This reduced dimensional state representation is then used as states for training a robust RL-based gait policy, eliminating the need for heuristic state selections or the use of template models for gait planning. The results demonstrate that the learned latent variables are disentangled and directly correspond to different gaits or speeds, such as moving forward, backward, or walking in place. Compared to traditional template model-based approaches, our framework exhibits superior performance and robustness in simulation. The trained policy effectively tracks a wide range of walking speeds and demonstrates good generalization capabilities to unseen scenarios. Guillermo A. Castillo, Bowen Weng, Wei Zhang 0013, Ayonga Hereid |
ICRA | 3 |
| 2024 | Multi-Resolution Planar Region Extraction for Uneven TerrainsabstractThis paper studies the problem of extracting planar regions in uneven terrains from unordered point cloud measurements. Such a problem is critical in various robotic applications such as robotic perceptive locomotion. While existing approaches have shown promising results in effectively extracting planar regions from the environment, they often suffer from issues such as low computational efficiency or loss of resolution. To address these issues, we propose a multi-resolution planar region extraction strategy in this paper that balances the accuracy in boundaries and computational efficiency. Our method begins with a pointwise classification preprocessing module, which categorizes all sampled points according to their local geometric properties to facilitate multi-resolution segmentation. Subsequently, we arrange the categorized points using an octree, followed by an in-depth analysis of nodes to finish multi-resolution plane segmentation. The efficiency and robustness of the proposed approach are verified via synthetic and real-world experiments, demonstrating our method’s ability to generalize effectively across various uneven terrains while maintaining real-time performance, achieving frame rates exceeding 35 FPS. Yinghan Sun, Linfang Zheng, Hua Chen 0007, Wei Zhang 0013 |
ICRA | 4 |
| 2024 | Task-Space Riccati Feedback based Whole Body Control for Underactuated Legged LocomotionabstractThis manuscript primarily aims to enhance the performance of whole-body controllers(WBC) for underactuated legged locomotion. We introduce a systematic parameter design mechanism for the floating-base feedback control within the WBC. The proposed approach involves utilizing the linearized model of unactuated dynamics to formulate a Linear Quadratic Regulator(LQR) and solving a Riccati gain while accounting for potential physical constraints through a second-order approximation of the log-barrier function. And then the user-tuned feedback gain for the floating base task is replaced by a new one constructed from the solved Riccati gain. Extensive simulations conducted in MuJoCo with a point bipedal robot, as well as real-world experiments performed on a quadruped robot, demonstrate the effectiveness of the proposed method. In the different bipedal locomotion tasks, compared with the user-tuned method, the proposed approach is at least 12% better and up to 50% better at linear velocity tracking, and at least 7% better and up to 47% better at angular velocity tracking. In the quadruped experiment, linear velocity tracking is improved by at least 3% and angular velocity tracking is improved by at least 23% using the proposed method. Shunpeng Yang, Zejun Hong, Patrick M. Wensing, Wei Zhang 0013, Hua Chen 0007 |
IROS | 5 |
| 2023 | HS-Pose: Hybrid Scope Feature Extraction for Category-level Object Pose EstimationabstractIn this paper, we focus on the problem of category-level object pose estimation, which is challenging due to the large intra-category shape variation. 3D graph convolution (3D-GC) based methods have been widely used to extract local geometric features, but they have limitations for complex shaped objects and are sensitive to noise. Moreover, the scale and translation invariant properties of 3D-GC restrict the perception of an object's size and translation information. In this paper, we propose a simple network structure, the HS-layer, which extends 3D-GC to extract hybrid scope latent features from point cloud data for category-level object pose estimation tasks. The proposed HS-layer: 1) is able to perceive local-global geometric structure and global information, 2) is robust to noise, and 3) can encode size and translation information. Our experiments show that the simple replacement of the 3D-GC layer with the proposed HS-layer on the baseline method (GPV-Pose) achieves a significant improvement, with the performance increased by 14.5% on 5°2cm metric and 10.3% on IoU75. Our method outperforms the state-of-the-art methods by a large margin (8.3% on 5°2cm, 6.9% on IoU75) on REAL275 dataset and runs in real-time (50 FPS)11Codeisavailable: https://github.com/Lynne-Zheng-Linfang/HS-Pose. Linfang Zheng, Chen Wang 0123, Yinghan Sun, Esha Dasgupta, Hua Chen 0007, Ales Leonardis, Wei Zhang 0013, Hyung Jin Chang |
CVPR | 7 |
| 2023 | Vision-based Six-Dimensional Peg-in-Hole for Practical Connector InsertionabstractWe study six-dimensional (6D) perceptive peg-in-hole problem for practical connector insertion task in this paper. To enable the manipulator system to handle different types of pegs in complex environment, we develop a perceptive robotic assembly system that utilizes an in-hand RGB-D camera for peg-in-hole with multiple types of pegs. The proposed framework addresses the critical hole detection and pose estimation problem through combining the learning-based detection with model-based pose estimation strategies. By exploiting the structure of the peg-in-hole task, we consider a rectangle-shape based characterization for modeling the candidate socket. Such a characterization allows us to design simple learning-based methods to detect and estimate the 6D pose of the target socket that balances between processing speed and accuracy. To validate our method, we test the performance of the proposed perceptive peg-in-hole solution using a KUKA iiwa7 robotic arm to accomplish the socket insertion task with two types of practical sockets (RJ45/HDMI). Without the need of additional search, our method achieves an acceptable success rate in the connector insertion tasks. The results confirm the reliability of our method and show that our method is suitable for real world application. Kun Zhang 0017, Chen Wang 0123, Hua Chen 0007, Jia Pan 0001, Michael Yu Wang, Wei Zhang 0013 |
ICRA | 6 |
| 2023 | Template Model Inspired Task Space Learning for Robust Bipedal LocomotionabstractThis work presents a hierarchical framework for bipedal locomotion that combines a Reinforcement Learning (RL)-based high-level (HL) planner policy for the online generation of task space commands with a model-based low-level (LL) controller to track the desired task space trajectories. Different from traditional end-to-end learning approaches, our HL policy takes insights from the angular momentum-based linear inverted pendulum (ALIP) to carefully design the observation and action spaces of the Markov Decision Process (MDP). This simple yet effective design creates an insightful mapping between a low-dimensional state that effectively captures the complex dynamics of bipedal locomotion and a set of task space outputs that shape the walking gait of the robot. The HL policy is agnostic to the task space LL controller, which increases the flexibility of the design and generalization of the framework to other bipedal robots. This hierarchical design results in a learning-based framework with improved performance, data efficiency, and robustness compared with the ALIP model-based approach and state-of-the-art learning-based frameworks for bipedal locomotion. The proposed hierarchical controller is tested in three different robots, Rabbit, a five-link underactuated planar biped; Walker2D, a seven-link fully-actuated planar biped; and Digit, a 3D humanoid robot with 20 actuated joints. The trained policy naturally learns human-like locomotion behaviors and is able to effectively track a wide range of walking speeds while preserving the robustness and stability of the walking gait even under adversarial conditions. Guillermo A. Castillo, Bowen Weng, Shunpeng Yang, Wei Zhang 0013, Ayonga Hereid |
IROS | 4 |
| 2023 | POMDP-Guided Active Force-Based Search for Robotic InsertionabstractIn robotic insertion tasks where the uncertainty exceeds the allowable tolerance, a good search strategy is essential for successful insertion and significantly influences efficiency. The commonly used blind search method is time-consuming and does not exploit the rich contact information. In this paper, we propose a novel search strategy that actively utilizes the information contained in the contact configuration and shows high efficiency. In particular, we formulate this problem as a Partially Observable Markov Decision Process (POMDP) with carefully designed primitives based on an in-depth analysis of the contact configuration's static stability. From the formulated POMDP, we can derive a novel search strategy. Thanks to its simplicity, this search strategy can be incorporated into a Finite-State-Machine (FSM) controller. The behaviors of the FSM controller are realized through a low-level Cartesian Impedance Controller. Our method is based purely on the robot's proprioceptive sensing and does not need visual or tactile sensors. To evaluate the effectiveness of our proposed strategy and control framework, we conduct extensive comparison experiments in simulation, where we compare our method with the baseline approach. The results demonstrate that our proposed method achieves a higher success rate with a shorter search time and search trajectory length compared to the baseline method. Additionally, we show that our method is robust to various initial displacement errors. Chen Wang 0123, Haoxiang Luo, Kun Zhang 0017, Hua Chen 0007, Jia Pan 0001, Wei Zhang 0013 |
IROS | 6 |
| 2023 | Quadruped Capturability and Push Recovery via a Switched-Systems Characterization of Dynamic BalanceabstractThis article studies capturability and push recovery for quadruped locomotion. Despite the rich literature on capturability analysis and push recovery for legged robots, existing tools have been developed mainly with the requirement of reaching static or quasi-static balance following a push. In practice, this requirement commonly restricts capturability analysis to cases with simple dynamics and fails to encode the time dependence of capturable states for legged locomotion with time-based gaits. To address these issues, we apply switched systems to model quadruped locomotion and extend capturability notions through a novel specification ofdynamic balance. We also provide an explicit model predictive control (EMPC) scheme to compute the dynamic balance and capturable tubes and offer a way of using the capturable tube to synthesize push recovery controllers. Such a generalization allows for a rigorous characterization of disturbance timing on the capturability of quadrupedal locomotion and opens the door of disturbance-timing-aware push recovery control strategies. Extensive simulation and hardware experiments illustrate the necessity of considering dynamic balance for quadrupedal push recovery, reveal how disturbance timing affects capturability, and demonstrate the significant improvement in disturbance rejection with the proposed strategy. Hardware experimental validations on a replica of the Mini Cheetah quadruped further verify that the proposed approach performs statistically better than the state-of-the-art baseline considered. Hua Chen 0007, Zejun Hong, Shunpeng Yang, Patrick M. Wensing, Wei Zhang 0013 |
IEEE Trans. Robotics | 5 |
| 2023 | On the Comparability and Optimal Aggressiveness of the Adversarial Scenario-Based Safety Testing of RobotsabstractThis article studies the class of scenario-based safety testing algorithms in the black-box safety testing configuration. For algorithms sharing the same state–action set coverage with different sampling distributions, it is commonly believed that prioritizing the exploration of high-risk states and actions leads to a better sampling efficiency. Our proposal disputes the above intuition by introducing an impossibility theorem that provably shows that all the safety testing algorithms of the aforementioned difference perform equally well with the same expected sampling efficiency. Moreover, for testing algorithms covering different sets of states and actions, the sampling efficiency criterion is no longer applicable as different algorithms do not necessarily converge to the same termination condition. We then propose a testing aggressiveness definition based on the almost safe set concept along with an unbiased and efficient algorithm that compares the aggressiveness between testing algorithms. Empirical observations from the safety testing of bipedal locomotion controllers and vehicle decision-making modules are also presented to support the proposed theoretical implications and methodologies. Bowen Weng, Guillermo A. Castillo, Wei Zhang 0013, Ayonga Hereid |
IEEE Trans. Robotics | 3 |
| 2022 | DynamicFilter: an Online Dynamic Objects Removal Framework for Highly Dynamic EnvironmentsabstractEmergence of massive dynamic objects will diversify spatial structures when robots navigate in urban environments. Therefore, the online removal of dynamic objects is critical. In this paper, we introduce a novel online removal framework for highly dynamic urban environments. The framework consists of the scan-to-map front-end and the map-to-map back-end modules. Both the front- and back-ends deeply integrate the visibility-based approach and map-based approach. The experiments validate the framework in highly dynamic simulation scenarios and real-world dataset. Tingxiang Fan, Bowen Shen, Hua Chen 0007, Wei Zhang 0013, Jia Pan 0001 |
ICRA | 4 |
| 2022 | On the Convergence of Multi-robot Constrained Navigation: A Parametric Control Lyapunov Function ApproachabstractThis paper studies the distributed multi-robot constrained navigation problem. While the multi-robot collision avoidance has been extensively studied in the literature with safety being the primary focus, the individual robot's destination convergence is not necessarily guaranteed. In particular, robots may get stuck in the local equilibria or periodic orbits of the multi-robot system, some of which are practically known as the deadlock and the livelock behaviors. Inspired by the combination of Control Lyapunov Function (CLF) and Control Barrier Function (CBF) for the nonlinear system's constrained stabilization, the authors present a guaranteed safe feedback control policy with improved convergence performance. The proposed Parametric CLF (PCLF) scheme adaptively determines the appropriate CLF parameterization within the in-stantaneous feasible action space. The algorithm also induces a conditional global asymptotic convergence guarantee for multi-robot system of single-integrator dynamics, and is empirically effective for nonlinear nonholonomic vehicle model. Empiri-cally, the proposed PCLF-CBF framework exhibits superior performance than state-of-the-art methods, including its de-generated counterpart of various CLF-CBF solutions. Bowen Weng, Hua Chen 0007, Wei Zhang 0013 |
ICRA | 3 |
| 2022 | TP-AE: Temporally Primed 6D Object Pose Tracking with Auto-EncodersabstractFast and accurate tracking of an object's motion is one of the key functionalities of a robotic system for achieving reliable interaction with the environment. This paper focuses on the instance-level six-dimensional (6D) pose tracking problem with a symmetric and textureless object under occlusion. We propose a Temporally Primed 6D pose tracking framework with Auto-Encoders (TP-AE) to tackle the pose tracking problem. The framework consists of a prediction step and a temporally primed pose estimation step. The prediction step aims to quickly and efficiently generate a guess on the object's real-time pose based on historical information about the target object's motion. Once the prior prediction is obtained, the temporally primed pose estimation step embeds the prior pose into the RGB-D input, and leverages auto-encoders to reconstruct the target object with higher quality under occlusion, thus improving the framework's performance. Extensive experiments show that the proposed 6D pose tracking method can accurately estimate the 6D pose of a symmetric and textureless object under occlusion, and significantly outperforms the state-of-the-art on T-LESS dataset while running in real-time at 26 FPS. Linfang Zheng, Ales Leonardis, Tze Ho Elden Tse, Nora Horanyi, Hua Chen 0007, Wei Zhang 0013, Hyung Jin Chang |
ICRA | 6 |
| 2022 | Three-Dimensional Dynamic Running with a Point-Foot Biped based on Differentially Flat SLIPabstractThis paper presents a novel framework for point- foot biped running in three-dimensional space. The proposed approach generates center of mass (CoM) reference trajectories based on a differentially flat spring-loaded inverted pendulum (SLIP) model. A foothold planner is used to select touch down location that renders optimal CoM trajectory for upcoming step in real time. Dynamically feasible trajectories of CoM and orientation are subsequently generated by a simplified single rigid body (SRB) model based model predictive control (MPC). A task-space controller is then applied online to compute whole- body joint torques which embeds these target dynamics into the robot. The proposed approach is evaluated on physical simulation of a 12 degree-of-freedom (DoF), 7.95 kg point-foot bipedal robot. The robot achieves stable running at at varying speeds with maximum value of 1.1 m/s. The proposed scheme is shown to be able to reject vertical disturbances of 8 N. s and lateral disturbance of 6.5 N. s applied at the robot base. Zejun Hong, Hua Chen 0007, Wei Zhang 0013 |
IROS | 3 |
| 2022 | On Safety Testing, Validation, and Characterization with Scenario-Sampling: A Case Study of Legged RobotsabstractThe dynamic response of the legged robot locomotion is non-Lipschitz and can be stochastic due to environmental uncertainties. To test, validate, and characterize the safety performance of legged robots, existing solutions on observed and inferred risk can be incomplete and sampling inefficient. Some formal verification methods suffer from the model precision and other surrogate assumptions. In this paper, we propose a scenario sampling based testing framework that characterizes the overall safety performance of a legged robot by specifying (i) where (in terms of a set of states) the robot is potentially safe, and (ii) how safe the robot is within the specified set. The framework can also help certify the commercial deployment of the legged robot in real-world environment along with human and compare safety performance among legged robots with different mechanical structures and dynamic properties. The proposed framework is further deployed to evaluate a group of state-of-the-art legged robot locomotion controllers from various model-based, deep neural network involved, and reinforcement learning based methods in the literature. Among a series of intended work domains of the studied legged robots (e.g. tracking speed on sloped surface, with abrupt changes on demanded velocity, and against adversarial push-over disturbances), we show that the method can adequately capture the overall safety characterization and the subtle performance insights. Many of the observed safety outcomes, to the best of our knowledge, have never been reported by the existing work in the legged robot literature. Bowen Weng, Guillermo A. Castillo, Wei Zhang 0013, Ayonga Hereid |
IROS | 3 |
| 2022 | Polytopic Planar Region Characterization of Rough Terrains for Legged LocomotionabstractThis paper studies the problem of constructing polytopic representations for planar regions from depth camera readings. This problem is of great importance for terrain mapping in complicated environment and has great potentials in legged locomotion applications. To address the polytopic planar region characterization problem, we propose a two-stage solution scheme. At the first stage, the planar regions embedded within a sequence of depth images are extracted individually and then merged to establish a terrain map containing only planar regions in a selected frame. To simplify the representations of the planar regions that are applicable to foothold planning for legged robots, we further approximate the extracted planar regions via convex polytopes at the second stage. With the polytopic representation, the proposed approach achieves a great balance between accuracy and simplicity. Experimental validations with RGB-D cameras are conducted to demonstrate the performance of the proposed scheme. The proposed scheme successfully characterizes the planar regions via polytopes with acceptable accuracy. More importantly, the run time of the overall scheme is less than 10ms (i.e., > 100Hz) throughout the tests, which strongly illustrates the advantages of our approach developed in this paper. Hua Chen 0007, Wei Zhang 0013 |
IROS | 4 |
| 2022 | Improved Task Space Locomotion Controller for a Quadruped Robot with Parallel MechanismsabstractIn this work, an advanced quadruped robot with abundant kinematic loops and passive joints is introduced. Due to the existence of many closed chains, the robot dynamic model is quite complex, and is derived using the Gauss's principle of least constraint. To explicitly consider the loop-closure constraints, we propose a task-space inverse dynamics based approach to obtain the robot locomotion controller. Besides, to meet the demand of high frequency (≥ 500Hz) in controller, an alternative method is provided. It uses the projected dynamics to find an analytical mapping from the desired contact force to the desired torque of actuators under full consideration of passive joints and loop-closure constraints. The effectiveness and efficiency of the proposed algorithms in this paper have been validated by simulation with a reliable physical engine MuJoCo. Shunpeng Yang, Wenchun Lin, Jaeho Noh, Bill Huang, Wei Zhang 0013, Hua Chen 0007 |
IROS | 6 |
| 2022 | Deterministic policy gradient: Convergence analysisabstractThe deterministic policy gradient (DPG) method proposed in Silver et al. [2014] has been demonstrated to exhibit superior performance particularly for applications with multi-dimensional and continuous action spaces. However, it remains unclear whether DPG converges, and if so, how fast it converges and whether it converges as efficiently as other PG methods. In this paper, we provide a theoretical analysis of DPG to answer those questions. We study the single timescale DPG (often the case in practice) in both on-policy and off-policy settings, and show that both algorithms attain an $\epsilon$-accurate stationary policy with a sample complexity of $\mathcal{O}(\epsilon^{-2})$. Moreover, we establish the convergence rate for DPG under Gaussian noise exploration, which is widely adopted in practice to improve the performance of DPG. To our best knowledge, this is the first non-asymptotic convergence characterization for DPG methods. Huaqing Xiong, Tengyu Xu, Lin Zhao 0009, Yingbin Liang, Wei Zhang 0013 |
UAI | 5 |
| 2021 | Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian SamplingabstractDespite the wide applications of Adam in reinforcement learning (RL), the theoretical convergence of Adam-type RL algorithms has not been established. This paper provides the first such convergence analysis for two fundamental RL algorithms of policy gradient (PG) and temporal difference (TD) learning that incorporate AMSGrad updates (a standard alternative of Adam in theoretical analysis), referred to as PG-AMSGrad and TD-AMSGrad, respectively. Moreover, our analysis focuses on Markovian sampling for both algorithms. We show that under general nonlinear function approximation, PG-AMSGrad with a constant stepsize converges to a neighborhood of a stationary point at the rate of O(1/T) (where T denotes the number of iterations), and with a diminishing stepsize converges exactly to a stationary point at the rate of O(log^2 T/√T). Furthermore, under linear function approximation, TD-AMSGrad with a constant stepsize converges to a neighborhood of the global optimum at the rate of O(1/T), and with a diminishing stepsize converges exactly to the global optimum at the rate of O(log T/√T). Our study develops new techniques for analyzing the Adam-type RL algorithms under Markovian sampling. Huaqing Xiong, Tengyu Xu, Yingbin Liang, Wei Zhang 0013 |
AAAI | 4 |
| 2021 | Belief Space Partitioning for Symbolic Motion PlanningabstractWe propose a memory-constrained partition-based method to extract symbolic representations of the belief state and its dynamics in order to solve planning problems in a partially observable Markov decision process (POMDP). Our K-means partitioning strategy uses a fixed number of symbols to represent the partitions of the belief space and ensures the parameterization of the belief dynamics does not grow exponentially as the system dimension increases. By casting our problem as a partitioning of the POMDP, we can then solve planning problems using traditional symbolic planning solvers (such as HTN or A* solvers). Our work is motivated by an autonomous underwater vehicle navigation problem where the vehicle is affected by uncertain flow conditions and receives severely limited position observations. Simulation experiments are provided to validate the performance of the proposed algorithms. Mengxue Hou, Tony X. Lin, Haomin Zhou 0001, Wei Zhang 0013, Catherine R. Edwards, Fumin Zhang 0001 |
ICRA | 4 |
| 2021 | Design of a deployable underwater robot for the recovery of autonomous underwater vehicles based on origami technique
Jisen Li, Yuliang Yang, Yongqi Li 0003, Qiujun Huang, Haibo Lu, Shengquan Li 0001, Wei Zhang 0013, Tao Mei 0001, Feng Wu 0001, Aidong Zhang 0002 |
ICRA | 10 |
| 2021 | Reachability-based Push Recovery for Humanoid Robots with Variable-Height Inverted PendulumabstractThis paper studies push recovery for humanoid robots based on a variable-height inverted pendulum (VHIP) model. We first develop an approach for treating zero-step capturability of the VHIP with a novel methodology based on Hamilton-Jacobi (HJ) reachability analysis. Such an approach uses the sub-zero level set of a value function to encode capturability of the VHIP, where the value function is obtained by numerically solving a HJ variational inequality offline. Based on this analysis, a simple and effective method for adjusting foothold locations is then devised for cases where the VHIP state is not zero-step capturable. In addition, the HJ reachability analysis naturally induces an optimal control law that allows for rapid planning with the VHIP during push recovery online. To enable use of the strategy with a position-controlled humanoid robot, an associated differential inverse kinematics based tracking controller is employed. The effectiveness of the overall framework is demonstrated with the UBTECH Walker robot in the MuJoCo simulator. Simulation validations show a significant improvement in push robustness as compared to the methods based on the classical linear inverted pendulum model. Shunpeng Yang, Hua Chen 0007, Zhefeng Cao, Patrick M. Wensing, Yizhang Liu, Jianxin Pang, Wei Zhang 0013 |
ICRA | 8 |
| 2021 | Robust Feedback Motion Policy Design Using Reinforcement Learning on a 3D Digit Bipedal RobotabstractIn this paper, a hierarchical and robust framework for learning bipedal locomotion is presented and successfully implemented on the 3D biped robot Digit built by Agility Robotics. We propose a cascade-structure controller that combines the learning process with intuitive feedback regulations. This design allows the framework to realize robust and stable walking with a reduced-dimensional state and action spaces of the policy, significantly simplifying the design and increasing the sampling efficiency of the learning method. The inclusion of feedback regulation into the framework improves the robustness of the learned walking gait and ensures the success of the sim-to-real transfer of the proposed controller with minimal tuning. We specifically present a learning pipeline that considers hardware-feasible initial poses of the robot within the learning process to ensure the initial state of the learning is replicated as close as possible to the initial state of the robot in hardware experiments. Finally, we demonstrate the feasibility of our method by successfully transferring the learned policy in simulation to the Digit robot hardware, realizing sustained walking gaits under external force disturbances and challenging terrains not incurred during the training process. To the best of our knowledge, this is the first time a learning-based policy is transferred successfully to the Digit robot in hardware experiments. Guillermo A. Castillo, Bowen Weng, Wei Zhang 0013, Ayonga Hereid |
IROS | 3 |
| 2021 | Quadruped Robot Hopping on Two LegsabstractThis paper presents a control strategy for quadruped robots to hop on their rear legs in three-dimensional space. The proposed approach generates nominal center of mass (CoM) trajectories based on a template spring-loaded inverted pendulum (SLIP) model. Tracking this reference remains a challenge due to the underactauted nature of balance with point feet. To address this challenge, a control-Lyapunov function based quadratic programming (CLF-QP) controller is proposed, which modulates nominal ground reaction forces (GRFs) to balance the torso while considering friction limits. The CLF construction is guided by a variational-based linearization (VBL) applied to a reduced-order single-rigid-body (SRB) model, and treats underactuation via solving a Riccati equation to obtain the CLF. A new balance control approach is presented that effectively decouples sagittal plane control (via re-planning) with lateral and rotational control (via the CLF and VBL). The proposed approach shows more robust balancing performance than the conventional CLF-QP approach. Simulations of the Mini Cheetah demonstrate in-place hopping with up to a 0.71m apex height. Shenggao Li 0001, Hua Chen 0007, Wei Zhang 0013, Patrick M. Wensing |
IROS | 3 |
| 2021 | Perceptive Autonomous Stair Climbing for Quadrupedal RobotsabstractThis paper studies autonomous stair climbing for quadrupedal robots with perception. Enabling quadrupeds to reliably climb staircases greatly expands their applicability in practical scenarios. For this structured task, we develop a simple yet effective perception and control framework for autonomous quadrupedal stair climbing. By exploiting the structural knowledge about the staircases, the proposed framework first extracts the geometric information about the staircase from measurements of the perception system. Then, the climbing velocity and associated foothold references during stair climbing are generated via simple optimization algorithms based on the geometric information about the staircase. Given these references, we use model predictive control based approach to generate input joint torques for controlling the quadruped to complete the whole stair climbing task. Simulation validations using the full dynamic model of the Unitree’s Aliengo quadruped with the MuJoCo simulator are performed, which demonstrate successful autonomous climbing of various staircases with different geometries. Effectiveness of the proposed strategy is further validated through hardware experiments on the real Aliengo robot with different real-world staircases. Shuhao Qi, Wenchun Lin, Zejun Hong, Hua Chen 0007, Wei Zhang 0013 |
IROS | 5 |
| 2021 | Encirclement Guaranteed Cooperative Pursuit with Robust Model Predictive ControlabstractThis paper studies a novel encirclement guaranteed cooperative pursuit problem involving N pursuers and a single evader in an unbounded two-dimensional game domain. Throughout the game, the pursuers are required to maintain encirclement of the evader, i.e., the evader should always stay inside the convex hull generated by all the pursuers, in addition to achieving the classical capture condition. To tackle this challenging cooperative pursuit problem, a robust model predictive control (RMPC) based formulation framework is first introduced, which simultaneously accounts for the encirclement and capture requirements under the assumption that the evader’s action is unavailable to all pursuers. Despite the reformulation, the resulting RMPC problem involves a bilinear constraint due to the encirclement requirement. To further handle such a bilinear constraint, a novel encirclement guaranteed partitioning scheme is devised that simplifies the original bilinear RMPC problem to a number of linear tube MPC (TMPC) problems solvable in a decentralized manner. Simulation experiments demonstrate the effectiveness of the proposed solution framework. Furthermore, comparisons with existing approaches show that the explicit consideration of the encirclement condition significantly improves the chance of successful capture of the evader in various scenarios. Chen Wang 0123, Hua Chen 0007, Jia Pan 0001, Wei Zhang 0013 |
IROS | 4 |
| 2021 | Force-feedback based Whole-body Stabilizer for Position-Controlled Humanoid RobotsabstractThis paper studies stabilizer design for position-controlled humanoid robots. Stabilizers are an essential part for position-controlled humanoids, whose primary objective is to adjust the control input sent to the robot to assist the tracking controller to better follow the planned reference trajectory. To achieve this goal, this paper develops a novel force-feedback based whole-body stabilizer that fully exploits the six-dimensional force measurement information and the whole-body dynamics to improve tracking performance. Relying on rigorous analysis of whole-body dynamics of position-controlled humanoids under unknown contact, the developed stabilizer leverages quadratic-programming based technique that allows cooperative consideration of both the center-of-mass tracking and contact force tracking. The effectiveness of the proposed stabilizer is demonstrated on the UBTECH Walker robot in the MuJoCo simulator. Simulation validations show a significant improvement in various scenarios as compared to commonly adopted stabilizers based on the zero-moment-point feedback and the linear inverted pendulum model. Shunpeng Yang, Hua Chen 0007, Zhen Fu, Wei Zhang 0013 |
IROS | 4 |
| 2021 | Finite-time theory for momentum Q-learningabstractExisting studies indicate that momentum ideas in conventional optimization can be used to improve the performance of Q-learning algorithms. However, the finite-time analysis for momentum-based Q-learning algorithms is only available for the tabular case without function approximation. This paper analyzes a class of momentum-based Q-learning algorithms with finite-time convergence guarantee. Specifically, we propose the MomentumQ algorithm, which integrates the Nesterov’s and Polyak’s momentum schemes, and generalizes the existing momentum-based Q-learning algorithms. For the infinite state-action space case, we establish the convergence guarantee for MomentumQ with linear function approximation under Markovian sampling. In particular, we characterize a finite-time convergence rate which is provably faster than the vanilla Q-learning. This is the first finite-time analysis for momentum-based Q-learning algorithms with function approximation. For the tabular case under synchronous sampling, we also obtain a finite-time convergence rate that is slightly better than the SpeedyQ (Azar et al., NIPS 2011). Finally, we demonstrate through various experiments that the proposed MomentumQ outperforms other momentum-based Q-learning algorithms. Bowen Weng, Huaqing Xiong, Lin Zhao 0009, Yingbin Liang, Wei Zhang 0013 |
UAI | 5 |
| 2020 | History-Gradient Aided Batch Size Adaptation for Variance Reduced AlgorithmsabstractVariance-reduced algorithms, although achieve great theoretical performance, can run slowly in practice due to the periodic gradient estimation with a large batch of data. Batch-size adaptation thus arises as a promising approach to accelerate such algorithms. However, existing schemes either apply prescribed batch-size adaption rule or exploit the information along optimization path via additional backtracking and condition verification steps. In this paper, we propose a novel scheme, which eliminates backtracking line search but still exploits the information along optimization path by adapting the batch size via history stochastic gradients. We further theoretically show that such a scheme substantially reduces the overall complexity for popular variance-reduced algorithms SVRG and SARAH/SPIDER for both conventional nonconvex optimization and reinforcement learning problems. To this end, we develop a new convergence analysis framework to handle the dependence of the batch size on history stochastic gradients. Extensive experiments validate the effectiveness of the proposed batch-size adaptation scheme. Kaiyi Ji, Zhe Wang 0021, Bowen Weng, Yi Zhou 0017, Wei Zhang 0013, Yingbin Liang |
ICML | 5 |
| 2020 | Hybrid Zero Dynamics Inspired Feedback Control Policy Design for 3D Bipedal Locomotion using Reinforcement LearningabstractThis paper presents a novel model-free reinforcement learning (RL) framework to design feedback control policies for 3D bipedal walking. Existing RL algorithms are often trained in an end-to-end manner or rely on prior knowledge of some reference joint trajectories. Different from these studies, we propose a novel policy structure that appropriately incorporates physical insights gained from the hybrid nature of the walking dynamics and the well-established hybrid zero dynamics approach for 3D bipedal walking. As a result, the overall RL framework has several key advantages, including lightweight network structure, short training time, and less dependence on prior knowledge. We demonstrate the effectiveness of the proposed method on Cassie, a challenging 3D bipedal robot. The proposed solution produces stable limit walking cycles that can track various walking speed in different directions. Surprisingly, without specifically trained with disturbances to achieve robustness, it also performs robustly against various adversarial forces applied to the torso towards both the forward and the backward directions. Guillermo A. Castillo, Bowen Weng, Wei Zhang 0013, Ayonga Hereid |
ICRA | 3 |
| 2020 | Improved Dynamic Window Approach for Dynamic Obstacle Avoidance of Quadruped RobotsabstractThis paper studies dynamic obstacle avoidance strategies of quadruped robots in an unknown and dynamically changing environment. Different from classical navigation problems with static environment, an unknown number of moving obstacles with unknown dynamics are present in the scenario considered in this paper. This unknown and rapidly changing environment prevents us from directly applying existing navigation approaches for quadruped robots. To address the additional challenges introduced by the dynamic obstacles, an improved dynamic window approach (DWA) is proposed in which both evaluation function and constraints are modified. To enable real-time implementation of the proposed algorithm, multi-layer techniques for processing camera point cloud data are applied for online extraction of environmental information. The proposed algorithm is tested both in simulation and on hardware with a quadruped robot equipped with a low-cost RGB-D camera. The results show that the proposed algorithm significantly increases the rate of success for avoiding dynamical obstacles and reduces the time spent for reaching the target. Hua Chen 0007, Wei Zhang 0013 |
IECON | 5 |
| 2020 | Analysis of Q-learning with Adaptation and Momentum Restart for Gradient DescentabstractExisting convergence analyses of Q-learning mostly focus on the vanilla stochastic gradient descent (SGD) type of updates. Despite the Adaptive Moment Estimation (Adam) has been commonly used for practical Q-learning algorithms, there has not been any convergence guarantee provided for Q-learning with such type of updates. In this paper, we first characterize the convergence rate for Q-AMSGrad, which is the Q-learning algorithm with AMSGrad update (a commonly adopted alternative of Adam for theoretical analysis). To further improve the performance, we propose to incorporate the momentum restart scheme to Q-AMSGrad, resulting in the so-called Q-AMSGradR algorithm. The convergence rate of Q-AMSGradR is also established. Our experiments on a linear quadratic regulator problem demonstrate that the two proposed Q-learning algorithms outperform the vanilla Q-learning with SGD updates. The two algorithms also exhibit significantly better performance than the DQN learning method over a batch of Atari 2600 games. Bowen Weng, Huaqing Xiong, Yingbin Liang, Wei Zhang 0013 |
IJCAI | 4 |
| 2020 | Velocity Regulation of 3D Bipedal Walking Robots with Uncertain Dynamics Through Adaptive Neural Network ControllerabstractThis paper presents a neural-network based adaptive feedback control structure to regulate the velocity of 3D bipedal robots under dynamics uncertainties. Existing Hybrid Zero Dynamics (HZD)-based controllers regulate velocity through the implementation of heuristic regulators that do not consider model and environmental uncertainties, which may significantly affect the tracking performance of the controllers. In this paper, we address the uncertainties in the robot dynamics from the perspective of the reduced dimensional representation of virtual constraints and propose the integration of an adaptive neural network-based controller to regulate the robot velocity in the presence of model parameter uncertainties. The proposed approach yields improved tracking performance under dynamics uncertainties. The shallow adaptive neural network used in this paper does not require training a priori and has the potential to be implemented on the real-time robotic controller. A comparative simulation study of a 3D Cassie robot is presented to illustrate the performance of the proposed approach under various scenarios. Guillermo A. Castillo, Bowen Weng, Terrence C. Stewart, Wei Zhang 0013, Ayonga Hereid |
IROS | 4 |
| 2020 | Finite-Time Analysis for Double Q-learningabstractAlthough Q-learning is one of the most successful algorithms for finding the best action-value function (and thus the optimal policy) in reinforcement learning, its implementation often suffers from large overestimation of Q-function values incurred by random sampling. The double Q-learning algorithm proposed in~\citet{hasselt2010double} overcomes such an overestimation issue by randomly switching the update between two Q-estimators, and has thus gained significant popularity in practice. However, the theoretical understanding of double Q-learning is rather limited. So far only the asymptotic convergence has been established, which does not characterize how fast the algorithm converges. In this paper, we provide the first non-asymptotic (i.e., finite-time) analysis for double Q-learning. We show that both synchronous and asynchronous double Q-learning are guaranteed to converge to an $\epsilon$-accurate neighborhood of the global optimum by taking $\tilde{\Omega}\left(\left( \frac{1}{(1-\gamma)^6\epsilon^2}\right)^{\frac{1}{\omega}} +\left(\frac{1}{1-\gamma}\right)^{\frac{1}{1-\omega}}\right)$ iterations, where $\omega\in(0,1)$ is the decay parameter of the learning rate, and $\gamma$ is the discount factor. Our analysis develops novel techniques to derive finite-time bounds on the difference between two inter-connected stochastic processes, which is new to the literature of stochastic approximation. Huaqing Xiong, Lin Zhao 0009, Yingbin Liang, Wei Zhang 0013 |
NeurIPS | 4 |
| 2019 | Reinforcement Learning Meets Hybrid Zero Dynamics: A Case Study for RABBITabstractThe design of feedback controllers for bipedal robots is challenging due to the hybrid nature of its dynamics and the complexity imposed by high-dimensional bipedal models. In this paper, we present a novel approach for the design of feedback controllers using Reinforcement Learning (RL) and Hybrid Zero Dynamics (HZD). Existing RL approaches for bipedal walking are inefficient as they do not consider the underlying physics, often requires substantial training, and the resulting controller may not be applicable to real robots. HZD is a powerful tool for bipedal control with local stability guarantees of the walking limit cycles. In this paper, we propose a non traditional RL structure that embeds the HZD framework into the policy learning. More specifically, we propose to use RL to find a control policy that maps from the robot’s reduced order states to a set of parameters that define the desired trajectories for the robot’s joints through the virtual constraints. Then, these trajectories are tracked using an adaptive PD controller. The method results in a stable and robust control policy that is able to track variable speed within a continuous interval. Robustness of the policy is evaluated by applying external forces to the torso of the robot. The proposed RL framework is implemented and demonstrated in OpenAI Gym with the MuJoCo physics engine based on the well-known RABBIT robot model. Guillermo A. Castillo, Bowen Weng, Ayonga Hereid, Zheng Wang 0002, Wei Zhang 0013 |
ICRA | 5 |
| 2019 | Model Predictive Tracking Control Design for a Robotic Fish with Controllable BarycentreabstractIn this paper, we present the dynamic modeling and model predictive tracking control for a fin-actuated robot with barycentre regulating mechanism in multiple motions. Specifically, a dynamic model for the robot is established firstly. Based on the dynamic model, a model predictive tracking control algorithm is proposed. And simulations of of tracking rectangle trajectory, sine-like trajectory, ascending trajectory, and spiral trajectory are conducted to validate the algorithm. The simulation results demonstrate that the proposed algorithm is able to implement trajectory tracking of the robot with small position error and orientation error. This paper contributes to trajectory tracking for an underwater robot with controllable barycentre in multiple motions, which has been rarely explored. Xingwen Zheng, Hua Chen 0007, Ouyang Jiao, Minglei Xiong, Wei Zhang 0013, Guangming Xie |
IECON | 5 |
| 2015 | A stochastic hybrid system approach to aggregated load modeling for demand responseabstractThis abstract presents an unified framework for aggregated modeling of various responsive loads. The general model consists of coupled partial differential equations derived by drawing the connection with stochastic hybrid systems. Lin Zhao 0009, Wei Zhang 0013 |
HSCC | 2 |
| 2012 | Communication scheduling for decentralized state estimationabstractThis paper considers decentralized state estimation subject to communication constraints. A group of agents measure the state of a process and obtain their state estimates by exchanging data with one another. Due to the communication constraint, only a few communication channels are available. The main objective of this paper is to allocate these channels among the agents so as to minimize their average estimation errors. Assuming the agents have the same sensing capability, we provide the optimal allocation strategy for both directed and un-directed communication channels. Chao Yang 0009, Junfeng Wu 0001, Ling Shi 0001, Wei Zhang 0013 |
ICARCV | 4 |
| 2012 | A hierarchical method for stochastic motion planning in uncertain environmentsabstractThis paper considers the problem of stochastic motion planning in uncertain environments, and extends existing chance constrained optimal control solutions. Due to the imperfect knowledge of the system state caused by motion uncertainty, sensor noise and environment uncertainty, the system constraints cannot be guaranteed to be satisfied and consequently must be considered probabilistically. To account for the uncertainty, the constraints are formulated as convex constraints on a random variable, known as chance constraints, with the violation probability of all the constraints guaranteed to be below a threshold. Standard chance constrained stochastic motion planning methods do not incorporate environmental sensing which typically leads to overly-conservative solutions. To address this, a novel hierarchical framework is proposed that consists of two main steps: an expected shortest path problem on an uncertain graph and a chance constrained motion planning problem. The first successful, real-time experimental demonstration of chance constrained control with uncertain constraint parameters and variables is also presented for a quadrotor equipped with a Kinect sensor navigating through an uncertain, cluttered 3D environment. Michael P. Vitus, Wei Zhang 0013, Claire J. Tomlin |
IROS | 2 |
| 2012 | A Hierarchical Flight Planning Framework for Air Traffic ManagementabstractThe continuous growth of air traffic demand, skyrocketing fuel price, and increasing concerns on safety and environmental impact of air transportation necessitate the modernization of the air traffic management (ATM) system in the United States. The design of such a large-scale networked system that involves complex interactions among automation and human operators poses new challenges for many engineering fields. This paper investigates several important facets of the future ATM system from a systems-level point of view. In particular, we develop a hierarchical decentralized decision architecture that can design 4-D (space +time) path plans for a large number of flights while satisfying weather and capacity constraints of the overall system. The proposed planning framework respects preferences of individual flights and encourages information sharing among different decision makers in the system, and thus has a great potential to reduce traffic delays and weather risks while maintaining safety standards. The framework is validated through a large-scale simulation based on real traffic data over the entire airspace of the contiguous United States. We envision that the hierarchical decentralization approach developed in this paper would also provide useful insights into the design of decision and information hierarchies for other large-scale infrastructure systems. Wei Zhang 0013, Maryam Kamgarpour, Dengfeng Sun, Claire J. Tomlin |
Proc. IEEE | 1 |
| 2011 | A differential game approach to planning in adversarial scenarios: A case study on capture-the-flagabstractCapture-the-flag is a complex, challenging game that is a useful proxy for many problems in robotics and other application areas. The game is adversarial, with multiple, potentially competing, objectives. This interplay between different factors makes the problem complex, even in the case of only two players. To make analysis tractable, previous approaches often make various limiting assumptions upon player actions. In this paper, we present a framework for analyzing and solving a two-player capture-the-flag game as a zero-sum differential game. Our problem formulation allows each player to make decisions rationally based upon the current player positions, assuming only an upper bound on the movement speeds. Using Hamilton-Jacobi reachability analysis, we compute winning regions for each player as subsets of the joint configuration space and derive the corresponding winning strategies. Simulation results are presented along with implications of the work as a tool for automation-aided decision-making for humans and mixed human-robot teams. Haomiao Huang, Jerry Ding, Wei Zhang 0013, Claire J. Tomlin |
ICRA | 3 |
| 2010 | A generating function approach to the stability of discrete-time switched linear systemsabstractExponential stability of switched linear systems under both arbitrary and proper switching is studied through two suitably defined families of functions called the strong and the weak generating functions. Various properties of the generating functions are established. It is found that the radii of convergence of the generating functions characterize the exponential growth rate of the trajectories of the switched linear systems. In particular, necessary and sufficient conditions for the exponential stability of the systems are derived based on these radii of convergence. Numerical algorithms for computing estimates of the generating functions are proposed and examples are presented for illustration purpose. Jianghai Hu, Jinglai Shen, Wei Zhang 0013 |
HSCC | 3 |
| 2009 | Stabilization of Discrete-Time Switched Linear Systems: A Control-Lyapunov Function Approach
Wei Zhang 0013, Alessandro Abate, Jianghai Hu |
HSCC | 1 |
| 2005 | Hiding privacy information in video surveillance systemabstractThis paper proposes a detailed framework of storing privacy information in surveillance video as a watermark. Authorized personnel is not only removed from the surveillance video as in J. Wickramasuriya et al. (2004) but also embedded into the video itself, which can only be retrieved with a secrete key. A perceptual-model-based compressed domain video watermarking scheme is proposed to deal with the huge payload problem in the proposed surveillance system. A signature is also embedded into the header of the video as in M. Pramateftakis et al. (2004) for authentication. Simulation results have shown that the proposed algorithm can embed all the privacy information into the video without affecting its visual quality. As a result, the proposed video surveillance system can monitor the unauthorized persons in a restricted environment, protect the privacy of the authorized persons but, at the same time, allow the privacy information to be revealed in a secure and reliable way. Wei Zhang 0013, Sen-Ching S. Cheung, Minghua Chen 0001 |
ICIP (3) | 1 |