VLDB 2026 Research / reviewers in the wild / expert
Zhibin Li 0001
dblp:89/6033-1
· DBLP profile ↗
42ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-6357-7419ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 6 first-author · 18 since 2021Systems, architecture and hardware · 34 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Generalist Robot Learning from Internet Video: A Survey (Abstract Reprint)abstractScaling deep learning to massive and diverse internet data has driven remarkable breakthroughs in domains such as video generation and natural language processing. Robot learning, however, has thus far failed to replicate this success and remains constrained by a scarcity of available data. Learning from Videos (LfV) methods aim to address this data bottleneck by augmenting traditional robot data with large-scale internet video. This video data provides foundational information regarding physical dynamics, behaviours, and tasks, and can be highly informative for general-purpose robots. This survey systematically examines the emerging field of LfV. We first outline essential concepts, including detailing fundamental LfV challenges such as distribution shift and missing action labels in video data. Next, we comprehensively review current methods for extracting knowledge from large-scale internet video, overcoming LfV challenges, and improving robot learning through video-informed training. The survey concludes with a critical discussion of future opportunities. Here, we emphasize the need for scalable foundation model approaches that can leverage the full range of available internet video and enhance the learning of robot policies and dynamics models. Overall, the survey aims to inform and catalyse future LfV research, driving progress towards general-purpose robots. Robert McCarthy, Daniel Tan 0001, Dominik Schmidt, Fernando Acero, Nathan Herr, Yilun Du, Thomas George Thuruthel, Zhibin Li 0001 |
AAAI | 8 |
| 2026 | eGAIT: Multi-Skilled Policy for Energy-Efficient Gait TransitionsabstractAchieving adaptive, multi-skilled, and energy-efficient locomotion is vital for advancing the operation of autonomous quadrupedal systems. This study presents eGAIT, a unified multi-skilled policy enabling energy-efficient and stable gait transitions across nine non-monotonic, velocity-optimized gaits, in response to dynamic velocity commands. The framework leverages a hybrid control architecture that integrates model-based and learning-based methods to address the entire locomotion pipeline. An MPC-based gait generator produces velocity-optimized trajectories, which are imitated through Proximal Policy Optimization (PPO), driven by a Adversarial Motion Prior (AMP) style reward to train distinct policies for specific velocity ranges. These policies are unified through a Hierarchical Reinforcement Learning (HRL) framework featuring a novel modified Deep Q-Network (eDQN) for real-time velocity-to-policy mapping. Training efficiency is enhanced by an auxiliary selector layer that guides velocity-policy mapping, while a sparsely activated stability reward mechanism ensures smooth gait transitions by incorporating geometric and rotational stability. Extensively validated in simulation and on a Unitree Go1 robot, eGAIT achieves a 100% success rate in velocity-to-policy mapping, a 35% improvement in energy efficiency, a 31% improvement in both velocity tracking and stability compared to the next best state-of-the-art method. This work advances autonomous quadrupedal locomotion, enabling longer, more efficient, and stable operations in dynamic environments. Supplementary materials and visualizations related to the paper can be found at: https://github.com/RPL-CS-UCL/egait/. Maria Stamatopoulou, Daniel Tan 0001, Rokas Bendikas, Valerio Modugno, Zhibin Li 0001, Dimitrios Kanoulas |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Towards Generalist Robot Learning from Internet Video: A SurveyabstractScaling deep learning to massive and diverse internet data has driven remarkable breakthroughs in domains such as video generation and natural language processing. Robot learning, however, has thus far failed to replicate this success and remains constrained by a scarcity of available data. Learning from Videos (LfV) methods aim to address this data bottleneck by augmenting traditional robot data with large-scale internet video. This video data provides foundational information regarding physical dynamics, behaviours, and tasks, and can be highly informative for general-purpose robots. This survey systematically examines the emerging field of LfV. We first outline essential concepts, including detailing fundamental LfV challenges such as distribution shift and missing action labels in video data. Next, we comprehensively review current methods for extracting knowledge from large-scale internet video, overcoming LfV challenges, and improving robot learning through video-informed training. The survey concludes with a critical discussion of future opportunities. Here, we emphasize the need for scalable foundation model approaches that can leverage the full range of available internet video and enhance the learning of robot policies and dynamics models. Overall, the survey aims to inform and catalyse future LfV research, driving progress towards general-purpose robots. Robert McCarthy, Daniel Tan 0001, Dominik Schmidt, Fernando Acero, Nathan Herr, Yilun Du, Thomas George Thuruthel, Zhibin Li 0001 |
J. Artif. Intell. Res. | 8 |
| 2025 | Efficient learning of robust multigait quadruped locomotion for minimizing the cost of transportabstractQuadruped robots are able to exhibit a range of gaits, each with its own traversability and energy efficiency characteristics. By actively coordinating between gaits in different scenarios, energy-efficient and adaptive locomotion can be achieved. This study investigates the performances of learned energy-efficient policies for quadrupedal gaits under different commands. We propose a training–synthesizing framework that integrates learned gait-conditioned locomotion policies into an efficient multiskill locomotion policy. The resulting control policy achieves low-cost smooth switching and controllable gaits. Our results of the learned multiskill policy demonstrate seamless gait transitions while maintaining energy optimality across all commands. Zhicheng Wang 0003, Meng Yee Chuah, Zhibin Li 0001, Jun Wu 0003, Qiuguo Zhu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | Erratum to: Efficient learning of robust multigait quadruped locomotion for minimizing the cost of transport
Zhicheng Wang 0003, Meng Yee Chuah, Zhibin Li 0001, Jun Wu 0003, Qiuguo Zhu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2024 | Intrinsic Language-Guided Exploration for Complex Long-Horizon Robotic Manipulation TasksabstractCurrent reinforcement learning algorithms struggle in sparse and complex environments, most notably in long-horizon manipulation tasks entailing a plethora of different sequences. In this work, we propose the Intrinsically Guided Exploration from Large Language Models (IGE-LLMs) framework. By leveraging LLMs as an assistive intrinsic reward, IGE-LLMs guides the exploratory process in reinforcement learning to address intricate long-horizon with sparse rewards robotic manipulation tasks. We evaluate our framework and related intrinsic learning methods in an environment challenged with exploration, and a complex robotic manipulation task challenged by both exploration and long-horizons. Results show IGE-LLMs (i) exhibit notably higher performance over related intrinsic methods and the direct use of LLMs in decision-making, (ii) can be combined and complement existing learning methods highlighting its modularity, (iii) are fairly insensitive to different intrinsic scaling parameters, and (iv) maintain robustness against increased levels of uncertainty and horizons. Eleftherios Triantafyllidis, Filippos Christianos, Zhibin Li 0001 |
ICRA | 3 |
| 2024 | TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image PairsabstractVideos capture the evolution of continuous dynamical systems over time in the form of discrete image sequences. Recently, video generation models have been widely used in robotic research. However, generating controllable videos from image-text pairs is an important yet underexplored research topic in both robotic and computer vision communities. This paper introduces an innovative and elegant framework named TiV-ODE, formulating this task as modeling the dynamical system in a continuous space. Specifically, our framework leverages the ability of Neural Ordinary Differential Equations (Neural ODEs) to model the complex dynamical system depicted by videos as a nonlinear ordinary differential equation. The resulting framework offers control over the generated videos’ dynamics, content, and frame rate, a feature not provided by previous methods. Experiments demonstrate the ability of the proposed method to generate highly controllable and visually consistent videos and its capability of modeling dynamical systems. Overall, this work is a significant step towards developing advanced controllable video generation models that can handle complex and dynamic scenes. Nanbo Li, Arushi Goel, Zonghai Yao, Zijian Guo 0002, Hamidreza Kasaei 0001, Seyed Mohammadreza Mohades Kasaei, Zhibin Li 0001 |
ICRA | 8 |
| 2024 | Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic RegressionabstractRecent advancements in reinforcement learning (RL) have led to remarkable achievements in robot locomotion capabilities. However, the complexity and "black-box" nature of neural network-based RL policies hinder their interpretability and broader acceptance, particularly in applications demanding high levels of safety and reliability. This paper introduces a novel approach to distill neural RL policies into more interpretable forms using Gradient Boosting Machines (GBMs), Explainable Boosting Machines (EBMs) and Symbolic Regression. By leveraging the inherent interpretability of generalized additive models, decision trees, and analytical expressions, we transform opaque neural network policies into more transparent "glass-box" models. We train expert neural network policies using RL and subsequently distill them into (i) GBMs, (ii) EBMs, and (iii) symbolic policies. To address the inherent distribution shift challenge of behavioral cloning, we propose to use the Dataset Aggregation (DAgger) algorithm with a curriculum of episode-dependent alternation of actions between expert and distilled policies, to enable efficient distillation of feedback control policies. We evaluate our approach on various robot locomotion gaits – walking, trotting, bounding, and pacing – and study the importance of different observations in joint actions for distilled policies using various methods. We train neural expert policies for 205 hours of simulated experience and distill interpretable policies with only 10 minutes of simulated interaction for each gait using the proposed method. Fernando Acero, Zhibin Li 0001 |
IROS | 2 |
| 2024 | DexSkills: Skill Segmentation Using Haptic Data for Learning Autonomous Long-Horizon Robotic Manipulation TasksabstractEffective execution of long-horizon tasks with dexterous robotic hands remains a significant challenge in real-world problems. While learning from human demonstrations has shown encouraging results, they require extensive data collection for training. Hence, decomposing long-horizon tasks into reusable primitive skills is a more efficient approach. To achieve so, we developed DexSkills, a novel supervised learning framework that addresses long-horizon dexterous manipulation tasks using primitive skills. DexSkills is trained to recognize and replicate a select set of skills using human demonstration data, which can then segment a demonstrated long-horizon dexterous manipulation task into a sequence of primitive skills to achieve one-shot execution by the robot directly. Significantly, DexSkills operates solely on proprioceptive and tactile data, i.e., haptic data. Our real-world robotic experiments show that DexSkills can accurately segment skills, thereby enabling autonomous robot execution of a diverse range of tasks. Xiaofeng Mao, Gabriele Giudici, Claudio Coppola, Kaspar Althoefer, Ildar Farkhatdinov, Zhibin Li 0001, Lorenzo Jamone |
IROS | 6 |
| 2024 | Efficient Tactile Sensing-based Learning from Limited Real-world Demonstrations for Dual-arm Fine Pinch-Grasp SkillsabstractImitation learning for robot dexterous manipulation, especially with a real robot setup, typically requires a large number of demonstrations. In this paper, we present a data-efficient learning from demonstration framework which exploits the use of rich tactile sensing data and achieves fine bimanual pinch grasping. Specifically, we employ a convolutional autoencoder network that can effectively extract and encode high-dimensional tactile information. Further, we develop a framework that achieves efficient multi-sensor fusion for imitation learning, allowing the robot to learn contact-aware sensorimotor skills from demonstrations. The ablation studies on encoded tactile features highlighted the effectiveness of incorporating rich contact information, which enabled dexterous bimanual grasping with active contact searching. Extensive experiments demonstrated the robustness of the fine pinch grasp policy directly learned from few-shot demonstration, including grasping of the same object with different initial poses, generalizing to ten unseen new objects, robust and firm grasping against external pushes, as well as contact-aware and reactive re-grasping in case of dropping objects under very large perturbations. Furthermore, the saliency map analysis method is used to describe weight distribution across various modalities during pinch grasping, confirming the effectiveness of our framework at leveraging multimodal information. The video is available online at: https://youtu.be/BlzxGgiKfck. Xiaofeng Mao, Ruoshi Wen, Seyed Mohammadreza Mohades Kasaei, Wanming Yu, Efi Psomopoulou, Nathan F. Lepora, Zhibin Li 0001 |
IROS | 8 |
| 2024 | Neural ODE-based Imitation Learning (NODE-IL): Data-Efficient Imitation Learning for Long-Horizon Multi-Skill Robot ManipulationabstractIn robotics, acquiring new skills through Imitation Learning (IL) is crucial for handling diverse complex tasks. However, model-free IL faces challenges of data inefficiency and prolonged training time, whereas model-based methods struggle to obtain accurate nonlinear models. To address these challenges, we developed Neural ODE-based Imitation Learning (NODE-IL), a novel model-based imitation learning framework that employs Neural Ordinary Differential Equations (Neural ODEs) for learning task dynamics and control policies. NODE-IL comprises (1) Dynamic-NODE for learning the continuous differentiable task’s transition dynamics model, and (2) Control-NODE for learning a long-horizon control policy in an MPC fashion, which are trained holistically. Extensively evaluated on challenging manipulation tasks, NODE-IL demonstrates significant advantages in data efficiency, requiring less than 70 samples to achieve robust performance. It outperforms Behavioral Cloning from Observation (BCO) and Gaussian Process Imitation Learning (GP-IL) methods, achieving 70% higher average success rate, and reducing translation errors for high-precision tasks, which demonstrates its robustness and accuracy, as an effective and efficient imitation learning approach for learning complex manipulation tasks. Shiyao Zhao, Seyed Mohammadreza Mohades Kasaei, Mohsen Khadem, Zhibin Li 0001 |
IROS | 5 |
| 2023 | Agile and Versatile Robot Locomotion via Kernel-based Residual LearningabstractThis work developed a kernel-based residual learning framework for quadrupedal robotic locomotion. Ini-tially, a kernel neural network is trained with data collected from an MPC controller. Alongside a frozen kernel network, a residual controller network is trained using reinforcement learning to acquire generalized locomotion skills and robust-ness against external perturbations. The proposed framework successfully learns a robust quadrupedal locomotion controller with high sample efficiency and controllability, which can provide omnidirectional locomotion at continuous velocities. We validated its versatility and robustness on unseen terrains that the expert MPC controller failed to traverse. Furthermore, the learned kernel can produce a range of functional locomotion behaviors and can generalize to unseen gaits. Milo Carroll, Zhaocheng Liu, Seyed Mohammadreza Mohades Kasaei, Zhibin Li 0001 |
ICRA | 4 |
| 2023 | Data-efficient Non-parametric Modelling and Control of an Extensible Soft ManipulatorabstractData-driven approaches have shown promising results in modeling and controlling robots, specifically soft and flexible robots where developing physics-based models are more challenging. However, these methods often require a large number of real data, and gathering such data is time-consuming and can damage the robot as well. This paper proposed a novel data-efficient and non-parametric approach to develop a continuous model using a small dataset of real robot demonstrations (only 25 points). To the best of our knowledge, the proposed approach is the most sample-efficient method for soft continuum robot. Furthermore, we employed this model to develop a controller to track arbitrary trajectories in the feasible kinematic space. To show the performance of the proposed approach, a set of trajectory-tracking experiments has been conducted. The results showed that the robot was able to track the references precisely even in presence of external loads (up to 25 grams). Moreover, fine object manipulation experiments were performed to demonstrate the effectiveness of the proposed method in real-world tasks. Finally, we compared its performance with common data-driven approaches in seen/useen-before trajectory tracking scenarios. The results validated that the proposed approach significantly outperformed the existing approaches in unseen-before scenarios and offered similar performance in seen-before scenarios. Seyed Mohammadreza Mohades Kasaei, Keyhan Kouhkiloui Babarahmati, Zhibin Li 0001, Mohsen Khadem |
ICRA | 3 |
| 2023 | Instance-wise Grasp Synthesis for Robotic GraspingabstractGenerating high-quality instance-wise grasp con-figurations provides critical information of how to grasp specific objects in a multi-object environment and is of high importance for robot manipulation tasks. This work proposed a novel Single-Stage Grasp (SSG) synthesis network, which performs high-quality instance-wise grasp synthesis in a single stage: instance mask and grasp configurations are generated for each object simultaneously. Our method outperforms state-of-the-art on robotic grasp prediction based on the OCID-Grasp dataset, and performs competitively on the JACQUARD dataset. The benchmarking results showed significant improvements compared to the baseline on the accuracy of generated grasp configurations. The performance of the proposed method has been validated through both extensive simulations and real robot experiments for three tasks including single object pick-and-place, grasp synthesis in cluttered environments and table cleaning task. Seyed Mohammadreza Mohades Kasaei, Hamidreza Kasaei 0001, Zhibin Li 0001 |
ICRA | 4 |
| 2023 | Modular Neural Network Policies for Learning In-Flight Object Catching with a Robot Hand-Arm SystemabstractWe present a modular framework designed to enable a robot hand-arm system to learn how to catch flying objects, a task that requires fast, reactive, and accurately-timed robot motions. Our framework consists of five core modules: (i) an object state estimator that learns object trajectory prediction, (ii) a catching pose quality network that learns to score and rank object poses for catching, (iii) a reaching control policy trained to move the robot hand to pre-catch poses, (iv) a grasping control policy trained to perform soft catching motions for safe and robust grasping, and (v) a gating network trained to synthesize the actions given by the reaching and grasping policy. The former two modules are trained via supervised learning and the latter three use deep reinforcement learning in a simulated environment. We conduct extensive evaluations of our framework in simulation for each module and the integrated system, to demonstrate high success rates of in-flight catching and robustness to perturbations and sensory noise. Whilst only simple cylindrical and spherical objects are used for training, the integrated system shows successful generalization to a variety of household objects that are not used in training. Fernando Acero, Eleftherios Triantafyllidis, Zhaocheng Liu, Zhibin Li 0001 |
IROS | 5 |
| 2023 | Run and Catch: Dynamic Object-Catching of Quadrupedal RobotsabstractQuadrupedal robots are performing increasingly more real-world capabilities, but are primarily limited to locomotion tasks. To expand their task-level abilities of object acquisition, i.e., run-to-catch as frisbee catching for dogs, this paper developed a control pipeline using stereo vision for legged robots which allows for dynamic catching balls while the robot is in motion. To achieve high-frame-rate tracking, we designed a ball that can actively emit homogeneous infrared (IR) light and then located the flying ball based on binocular vision positioning using the onboard RealSense D450 camera with an additional IR bandpass filter. The camera was mounted on top of a 2-DoF head to gain a full view of the target ball. A state estimation module was developed to fuse the vision positioning, camera motor readings, localization result of RealSense T265 equipped on the back, and the legged odometry output altogether. With the use of a ballistic model, we achieved a robust estimation of both the ball and robot positions in an inertial coordinate. Additionally, we developed a close-loop catching strategy and employed trajectory prediction so that tracking and run-to-catch were performed simultaneously, which is critical for such drastically dynamic and precise tasks. The proposed approach was validated through both static testing and dynamic catch experiments conducted on the CyberDog robot with a high success rate. Yangwei You, Tianlin Liu, Xiaowei Liang, Mingliang Zhou 0003, Zhibin Li 0001, Shiwu Zhang |
IROS | 6 |
| 2022 | Robust Impedance Control for Dexterous Interaction Using Fractal Impedance Controller with IK-OptimisationabstractRobust dynamic interactions are required to move robots in daily environments alongside humans. Optimisation and learning methods have been used to mimic and reproduce human movements. However, they are often not robust and their generalisation is limited. This work proposed a hierarchical control architecture for robot manipulators and provided capabilities of reproducing human-like motions during unknown interaction dynamics. Our results show that the reproduced end-effector trajectories can preserve the main characteristics of the initial human motion recorded via a motion capture system, and are robust against external perturbations. The data indicate that some detailed movements are hard to reproduce due to the physical limits of the hardware that cannot reach the same velocity recorded in human movements. Nevertheless, these technical problems can be addressed by using better hardware and our proposed algorithms can still be applied to produce imitated motions. Carlo Tiseo, Quentin Rouxel, Zhibin Li 0001, Michael N. Mistry |
ICRA | 3 |
| 2022 | Accessibility-Based Clustering for Efficient Learning of Locomotion SkillsabstractFor model-free deep reinforcement learning of quadruped locomotion, the initialization of robot configurations is crucial for data efficiency and robustness. This work focuses on algorithmic improvements of data efficiency and robustness simultaneously through automatic discovery of initial states, which is achieved by our proposed K-Access algorithm based on accessibility metrics. Specifically, we formulated accessibility metrics to measure the difficulty of transitions between two arbitrary states, and proposed a novel K-Access algorithm for state-space clustering that automatically discovers the centroids of the static-pose clusters based on the accessibility metrics. By using the discovered centroidal static poses as the initial states, we can improve data efficiency by reducing redundant explorations, and enhance the robustness by more effective explorations from the centroids to sampled poses. Focusing on fall recovery as a very hard set of locomotion skills, we validated our method extensively using an 8-DoF quadrupedal robot Bittle. Compared to the baselines, the learning curve of our method converges much faster, requiring only 60% of training episodes. With our method, the robot can successfully recover to standing poses within 3 seconds in 99.4% of the test cases. Moreover, the method can generalize to other difficult skills successfully, such as backflipping. Wanming Yu, Zhibin Li 0001 |
ICRA | 3 |
| 2022 | Real-time Digital Double Framework to Predict Collapsible Terrains for Legged RobotsabstractInspired by the digital twinning systems, a novel real-time digital double framework is developed to enhance robot perception of the terrain conditions. Based on the very same physical model and motion control, this work exploits the use of such simulated digital double synchronized with a real robot to capture and extract discrepancy information between the two systems, which provides high dimensional cues in multiple physical quantities to represent differences between the modelled and the real world. Soft, non-rigid terrains cause common failures in legged locomotion, whereby visual perception solely is insufficient in estimating such physical properties of terrains. We used digital double to develop the estimation of the collapsibility, which addressed this issue through physical interactions during dynamic walking. The discrepancy in sensory measurements between the real robot and its digital double are used as input of a learning-based algorithm for terrain collapsibility analysis. Although trained only in simulation, the learned model can perform collapsibility estimation successfully in both simulation and real world. Our evaluation of results showed the generalization to different scenarios and the advantages of the digital double to reliably detect nuances in ground conditions. Garen Haddeler, Hari P. Palanivelu, Yung Chuen Ng, Fabien Colonnier, Albertus Hendrawan Adiwahono, Zhibin Li 0001, Chee-Meng Chew, Meng Yee Chuah |
IROS | 6 |
| 2021 | Robust High-Transparency Haptic Exploration for Dexterous TelemanipulationabstractRobotic teleoperation provides human-in-the-loop capabilities of complex manipulation tasks in dangerous or remote environments, such as for planetary exploration or nuclear decommissioning. This work proposes a novel telemanipulation architecture using a passive Fractal Impedance Controller (FIC), which does not depend upon an active viscous component for guaranteeing stability. Compared to a traditional impedance controller in ideal conditions (no delays and maximum communication bandwidth), our proposed method yields higher transparency in interaction and demonstrates superior dexterity and capability in our telemanipulation test scenarios. We also validate its performance with extreme delays up to 1 s and communication bandwidths as low as 10 Hz. All results validate a consistent stability when using the proposed controller in challenging conditions, regardless of operator expertise. Keyhan Kouhkiloui Babarahmati, Carlo Tiseo, Quentin Rouxel, Zhibin Li 0001, Michael N. Mistry |
ICRA | 4 |
| 2021 | Meta-Learning for Fast Adaptive Locomotion with Uncertainties in Environments and Robot DynamicsabstractThis work developed meta-learning control policies to achieve fast online adaptation to different changing conditions, which generate diverse and robust locomotion. The proposed method updates the interaction model constantly, samples feasible sequences of actions of estimated state-action trajectories, and then applies the optimal actions to maximize the reward. To achieve online model adaptation, our proposed method learns different latent vectors of each training condition, which is selected online based on newly collected data from the past 10 samples within 0.2s. Our work designs appropriate state space and reward functions, and optimizes feasible actions in an MPC fashion which are sampled directly in the joint space with constraints, hence requiring no prior design or training of specific gaits. We further demonstrated the robot’s capability of detecting unexpected changes during the interaction and adapting the control policy in less than 0.2s. The extensive validation on the SpotMicro robot in a physics simulation shows adaptive and robust locomotion skills under changing ground friction, external pushes, and different robot dynamics including motor failures and the whole leg amputation. Timothée Anne, Jack Wilkinson, Zhibin Li 0001 |
IROS | 3 |
| 2020 | Unified Push Recovery Fundamentals: Inspiration from Human StudyabstractCurrently for balance recovery, humans outperform humanoid robots which use hand-designed controllers in terms of the diverse actions. This study aims to close this gap by finding core control principles that are shared across ankle, hip, toe and stepping strategies by formulating experiments to test human balance recoveries and define criteria to quantify the strategy in use. To reveal fundamental principles of balance strategies, our study shows that a minimum jerk controller can accurately replicate comparable human behaviour at the Centre of Mass level. Therefore, we formulate a general Model-Predictive Control (MPC) framework to produce recovery motions in any system, including legged machines, where the framework parameters are tuned for time-optimal performance in robotic systems. Christopher McGreavy, Daniel F. N. Gordon, Kang Tan, Wouter Wolfslag, Sethu Vijayakumar, Zhibin Li 0001 |
ICRA | 7 |
| 2020 | Learning Pregrasp Manipulation of Objects from Ungraspable PosesabstractIn robotic grasping, objects are often occluded in ungraspable configurations such that no feasible grasp pose can be found, e.g. large flat boxes on the table that can only be grasped once lifted. Inspired by human bimanual manipulation, e.g. one hand to lift up things and the other to grasp, we address this type of problems by introducing pregrasp manipulation – push and lift actions. We propose a model-free Deep Reinforcement Learning framework to train feedback control policies that utilize visual information and proprioceptive states of the robot to autonomously discover robust pregrasp manipulation. The robot arm learns to push the object first towards a support surface and then lift up one side of the object, creating an object-table clearance for possible grasping solutions. Furthermore, we show the robustness of the proposed learning framework in training pregrasp policies that can be directly transferred to a real robot. Lastly, we evaluate the effectiveness and generalization ability of the learned policy in real-world experiments, and demonstrate pregrasp manipulation of objects with various sizes, shapes, weights, and surface friction. Zhaole Sun, Chuanyu Yang, Zhibin Li 0001 |
ICRA | 5 |
| 2020 | Optimisation of Body-ground Contact for Augmenting the Whole-Body Loco-manipulation of Quadruped RobotsabstractLegged robots have great potential to perform complex loco-manipulation tasks, yet it is challenging to keep the robot balanced while it interacts with the environment. In this paper we investigated the use of additional contact points for maximising the robustness of loco-manipulation motions. Specifically, body-ground contact was studied for its ability to enhance robustness and manipulation capabilities of quadrupedal robots. We proposed equipping the robot with prongs: small legs rigidly attached to the body which create body-ground contact at controllable point-contacts. The effect of these prongs on robustness was quantified by computing the Smallest Unrejectable Force (SUF), a measure of robustness related to Feasible Wrench Polytopes. We applied the SUF to evaluate the robustness of the system, and proposed an effective approximation of the SUF that can be computed at near-real-time speed. We developed a hierarchical quadratic programming based whole-body controller that can control stable interaction when the prongs are in contact with the ground. This novel prong concept and complementary control framework were implemented on hardware to validate their effectiveness by showing increased robustness and newly enabled loco-manipulation tasks, such as obstacle clearance and manipulation of a large object. Wouter Wolfslag, Christopher McGreavy, Guiyang Xin, Carlo Tiseo, Sethu Vijayakumar, Zhibin Li 0001 |
IROS | 6 |
| 2018 | Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain ChallengeabstractThis paper presents a deep learning framework that is capable of solving partially observable locomotion tasks based on our novel interpretation of Recurrent Deterministic Policy Gradient (RDPG). We study on bias of sampled error measure and its variance induced by the partial observability of environment and subtrajectory sampling, respectively. Three major improvements are introduced in our RDPG based learning framework: tail-step bootstrap of temporal difference, initialisation of hidden state using past subtrajectory, truncation of temporal backpropagation, and injection of external experiences learned by other agents. The proposed learning framework was implemented to solve the Bipedal-Walker challenge in OpenAI's gym simulation environment where only partial state information is available. Our simulation study shows that the autonomous behaviors generated by the RDPG agent are highly adaptive to a variety of obstacles and enables the agent to effectively traverse rugged terrains for long distance with higher success rate than leading contenders. Doo Re Song, Chuanyu Yang, Christopher McGreavy, Zhibin Li 0001 |
ICARCV | 4 |
| 2018 | Comparison Study of Nonlinear Optimization of Step Durations and Foot Placement for Dynamic WalkingabstractThis paper studies bipedal locomotion as a nonlinear optimization problem based on continuous and discrete dynamics, by simultaneously optimizing the remaining step duration, the next step duration and the foot location to achieve robustness. The linear inverted pendulum as the motion model captures the center of mass dynamics and its low-dimensionality makes the problem more tractable. We first formulate a holistic approach to search for optimality in the three-dimensional parametric space and use these results as baseline. To further improve computational efficiency, our study investigates a sequential approach with two stages of customized optimization that first optimizes the current step duration, and subsequently the duration and location of the next step. The effectiveness of both approaches is successfully demonstrated in simulation by applying different perturbations. The comparison study shows that these two approaches find mostly the same optimal solutions, but the latter requires considerably less computational time, which suggests that the proposed sequential approach is well suited for real-time implementation with a minor trade-off in optimality. Iordanis Chatzinikolaidis, Zhibin Li 0001 |
ICRA | 4 |
| 2018 | An Improved Formulation for Model Predictive Control of Legged Robots for Gait Planning and Feedback ControlabstractPredictive control methods for walking commonly use low dimensional models, such as a Linear Inverted Pendulum Model (LIPM), for simplifying the complex dynamics of legged robots. This paper identifies the physical limitations of the modeling methods that do not account for external disturbances, and then analyzes the issues of numerical stability of Model Predictive Control (MPC)using different models with variable receding horizons. We propose a new modeling formulation that can be used for both gait planning and feedback control in an MPC scheme. The advantages are the improved numerical stability for long prediction horizons and the robustness against various disturbances. Benchmarks were rigorously studied to compare the proposed MPC scheme with the existing ones in terms of numerical stability and disturbance rejection. The effectiveness of the controller is demonstrated in both MATLAB and Gazebo simulations. Zhibin Li 0001 |
IROS | 2 |
| 2017 | A study of nonlinear forward models for dynamic walkingabstractThis paper offers a novel insight of using nonlinear models for the control to produce more robust and natural walking gaits for humanoid robots. The sagittal and lateral gait control needs to be treated differently, hence, we proposed two types of suitable nonlinear models, which allow forward simulations to look ahead, and thus, predict accurately the future trajectory/state at the end of the current step. Subsequently, by performing multiple forward simulations in a similar manner for the next step and using the gradient descent method, an appropriate foot placement can be found to achieve precise walking speed. By doing this two-step lookahead, all trajectories of the support and the swing leg can be generated. Our proposed controller can plan trajectories at the beginning of each step or actively re-plan according to task state errors. It is validated effectively in simulations performed in both ADAMS and Open Dynamic Engine. The robot can successfully traverse up/down a stair and recover from pushes with more natural looking gaits compared to the conventional bent-knee style. The reasonable computational time also indicates the feasibility of real-time implementation on real robots. Yangwei You, Chengxu Zhou, Zhibin Li 0001, Nikolaos G. Tsagarakis |
ICRA | 3 |
| 2017 | Humanoid Balancing Behavior Featured by Underactuated Foot MotionabstractA novel control synthesis is proposed for humanoids to demonstrate unique foot-tilting behaviors that are comparable to humans in balance recovery. Our study of model-based behaviors explains the underlying mechanism and the significance of foot tilting well. Our main algorithms are composed of impedance control at the center of mass, virtual stoppers that prevent overtilting of the feet, and postural control for the torso. The proof of concept focuses on the sagittal scenario and the proposed control is effective to produce human-like balancing behaviors characterized by active foot tilting. The successful replication of this behavior on a real humanoid proves the feasibility of deliberately controlled underactuation. The experimental validation was rigorously performed, and the data from the submodules and the entire control were presented and analyzed. Zhibin Li 0001, Chengxu Zhou, Qiuguo Zhu, Rong Xiong |
IEEE Trans. Robotics | 1 |
| 2015 | Fall Prediction of legged robots based on energy state and its implication of balance augmentation: A study on the humanoidabstractIn this paper, we propose an Energy based Fall Prediction (EFP) which observes the real-time balance status of a humanoid robot during standing. The EFP provides an analytic and quantitative measure of the level of balance. Both simulation and experimental studies were conducted and compared with the previously proposed indicators, such as Capture Point (CP) and Foot Rotation Indicator (FRI). The EFP also suggests the balance augmentation by active foot tilting to create larger potential barriers. As a proof of concept, a hybrid balance controller was designed to stabilize the robot including under-actuation phases so the robot can also balance with shoes. Our study reveals that both EFP and CP successfully predict falling about 0.2s in advance for the tested robot, while the FRI fails due to the light weight of the foot and limited resolution of the force/torque measurement. Zhibin Li 0001, Chengxu Zhou, Juan Alejandro Castano, Xin Wang 0041, Francesca Negrello, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
ICRA | 1 |
| 2015 | Active control of under-actuated foot tilting for humanoid push recoveryabstractWe propose a novel control framework to demonstrate a unique foot tilting maneuver based on ankle torque control for humanoid balance recovery. The framework consists of the variable impedance regulation at the center of mass of the robot based on the ankle torque control, the virtual stoppers to prevent over tilting of the feet, and the body attitude control. The scope of our paper focuses on the sagittal scenario as the first proof of concept on the balance recovery by means of active foot tilting without losing stability. Our study demonstrates the success of the control implementation for the humanoid push recovery and the feasibility of having actively controlled foot tilting. The experimental data are presented and analyzed. Zhibin Li 0001, Chengxu Zhou, Qiuguo Zhu, Rong Xiong, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
IROS | 1 |
| 2015 | From one-legged hopping to bipedal running and walking: A unified foot placement control based on regression analysisabstractThis paper aims at developing a unified and adaptive foot placement control for legged robots. The locomotion control of legged robots can be classified into three parts as body height control, body attitude control, and forward velocity control. In our study, the body attitude is controlled at stance phase by the hip actuator, and the height is controlled by the motion of the stance leg. In this case, the foot placement has a nearly linear correlation with forward velocity. Hereby, a generic foot placement controller is developed to control the forward velocity based on the online linear regression analysis of their coupled correlation. Our proposed algorithm is capable of adjusting the control parameters automatically, and is featured by good adaptability and higher control accuracy that outperforms the empirical tuning. The very same controller is able to produce stable hopping with accurate forward velocity tracking even with unknown mass offset, as well as stable bipedal running and walking with accurate velocity tracking. Yangwei You, Zhibin Li 0001, Darwin G. Caldwell, Nikolaos G. Tsagarakis |
IROS | 2 |
| 2015 | Exploiting the redundancy for humanoid robots to dynamically step over a large obstacleabstractIn this paper, we resolve the issue of stepping over a large obstacle by exploiting the redundancy of pelvis rotation and the versatility of foot trajectories for the humanoids. The control framework consists of a motion pattern that exploits the redundancy of pelvis rotation to enlarge the kinematic workspace, a generic foot trajectory generation which can be modified by a parametric interface to adapt to a specific task as well as utilizing the hip abduction to avoid obstacle collision. Moreover, the compensation strategies are also presented for reducing the discrepancies to implement the dynamic stepping motion on a real robot. The effectiveness is validated by COMAN's capability of dynamically stepping over a large obstacle of 10cm height by 5cm width which is almost 20% of its leg length in both simulation and experiment. Chengxu Zhou, Xin Wang 0041, Zhibin Li 0001, Darwin G. Caldwell, Nikolaos G. Tsagarakis |
IROS | 3 |
| 2014 | A passivity based compliance stabilizer for humanoid robotsabstractThis paper presents a passivity based compliance stabilizer for humanoid robots. The proposed stabilizer is an admittance controller that uses the force/torque sensing in feet to actively regulate the compliance for the position controlled system. The low stiffness provided by the stabilizer permits compliant interaction with external forces, and the active damping control guarantees the passivity by dissipating the excessive energy delivered by disturbances. Both the theoretical work and simulation validations are presented. The effectiveness of the stabilizer is demonstrated by the simulations of a simplified cart-table model and the multi-body model of a humanoid under impulsive/periodic force perturbations during standing and walking in place. Simulation data show the quantitative evaluation of the stabilization effect by comparing the responses of body attitude, center of mass, center of pressure without and with the stabilizer. Chengxu Zhou, Zhibin Li 0001, Juan Alejandro Castano, Houman Dallali, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
ICRA | 2 |
| 2013 | COMpliant huMANoid COMAN: Optimal joint stiffness tuning for modal frequency controlabstractThe incorporation of passive compliance in robotic systems could improve their performance during interactions and impacts, for energy storage and efficiency, and for general safety for both the robots and humans. This paper presents the recently developed COMpliant huMANoid COMAN. COMAN is actuated by passive compliance actuators based on the series elastic actuation principle (SEA). The design and implementation of the overall body of the robot is discussed including the realization of the different body segments and the tuning of the joint distributed passive elasticity. This joint stiffness tuning is a critical parameter in the performance of compliant systems. A novel systematic method to optimally tune the joint elasticity of multi-dof SEA robots based on resonance analysis and energy storage maximization criteria forms one of the key contributions of this work. The paper will show this method being applied to the selection of the passive elasticity of COMAN legs. The first completed robot prototype is presented accompanied by experimental walking trials to demonstrate its operation. Nikolaos G. Tsagarakis, Stephen Morfey, Gustavo A. Medrano-Cerda, Zhibin Li 0001, Darwin G. Caldwell |
ICRA | 4 |
| 2013 | Stabilizing humanoids on slopes using terrain inclination estimationabstractThis paper presents an integrated control framework for balancing humanoids on uneven terrains combining stabilization control and terrain inclination estimation. The stabilization is realized by passivity based admittance control that utilizes the force/torque feedback in feet to actively regulate the compliance. The logic-based terrain estimation algorithm exploits feet to probe the terrain inclination and deals with underactuation when feet tilt on the contact surface. The equilibrium position in the admittance control is thereby adapted for recovering balance on the slope. Both the theoretical work and experimental validation are presented. The method is implemented and validated on the real humanoid by demonstrating the capability of estimating terrain inclination, balancing on the slope with varying gradient, and maintaining upright posture in the meantime. Experimental data such as inclination estimation in the comparison study, center of pressure measurement, and body attitude compensation are presented and analyzed. Zhibin Li 0001, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
IROS | 1 |
| 2013 | Optimal ankle compliance regulation for humanoid balancing controlabstractKeeping balance is the main concern for humanoids in standing and walking tasks. This paper endeavors to acquire optimal ankle stabilization methods for humanoids with passive and active compliance and explain ankle balancing strategy from the compliance regulation perspective. Unlike classical stiff humanoids, the compliant ones can control both impedance and position during task operation. Optimal compliance regulation is resolved to maximize the stability of the humanoids. The linearized model is proposed to obtain the optimal ankle impedance for stabilizing against impacts. The nonlinear model is proposed as well and compared with the linear one. The proposed methods are validated by experiments on an intrinsically compliant humanoid using passivity based admittance and impedance controllers both in joint and Cartesian space. Mohamad Mosadeghzad, Zhibin Li 0001, Nikolaos G. Tsagarakis, Gustavo A. Medrano-Cerda, Houman Dallali, Darwin G. Caldwell |
IROS | 2 |
| 2012 | Walking trajectory generation for humanoid robots with compliant joints: Experimentation with COMAN humanoidabstractThis work introduces a walking pattern generator suitable for humanoids with inherent joint compliance. The proposed walking pattern generator computes the desired center of mass (COM) references on-line based on the COM state feedback. The position and velocity of the COM are the feedback variables, and the constraint ground reaction force (GRF), which is limited by the support polygon, is the control effort to drive the COM states to track the desired ones. The zero moment point (ZMP) is obtained naturally as a result of GRF interaction with robot feet. The proposed COM tracking scheme demands a lower bandwidth from the controller compared to the ZMP tracking schemes. Experimental data of the real compliant humanoid, such as ZMP, COM motion, and GRF are presented to demonstrate the validation of the proposed gait generation method. Zhibin Li 0001, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
ICRA | 1 |
| 2012 | Stabilization for the compliant humanoid robot COMAN exploiting intrinsic and controlled complianceabstractThe work presents the standing stabilization of a compliant humanoid robot against external force disturbances and variations of the terrain inclination. The novel contribution is the proposed control scheme which consists of three strategies named compliance control in the transversal plane, body attitude control, and potential energy control, all combined with the intrinsic passive compliance in the robot. The physical compliant elements of the robot are exploited to react at the first instance of the impact while the active compliance control is applied to further absorb the impact and dissipate the elastic energy stored in springs preventing the high rate of spring recoil. The body attitude controller meanwhile regulates the spin angular momentum to provide more agile reactions by changing body inclination. The potential energy control module constrains the robot center of mass (COM) in a virtual slope to convert the excessive kinetic energy into potential energy to prevent falling. Experiments were carried out with the proposed balance stabilization control demonstrating superior balance performance. The compliant humanoid was capable of recovering from external force disturbances and moderate or even abrupt variations of the terrain inclination. Experimental data such as the impulse forces, real COM, center of pressure (COP) and the spring elastic energy are presented and analyzed. Zhibin Li 0001, Bram Vanderborght, Nikolaos G. Tsagarakis, Luca Colasanto, Darwin G. Caldwell |
ICRA | 1 |
| 2012 | Internal model control for improving the gait tracking of a compliant humanoid robotabstractThis paper reports on the modelling and trajectory generation of an intrinsically compliant humanoid robot. To achieve adequate gait tracking performance in a compliant robot is not trivial and cannot be addressed with the traditional control approaches used for stiff robots. To permit the development of effective gait generators which take into account the additional dynamic effects due to intrinsic compliance, an appropriate model which can predict the robot motion dynamics is required. In this work, we propose a model which combines the inverted pendulum model approach with a compliant model (Cartesian) at the level of the COM. Based on this model which permits to predict the motion of the centre of mass (COM) of the compliant robot an Internal Model Control strategy is adopted to improve the gait tracking performance. The derivation of the model is introduced followed by experimental validation which demonstrates the tracking performance achieved by the proposed reduced model. The Internal Model Control is subsequently discussed and validated on the COmpliant huMANoid COMAN using a series of ZMP based walking gaits. Luca Colasanto, Nikolaos G. Tsagarakis, Zhibin Li 0001, Darwin G. Caldwell |
IROS | 3 |
| 2011 | The design of the lower body of the compliant humanoid robot "cCub"abstractThe “iCub ”is a robotic platform that was developed by the RobotCub [1] consortium to provide the cognition research community with an open “child-like ”humanoid platform for understanding and development of cognitive systems [1]. In this paper we present the mechanical realization of the lower body developed for the “cCub ”humanoid robot, a derivative of the original “iCub”, which has passive compliance in the major joints of the legs. It is hypothesized that this will give to the robot high versatility to cope with unpredictable disturbance ranging from small uneven terrain variations to unexpected collisions or even accidental falls. As part of the AMARSI European project, the passive compliance of this newly developed robot will be exploited for safer interaction, energy efficient and more aggressive damage-safe learning. The passive compliant actuation module used is a compact unit based on the series elastic actuator principle (SEA). In addition to the passive compliance the “cCub ”design includes other significant updates over the original prototype such as full joint state sensing including joint torque sensing and improved range of motion and torque capabilities. In this paper, the new leg mechanisms of the “cCub ”robot are introduced. Nikolaos G. Tsagarakis, Zhibin Li 0001, Jody Alessandro Saglia, Darwin G. Caldwell |
ICRA | 2 |
| 2010 | Trajectory generation of straightened knee walking for humanoid robot iCubabstractMost humanoid robots walk with bent knees, which particularly requires high motor torques at knees and gives an unnatural walking manner. It is therefore essential to design a control method that produces a motion which is more energy efficient and natural comparable to those performed by humans. In this paper, we address this issue by modeling the virtual spring-damper based on the cart-table model. This strategy utilizes the preview control, which generates the desired horizontal motion of the center of mass (COM), and the virtual spring-damper for generating the vertical COM motion. The theoretical feasibility of this hybrid strategy is demonstrated in Matlab simulation of a multi-body bipedal model. Knee joint patterns, ground reaction force (GRF) patterns, COM trajectories are presented. The successful walking gaits of the child humanoid "iCub" in the dynamic simulator validate the proposed scheme. The joint torques required by the proposed strategy are reduced, compared with the one required by the cart-table model. Zhibin Li 0001, Nikolaos G. Tsagarakis, Darwin G. Caldwell, Bram Vanderborght |
ICARCV | 1 |