Hang Lai

dblp:269/9405 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-1000-3232ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning Embodied Quadruped Agents for Posture-Aware Locomotion
abstract
Recent advances in embodied intelligence have opened new directions for autonomous agents to operate in complex physical environments. In this work, we propose a Posture-Aware Locomotion Agent (PALA) trained via deep reinforcement learning. Unlike traditional quadruped agents that focus solely on velocity tracking, our agent learns to track task-oriented 6D motion commands, including linear and angular velocities, as well as desired body posture (height, pitch, and roll), in real time using only proprioceptive sensing and external commands. To improve robustness and terrain adaptability, we introduce two key heuristic designs: a progressive reward curriculum and an orientation command resampling strategy. Combined with asymmetric actor-critic training, adversarial motion priors, and domain randomization, these components enable a single policy to generalize zero-shot across diverse and challenging environments. Additionally, we extend PALA to handle high-level instructions by integrating it with an autonomous agent powered by a large language model (LLM), enabling natural language task descriptions to be directly translated into executable 6D commands. Extensive simulation and real-world experiments demonstrate the responsiveness and accuracy of our posture-aware locomotion agent, underscoring its potential as a core component for embodied systems operating in unstructured settings.
Xiangyu Miao, Hang Lai, Xinpeng Di, Jiahang Cao, Yong Yu 0001, Weinan Zhang 0001
DAI3
2025 World Model-Based Perception for Visual Legged Locomotion
abstract
Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often data-inefficient and intricate. To address this issue, traditional methods attempt to learn a teacher policy with access to privileged information first and then learn a student policy to imitate the teacher's behavior with visual input. Despite some progress, this imitation framework prevents the student policy from achieving optimal performance due to the information gap between inputs. Furthermore, the learning process is unnatural since animals intuitively learn to traverse different terrains based on their understanding of the world without privileged knowledge. Inspired by this natural ability, we propose a simple yet effective method, World Model-based Perception (WMP), which builds a world model of the environment and learns a policy based on the world model. We illustrate that though completely trained in simulation, the world model can make accurate predictions of real-world trajectories, thus providing informative signals for the policy controller. Extensive simulated and real-world experiments demonstrate that WMP outperforms state-of-the-art baselines in traversability and robustness. Videos and Code are available at: https://wmp-loco.github.io/.
Hang Lai, Jiahang Cao, Jiafeng Xu, Yunfeng Lin, Tao Kong, Yong Yu 0001, Weinan Zhang 0001
ICRA1
2025 LoopSR: Looping Sim-and-Real for Lifelong Policy Adaptation of Legged Robots
abstract
Reinforcement Learning (RL) has shown its remarkable and generalizable capability in legged locomotion through sim-to-real transfer. However, while adaptive methods like domain randomization are expected to enhance policy robustness across diverse environments, they potentially compromise the policy’s performance in any specific environment, leading to suboptimal real-world deployment due to the No Free Lunch Theorem. To address this, we propose LoopSR, a lifelong policy adaptation framework that continuously refines RL policies in the post-deployment stage. LoopSR employs a transformer-based encoder to map real-world trajectories into a latent space and reconstruct a digital twin of the real world for further improvement. Autoencoder architecture and contrastive learning methods are adopted to enhance feature extraction of real-world dynamics. Simulation parameters for continual training are derived by combining predicted values from the decoder with retrieved parameters from a pre-collected simulation trajectory dataset. By leveraging simulated continual training, LoopSR achieves superior data efficiency compared with strong baselines, yielding eminent performance with limited data in both sim-to-sim and sim-to-real experiments.
Weiji Xie, Jiahang Cao, Hang Lai, Weinan Zhang 0001
IROS4
2024 A survey on model-based reinforcement learning
Fan-Ming Luo, Tian Xu 0003, Hang Lai, Xiong-Hui Chen, Weinan Zhang 0001, Yang Yu 0001
Sci. China Inf. Sci.3
2023 Adversarially Trained Environment Models Are Effective Policy Evaluators and Improvers - An Application to Information Retrieval
abstract
The essence of information retrieval (IR) is to find the most useful information items (or documents) according to the user’s information need and present the items to the users in the form of a ranking list. The widely used evaluation metrics for a ranking list are NDCG, MAP, hit ratio etc., which are based on strong assumptions on the users’ examining and click behaviors when interacting with the ranking list. In modern IR scenarios, it has been shown that users’ behavior can be highly personalized, diverse and dynamic, which leads to the failure of those assumptions. Click models (CMs) are proposed to learn such complex user behaviors. However, most of existing works on CM still focus on fitting the logged behavior data while little attention has been paid to how CMs can both evaluate and improve rankers effectively. In this paper, we perform an in-depth investigation into how a CM could simultaneously evaluate and improve the ranking policy for IR. Specifically, we first make a theoretical analysis on the discrepancy of evaluated performance between a learned click model and the real user. Then based on the analysis, we discuss the principles of learning a good CM that could reduce such a discrepancy and accordingly propose a novel rank-oriented click model (RankCM). Furthermore, we conducted extensive experiments on how a CM can both evaluate and improve a ranker effectively based on four tasks, namely data fitting, click simulation, policy ranking and policy improvement, where RankCM demonstrates its comprehensive effectiveness and superiority over the compared CMs. Our study claims that adversarially trained environment models are effective policy evaluators and improvers, and this work could serve as an application to information retrieval. The innovations from the experiments as well as the future promising investigations are further discussed.
Yifan Liu 0008, Xinyi Dai, Jianghao Lin, Hang Lai, Yunfei Liu 0002, Yong Yu 0001
DAI5
2023 Adaptive Control Strategy for Quadruped Robots in Actuator Degradation Scenarios
abstract
Quadruped robots have strong adaptability to extreme environments but may also experience faults. Once these faults occur, robots must be repaired before returning to the task, reducing their practical feasibility. One prevalent concern among these faults is actuator degradation, stemming from factors like device aging or unexpected operational events. Traditionally, addressing this problem has relied heavily on intricate fault-tolerant design, which demands deep domain expertise from developers and lacks generalizability. Learning-based approaches offer effective ways to mitigate these limitations, but a research gap exists in effectively deploying such methods on real-world quadruped robots. This paper introduces a pioneering teacher-student framework rooted in reinforcement learning, named Actuator Degeneration Adaptation Transformer (Adapt), aimed at addressing this research gap. This framework produces a unified control strategy, enabling the robot to sustain its locomotion and perform tasks despite sudden joint actuator faults, relying exclusively on its internal sensors. Empirical evaluations on the Unitree A1 platform validate the deployability and effectiveness of Adapt on real-world quadruped robots, and affirm the robustness and practicality of our approach.
Hang Lai, Yong Yu 0001, Ying Wen 0001
DAI3
2023 Sim-to-Real Transfer for Quadrupedal Locomotion via Terrain Transformer
abstract
Deep reinforcement learning has recently emerged as an appealing alternative for legged locomotion over multiple terrains by training a policy in physical simulation and then transferring it to the real world (i.e., sim-to-real transfer). Despite considerable progress, the capacity and scalability of traditional neural networks are still limited, which may hinder their applications in more complex environments. In contrast, the Transformer architecture has shown its superiority in a wide range of large-scale sequence modeling tasks, including natural language processing and decision-making problems. In this paper, we propose Terrain Transformer (TERT), a high-capacity Transformer model for quadrupedal locomotion control on various terrains. Furthermore, to better leverage Transformer in sim-to-real scenarios, we present a novel two-stage training framework consisting of an offline pretraining stage and an online correction stage, which can naturally integrate Transformer with privileged training. Extensive experiments in simulation demonstrate that TERT outperforms state-of-the-art baselines on different terrains in terms of return, energy consumption and control smoothness. In further real-world validation, TERT successfully traverses nine challenging terrains, including sand pit and stair down, which can not be accomplished by strong baselines.
Hang Lai, Weinan Zhang 0001, Xialin He, Zheng Tian 0002, Yong Yu 0001, Jun Wang 0012
ICRA1
2023 Multi-embodiment Legged Robot Control as a Sequence Modeling Problem
abstract
Robots are traditionally bounded by a fixed embodiment during their operational lifetime, which limits their ability to adapt to their surroundings. Co-optimizing control and morphology of a robot, however, is often inefficient due to the complex interplay between the controller and morphology. In this paper, we propose a learning-based control method that can inherently take morphology into consideration such that once the control policy is trained in the simulator, it can be easily deployed to real robots with different embodiments. In particular, we present the Embodiment-aware Transformer (EAT), an architecture that casts this control problem as conditional sequence modeling. EAT outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired robot embodiment, past states, and actions, our EAT model can generate future actions that best fit the current robot embodiment. Experimental results show that EAT can outperform all other alternatives in embodiment-varying tasks, and succeed in an example of real-world evolution tasks: stepping down a stair through updating the morphology alone. We hope that EAT will inspire a new push toward real-world evolution across many domains, where algorithms like EAT can blaze a trail by bridging the field of evolutionary robotics and big data sequence modeling.
Weinan Zhang 0001, Hang Lai, Zheng Tian 0002, Laurent Kneip, Jun Wang 0012
ICRA3
2023 Adaptation Augmented Model-based Policy Optimization
abstract
Compared to model-free reinforcement learning (RL), model-based RL is often more sample efficient by leveraging a learned dynamics model to help decision making. However, the learned model is usually not perfectly accurate and the error will compound in multi-step predictions, which can lead to poor asymptotic performance. In this paper, we first derive an upper bound of the return discrepancy between the real dynamics and the learned model, which reveals the fundamental problem of distribution shift between simulated data and real data. Inspired by the theoretical analysis, we propose an adaptation augmented model-based policy optimization (AMPO) framework to address the distribution shift problem from the perspectives of feature learning and instance re-weighting, respectively. Specifically, the feature-based variant, namely FAMPO, introduces unsupervised model adaptation to minimize the integral probability metric (IPM) between feature distributions from real and simulated data, while the instance-based variant, termed as IAMPO, utilizes importance sampling to re-weight the real samples used to train the model. Besides model learning, we also investigate how to improve policy optimization in the model usage phase by selecting simulated samples with different probability according to their uncertainty. Extensive experiments on challenging continuous control tasks show that FAMPO and IAMPO, coupled with our model usage technique, achieves superior performance against baselines, which demonstrates the effectiveness of the proposed methods.
Jian Shen 0003, Hang Lai, Minghuan Liu, Han Zhao 0002, Yong Yu 0001, Weinan Zhang 0001
J. Mach. Learn. Res.2
2021 On Effective Scheduling of Model-based Reinforcement Learning
abstract
Model-based reinforcement learning has attracted wide attention due to its superior sample efficiency. Despite its impressive success so far, it is still unclear how to appropriately schedule the important hyperparameters to achieve adequate performance, such as the real data ratio for policy optimization in Dyna-style model-based algorithms. In this paper, we first theoretically analyze the role of real data in policy training, which suggests that gradually increasing the ratio of real data yields better performance. Inspired by the analysis, we propose a framework named AutoMBPO to automatically schedule the real data ratio as well as other hyperparameters in training model-based policy optimization (MBPO) algorithm, a representative running case of model-based methods. On several continuous control tasks, the MBPO instance trained with hyperparameters scheduled by AutoMBPO can significantly surpass the original one, and the real data ratio schedule found by AutoMBPO shows consistency with our theoretical analysis.
Hang Lai, Jian Shen 0003, Weinan Zhang 0001, Ruiming Tang, Yong Yu 0001, Zhenguo Li
NeurIPS1
2020 Bidirectional Model-based Policy Optimization
abstract
Model-based reinforcement learning approaches leverage a forward dynamics model to support planning and decision making, which, however, may fail catastrophically if the model is inaccurate. Although there are several existing methods dedicated to combating the model error, the potential of the single forward model is still limited. In this paper, we propose to additionally construct a backward dynamics model to reduce the reliance on accuracy in forward model predictions. We develop a novel method, called Bidirectional Model-based Policy Optimization (BMPO) to utilize both the forward model and backward model to generate short branched rollouts for policy optimization. Furthermore, we theoretically derive a tighter bound of return discrepancy, which shows the superiority of BMPO against the one using merely the forward model. Extensive experiments demonstrate that BMPO outperforms state-of-the-art model-based methods in terms of sample efficiency and asymptotic performance.
Hang Lai, Jian Shen 0003, Weinan Zhang 0001, Yong Yu 0001
ICML1