Weiji Xie

dblp:371/8889 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 67% Legged, aerial and field robots · 14% Motion planning and robot control · 14%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning › self-play
fictitious self-play
1.012026
Offline Fictitious Self-Play for Competitive Games · AAAI 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
Offline Fictitious Self-Play for Competitive Games · AAAI 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning
offline multi-agent reinforcement learning
1.012026
Offline Fictitious Self-Play for Competitive Games · AAAI 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play
1.012026
Offline Fictitious Self-Play for Competitive Games · AAAI 2026
Robotics › Legged, aerial and field robots › legged robots
humanoid locomotion
0.912025
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills · NeurIPS 2025
Robotics › Motion planning and robot control
whole-body control
0.912025
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills · NeurIPS 2025
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium
0.312026
Offline Fictitious Self-Play for Competitive Games · AAAI 2026
Robotics › Robot manipulation › learning from demonstration
motion imitation
0.312025
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

importance sampling · 2.0best response learning · 2.0motion retargeting · 0.9bi-level optimization · 0.9asymmetric actor-critic · 0.9
YearPublicationVenuePosition
2026 Offline Fictitious Self-Play for Competitive Games
abstract
Offline Reinforcement Learning (RL) enables policy improvement from fixed datasets without online interactions, making it highly suitable for real-world applications lacking efficient simulators. Despite its success in the single-agent setting, offline multi-agent RL remains a challenge, especially in competitive games. Firstly, unaware of the game structure, it is impossible to interact with the opponents and conduct a major learning paradigm, self-play, for competitive games. Secondly, real-world datasets cannot cover all the state and action space in the game, resulting in barriers to identifying Nash equilibrium (NE). To address these issues, this paper introduces Off-FSP, the first practical model-free offline RL algorithm for competitive games. We start by simulating interactions with various opponents by adjusting the weights of the fixed dataset with importance sampling. This technique allows us to learn the best responses to different opponents and employ the Offline Self-Play learning framework. To overcome the challenge of partial coverage, we combine the single-agent offline RL method with Fictitious Self-Play (FSP) to approximate NE by constraining the approximate best responses away from out-of-distribution actions. Experiments on matrix games, extensive-form poker, and board games demonstrate that Off-FSP achieves significantly lower exploitability than state-of-the-art baselines. Finally, we validate Off-FSP on a real-world human-robot competitive task, demonstrating its potential for solving complex, hard-to-simulate real-world problems.
Jingxiao Chen, Weiji Xie, Weinan Zhang 0001, Yong Yu 0001, Ying Wen 0001
AAAI2
2025 LoopSR: Looping Sim-and-Real for Lifelong Policy Adaptation of Legged Robots
abstract
Reinforcement Learning (RL) has shown its remarkable and generalizable capability in legged locomotion through sim-to-real transfer. However, while adaptive methods like domain randomization are expected to enhance policy robustness across diverse environments, they potentially compromise the policy’s performance in any specific environment, leading to suboptimal real-world deployment due to the No Free Lunch Theorem. To address this, we propose LoopSR, a lifelong policy adaptation framework that continuously refines RL policies in the post-deployment stage. LoopSR employs a transformer-based encoder to map real-world trajectories into a latent space and reconstruct a digital twin of the real world for further improvement. Autoencoder architecture and contrastive learning methods are adopted to enhance feature extraction of real-world dynamics. Simulation parameters for continual training are derived by combining predicted values from the decoder with retrieved parameters from a pre-collected simulation trajectory dataset. By leveraging simulated continual training, LoopSR achieves superior data efficiency compared with strong baselines, yielding eminent performance with limited data in both sim-to-sim and sim-to-real experiments.
Weiji Xie, Jiahang Cao, Hang Lai, Weinan Zhang 0001
IROS2
2025 Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement Learning
abstract
Humans possess delicate dynamic balance mechanisms that enable them to maintain stability across diverse terrains and under extreme conditions. However, despite significant advances recently, existing locomotion algorithms for humanoid robots are still struggle to traverse extreme environments, especially in cases that lack external perception (e.g., vision or LiDAR). This is because current methods often rely on gait-based or perception-condition rewards, lacking effective mechanisms to handle unobservable obstacles and sudden balance loss. To address this challenge, we propose a novel whole-body locomotion algorithm based on dynamic balance and Reinforcement Learning (RL) that enables humanoid robots to traverse extreme terrains, particularly narrow pathways and unexpected obstacles, using only proprioception. Specifically, we introduce a dynamic balance mechanism by leveraging a novel Zero Moment Point (ZMP)-driven reward and task-driven rewards in a whole-body actor-critic framework, aiming to achieve coordinated actions of the upper and lower limbs for robust locomotion. Experiments conducted on a full-sized Unitree H1-2 robot verify the ability of our method to maintain balance on extremely narrow terrains and under external disturbances, demonstrating its effectiveness in enhancing the robot's adaptability to complex environments. The videos are given at https://whole-body-loco.github.io.
Weiji Xie, Chenjia Bai, Jiyuan Shi, Yunfei Ge, Weinan Zhang 0001, Xuelong Li 0001
IROS1
2025 KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
abstract
Humanoid robots are promising to acquire various skills by imitating human behaviors. However, existing algorithms are only capable of tracking smooth, low-speed human motions, even with delicate reward and curriculum design. This paper presents a physics-based humanoid control framework, aiming to master highly-dynamic human behaviors such as Kungfu and dancing through multi-steps motion processing and adaptive motion tracking. For motion processing, we design a pipeline to extract, filter out, correct, and retarget motions, while ensuring compliance with physical constraints to the maximum extent. For motion imitation, we formulate a bi-level optimization problem to dynamically adjust the tracking accuracy tolerance based on the current tracking error, creating an adaptive curriculum mechanism. We further construct an asymmetric actor-critic framework for policy training. In experiments, we train whole-body control policies to imitate a set of highly dynamic motions. Our method achieves significantly lower tracking errors than existing approaches and is successfully deployed on the Unitree G1 robot, demonstrating stable and expressive behaviors. The project page is https://kungfubot.github.io.
Weiji Xie, Jinrui Han, Jiakun Zheng, Huanyu Li 0014, Jiyuan Shi, Weinan Zhang 0001, Chenjia Bai, Xuelong Li 0001
NeurIPS1