Yizheng Zhang

dblp:58/10545 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Motion planning and robot control · 32% Legged, aerial and field robots · 27% Robot manipulation · 27%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control
robot learning
1.522024
TossNet: Learning to Accurately Measure and Predict Robot Throwing of Arbitrary Objects in Real Time With Proprioceptive Sensing · IEEE Trans. Robotics 2024
Learning Highly Dynamic Behaviors for Quadrupedal Robots · ICRA 2024
Robotics › Robot manipulation › nonprehensile manipulation
dynamic manipulation
0.812024
TossNet: Learning to Accurately Measure and Predict Robot Throwing of Arbitrary Objects in Real Time With Proprioceptive Sensing · IEEE Trans. Robotics 2024
Robotics › Motion planning and robot control
dynamic modeling
0.812024
Relative Policy-Transition Optimization for Fast Policy Transfer · AAAI 2024
Robotics › Legged, aerial and field robots
humanoid robot
0.812024
HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid · NeurIPS 2024
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion
0.812024
Learning Highly Dynamic Behaviors for Quadrupedal Robots · ICRA 2024
Robotics › Robot manipulation
object rearrangement
0.812024
HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid · NeurIPS 2024
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
0.812024
Relative Policy-Transition Optimization for Fast Policy Transfer · AAAI 2024
Robotics › Robot manipulation › robot sensing
proprioceptive sensing
0.812024
TossNet: Learning to Accurately Measure and Predict Robot Throwing of Arbitrary Objects in Real Time With Proprioceptive Sensing · IEEE Trans. Robotics 2024
Robotics › Legged, aerial and field robots › legged robots › legged robot locomotion
quadruped locomotion
0.812024
Learning Highly Dynamic Behaviors for Quadrupedal Robots · ICRA 2024
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
0.212024
HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid · NeurIPS 2024
Robotics › Motion planning and robot control › robot control
learning control
0.212024
Learning Highly Dynamic Behaviors for Quadrupedal Robots · ICRA 2024
Robotics › Motion planning and robot control
robot control
0.212024
Learning Highly Dynamic Behaviors for Quadrupedal Robots · ICRA 2024
Robotics › Autonomous driving
trajectory prediction
0.212024
TossNet: Learning to Accurately Measure and Predict Robot Throwing of Arbitrary Objects in Real Time With Proprioceptive Sensing · IEEE Trans. Robotics 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.5teacher-student distillation · 0.8relative transition optimization · 0.8relative policy optimization · 0.8proprioceptive sensing · 0.8motion capture data · 0.8imitation learning · 0.8end-to-end learning · 0.8behavior cloning · 0.8adversarial motion priors · 0.8
YearPublicationVenuePosition
2024 Relative Policy-Transition Optimization for Fast Policy Transfer
abstract
We consider the problem of policy transfer between two Markov Decision Processes (MDPs). We introduce a lemma based on existing theoretical results in reinforcement learning to measure the relativity gap between two arbitrary MDPs, that is the difference between any two cumulative expected returns defined on different policies and environment dynamics. Based on this lemma, we propose two new algorithms referred to as Relative Policy Optimization (RPO) and Relative Transition Optimization (RTO), which offer fast policy transfer and dynamics modelling, respectively. RPO transfers the policy evaluated in one environment to maximize the return in another, while RTO updates the parameterized dynamics model to reduce the gap between the dynamics of the two environments. Integrating the two algorithms results in the complete Relative Policy-Transition Optimization (RPTO) algorithm, in which the policy interacts with the two environments simultaneously, such that data collections from two environments, policy and transition updates are completed in one closed loop to form a principled learning framework for policy transfer. We demonstrate the effectiveness of RPTO on a set of MuJoCo continuous control tasks by creating policy transfer problems via variant dynamics.
Yizheng Zhang, Baoxiang Wang 0001, Lei Han 0001
AAAI3
2024 Learning Highly Dynamic Behaviors for Quadrupedal Robots
abstract
Learning highly dynamic behaviors for robots has been a longstanding challenge. Traditional approaches have demonstrated robust locomotion, but the exhibited behaviors lack diversity and agility. They employ approximate models, which lead to compromises in performance. Data-driven approaches have been shown to reproduce agile behaviors of animals, but typically have not been able to learn highly dynamic behaviors. In this paper, we propose a learning-based approach to enable robots to learn highly dynamic behaviors from animal motion data. The learned controller is deployed on a quadrupedal robot and the results show that the controller is able to reproduce highly dynamic behaviors including sprinting, jumping and sharp turning. Various behaviors can be activated through human interaction using a stick with markers attached to it. Based on the motion pattern of the stick, the robot exhibits walking, running, sitting and jumping, much like the way humans interact with a pet.
Jiapeng Sheng, Tingguang Li, Qingxu Zhu 0001, Yizheng Zhang, Lei Han 0001
ICRA8
2024 HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid
abstract
Physical Human-Scene Interaction (HSI) plays a crucial role in numerous applications. However, existing HSI techniques are limited to specific object dynamics and privileged information, which prevents the development of more comprehensive applications. To address this limitation, we introduce HumanVLA for general object rearrangement directed by practical vision and language. A teacher-student framework is utilized to develop HumanVLA. A state-based teacher policy is trained first using goal-conditioned reinforcement learning and adversarial motion prior. Then, it is distilled into a vision-language-action model via behavior cloning. We propose several key insights to facilitate the large-scale learning process. To support general object rearrangement by physical humanoid, we introduce a novel Human-in-the-Room dataset encompassing various rearrangement tasks. Through extensive experiments and analysis, we demonstrate the effectiveness of our approach.
Yizheng Zhang, Yong-Lu Li 0001, Lei Han 0001, Cewu Lu
NeurIPS2
2024 TossNet: Learning to Accurately Measure and Predict Robot Throwing of Arbitrary Objects in Real Time With Proprioceptive Sensing
abstract
Accurate measuring and modeling of dynamic robot manipulation (e.g., tossing and catching) is particularly challenging, due to the inherent nonlinearity, complexity, and uncertainty in high-speed robot motions and highly dynamic robot–object interactions happening in very short distances and times. Most studies leverage extrinsic sensors such as visual and tactile feedback toward task or object-centric modeling of manipulation dynamics, which, however, may hit bottleneck due to the significant cost and complexity, e.g., the environmental restrictions. In this work, we investigate whether using solely the on-board proprioceptive sensory modalities can effectively capture and characterize dynamic manipulation processes. In particular, we present an object-agnostic strategy to learn the robot toss dynamics of arbitrary unknown objects from the spatio-temporal variations of robot toss movements and wrist-force/torque (F/T) observations. We then propose TossNet, an end-to-end formulation that jointly measures the robot toss dynamics and predicts the resulting flying trajectories of the tossed objects. Experimental results in both simulation and real-world scenarios demonstrate that our methods can accurately model the robot toss dynamics of both seen and unseen objects, and predict their flying trajectories with superior prediction accuracy in nearly real-time. Ablative results are also presented to demonstrate the effectiveness of each proprioceptive modality and their correlations in modeling the toss dynamics. Case studies show that TossNet can be applied on various real robot platforms for challenging tossing-centric robot applications, such as blind juggling and high-precise robot pitching.
Lipeng Chen, Weifeng Lu, Kun Zhang 0017, Yizheng Zhang, Yu Zheng 0001
IEEE Trans. Robotics4
2023 Learning Terrain-Adaptive Locomotion with Agile Behaviors by Imitating Animals
abstract
In this paper, we present a general learning framework for controlling a quadruped robot that can mimic the behavior of real animals and traverse challenging terrains. Our method consists of two steps: an imitation learning step to learn from motions of real animals, and a terrain adaptation step to enable generalization to unseen terrains. We capture motions from a Labrador on various terrains to facilitate terrain adaptive locomotion. Our experiments demonstrate that our policy can traverse various terrains and produce a natural-looking behavior. We deployed our method on the real quadruped robot$\boldsymbol{Max}$[1] via zero-shot simulation-to-reality transfer, achieving a speed of 1.1 m/s on stairs climbing.
Tingguang Li, Yizheng Zhang, Qingxu Zhu 0001, Jiapeng Sheng, Wanchao Chi, Lei Han 0001
IROS2
2022 RECCraft System: Towards Reliable and Efficient Collective Robotic Construction
abstract
This research presents a novel Collective Robotic Construction (CRC) system named RECCraft. The RECCraft hardware system is composed of the mobile manipulation vehicles, the cubic blocks, and the folding ramp blocks. Solid connection and easy removal of the blocks are achieved by an electropermanent magnet and silicon steel sheets. With one degree of freedom (DOF) lifting manipulator, the robot can carry a block 3.7 times its volume. An active folding ramp block can provide a robust passage to the upper level for the robot. Our study focuses on systemic improvement of the construction speed and reliability of the robotic construction system. Visual perception system realized by Apritag is adopted, featured by convenient deployment and high precision, to provide a reliable guarantee for robotic construction. RL-based planner provides end-to-end solution for planning tasks of building multi-layer constructions, which is validated by simulation platform and real prototype. Compared with construction speed of existing robotic construction systems, our proposed RECCraft system achieves state-of-the-art level. The robot builds a 2-layer construction by RL-based planner in 4 minutes and 16 seconds, which achieves construction volumetric throughput of 6.7×105mm3/s.
Qiwei Xu, Yizheng Zhang, Shenghao Zhang 0001, Zhuoxing Wu, Xiong Li 0001, Jiahong Chen, Zengjun Zhao, Luyang Tang, Zhengyou Zhang, Lei Han 0001
IROS2
2022 Time-of-Use Scheduling Problem with Equal-Length Jobs
Vincent Chau, Chenchen Fu, Weiwei Wu 0001, Yizheng Zhang
TAMC5