Haoqi Yuan

dblp:254/2084 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 54% Robot manipulation · 26% Language models and text generation · 10%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
dexterous manipulation
1.922026
Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations · AAAI 2026
Cross-Embodiment Dexterous Grasping with Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
hierarchical reinforcement learning
1.522024
RL-GPT: Integrating Reinforcement Learning and Code-as-policy · NeurIPS 2024
Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024
Robotics › Robot manipulation › dexterous manipulation
bimanual dexterous manipulation
1.012026
Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations · AAAI 2026
Robotics › Robot manipulation › grasping
multifingered grasping
0.912025
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping · ICLR 2025
Machine learning › Reinforcement learning › policy optimization
residual policy learning
0.912025
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping · ICLR 2025
Robotics › Motion planning and robot control
robot learning
0.912025
Cross-Embodiment Dexterous Grasping with Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
exploration
0.812024
Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online Adaptation · NeurIPS 2024
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy
0.812024
Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.812024
Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online Adaptation · NeurIPS 2024
Natural language and speech › Language models and text generation
LLM agents
0.812024
RL-GPT: Integrating Reinforcement Learning and Code-as-policy · NeurIPS 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online Adaptation · NeurIPS 2024
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization
0.812024
Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online Adaptation · NeurIPS 2024
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.812024
Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning · ICML 2022
Machine learning › Reinforcement learning
meta-reinforcement learning
0.612022
Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning · ICML 2022
Machine learning › Reinforcement learning › meta-reinforcement learning
offline meta-reinforcement learning
0.612022
Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning · ICML 2022
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
task representation learning
0.612022
Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning · ICML 2022
Robotics › Robot manipulation
grasping
0.312025
Cross-Embodiment Dexterous Grasping with Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.312025
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping · ICLR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning
0.212024
RL-GPT: Integrating Reinforcement Learning and Code-as-policy · NeurIPS 2024
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning
0.212024
Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.6teacher-student policy learning · 1.0knowledge distillation · 1.0vision-based policy · 0.9retargeting mapping · 0.9residual policy learning · 0.9mixture of experts · 0.9eigengrasp action space · 0.9goal clustering · 0.8behavior regularization · 0.8
YearPublicationVenuePosition
2026 Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations
abstract
Bimanual dexterous manipulation is a critical yet underexplored area in robotics. Its high-dimensional action space and inherent task complexity present significant challenges for policy learning, and the limited task diversity in existing benchmarks hinders general-purpose skill development. Existing approaches largely depend on reinforcement learning, often constrained by intricately designed reward functions tailored to a narrow set of tasks. In this work, we present a novel approach for efficiently learning diverse bimanual dexterous skills from abundant human demonstrations. Specifically, we introduce BiDexHD, a framework that unifies task construction from existing bimanual datasets and employs teacher-student policy learning to address all tasks. The teacher learns state-based policies using a general two-stage reward function across tasks with shared behaviors, while the student distills the learned multi-task policies into a vision-based policy. With BiDexHD, scalable learning of numerous bimanual dexterous skills from auto-constructed tasks becomes feasible, offering promising advances toward universal bimanual dexterous manipulation. Experiments on TACO tool-using dataset spanning 141 tasks across 6 categories demonstrate a task fulfillment rate of 74.59% on trained tasks and 51.07% on unseen tasks. We further transfer BiDexHD to 11 ARCTIC collaborative tasks and achieve an average of 80.49% task fulfillment rate on trained tasks and 65.99% on unseen task. All empirical results demonstrate the effectiveness and competitive zero-shot generalization capabilities of BiDexHD.
Haoqi Yuan, Yuhui Fu 0004, Zongqing Lu 0002
AAAI2
2025 Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping
abstract
Universal dexterous grasping across diverse objects presents a fundamental yet formidable challenge in robot learning. Existing approaches using reinforcement learning (RL) to develop policies on extensive object datasets face critical limitations, including complex curriculum design for multi-task learning and limited generalization to unseen objects. To overcome these challenges, we introduce ResDex, a novel approach that integrates residual policy learning with a mixture-of-experts (MoE) framework. ResDex is distinguished by its use of geometry-agnostic base policies that are efficiently acquired on individual objects and capable of generalizing across a wide range of unseen objects. Our MoE framework incorporates several base policies to facilitate diverse grasping styles suitable for various objects. By learning residual actions alongside weights that combine these base policies, ResDex enables efficient multi-task RL for universal dexterous grasping. ResDex achieves state-of-the-art performance on the DexGraspNet dataset comprising 3,200 objects with an 88.8% success rate. It exhibits no generalization gap with unseen objects and demonstrates superior training efficiency, mastering all tasks within only 12 hours on a single GPU. For further details and videos, visit our project page.
Ziye Huang, Haoqi Yuan, Yuhui Fu 0004, Zongqing Lu 0002
ICLR2
2025 Cross-Embodiment Dexterous Grasping with Reinforcement Learning
abstract
Dexterous hands exhibit significant potential for complex real-world grasping tasks. While recent studies have primarily focused on learning policies for specific robotic hands, the development of a universal policy that controls diverse dexterous hands remains largely unexplored. In this work, we study the learning of cross-embodiment dexterous grasping policies using reinforcement learning (RL). Inspired by the capability of human hands to control various dexterous hands through teleoperation, we propose a universal action space based on the human hand's eigengrasps. The policy outputs eigengrasp actions that are then converted into specific joint actions for each robot hand through a retargeting mapping. We simplify the robot hand's proprioception to include only the positions of fingertips and the palm, offering a unified observation space across different robot hands. Our approach demonstrates an 80\% success rate in grasping objects from the YCB dataset across four distinct embodiments using a single vision-based policy. Additionally, our policy exhibits zero-shot generalization to two previously unseen embodiments and significant improvement in efficient finetuning. For further details and videos, visit our project page (https://sites.google.com/view/crossdex).
Haoqi Yuan, Yuhui Fu 0004, Zongqing Lu 0002
ICLR1
2025 Creative Agents: Empowering Agents with Imagination for Creative Tasks
abstract
We study building embodied agents for open-ended creative tasks. While existing methods build instruction-following agents that can perform diverse open-ended tasks, none of them demonstrates creativity – the ability to give novel and diverse solutions implicit in the language instructions. This limitation comes from their inability to convert abstract language instructions into concrete goals and perform long-horizon planning for such complicated goals. Given the observation that humans perform creative tasks with imagination, we propose a class of solutions, where the controller is enhanced with an imaginator generating detailed imaginations of task outcomes conditioned on language instructions. We introduce several approaches to implementing the components of creative agents. We implement the imaginator with either a large language model for textual imagination or a diffusion model for visual imagination. The controller can either be a behavior-cloning policy or a pre-trained foundation model generating executable codes in the environment. We benchmark creative tasks with the challenging open-world game Minecraft, where the agents create diverse buildings given free-form language instructions. We propose novel evaluation metrics for open-ended creative tasks utilizing GPT-4V, which holds many advantages over existing metrics. We perform a detailed experimental analysis of creative agents, showing that creative agents are the first AI agents accomplishing diverse building creation in the survival mode of Minecraft. Our benchmark and models are open-source for future research on creative agents (https://github.com/PKU-RL/Creative-Agents).
Penglin Cai, Yuhui Fu 0004, Haoqi Yuan, Zongqing Lu 0002
UAI4
2024 Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning
abstract
Pre-training on task-agnostic large datasets is a promising approach for enhancing the sample efficiency of reinforcement learning (RL) in solving complex tasks. We present PTGM, a novel method that pre-trains goal-based models to augment RL by providing temporal abstractions and behavior regularization. PTGM involves pre-training a low-level, goal-conditioned policy and training a high-level policy to generate goals for subsequent RL tasks. To address the challenges posed by the high-dimensional goal space, while simultaneously maintaining the agent's capability to accomplish various skills, we propose clustering goals in the dataset to form a discrete high-level action space. Additionally, we introduce a pre-trained goal prior model to regularize the behavior of the high-level policy in RL, enhancing sample efficiency and learning stability. Experimental results in a robotic simulation environment and the challenging open-world environment of Minecraft demonstrate PTGM’s superiority in sample efficiency and task performance compared to baselines. Moreover, PTGM exemplifies enhanced interpretability and generalization of the acquired low-level skills.
Haoqi Yuan, Zhancun Mu, Feiyang Xie, Zongqing Lu 0002
ICLR1
2024 RL-GPT: Integrating Reinforcement Learning and Code-as-policy
abstract
Large Language Models (LLMs) have demonstrated proficiency in utilizing various tools by coding, yet they face limitations in handling intricate logic and precise control. In embodied tasks, high-level planning is amenable to direct coding, while low-level actions often necessitate task-specific refinement, such as Reinforcement Learning (RL). To seamlessly integrate both modalities, we introduce a two-level hierarchical framework, RL-GPT, comprising a slow agent and a fast agent. The slow agent analyzes actions suitable for coding, while the fast agent executes coding tasks. This decomposition effectively focuses each agent on specific tasks, proving highly efficient within our pipeline. Our approach outperforms traditional RL methods and existing GPT agents, demonstrating superior efficiency. In the Minecraft game, it rapidly obtains diamonds within a single day on an RTX3090. Additionally, it achieves SOTA performance across all designated MineDojo tasks.
Shaoteng Liu, Haoqi Yuan, Minda Hu, Yukang Chen, Shu Liu 0005, Zongqing Lu 0002, Jiaya Jia
NeurIPS2
2024 Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online Adaptation
abstract
Efficiently solving unseen tasks remains a challenge in reinforcement learning (RL), especially for long-horizon tasks composed of multiple subtasks. Pre-training policies from task-agnostic datasets has emerged as a promising approach, yet existing methods still necessitate substantial interactions via RL to learn new tasks. We introduce MGPO, a method that leverages the power of Transformer-based policies to model sequences of goals, enabling efficient online adaptation through prompt optimization. In its pre-training phase, MGPO utilizes hindsight multi-goal relabeling and behavior cloning. This combination equips the policy to model diverse long-horizon behaviors that align with varying goal sequences. During online adaptation, the goal sequence, conceptualized as a prompt, is optimized to improve task performance. We adopt a multi-armed bandit framework for this process, enhancing prompt selection based on the returns from online trajectories. Our experiments across various environments demonstrate that MGPO holds substantial advantages in sample efficiency, online adaptation performance, robustness, and interpretability compared with existing methods.
Haoqi Yuan, Yuhui Fu 0004, Feiyang Xie, Zongqing Lu 0002
NeurIPS1
2022 Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning
abstract
We study offline meta-reinforcement learning, a practical reinforcement learning paradigm that learns from offline data to adapt to new tasks. The distribution of offline data is determined jointly by the behavior policy and the task. Existing offline meta-reinforcement learning algorithms cannot distinguish these factors, making task representations unstable to the change of behavior policies. To address this problem, we propose a contrastive learning framework for task representations that are robust to the distribution mismatch of behavior policies in training and test. We design a bi-level encoder structure, use mutual information maximization to formalize task representation learning, derive a contrastive learning objective, and introduce several approaches to approximate the true distribution of negative pairs. Experiments on a variety of offline meta-reinforcement learning benchmarks demonstrate the advantages of our method over prior methods, especially on the generalization to out-of-distribution behavior policies.
Haoqi Yuan, Zongqing Lu 0002
ICML1
2021 DMotion: Robotic Visuomotor Control with Unsupervised Forward Model Learned from Videos
abstract
Learning an accurate model of the environment is essential for model-based control tasks. Existing methods in robotic visuomotor control usually learn from data with heavily labelled actions, object entities or locations, which can be demanding in many cases. To cope with this limitation, we propose a method, dubbed DMotion, that trains a forward model from video data only, via disentangling the motion of controllable agent to model the transition dynamics. An object extractor and an interaction learner are trained in an end-to-end manner without supervision. The agent’s motions are explicitly represented using spatial transformation matrices containing physical meanings. In the experiments, DMotion achieves superior performance on learning an accurate forward model in a Grid World environment, as well as a more realistic robot control environment in simulation. With the accurate learned forward models, we further demonstrate their usage in model predictive control as an effective approach for robotic manipulations. Code, video and more materials are available at: https://hyperplane-lab.github.io/dmotion.
Haoqi Yuan, Ruihai Wu, Andrew Zhao, Haipeng Zhang 0006, Hao Dong 0003
IROS1