Haonan Duan 0001

dblp:273/7767-1 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-3608-4467ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Robot manipulation · 38% Transfer learning and domain adaptation · 15% Motion planning and robot control · 13%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
grasping
1.022025
Learning Realistic and Reasonable Grasps for Anthropomorphic Hand in Cluttered Scenes · ICRA 2024
Reactive Human-to-Robot Dexterous Handovers for Anthropomorphic Hand · IEEE Trans. Robotics 2025
Machine learning › Transfer learning and domain adaptation › domain generalization
cross-area generalization
1.012026
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy · AAAI 2026
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.012026
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy · AAAI 2026
Machine learning › Generative modeling
diffusion model
0.912025
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy · ICCV 2025
Robotics › Motion planning and robot control
robot learning
0.912025
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy · ICCV 2025
Human-robot interaction › physical human-robot interaction
object handover
0.912025
Reactive Human-to-Robot Dexterous Handovers for Anthropomorphic Hand · IEEE Trans. Robotics 2025
Human-robot interaction
physical human-robot interaction
0.912025
Reactive Human-to-Robot Dexterous Handovers for Anthropomorphic Hand · IEEE Trans. Robotics 2025
Computer vision › 3D vision
shape matching
0.812024
Learning Human-Like Functional Grasping for Multifinger Hands From Few Demonstrations · IEEE Trans. Robotics 2024
Robotics › Robot manipulation › grasping › grasp planning
grasp selection
0.312025
Reactive Human-to-Robot Dexterous Handovers for Anthropomorphic Hand · IEEE Trans. Robotics 2025

Methods — techniques the papers use, named apart from their topics

motion planning · 1.7collision detection · 1.7closed-loop control · 1.7coordinate frame transformation · 1.0camera extrinsic calibration · 1.0in-context conditioning · 0.9diffusion transformer · 0.9point cloud learning · 0.8grasp optimization · 0.8contact modeling · 0.8
YearPublicationVenuePosition
2026 Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
abstract
Vision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces. Although training data are collected from diverse camera perspectives, the models typically predict end-effector poses within the robot base coordinate frame, resulting in spatial inconsistencies. To mitigate this limitation, we introduce the Observation-Centric VLA (OC-VLA) framework, which grounds action predictions directly in the camera observation space. Leveraging the camera’s extrinsic calibration matrix, OC-VLA transforms end-effector poses from the robot base coordinate system into the camera coordinate system, thereby unifying prediction targets across heterogeneous viewpoints. This lightweight, plug-and-play strategy ensures robust alignment between perception and action, substantially improving model resilience to camera viewpoint variations. The proposed approach is readily compatible with existing VLA architectures, requiring no substantial modifications. Comprehensive evaluations on both simulated and real-world robotic manipulation tasks demonstrate that OC-VLA accelerates convergence, enhances task success rates, and improves cross-view generalization.
Haonan Duan 0001, Haoran Hao 0003, Yu Qiao 0001, Jifeng Dai, Zhi Hou
AAAI2
2025 Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
abstract
While recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or continuous actions constrains adaptability to heterogeneous action spaces. We present Dita, a scalable framework that leverages Transformer architectures to directly denoise continuous action sequences through a unified multimodal diffusion process. Departing from prior methods that condition denoising on fused embeddings via shallow networks, Dita employs in-context conditioning -- enabling fine-grained alignment between denoised actions and raw visual tokens from historical observations. This design explicitly models action deltas and environmental nuances. By scaling the diffusion action denoiser alongside the Transformer's scalability, Dita effectively integrates cross-embodiment datasets across diverse camera perspectives, observation scenes, tasks, and action spaces. Such synergy enhances robustness against various variances and facilitates the successful execution of long-horizon tasks. Evaluations across extensive benchmarks demonstrate state-of-the-art or comparative performance in simulation. Notably, Dita achieves robust real-world adaptation to environmental variances and complex long-horizon tasks through 10-shot finetuning, using only third-person camera inputs. The architecture establishes a versatile, lightweight and open-source baseline for generalist robot policy learning. Project Page: https://robodita.github.io.
Zhi Hou, Yuwen Xiong, Haonan Duan 0001, Hengjun Pu, Ronglei Tong, Chengyang Zhao, Xizhou Zhu, Yu Qiao 0001, Jifeng Dai, Yuntao Chen
ICCV4
2025 Reactive Human-to-Robot Dexterous Handovers for Anthropomorphic Hand
abstract
Human-robot object handovers are essential for robots to effectively serve human needs in various domains of human–robot interaction and collaboration, yet remain a significant challenge. Remarkable progress has been made by parallel-jaw gripper robots in grasp generation and motion planning for handovers, while few studies address this issue using anthropomorphic hands, necessitating the ability to handle higher collision probabilities and lower approaching space under occluded situations. In this article, we present a reactive human-to-robot dexterous handover framework for anthropomorphic hands. The closed-loop framework employs an effective collision detection and grasp selection approach to ensure safe and smooth motion in unstructured environments. We implement a handover system using a UR5 robot arm and a Schunk SVH Hand based on the presented framework, which can react to human motion during the handover process, generalize to diverse objects with different 6-DoF poses, and execute suitable grasp configurations. The generalizability, reliability, and efficiency of our method are demonstrated through the handover of 30 novel objects, a system ablation study for submodule evaluation, and a user study assessment involving eight participants.
Haonan Duan 0001, Peng Wang 0024, Daheng Li, Wei Wei 0062, Yongkang Luo 0001, Guoqiang Deng
IEEE Trans. Robotics1
2024 Learning Realistic and Reasonable Grasps for Anthropomorphic Hand in Cluttered Scenes
abstract
Grasping is one of the most fundamental skills for humans to interact with objects. However, it remains a challenging problem for anthropomorphic hands, due to the lack of object affordance understanding and high-dimensional grasp planning. In this work, we propose an anthropomorphic hand grasping framework to learn realistic and reasonable grasps in cluttered scenes, which tackles the problem in three items: 1) graspable point segmentation; 2) hand grasp generation and 3) grasp optimization. Specifically, our method generates high-quality hand grasps efficiently without complete object models by learning graspable points, associated grasp configurations from observed point cloud in a parallel manner and optimizing predicted grasps based on hand-object contacts. Simulation experiments show that our model generates physical plausible grasps for the anthropomorphic hand effectively with over 70% success rate. Real-world experiments demonstrate that the model trained in simulation performs satisfactorily in real-world scenarios for unseen objects.
Haonan Duan 0001, Daheng Li, Wei Wei 0062, Yayu Huang, Peng Wang 0024
ICRA1
2024 Learning Human-Like Functional Grasping for Multifinger Hands From Few Demonstrations
abstract
This article investigates the challenge of enabling multifinger hands to perform human-like functional grasping for various intentions. However, accomplishing functional grasping in real robot hands present many challenges, including handling generalization ability for kinematically diverse robot hands, generating intention-conditioned grasps for a large variety of objects, and incomplete perception from a single-view camera. In this work, we first propose a six-step functional grasp synthesis algorithm based on fine-grained contact modeling. With the fine-grained contact-based optimization and learned dense shape correspondence, the algorithm is adaptable to various objects of the same category and a wide range of multifinger hands using few demonstrations. Second, over 10 k functional grasps are synthesized to train our neural network, named DexFG-Net, which generates intention-conditioned grasps based on reconstructed object. Extensive experiments in the simulation and physical grasps indicate that the grasp synthesis algorithm can produce human-like functional grasp with robust stability and functionality, and the DexFG-Net can generate plausible and human-like intention-conditioned grasping postures for anthropomorphic hands.
Wei Wei 0062, Peng Wang 0024, Yongkang Luo 0001, Wanyi Li 0002, Daheng Li, Yayu Huang, Haonan Duan 0001
IEEE Trans. Robotics8