Anqing Duan

dblp:212/2457 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-9666-018XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Safe online reinforcement learning with diffusion world model and Langevin dynamics
Yuanda Wang, Changyin Sun 0001, Anqing Duan
Expert Syst. Appl.4
2025 Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile Manipulation
abstract
Enabling humanoid robots to perform long-horizon mobile manipulation planning in real-world environments based on embodied perception and comprehension abilities has been a longstanding challenge. With the recent rise of large language models (LLMs), there has been a notable increase in the development of LLM-based planners. These approaches either utilize human-provided textual representations of the real world or heavily depend on prompt engineering to extract such representations, lacking the capability to quantitatively understand the environment, such as determining the feasibility of manipulating objects. To address these limitations, we present the Instruction-Augmented Long-Horizon Planning (IALP) system, a novel framework that employs LLMs to generate feasible and optimal actions based on real-time sensor feedback, including grounded knowledge of the environment, in a closed-loop interaction. Distinct from prior works, our approach augments user instructions into PDDL problems by leveraging both the abstract reasoning capabilities of LLMs and grounding mechanisms. By conducting various real-world long-horizon tasks, each consisting of seven distinct manipulatory skills, our results demonstrate that the IALP system can efficiently solve these tasks with an average success rate exceeding 80%. Our proposed method can operate as a high-level planner, equipping robots with substantial autonomy in unstructured environments through the utilization of multi-modal sensor inputs.
Fangyuan Wang 0002, Shipeng Lyu, Peng Zhou 0018, Anqing Duan, Guodong Guo, David Navarro-Alarcon
AAAI4
2025 Towards Safe Imitation Learning via Potential Field-Guided Flow Matching
abstract
Deep generative models, particularly diffusion and flow matching models, have recently shown remarkable potential in learning complex policies through imitation learning. However, the safety of generated motions remains overlooked, particularly in complex environments with inherent obstacles. In this work, we address this critical gap by proposing Potential Field-Guided Flow Matching Policy (PF2MP), a novel approach that simultaneously learns task policies and extracts obstacle-related information, represented as a potential field, from the same set of successful demonstrations. During inference, PF2MP modulates the flow matching vector field via the learned potential field, enabling safe motion generation. By leveraging these complementary fields, our approach achieves improved safety without compromising task success across diverse environments, such as navigation tasks and robotic manipulation scenarios. We evaluate PF2MP in both simulation and real-world settings, demonstrating its effectiveness in task space and joint space control. Experimental results demonstrate that PF2MP enhances safety, achieving a significant reduction of collisions compared to baseline policies. This work paves the way for safer motion generation in unstructured and obstacle-rich environments.
Anqing Duan, Zezhou Sun, Leonel Rozo, Noémie Jaquier, Dezhen Song, Yoshihiko Nakamura
IROS2
2025 Human-in-the-Loop Robot Learning for Smart Manufacturing: A Human-Centric Perspective
abstract
Robot learning has attracted an ever-increasing attention by automating complex tasks, reducing errors, and increasing production speed and flexibility, which leads to significant advancements in manufacturing intelligence. However, its low training efficiency, limited real-time feedback, and challenges in adapting to untrained scenarios hinder its applications in smart manufacturing. Introducing a human role in the training loop, a practice known as human-in-the-loop (HITL) robot learning, can improve the performance of robots by leveraging human prior knowledge. Nonetheless, the exploration of HITL robot learning within the context of human-centric smart manufacturing remains in its infancy. This study provides a holistic literature review for understanding HITL robot learning within an industrial context from a human-centric perspective. A united structure is presented to encompass different aspects of human intelligence in HITL robot learning, highlighting perception, cognition, behavior, and notably, empathy. Then, the typical applications in manufacturing scenarios are analyzed to expand the research landscape for smart manufacturing. Finally, it introduces the empirical challenges and future directions for HITL robot learning in the next industrial revolution era.
Hongpeng Chen, Junming Fan, Anqing Duan, Chenguang Yang 0001, David Navarro-Alarcon, Pai Zheng
IEEE Trans Autom. Sci. Eng.4
2025 Safe Learning by Constraint-Aware Policy Optimization for Robotic Ultrasound Imaging
abstract
Ultrasound-based medical examination usually requires establishing proper contact between an ultrasound probe and a human body that ensures the quality of ultrasound images. The scanning skills are quite challenging for a robot to learn primarily due to the complex coupling between the applied force profile and the resulting ultrasound image quality. While reinforcement learning appears as a powerful tool for learning complex robot skills, the deployment of these algorithms in medical robots demands special attention due to the evident safety concerns that arise from physical probe-tissue interactions. In this paper, we explicitly consider external constraints on the force magnitude when searching for the optimal policy parameters to enhance safety during ultrasound-guided robotic interventions. In particular, we study policy optimization under the framework of a constrained Markov decision process. The resulting gradient-based policy update is then subject to the involved constraints, which can be readily addressed by the primal-dual interior-point technique. In addition, upon the observation that policy update requires consecutive policies to be close to each other to have stable and robust performance with reinforcement learning algorithms, we design the learning rate of policy gradient from an imitation perspective. The performance of the proposed constraint-aware policy optimization method is validated with experiments of robotic ultrasound imaging for spinal diagnosisNote to Practitioners—This paper was motivated by the problem of safely learning the optimal interaction force strategy to facilitate robotic ultrasound imaging. Existing approaches to robotic ultrasound imaging usually empirically set a constant value for the scanning force, despite the fact the force strategy plays an important role in the quality of the ultrasound images. This paper suggests the usage of reinforcement learning to identify the optimal interaction force due to the complex acoustic coupling between the force and the ultrasound image quality. Specifically, we propose constraint-aware reinforcement learning in view of the safety-critical issues as a result of physical human-probe interaction. We then conduct a theoretical analysis of the proposed safe reinforcement learning, including monotonic improvement and policy value bound under mild assumptions. Preliminary real experiments on ultrasound imaging of the spine of a phantom for scoliosis assessment suggest that the proposed approach can safely learn the optimal scanning force without violating the prescribed force threshold. In the future, we would like to apply our approach to learning the optimal scanning force on different organs of interest of human subjects.
Anqing Duan, Chenguang Yang 0001, Shengzeng Huo, Peng Zhou 0018, Wanyu Ma, David Navarro-Alarcon
IEEE Trans Autom. Sci. Eng.1
2025 Explicit-Implicit Subgoal Planning for Long-Horizon Tasks With Sparse Rewards
abstract
The challenges inherent in long-horizon tasks in robotics persist due to the typical inefficient exploration and sparse rewards in traditional reinforcement learning approaches. To address these challenges, we have developed a novel algorithm, termed hlexplicit-implicit subgoal planning (EISP), designed to tackle long-horizon tasks through a divide-and-conquer approach. We utilize two primary criteria, feasibility and optimality, to ensure the quality of the generated subgoals. EISP consists of three components: a hybrid subgoal generator, a hindsight sampler, and a value selector. The hybrid subgoal generator uses an explicit model to infer subgoals and an implicit model to predict the final goal, inspired by way of human thinking that infers subgoals by using the current state and final goal as well as reason about the final goal conditioned on the current state and given subgoals. Additionally, the hindsight sampler selects valid subgoals from an offline dataset to enhance the feasibility of the generated subgoals. While the value selector utilizes the value function in reinforcement learning to filter the optimal subgoals from subgoal candidates. To validate our method, we conduct four long-horizon tasks in both simulation and the real world. The obtained quantitative and qualitative data indicate that our approach achieves promising performance compared to other baseline methods. These experimental results can be seen on the website https://sites.google.com/view/vaesi.
Fangyuan Wang 0002, Anqing Duan, Peng Zhou 0018, Shengzeng Huo, Guodong Guo, Chenguang Yang 0001, David Navarro-Alarcon
IEEE Trans Autom. Sci. Eng.2
2025 Human-Aware Reactive Task Planning of Sequential Robotic Manipulation Tasks
abstract
The recent emergence of Industry 5.0 underscores the need for increased autonomy in human–robot interaction (HRI), presenting both motivation and challenges in achieving resilient and energy-efficient production systems. To address this, in this article, we introduce a strategy for seamless collaboration between humans and robots in manufacturing and maintenance tasks. Our method enables smooth switching between temporary HRI (human-aware mode) and long-horizon automated manufacturing (fully automatic mode), effectively solving the human–robot coexistence problem. We develop a task progress monitor that decomposes complex tasks into robot-centric action sequences, further divided into three-phase subtasks. A trigger signal orchestrates mode switches based on detected human actions and their contribution to the task. In addition, we introduce a human agent coefficient matrix, computed using selected environmental features, to determine cut-points for reactive execution by each robot. To validate our approach, we conducted extensive experiments involving robotic manipulators performing representative manufacturing tasks in collaboration with humans. The results show promise for advancing HRI, offering pathways to enhancing sustainability within Industry 5.0. Our work lays the foundation for intelligent manufacturing processes in future societies, marking a pivotal step toward realizing the full potential of human–robot collaboration.
Wanyu Ma, Anqing Duan, Hoi-Yin Lee, Pai Zheng, David Navarro-Alarcon
IEEE Trans. Ind. Informatics2
2025 Learning Rhythmic Trajectories With Geometric Constraints for Laser-Based Skincare Procedures
abstract
The increasing deployment of robots has significantly enhanced the automation levels across a wide and diverse range of industries. This article investigates the automation challenges of laser-based dermatology procedures in the beauty industry. This group of related manipulation tasks involves delivering energy from a cosmetic laser onto the skin with repetitive patterns. To automate this procedure, we propose to use a robotic manipulator and endow it with the dexterity of a skilled dermatology practitioner through a learning-from-demonstration framework. To ensure that the cosmetic laser can properly deliver the energy onto the skin surface of an individual, we develop a novel structured prediction-based imitation learning algorithm with the merit of handling geometric constraints. Notably, our proposed algorithm effectively tackles the imitation challenges associated with quasi-periodic motions, a common feature of many laser-based cosmetic tasks. The conducted real-world experiments illustrate the performance of our robotic beautician in mimicking realistic dermatological procedures. Our new method is shown to not only replicate the rhythmic movements from the provided demonstrations but also to adapt the acquired skills to previously unseen scenarios and subjects.
Anqing Duan, Wanli Liuchen, Raffaello Camoriano, Lorenzo Rosasco, David Navarro-Alarcon
IEEE Trans. Robotics1
2024 Imitating Tool-Based Garment Folding From a Single Visual Observation Using Hand-Object Graph Dynamics
abstract
Garment folding is a ubiquitous domestic task that is difficult to automate due to the highly deformable nature of fabrics. In this article, we propose a novel method of learning from demonstrations that enables robots to autonomously manipulate an assistive tool to fold garments. In contrast to traditional methods (that rely on low-level pixel features), our proposed solution uses a dense visual descriptor to encode the demonstration into a high-levelhand-object graph(HoG) that allows to efficiently represent the interactions between the manipulated tool and robots. With that, we leverage graph neural network to autonomously learn the forward dynamics model from HoGs, then, given only a single demonstration, the imitation policy is optimized with a model predictive controller to accomplish the folding task. To validate the proposed approach, we conducted a detailed experimental study on a robotic platform instrumented with vision sensors and a custom-made end-effector that interacts with the folding board.
Peng Zhou 0018, Jiaming Qi, Anqing Duan, Shengzeng Huo, David Navarro-Alarcon
IEEE Trans. Ind. Informatics3
2023 Fourier-Based Multi-Agent Formation Control to Track Evolving Closed Boundaries
abstract
The automatic monitoring/tracking of environmental boundaries by multi-agent systems is a fundamental problem that has many practical applications. In this paper, we address this problem with formation control techniques based on parame tric curves that represent the boundary’s feedback shape. For that, we approximate the curve with truncated Fourier series, whose finite coefficients are utilized to characterize the curve’s shape and to automatically distribute the agents along it. These feedback Fourier coefficients are exploited to design a new type of formation controller that drives the agents to form desired curves. A detailed stability analysis is provided for the proposed control methodology, considering both fixed and switching multi-agent topologies. The reported numerical simulation and experimental studies demonstrate the performance and feasibility of our new method to track closed boundaries of different shapes.
José Guadalupe Romero, Luiza Labazanova, Anqing Duan, Xiang Li 0009, David Navarro-Alarcon
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 A Multisensor Interface to Improve the Learning Experience in Arc Welding Training Tasks
abstract
This article presents the development of a multisensor user interface to facilitate the instruction of arc welding tasks. Traditional methods to acquire hand-eye coordination skills are typically conducted through one-to-one instruction, where trainees must wear protective helmets and conduct several tests. These approaches are inefficient as the harmful light emitted from the electric arc impedes the close monitoring of the process. Practitioners can only observe a small bright spot. To tackle these problems, recent training approaches have leveraged virtual reality to safely simulate the process and visualize the geometry of the workpieces. However, the synthetic nature of these types of simulation platforms reduces their effectiveness as they fail to comprise actual welding interactions with the environment, which hinders the trainees' learning process. To provide users with a real welding experience, we have developed a new multisensor extended reality platform for arc welding training. Our system is composed of: 1) An HDR camera, monitoring the real welding spot in real time. 2) A depth sensor, capturing the 3-D geometry of the scene; and 3) A head-mounted VR display, visualizing the process safely. Our innovative platform provides users with a “bot trainer,” virtual cues of the seam geometry, automatic spot tracking, and performance scores. To validate the platform's feasibility, we conduct extensive experiments with several welding training tasks. We show that compared with the traditional training practice and recent virtual reality approaches, our automated multisensor method achieves better performances in terms of accuracy, learning curve, and effectiveness.
Hoi-Yin Lee, Peng Zhou 0018, Anqing Duan, Jiangliu Wang, Victor Wu, David Navarro-Alarcon
IEEE Trans. Hum. Mach. Syst.3
2019 Learning to Sequence Multiple Tasks with Competing Constraints
abstract
Imitation learning offers a general framework where robots can efficiently acquire novel motor skills from demonstrations of a human teacher. While many promising achievements have been shown, the majority of them are only focused on single-stroke movements, without taking into account the problem of multi-tasks sequencing. Conceivably, sequencing different atomic tasks can further augment the robot's capabilities as well as avoid repetitive demonstrations. In this paper, we propose to address the issue of multi-tasks sequencing with emphasis on handling the so-called competing constraints, which emerge due to the existence of the concurrent constraints from Cartesian and joint trajectories. Specifically, we explore the null space of the robot from an information-theoretic perspective in order to maintain imitation fidelity during transition between consecutive tasks. The effectiveness of the proposed method is validated through simulated and real experiments on the iCub humanoid robot.
Anqing Duan, Raffaello Camoriano, Diego Ferigo, Daniele Calandriello, Lorenzo Rosasco, Daniele Pucci
IROS1