Yiwei Wang 0002

dblp:50/5889-2 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-2135-5359ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TOOL-CURE: Tool Selection via Curriculum-Enhanced Reinforcement Learning with Sample Screening for LLMs
abstract
Large language models (LLMs) are increasingly deployed as intelligent agents capable of executing complex real-world tasks through external tool interactions, but effective tool selection remains challenging due to the inherent limitations of real-world training data. These datasets suffer from severe tool imbalance following long-tail distributions, data scarcity for specialized tools, logic conflicts between user queries and available tools, rapidly evolving toolsets, and the presence of subpar samples including partially correct and dirty examples. Existing supervised fine-tuning (SFT) approaches struggle with these multifaceted challenges as they require abundant high-quality data, treat all labeled examples as ground truth regardless of quality, and lack the flexibility to generalize beyond specific query-tool pairings seen during training. While reinforcement learning (RL) offers a promising alternative through outcome-based learning, vanilla approaches like Group Relative Policy Optimization (GRPO) suffer from training instability due to conflicting reward signals and inefficient learning from weak signals. To address these issues, we propose TOOL-CURE, a novel method with two key improvements to GRPO: Proficiency-Scaled Curriculum Learning (PSCL), which organizes training into a two-stage curriculum that builds foundational skills on easier samples before progressing to harder ones, and Online Policy Guarding via Sample Screening (OPGSS), which continuously assesses rollout quality and masks dirty samples to prevent noisy gradients from destabilizing policy updates. Our approach enables stable and efficient learning from heterogeneous real-world data, resulting in a robust tool-selection agent that demonstrates significant improvements in accuracy and generalization capability. Our code is available at: https://github.com/einnullnull/TOOL-CURE.git.
Jie Zhang 0115, Dongsheng Bi, Tao Sun 0018, Jian Wang 0108, Yiwei Wang 0002
WSDM6
2026 Task-Adaptive Analytical Affordance Estimation for Feature-Based Manipulation of Soft Tissues in Robotic Surgery
abstract
Robotic soft tissue manipulation in surgery presents significant challenges due to the tissue’s high deformability and the spatial constraints of the surgical environment. While data-driven methods for affordance estimation are common, they often face challenges with data requirements and generalization in surgical scenarios. To address these challenges, a novel framework is proposed that combines a deformation model-based shape controller with an analytical affordance estimation approach for multi-contact scenarios by employing a manipulability metric. The method leverages the deformation Jacobian matrix derived from a linearized deformation model to evaluate the manipulability of candidate multi-contact points, providing a robust and data-efficient solution. For tissue manipulation, a differentiable deformation model is employed to efficiently compute forward and backward deformation processes in real-time. Point-based visual features are constructed to represent and track the tissue deformation, enabling precise control through visual feedback. The proposed framework has been validated through simulations and physical experiments. These results demonstrate the real-time (∼ 60Hz) ability to achieve targeted configuration with high accuracy (< 1mm RMSE for each marker position) and a strong correlation (near-zero p-value) between the predicted affordance and the observed manipulation efficiency, confirming the effectiveness of our approach in soft tissue manipulation and affordance estimation.
Sihang Yang, Yiwei Wang 0002, Huan Zhao 0001, Han Ding 0002
IEEE Trans Autom. Sci. Eng.3
2025 Robust Robotic Breast Ultrasound Scanning and Real-Time Lesion Localization
abstract
The inherent flexibility and real-time deformation of breast tissue pose significant challenges for achieving full coverage and accurate lesion localization in autonomous breast ultrasound scanning. This paper introduces a robust finite state machine-based framework that mimics the decision-making process of an experienced physician, dynamically transitioning between the global breast scan and the fine lesion scan. An autonomous radial and anti-radial global scan pattern ensures comprehensive breast coverage. To avoid lesion misidentification caused by soft tissue movement, a real-time lesion fine scan method is proposed for lesion detection and localization. Experimental results demonstrate that the system in full coverage tests achieves 7 identified lesions out of 7 existing lesions and maintains a robust localization accuracy of$\mathbf{3. 2 3 ~ m m}$across phantoms with varying stiffnesses.
Zhiyan Cao, Yiwei Wang 0002, Huan Zhao 0001, Han Ding 0001
ICRA2
2025 Autonomous Bimanual Manipulation of Deformable Objects Using Deep Reinforcement Learning Guided Adaptive Control
abstract
Deformable object manipulation (DOM) which is a common subtask in various surgical procedures represents an inevitable challenge in robot-assisted surgery (RAS) due to complex nonlinear deformation. This paper proposes a deep reinforcement learning guided adaptive control (RLAC) modelfree framework, which combines learning-based and Jacobianbased methods. To complement each other for optimized performance, we harness the sampling of deep reinforcement learning (DRL) policy explored in simulations to solve a reasonable estimation of the initial deformation Jacobian. In early control iterations, the actions suggested by the DRL agent are adopted until the estimated real-time Jacobian approximates the actual deformation model. Subsequently, the independent Jacobianbased adaptive control (AC) with sufficient initial deformation awareness begins execution to achieve precise internal feature manipulation on deformable objects. Experimental results demonstrate that our method enables more efficient positioning and exhibits near-optimal positioning paths. RLAC with robust sim-to-real performance provides a feasible approach for the complex autonomous DOM in the real world.
Sihang Yang, Yiwei Wang 0002, Huan Zhao 0001, Han Ding 0001
ICRA3
2025 Leveraging Surgical Activity Grammar for Primary Intention Prediction in Laparoscopy Procedures
abstract
Surgical procedures are inherently complex and dynamic, with intricate dependencies and various execution paths. Accurate identification of the intentions behind critical actions, referred to as Primary Intentions (PIs), is crucial to understanding and planning the procedure. This paper presents a novel framework that advances PI recognition in instructional videos by combining top-down grammatical structure with bottom-up visual cues. The grammatical structure is based on a rich corpus of surgical procedures, offering a hierarchical perspective on surgical activities. A grammar parser, utilizing the surgical activity grammar, processes visual data obtained from laparoscopic images through surgical action detectors, ensuring a more precise interpretation of the visual information. Experimental results on the benchmark dataset demonstrate that our method outperforms existing surgical activity detectors that rely solely on visual features. Our research provides a promising foundation for developing advanced robotic surgical systems with enhanced planning and automation capabilities.
Jie Zhang 0115, Song Zhou, Yiwei Wang 0002, Chidan Wan, Huan Zhao 0001, Xiong Cai, Han Ding 0001
ICRA3
2025 Knowledge-Driven Framework for Anatomical Landmark Annotation in Laparoscopic Surgery
abstract
Accurate and reliable annotation of anatomical landmarks in laparoscopic surgery remains a challenge due to varying degrees of landmark visibility and changing shapes of human tissues during a surgical procedure in videos. In this paper, we propose a knowledge-driven framework that integrates prior surgical expertise with visual data to address this problem. Inspired by visual reasoning knowledge of tool-anatomy interactions, our framework models a spatio-temporal graph to represent the static topology of tool and tissue and dynamic transitions of landmarks' temporal behavior. By assigning explainable features of the surgical scene as node attributes in the graph, the surgical context is incorporated into the knowledge space. An attention-guided message passing mechanism across the graph dynamically adjusts the focus in different scenarios, enabling robust tracking of landmark states throughout the surgical process. Evaluations on the clinical dataset demonstrate the framework's ability to effectively use the inductive bias of explainable features to label landmarks, showing its potential in tackling intricate surgical tasks with improved stability and reliability.
Jie Zhang 0115, Song Zhou, Yiwei Wang 0002, Huan Zhao 0001, Han Ding 0001
IEEE Trans. Medical Imaging3
2025 Human-Like Robot Action Policy Through Game-Theoretic Intent Inference for Human-Robot Collaboration
abstract
Harmonious human-robot collaboration requires the robot to behave like a human partner, which raises the critical question of what factors make the robot do so. This paper proposes a series of policies based on empathetic and non-empathetic intent inference, proactive and reactive action planning, and ego and non-ego action styles to examine which modules enable robots to exhibit human-like behaviors. Two series of experiments are conducted with human subjects to test the performance of the proposed controllers. In Experiment 1, the participant must identify whether the collaborating partner is a human, similar to a Turing test. The classification results empirically verify that the designed empathetic proactive policies enable the robot to exhibit human-like behaviors. Experiment 2 indicates that the proposed policy can be applied to complex collaborative tasks, and this result is consistent with the findings of Experiment 1. From empirical evidence from the experiments, we believe that empathy and proactive policies are essential elements to enable robots to perform human-like actions.
Yubo Sheng, Yiwei Wang 0002, Haoyuan Cheng, Huan Zhao 0001, Han Ding 0001
IEEE Trans. Robotics2
2024 Vascular Centerline-Guided Autonomous Navigation Methods for Robot-Lead Endovascular Interventions
abstract
In minimally invasive endovascular interventional surgery, guidewire navigation is an indispensable process. However, even experienced physicians often encounter difficulties in manually manipulating the guidewire for branch selection, while also facing the risk of radiation exposure. In this study, we investigated robotic autonomous guidewire navigation methods. An electromagnetic system was used to track the real-time position and orientation of the guidewire tip, and a state space representing the guidewire within the vascular environment was constructed to guide the robot in precise guidewire manipulation. Experimental results demonstrated that the proposed trial-and-error and centerline-guided methods successfully completed navigation tasks in a static environment, outperforming human navigation performance in terms of trajectory smoothness, trajectory length, and incorrect branch entry counts. For dynamic environment navigation, dynamic time warping (DTW), a technique for measuring the similarity between two temporal sequences, was integrated into the centerline-guided method. The proposed approaches eliminate the need for visual feedback and thereby minimizing the risk of radiation exposure for both patients and medical staff present in the operating room during the procedure.
Naner Li, Yiwei Wang 0002, Haoyuan Cheng, Huan Zhao 0001, Han Ding 0001
ICRA2
2023 Laparoscopic Image-Based Critical Action Recognition and Anticipation With Explainable Features
abstract
Surgical workflow analysis integrates perception, comprehension, and prediction of the surgical workflow, which helps real-time surgical support systems provide proper guidance and assistance for surgeons. This article promotes the idea of critical actions, which refer to the essential surgical actions that progress towards the fulfillment of the operation. Fine-grained workflow analysis involves recognizing current critical actions and previewing the moving tendency of instruments in the early stage of critical actions. Aiming at this, we propose a framework that incorporates operational experience to improve the robustness and interpretability of action recognition in in-vivo situations. High-dimensional images are mapped into an experience-based explainable feature space with low dimensions to achieve critical action recognition through a hierarchical classification structure. To forecast the instrument's motion tendency, we model the motion primitives in the polar coordinate system (PCS) to represent patterns of complex trajectories. Given the laparoscopy variance, the adaptive pattern recognition (APR) method, which adapts to uncertain trajectories by modifying model parameters, is designed to improve prediction accuracy. The in-vivo dataset validations show that our framework fulfilled the surgical awareness tasks with exceptional accuracy and real-time performance.
Jie Zhang 0115, Song Zhou, Yiwei Wang 0002, Shenchao Shi, Chidan Wan, Huan Zhao 0001, Xiong Cai, Han Ding 0001
IEEE J. Biomed. Health Informatics3
2022 Bounded Rational Game-theoretical Modeling of Human Joint Actions with Incomplete Information
abstract
As humans and robots start to collaborate in close proximity, robots are tasked to perceive, comprehend, and anticipate human partners' actions, which demands a predictive model to describe how humans collaborate with each other in joint actions. Previous studies either simplify the collaborative task as an optimal control problem between two agents or do not consider the learning process of humans during repeated interaction. This idyllic representation is thus not able to model human rationality and the learning process. In this paper, a bounded-rational and game-theoretical human cooperative model is developed to describe the cooperative behaviors of the human dyad. An experiment of a joint object pushing collaborative task was conducted with 30 human subjects using haptic interfaces in a virtual environment. The proposed model uses inverse optimal control (IOC) to model the reward parameters in the collaborative task. The collected data verified the accuracy of the predicted human trajectory generated from the bounded rational model excels the one with a fully rational model. We further provide insight from the conducted experiments about the effects of leadership on the performance of human collaboration.
Yiwei Wang 0002, Pallavi Shintre, Sunny Amatya
IROS1
2022 Automatic Keyframe Detection for Critical Actions from the Experience of Expert Surgeons
abstract
Robot-Assisted Minimally Invasive Surgery (RAMIS), which introduced robot-actuated invasive tools to increase the dexterity and efficiency of traditional MIS, has become popular. Investigations on how to achieve autonomy in RAMIS have drawn vast intention recently, which urges further insights into the process of the surgical procedures. In this paper, the definition of critical actions, which discriminates the essential stages from regular surgical actions, is proposed to help decompose the complicated surgical processes. A critical intra-operative moment of the surgical workflow, which is called the keyframe, is introduced to indicate the beginning or ending moments of the critical actions. A keyframe detection method is proposed for critical action identification based on a new in-vivo dataset labeled by expert surgeons. Surgeons' criteria for critical actions are captured by the explainable features, which can be extracted from the raw laparoscopic images with a two-stage network. Motivated by the surgeon's decision process of keyframes, a hierarchical structure is designed for keyframe identification by checking the spatial-temporal characteristics of the explainable features. Experimental results show that the reliability of the proposed method for keyframe detection achieves unanimous agreement by expert surgeons.
Jie Zhang 0115, Shenchao Shi, Yiwei Wang 0002, Chidan Wan, Huan Zhao 0001, Xiong Cai, Han Ding 0001
IROS3
2019 How Shall I Drive? Interaction Modeling and Motion Planning towards Empathetic and Socially-Graceful Driving
abstract
While intelligence of autonomous vehicles (AVs) has significantly advanced in recent years, accidents involving AVs suggest that these autonomous systems lack gracefulness in driving when interacting with human drivers. In the setting of a two-player game, we propose model predictive control based on social gracefulness, which is measured by the discrepancy between the actions taken by the AV and those that could have been taken in favor of the human driver. We define social awareness as the ability of an agent to infer such favorable actions based on knowledge about the other agent's intent, and further show that empathy, i.e., the ability to understand others' intent by simultaneously inferring others' understanding of the agent's self intent, is critical to successful intent inference. Lastly, through an intersection case, we show that the proposed gracefulness objective allows an AV to learn more sophisticated behavior, such as passive-aggressive motions that gently force the other agent to yield.
Steven Elliott, Yiwei Wang 0002, Yezhou Yang
ICRA3