Huxin Gao

dblp:294/2400 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0002-8700-560XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-Assisted Endoscopic Submucosal Dissection
abstract
Robot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual operation, thereby enhancing dissection efficiency and accuracy. Accurate prediction of dissection trajectories is crucial for better decision-making, reducing intraoperative errors, and improving surgical training. Nevertheless, predicting these trajectories is challenging due to variable tumor margins and dynamic visual conditions. To address this issue, we create the ESD Trajectory and Confidence Map-based Safety Margin (ETSM) dataset with 1849 short clips, focusing on submucosal dissection with a dual-arm robotic system. We also introduce a framework that combines optimal dissection trajectory prediction with a confidence map-based safety margin, providing a more secure and intelligent decision-making tool to minimize surgical risks for ESD procedures. Additionally, we propose the Regression-based Confidence Map Prediction Network (RCMNet), which utilizes a regression approach to predict confidence maps for dissection areas, thereby delineating various levels of safety margins. We evaluate our RCMNet using three distinct experimental setups: in-domain evaluation, robustness assessment, and out-of-domain evaluation. Experimental results show that our approach excels in the confidence map-based safety margin prediction task, achieving a mean absolute error (MAE) of only 3.18. To the best of our knowledge, this is the first study to apply a regression approach for visual guidance concerning delineating varying safety levels of dissection areas. Our approach bridges gaps in current research by improving prediction accuracy and enhancing the safety of the dissection process, showing great clinical significance in practice. The dataset and code are available at https://github.com/FrankMOWJ/RCMNet.
Mengya Xu, Wenjin Mo, Guankun Wang, Huxin Gao, An Wang 0007, Long Bai 0008, Chaoyang Lyu, Xiaoxiao Yang, Zhen Li 0026, Hongliang Ren 0001
ICRA4
2025 Three-Dimension Tip Force Perception and Axial Contact Location Identification for Flexible Endoscopy Using Tissue-Compliant Soft Distal Attachment Cap Sensors
abstract
In endoluminal surgeries, inserting a flexible endo-scope is one of the fundamental procedures. During this process, vision remains the primary feedback, while the perception of tactile magnitude and location is insufficient. This limitation can hinder the clinician's efficiency when navigating the endoscope through various segments of the natural lumens. To address this issue, we propose a fiber Bragg grating (FBG)-based tissue-compliant sensor cap with multi-mode sensing capabilities, including contact location identification at the terminal surface and the three-dimensional contact force perception at the tip. The soft sensor cap can be affixed to the standard endoscope tip, like a distal attachment cap, for easy installation. Utilizing the relative contact location information, operators can adjust the steerable segment of the endoscope when transitioning from one segment of a natural orifice to a narrower segment, which may be obstructed by constricted lumens. A finite element analysis simulation and the corresponding calibration process based on learning-based approaches have been carried out. The FBG-based sensor can perceive the tip contact force and identify the axial contact location with high precision, where the force perception error is less than 3%, and the contact location identification accuracy is 98.8%. The experimental results demonstrate the potential of the proposed sensing mechanism to be applied in surgeries requiring endoscope insertions.
Yang Yang 0165, Yang Yang 0164, Huxin Gao, Jiewen Lai, Hongliang Ren 0001
ICRA4
2025 Head-mounted Robotic Needle Positioning: Learning from Augmented Reality Demonstration of Neuronavigation and Planning
abstract
Abstract— Robotic needle positioning tasks in neurosurgery often face challenges due to insufficient perception of planar guidance images during surgery. In this work, we propose an Augmented Reality (AR) interface to help perform the robotic needle positioning tasks by learning from demonstration (LfD). Enhanced immersion in the workflow is achieved by displaying surgical scenes and calculated navigation information. The framework utilizes mixed interactive interfaces in virtual and real environments, enhancing demonstration efficiency and quality. A head-mounted display and an optical tracking system are utilized to perform the visualization and needle tracking. Gaussian Mixture Model (GMM) and Gaussian Mixture Regression (GMR) are employed to learn a robust and smooth trajectory policy from demonstrations. Experiments on robot reproduction of the needle positioning task achieved a final positioning error of 0.6 mm and an average trajectory error of 1.07 mm. Comparative user studies with haptic device-based teleoperation exhibit a low completion time of 62.76 s and reduced workload of the proposed system.
Zhiwei Fang, Hok Man Hung, Huxin Gao, Hongliang Ren 0001
IROS3
2025 CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
Guankun Wang, Han Xiao 0010, Renrui Zhang, Huxin Gao, Long Bai 0008, Xiaoxiao Yang, Zhen Li 0026, Hongsheng Li 0001, Hongliang Ren 0001
ACM Multimedia4
2025 Enhancing Anti-Interference of Magnetic Tracking: A MagRobustNet-Based Framework With Self-Supervised Anomaly Detection and Measurements Recovery
abstract
Magnetic tracking technology shows great promise for applications in medicine and industry. However, it often suffers from diverse and unpredictable interferences in practical applications, such as hard-/soft-iron interferences and sensor saturation, leading to reduced localization accuracy or even tracking failure. Thus, we propose a two-step framework to mitigate the impact of interferences based on MagRobustNet, a UNet-like autoencoder network. In the first step, disjoint mask sets are used in conjunction with MagRobustNet to detect anomalous measurements subject to disturbances. In the second step, the interfered regions are masked, and MagRobustNet is applied again to recover their expected measurements from neighboring normal data. Experimental results from testing in four interference scenarios showed that the proposed method improved the average position accuracy by 76.2%, enhancing the tracking system's anti-interference capability. In addition, the proposed method can indicate the interfered regions, thereby prompting the adjustment of the magnetometer array to an interference-free location and offering a new potential diagnostic method for localizing ingested foreign bodies in clinical practice.
Shijian Su, Huxin Gao, Houde Dai, Hongliang Ren 0001
IEEE Trans. Ind. Informatics3
2024 OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted Surgery
abstract
In the realm of automated robotic surgery and computer-assisted interventions, understanding robotic surgical activities stands paramount. Existing algorithms dedicated to surgical activity recognition predominantly cater to pre-defined closed-set paradigms, ignoring the challenges of real-world open-set scenarios. Such algorithms often falter in the presence of test samples originating from classes unseen during training phases. To tackle this problem, we introduce an innovative Open-Set Surgical Activity Recognition (OSSAR) framework. Our solution leverages the hyperspherical reciprocal point strategy to enhance the distinction between known and unknown classes in the feature space. Additionally, we address the issue of over-confidence in the closed set by refining model calibration, avoiding misclassification of unknown classes as known ones. To support our assertions, we establish an open-set surgical activity benchmark utilizing the public JIGSAWS dataset. Besides, we also collect a novel dataset on endoscopic submucosal dissection for surgical activity tasks. Extensive comparisons and ablation experiments on these datasets demonstrate the significant outperformance of our method over existing state-of-the-art approaches. Our proposed solution can effectively address the challenges of real-world surgical scenarios. Our code is publicly accessible at github.com/longbai1006/OSSAR.
Long Bai 0008, Guankun Wang, Jie Wang 0097, Xiaoxiao Yang, Huxin Gao, An Wang 0007, Mobarakol Islam, Hongliang Ren 0001
ICRA5
2024 Inconstant curvature kinematics of parallel continuum robot without static model
abstract
In the study of minimally invasive surgical robots, a mini parallel continuum robot has shown motion advantage after passing through a long and winding working channel. However, due to the interaction force between the elastic wires of the parallel robots during motion generation processes, the constant curvature assumption has shown modeling errors. This causes the current geometric kinematic model to become unreliable. Therefore, there is a need for a more accurate kinematic model in the absence of a complicated static model. This paper aims to solve this issue. The simulation in ANSYS is carried out, and the shape of one of the driving wires, when bending, is fitted by a two-segment polynomial curve. Then, the position of the distal wrist tip can be calculated based on the curve shape. To verify the accuracy of the proposed model, bending simulation and experiment are carried out. The accuracy of the proposed model is compared with that of the kinematic model based on constant curvature assumption. The result shows that the proposed model can get more accurate results, especially when the driving wire displacement increases. For a 10 mm parallel robot, when the displacements of the two pairs of wires are both 3.0 mm, the errors of the two models are 0.42 mm and 5.79 mm (4.2% and 57.9%), respectively.
Huxin Gao, Hongliang Ren 0001
ICRA2
2024 Head-Mounted Hydraulic Needle Driver for Targeted Interventions in Neurosurgery
abstract
Needle interventions are crucial in neurosurgery, requiring high precision and stability. This paper presents a 5-DoF head-mounted hydraulic needle robot designed for accurate and targeted needle insertion and neuroimaging in the deep brain. The robot is compact and lightweight by utilizing a hydraulic pipe transmission to connect the needle driver and actuator. The syringe pistons serve as the actuator and executor, enabling synchronized motion, minimal hysteresis, and high-accuracy insertion. The hydraulic transmission system exhibits hysteresis of less than 0.8 mm, with bidirectional insertion accuracy of approximately 0.05 mm. The resulting needle driver features a compact structure measuring 48 mm × 25 mm × 9 mm, accompanied by a 70-mm-long needle guide. The needle driver is mainly 3D printed, while the hydraulic transmission ensures full compatibility with magnetic resonance imaging (MRI) by isolating all electromagnetic parts from the executor. This compact and lightweight robot-assisted needle intervention system significantly enhances the safety, accuracy, and effectiveness of deep-brain neuroimaging. The feasibility of precise positioning and insertion is further demonstrated by deploying an optical coherence tomography (OCT) microneedle in a rat brain.
Zhiwei Fang, Chao Xu 0008, Huxin Gao, Danny Tat-Ming Chan, Wu Yuan 0001, Hongliang Ren 0001
IROS3
2024 LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion
Tong Chen 0011, Qingcheng Lyu, Long Bai 0008, Erjian Guo, Huxin Gao, Xiaoxiao Yang, Hongliang Ren 0001, Luping Zhou
MICCAI (6)5
2023 SAVAnet: Surgical Action-Driven Visual Attention Network for Autonomous Endoscope Control
abstract
An endoscope holder must understand the detailed surgical actions and the surgeons’ visual attention to keep important targets in the field of endoscopic view during operations. From an intensive analysis of the surgeons’ attention mechanism, we included that surgical actions, like cutting, suturing, etc., play an important role in determining the positions and weights of visual attention points during a dynamic surgical scene. To perform this process, this work proposes a Surgical Action-driven Visual Attention network (SAVAnet) and applies the network in autonomous endoscope control. Four scenarios are constructed in the da Vinci V-rep simulator: pick&place and needle exercise in a general laparoscopic training environment, needle driving with and without obstacle removal in an abdominal cavity, to create datasets for network training. The results show that the network has an outstanding performance in surgical action prediction with a high average accuracy of over 91%. Additionally, with surgical action guidance, the attention point prediction has higher accuracy and accords with surgeons’ visual attention. Finally, the acquired attention points are utilized to execute visual servoing in simulation. The results verify that the SAVAnet is feasible for autonomous endoscope control in real-time and lays a theoretical foundation for future sim-to-real execution. Note to Practitioners—This paper was motivated by the problem of endowing an endoscope with surgeons’ visual attention mechanism, which is affected by surgical actions, for autonomous endoscope control. An eye-tracking device has been utilized to detect surgeon’s visual attention in real-time and then control the endoscope to follow what the surgeon is looking at. However, this approach is susceptible to the surgical environment. Besides, many instrument detection and segmentation algorithms are developed for automatic surgical instrument tracking. However, surgeons’ visual attention does not always focus on the instruments during operations. In this work, we propose a novel SAVAnet to determine visual attention based on surgical actions. We prove from many qualitative and quantitative experiments that surgical actions play a significant role in determining visual attention. The designed SAVAnet can predict surgical actions correctly and then effectively guide the choice of visual attention. Finally, the simulation results show that the SAVAnet can endow endoscope with surgeons’ visual attention to perform self-control in real time. In future research, we will train the SAVAnet using real datasets and conduct more physical experiments on real surgical robots.
Huxin Gao, Weichen Fan, Liang Qiu 0002, Xiaoxiao Yang, Zhen Li 0026, Xiuli Zuo, Max Q.-H. Meng, Hongliang Ren 0001
IEEE Trans Autom. Sci. Eng.1
2023 AMagPoseNet: Real-Time Six-DoF Magnet Pose Estimation by Dual-Domain Few-Shot Learning From Prior Model
abstract
Traditional magnetic tracking approaches based on mathematical models and optimization algorithms are computationally intensive, depend on initial guesses, and do not guarantee convergence to a global optimum. Although fully supervised data-driven deep learning can solve the above issues, the demand for a comprehensive dataset hampers its applicability in magnetic tracking. Thus, we propose an annular magnet pose estimation network (called AMagPoseNet) based on dual-domain few-shot learning from a prior mathematical model, which consists of two subnetworks: PoseNet and CaliNet. PoseNet learns to estimate the magnet pose from the prior mathematical model, and CaliNet is designed to narrow the gap between the mathematical model domain and the real-world domain. Experimental results reveal that the AMagPoseNet outperforms the optimization-based method regarding localization accuracy (1.87$\pm$1.14 mm, 1.89$\pm \text{0.81}^{\circ }$), robustness (nondependence on initial guesses), and computational latency (2.08$\pm$0.02 ms). In addition, the six-degree-of-freedom pose of the magnet could be estimated when discriminative magnetic field features are provided. With the assistance of the mathematical model, the AMagPoseNet requires only a few real-world samples and has excellent performance, showing great potential for practical biomedical and industrial applications.
Shijian Su, Sishen Yuan, Mengya Xu, Huxin Gao, Xiaoxiao Yang, Hongliang Ren 0001
IEEE Trans. Ind. Informatics4
2023 Federated Semi-Supervised Learning for Medical Image Segmentation via Pseudo-Label Denoising
abstract
Distributed big data and digital healthcare technologies have great potential to promote medical services, but challenges arise when it comes to learning predictive model from diverse and complex e-health datasets. Federated Learning (FL), as a collaborative machine learning technique, aims to address the challenges by learning a joint predictive model across multi-site clients, especially for distributed medical institutions or hospitals. However, most existing FL methods assume that clients possess fully labeled data for training, which is often not the case in e-health datasets due to high labeling costs or expertise requirement. Therefore, this work proposes a novel and feasible approach to learn a Federated Semi-Supervised Learning (FSSL) model from distributed medical image domains, where a federated pseudo-labeling strategy for unlabeled clients is developed based on the embedded knowledge learned from labeled clients. This greatly mitigates the annotation deficiency at unlabeled clients and leads to a cost-effective and efficient medical image analysis tool. We demonstrated the effectiveness of our method by achieving significant improvements compared to the state-of-the-art in both fundus image and prostate MRI segmentation tasks, resulting in the highest Dice scores of 89.23% and 91.95% respectively even with only a few labeled clients participating in model training. This reveals the superiority of our method for practical deployment, ultimately facilitating the wider use of FL in healthcare and leading to better patient outcomes.
Liang Qiu 0002, Jierong Cheng, Huxin Gao, Wei Xiong 0001, Hongliang Ren 0001
IEEE J. Biomed. Health Informatics3
2022 GESRsim: Gastrointestinal Endoscopic Surgical Robot Simulator
abstract
Robot-assisted gastrointestinal endoscopic surgery (GES) as a kind of natural orifice transluminal endoscopic surgery (NOTES) is the next-generation minimally invasive surgery (MIS). Besides, rendering certain autonomy to a Gas-trointestinal Endoscopic Surgical Robot (GESR) is promising but highly challenging. Therefore, to accelerate the development and augment the autonomy of GESR, we use CoppeliaSim to develop the first robotic simulator for the GESR system (GESRsim) based on our previous design. The GESRsim provides several 3D models and kinematics of our designed manipulators and endoscopic snake bone. Additionally, we build several scenes for robotic GES training and then utilize different programming interfaces to perform teleoperation. Furthermore, several advanced control algorithms, including visual servoing (VS) and deep reinforcement learning (DRL), are implemented to verify the performance of the GESRsim.
Huxin Gao, Zedong Zhang, Xiao Xiao 0006, Liang Qiu 0002, Xiaoxiao Yang, Ruoyi Hao, Xiuli Zuo, Hongliang Ren 0001
IROS1
2021 Remote-Center-of-Motion Recommendation toward Brain Needle Intervention Using Deep Reinforcement Learning
abstract
Brain needle intervention is a specific diagnosis and therapy procedure in brain disorders, such as brain tumors and Parkinson’s disease. Preoperative needle path planning is a vital step to guarantee the patient’s safety and reduce lesions. For positioning accuracy in the CT/MRI environment, we have developed a novel needle intervention robot in our previous work. Because the robot is currently designed for the rigid needle, the task of preoperative path-planning is to search for an optimal Remote Center of Motion (RCM) for needle insertion. Therefore, this work proposes an RCM recommendation system using deep reinforcement learning. Considering the robot kinematics, this system takes the following criteria/constraints into consideration: clinical obstacle (blood vessels, tissues) avoidance (COA), mechanically inverse kinematics (MIK) and mechanically less motion (MLM) for the robot. We design a reward function to combine the above three criteria based on their corresponding importance level and utilize proximal policy optimization (PPO) as the main agent of reinforcement learning (RL). RL methods are proved to be competent in searching the RCM, which satisfies the above criteria simultaneously. On the one hand, the results present that RL agents obtain the success rate of finishing the designed task at 93%, which has reached the human level in the tests. On the other hand, the RL agents have the remarkable capability of combining more complex criteria/constraints in future work.
Huxin Gao, Xiao Xiao 0006, Liang Qiu 0002, Max Q.-H. Meng, Nicolas Kon Kam King, Hongliang Ren 0001
ICRA1
2021 Magnetically-Connected Modular Reconfigurable Mini-robotic System with Bilateral Isokinematic Mapping and Fast On-site Assembly towards Minimally Invasive Procedures
abstract
This paper presents a modular and reconfigurable mini-robotic system with 5 degrees of freedom (DoFs) towards minimally invasive surgery (MIS). The mini-robotic system consists of two modules, a 2-DoFs rotational end-effector, and a 3-DoFs positioning platform. The 2-DoFs rotational end-effector is based on a spring-spherical joint mechanism, whose rotation is controlled by Bowden-cable. The 3-DoFs positioning platform is based on the linear Delta parallel mechanism. Magnetic spherical joints are adopted to replace the traditional spherical joint. The magnetic joint connections enable fast assembling and disassemble of the end platform and kinematic chains. Different surgical instruments can be installed without changing the driver and control system. A flexible shaft actuates the 3-DoFs positioning platform to arrange the motors away from the manipulator side. Based on these structure characteristics, the 3-DoFs positioning platform’s size is dramatically reduced. The outer diameter of the current prototype is 32.5 mm. The single-axis positioning accuracy of the 3-DoFs positioning platform is within -1 mm to 0.85 mm. Three axes tracking experiments are also carried out, with the positioning errors of ± 1.2 mm for cylindrical curves and -1.5 mm to 2 mm for spherical helix curves. Static and dynamic load capabilities are also tested. Finally, the feasibility of the proposed system is demonstrated.
Xiao Xiao 0006, Shilei Xu, Huxin Gao, Max Q.-H. Meng, Hongliang Ren 0001
ICRA5