VLDB 2026 Research / reviewers in the wild / expert
Hongliang Ren 0001
dblp:44/3343 · also Hong Liang Ren 0001
· DBLP profile ↗
140ranked-venue papers
12as first author
85since 2021 · last 2026
0000-0002-6488-1551ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 72 · 4 first-author · 48 since 2021Artificial intelligence and machine learning · 58 · 5 first-author · 34 since 2021Systems, architecture and hardware · 46 · 5 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 21 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionabstractEnabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexplored in surgical videos, with existing approaches struggling to generalize due to reliance on static visual cues and predefined instrument names. In this work, we introduce SurgRef, a novel motion-guided framework that grounds free-form language expressions in instrument motion, capturing how tools move and interact across time, rather than what they look like. This allows models to understand and segment instruments even under occlusion, ambiguity, or unfamiliar terminology. To train and evaluate SurgRef, we present Ref-IMotion, a diverse, multi-institutional video dataset with dense spatiotemporal masks and rich motion-centric expressions. SurgRef achieves state-of-the-art accuracy and generalization across surgical procedures, setting a new benchmark for robust, language-driven surgical video segmentation. Kun Yuan 0004, Long Bai 0008, Nassir Navab, Hongliang Ren 0001, Hong Joo Lee 0001, Tom Vercauteren, Nicolas Padoy |
AAAI | 7 |
| 2026 | Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and BenchmarkabstractAccurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate information, they lack the semantic context necessary to understand tracking failure mechanisms. We introduce VL-SurgPT, the first large-scale multimodal dataset that bridges visual tracking with textual descriptions of point status in surgical scenes. The dataset comprises 908 in vivo video clips, including 754 for tissue tracking (17,171 annotated points across five challenging scenarios) and 154 for instrument tracking (covering seven instrument types with detailed keypoint annotations). We establish comprehensive benchmarks using eight state-of-the-art tracking methods and propose TG-SurgPT, a text-guided tracking approach that leverages semantic descriptions to improve robustness in visually challenging conditions. Experimental results demonstrate that incorporating point status information significantly improves tracking accuracy and reliability, particularly in adverse visual scenarios where conventional vision-only methods struggle. By bridging visual and linguistic modalities, VL-SurgPT enables the development of context-aware tracking systems crucial for advancing computer-assisted surgery applications that can maintain performance even under challenging intraoperative conditions. Rulin Zhou, Wenlong He, An Wang 0007, Jianhang Zhang, Xuanhui Zeng, Chaowei Zhu, Haijun Hu, Hongliang Ren 0001 |
AAAI | 9 |
| 2026 | EndoControlMag: Robust endoscopic vascular motion magnification with periodic reference resetting and hierarchical tissue-aware dual-mask controlabstractAccurate visualization of subtle vascular dynamics is a knowledge-intensive challenge in minimally invasive surgery. Conventional imaging systems struggle to reveal these imperceptible motions amidst the dynamic complexity of surgical scenes, limiting decision-making reliability. We introduce EndoControlMag , a Lagrangian framework that employs mask-conditioned magnification to selectively enhance vascular motion while preserving the structural integrity of surrounding tissues in endoscopic videos. Our approach integrates two key designs: Periodic Reference Resetting (PRR) , which divides videos into short overlapping clips with dynamically updated reference frames to alleviate error accumulation while maintaining temporal coherence, and Hierarchical Tissue-aware Magnification (HTM) , which combines pretrained visual tracking for accurate vessel localization with dual-mode adaptive softening strategies. HTM employs either motion-based softening that modulates magnification strength proportional to observed tissue displacement, or distance-based exponential decay that simulates biomechanical force attenuation. This strategy enables robust performance across diverse surgical scenarios where motion-based softening excels with complex tissue deformations and distance-based softening provides stability under unreliable optical flow conditions. To validate generality and scalability, we construct EndoVMM24, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Extensive quantitative metrics, qualitative assessments, and expert surgeon evaluations demonstrate that EndoControlMag significantly outperforms existing methods in magnification accuracy, image quality, and robustness. This work advances engineering informatics for surgical vision by providing a reproducible, context-aware framework that supports reliable decision-making in minimally invasive procedures. The code, dataset, and video results are available at https://cho-haz.github.io/EndoControlMag/ . An Wang 0007, Rulin Zhou, Mengya Xu, Yiru Ye, Longfei Gou, Yiting Chang, Hao Chen 0011, Chwee Ming Lim, Jiankun Wang 0001, Hongliang Ren 0001 |
Adv. Eng. Informatics | 10 |
| 2026 | Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challengeabstractReliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding. Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm |
Medical Image Anal. | 37 |
| 2026 | EndoChat: Grounded multimodal large language model for endoscopic surgeryabstractRecently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a deficiency of MLLMs specialized for surgical scene understanding in endoscopic procedures. To this end, we present EndoChat, an MLLM tailored to address various dialogue paradigms and subtasks in understanding endoscopic procedures. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and seven surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, who provide positive feedback on the majority of conversation cases generated by EndoChat. Overall, these results demonstrate that EndoChat has the potential to advance training and automation in robotic-assisted surgery. Our dataset and model are publicly available at https://github.com/gkw0010/EndoChat. Guankun Wang, Long Bai 0008, Kun Yuan 0004, Zhen Li 0026, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001, Fan Zhang 0016, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
Medical Image Anal. | 16 |
| 2026 | Uncertainty-Aware Cross-Scale Hand-Eye Calibration of 2-D Optical Coherence Tomography Using a Plane Target
Haitian Lyu, Jiewen Lai, Ruiyang Zhang, Wu Yuan 0001, Hongliang Ren 0001 |
IEEE Trans. Robotics | 6 |
| 2025 | SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline InferenceabstractSurgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging preceding frames to predict the current frame. Despite great progress, they formulated the task as a series of frame-wise classification, which resulted in a lack of global context of the entire procedure and incoherent predictions. Moreover, besides online analysis, accurate offline surgical phase recognition is also in significant clinical need for retrospective analysis, and existing online algorithms do not fully analyze the entire video, thereby limiting accuracy in offline analysis. To over-come these challenges and enhance both online and offline inference capabilities, we propose a universal Surgical Phase LocalizAtion Network, named SurgPLAN++, with the principle of temporal detection. To ensure a global understanding of the surgical procedure, we devise a phase localization strategy for SurgPLAN ++ to predict phase segments across the entire video through phase proposals. For online analysis, to generate high-quality phase proposals, SurgPLAN++ incorporates a data augmentation strategy to extend the streaming video into a pseudo-complete video through mirroring, center-duplication, and down-sampling. For offline analysis, SurgPLAN++ capi-talizes on its global phase prediction framework to continu-ously refine preceding predictions during each online inference step, thereby significantly improving the accuracy of phase recognition. We perform extensive experiments to validate the effectiveness, and our SurgPLAN++ achieves remarkable performance in both online and offline modes, which outper-forms state-of-the-art methods. The source code is available at https://github.com/franciszchenlSurgPLAN-Plus. Zhen Chen 0018, Xingjian Luo, Jinlin Wu, Long Bai 0008, Zhen Lei 0001, Hongliang Ren 0001, Sébastien Ourselin, Hongbin Liu 0001 |
ICRA | 6 |
| 2025 | Variable-Stiffness Nasotracheal Intubation Robot with Passive Buffering: A Modular Platform in Mannequin Studies
Ruoyi Hao, Jiewen Lai, Wenqi Zhong, Dihong Xie, Yang Zhang 0053, Catherine Po Ling Chan, Jason Ying-Kuen Chan, Hongliang Ren 0001 |
ICRA | 10 |
| 2025 | Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and MappingabstractSimultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems struggle with accurate depth and surface reconstruction due to multi-view inconsistencies. Simply incorporating SLAM and 3DGS leads to mismatches between the reconstructed frames. In this work, we present Endo-2DTAM, a real-time endoscopic SLAM system with 2D Gaussian Splatting (2DGS) to address these challenges. Endo-2DTAM incorporates a surface normal-aware pipeline, which consists of tracking, mapping, and bundle adjustment modules for geometrically accurate reconstruction. Our robust tracking module combines point-topoint and point-to-plane distance metrics, while the mapping module utilizes normal consistency and depth distortion to enhance surface reconstruction quality. We also introduce a pose-consistent strategy for efficient and geometrically coherent keyframe sampling. Extensive experiments on public endoscopic datasets demonstrate that Endo-2DTAM achieves an RMSE of$1.87 \pm 0.63 \mathbf{m m}$for depth reconstruction of surgical scenes while maintaining computationally efficient tracking, high-quality visual appearance, and real-time rendering. Our code will be released at github.com/lastbasket/Endo-2DTAM. Yiming Huang 0007, Beilei Cui, Long Bai 0008, Zhen Chen 0018, Jinlin Wu, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
ICRA | 8 |
| 2025 | Minimally Invasive Endotracheal Inside-Out Flexible Needle Driving System Towards Microendoscope-Guided Robotic TracheostomyabstractOpen tracheostomy (OT) is considered the traditional way and golden standard for treating airway obstruction patients. However, OT has many unavoidable drawbacks, including strict performing scenarios, significant scarring, and the risk of surgeon infection. Percutaneous dilation tracheostomy (PDT) emerges, with advantages including a lower cost, smaller scarring, and better protection of surgeons from inflecting by aerosol. However, the outside-in puncture manner of PDT has a risk of piercing the post-tracheal wall and the esophagus with uncontrolled force. Additionally, locating tracheal rings and determining the puncture site externally can be challenging for certain patients, such as those who are obese or have undergone neck surgery, while this procedure typically relies on palpation and the surgeon's expertise. Hence, to improve the safety and simplicity of tracheostomy, a minimally-invasive endotracheal inside-out flexible needle-driving system towards microendoscope-guided robotic tracheostomy (MERT) has been proposed in this paper. Guided by an optical coherence tomography (OCT) probe and a microendoscope, the robot inserts into the trachea and performs an inside-out puncture using a flexible needle. The robot can work through a standard endotracheal tube (ETT), and the puncture direction of the flexible needle is variable. Kinematics and statics models of the flexible needle have been derived, and the minimum position errors generated in the kinematics and statics validation experiments are$0.57 \pm 0.21 \mathbf{~ m m}$and$0.27 \pm 0.21 \mathbf{~ m m}$. Finally, a porcine trachea puncture experiment is carried out, and the feasibility of the proposed system is verified. Botao Lin, Sishen Yuan, Tinghua Zhang, Ruoyi Hao, Wu Yuan 0001, Chwee Ming Lim, Hongliang Ren 0001 |
ICRA | 8 |
| 2025 | Improving Efficiency in Path Planning: Tangent Line Decomposition AlgorithmabstractThis paper introduces a tangent line decomposition (TLD) algorithm that efficiently finds collision-free paths close to optimal in both 2D and 3D environments. Compared with the existing visibility line-based algorithms, the proposed algorithm innovatively proposed the concept of tangent line decomposition, which decomposes complicated planning into many simple steps. For each step, only one key obstacle is taken into consideration. Besides, instead of constructing a complete graph, a best-first search algorithm is used to avoid searching redundant edges. The path planned by the algorithm is not the optimal path. However, following the idea of the informed RRT* algorithm, the path length planned by TLD can be used as a precondition for other optimal algorithms. In this way, the overall efficiency can be significantly improved. The simulations show that the proposed methods outperform existing methods regarding planning efficiency and solution quality. Hongliang Ren 0001 |
ICRA | 2 |
| 2025 | ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-Assisted Endoscopic Submucosal DissectionabstractRobot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual operation, thereby enhancing dissection efficiency and accuracy. Accurate prediction of dissection trajectories is crucial for better decision-making, reducing intraoperative errors, and improving surgical training. Nevertheless, predicting these trajectories is challenging due to variable tumor margins and dynamic visual conditions. To address this issue, we create the ESD Trajectory and Confidence Map-based Safety Margin (ETSM) dataset with 1849 short clips, focusing on submucosal dissection with a dual-arm robotic system. We also introduce a framework that combines optimal dissection trajectory prediction with a confidence map-based safety margin, providing a more secure and intelligent decision-making tool to minimize surgical risks for ESD procedures. Additionally, we propose the Regression-based Confidence Map Prediction Network (RCMNet), which utilizes a regression approach to predict confidence maps for dissection areas, thereby delineating various levels of safety margins. We evaluate our RCMNet using three distinct experimental setups: in-domain evaluation, robustness assessment, and out-of-domain evaluation. Experimental results show that our approach excels in the confidence map-based safety margin prediction task, achieving a mean absolute error (MAE) of only 3.18. To the best of our knowledge, this is the first study to apply a regression approach for visual guidance concerning delineating varying safety levels of dissection areas. Our approach bridges gaps in current research by improving prediction accuracy and enhancing the safety of the dissection process, showing great clinical significance in practice. The dataset and code are available at https://github.com/FrankMOWJ/RCMNet. Mengya Xu, Wenjin Mo, Guankun Wang, Huxin Gao, An Wang 0007, Long Bai 0008, Chaoyang Lyu, Xiaoxiao Yang, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 10 |
| 2025 | Three-Dimension Tip Force Perception and Axial Contact Location Identification for Flexible Endoscopy Using Tissue-Compliant Soft Distal Attachment Cap SensorsabstractIn endoluminal surgeries, inserting a flexible endo-scope is one of the fundamental procedures. During this process, vision remains the primary feedback, while the perception of tactile magnitude and location is insufficient. This limitation can hinder the clinician's efficiency when navigating the endoscope through various segments of the natural lumens. To address this issue, we propose a fiber Bragg grating (FBG)-based tissue-compliant sensor cap with multi-mode sensing capabilities, including contact location identification at the terminal surface and the three-dimensional contact force perception at the tip. The soft sensor cap can be affixed to the standard endoscope tip, like a distal attachment cap, for easy installation. Utilizing the relative contact location information, operators can adjust the steerable segment of the endoscope when transitioning from one segment of a natural orifice to a narrower segment, which may be obstructed by constricted lumens. A finite element analysis simulation and the corresponding calibration process based on learning-based approaches have been carried out. The FBG-based sensor can perceive the tip contact force and identify the axial contact location with high precision, where the force perception error is less than 3%, and the contact location identification accuracy is 98.8%. The experimental results demonstrate the potential of the proposed sensing mechanism to be applied in surgeries requiring endoscope insertions. Yang Yang 0165, Yang Yang 0164, Huxin Gao, Jiewen Lai, Hongliang Ren 0001 |
ICRA | 6 |
| 2025 | Head-mounted Robotic Needle Positioning: Learning from Augmented Reality Demonstration of Neuronavigation and PlanningabstractAbstract— Robotic needle positioning tasks in neurosurgery often face challenges due to insufficient perception of planar guidance images during surgery. In this work, we propose an Augmented Reality (AR) interface to help perform the robotic needle positioning tasks by learning from demonstration (LfD). Enhanced immersion in the workflow is achieved by displaying surgical scenes and calculated navigation information. The framework utilizes mixed interactive interfaces in virtual and real environments, enhancing demonstration efficiency and quality. A head-mounted display and an optical tracking system are utilized to perform the visualization and needle tracking. Gaussian Mixture Model (GMM) and Gaussian Mixture Regression (GMR) are employed to learn a robust and smooth trajectory policy from demonstrations. Experiments on robot reproduction of the needle positioning task achieved a final positioning error of 0.6 mm and an average trajectory error of 1.07 mm. Comparative user studies with haptic device-based teleoperation exhibit a low completion time of 62.76 s and reduced workload of the proposed system. Zhiwei Fang, Hok Man Hung, Huxin Gao, Hongliang Ren 0001 |
IROS | 4 |
| 2025 | CapsDT: Diffusion-Transformer for Capsule Robot ManipulationabstractVision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endoscopy robotics, particularly endoscopy capsule robots that perform actions within the digestive system, remains unexplored. The integration of VLA models into endoscopy robots allows more intuitive and efficient interactions between human operators and medical devices, improving both diagnostic accuracy and treatment outcomes. In this work, we design CapsDT, a Diffusion Transformer model for capsule robot manipulation in the stomach. By processing interleaved visual inputs, and textual instructions, CapsDT can infer corresponding robotic control signals to facilitate endoscopy tasks. In addition, we developed a capsule endoscopy robot system, a capsule robot controlled by a robotic arm-held magnet, addressing different levels of four endoscopy tasks and creating corresponding capsule robot datasets within the stomach simulator. Comprehensive evaluations on various robotic tasks indicate that CapsDT can serve as a robust vision-language generalist, achieving state-of-the-art performance in various levels of endoscopy tasks while achieving a 26.25% success rate in real-world simulation manipulation. Xiting He, Mingwu Su, Xinqi Jiang, Long Bai 0008, Hongliang Ren 0001 |
IROS | 5 |
| 2025 | Adjusting Tissue Puncture Omnidirectionally In Situ with Pneumatic Rotatable Biopsy Mechanism and Hierarchical Airflow Management in Tortuous Luminal PathwaysabstractIn situ tissue biopsy with an endoluminal catheter is an efficient approach for disease diagnosis, featuring low invasiveness and few complications. However, the endoluminal catheter struggles to adjust the biopsy direction by distal endoscope bending or proximal twisting for tissue sampling within the tortuous luminal organs, due to friction-induced hysteresis and narrow spaces. Here, we propose a pneumatically-driven robotic catheter enabling the adjustment of the sampling direction without twisting the catheter for an accurate in situ omnidirectional biopsy. The distal end of the robotic catheter consists of a pneumatic bending actuator for the catheter’s deployment in torturous luminal organs and a pneumatic rotatable biopsy mechanism (PRBM). By hierarchical airflow control, the PRBM can adjust the biopsy direction under low airflow and deploy the biopsy needle with higher airflow, allowing for rapid omnidirectional sampling of tissue in situ. This paper describes the design, modeling, and characterization of the proposed robotic catheter, including repeated deployment assessments of the biopsy needle, puncture force measurement, and validation via phantom tests. The PRBM prototype has six sampling directions evenly distributed across 360 degrees when actuated by a positive pressure of 0.3 MPa. The pneumatically-driven robotic catheter provides a novel biopsy strategy, potentially facilitating in situ multidirectional biopsies in tortuous luminal organs with minimum invasiveness. Botao Lin, Tinghua Zhang, Sishen Yuan, Jiaole Wang, Wu Yuan 0001, Hongliang Ren 0001 |
IROS | 7 |
| 2025 | Exploring Stiffness Gradient Effects in Magnetically Induced Metamorphic Materials via Continuum Simulation and ValidationabstractMagnetic soft continuum robots are capable of bending with remote control in confined space environments, and they have been applied in various bioengineering contexts. As one type of ferromagnetic soft continuums, the Magnetically Induced Metamorphic Materials (MIMMs)-based continuum (MC) exhibits similar bending behaviors. Based on the characteristics of its base material, MC is flexible in modifying unit stiffness and convenient in molding fabrication. However, recent studies on magnetic continuum robots have primarily focused on one or two design parameters, limiting the development of a comprehensive magnetic continuum bending model. In this work, we constructed graded-stiffness MCs (GMCs) and developed a numerical model for GMCs’ bending performance, incorporating four key parameters that determine their performance. The simulated bending results were validated with real bending experiments in four different categories: varying magnetic field, cross-section, unit stiffness, and unit length. The graded-stiffness design strategy applied to GMCs prevents sharp bending at the fixed end and results in a more circular curvature. We also trained an expansion model for GMCs’ bending performance that is highly efficient and accurate compared to the simulation process. An extensive library of bending prediction for GMCs was built using the trained model. Yang Yang 0164, Yiming Huang 0007, Hongliang Ren 0001 |
IROS | 4 |
| 2025 | Learning to Perform Low-Contact Autonomous Nasotracheal Intubation by Recurrent Action-Confidence Chunking with TransformerabstractNasotracheal intubation (NTI) is critical for establishing artificial airways in clinical anesthesia and critical care. Current manual methods face significant challenges, including cross-infection, especially during respiratory infection care, and insufficient control of endoluminal contact forces, increasing the risk of mucosal injuries. While existing studies have focused on automated endoscopic insertion, the automation of NTI remains unexplored despite its unique challenges: Nasotracheal tubes exhibit greater diameter and rigidity than standard endoscopes, substantially increasing insertion complexity and patient risks. We propose a novel autonomous NTI system with two key components to address these challenges. First, an autonomous NTI system is developed, incorporating a prosthesis embedded with force sensors, allowing for safety assessment and data filtering. Then, the Recurrent Action-Confidence Chunking with Transformer (RACCT) model is developed to handle complex tube-tissue interactions and partial visual observations. Experimental results demonstrate that the RACCT model outperforms the ACT model in all aspects and achieves a 66% reduction in average peak insertion force compared to manual operations while maintaining equivalent success rates. This validates the system’s potential for reducing infection risks and improving procedural safety. Ruoyi Hao, Yiming Huang 0007, Dihong Xie, Catherine Po Ling Chan, Jason Ying-Kuen Chan, Hongliang Ren 0001 |
IROS | 7 |
| 2025 | Jacobian Exploratory Dual-Phase Reinforcement Learning for Dynamic Endoluminal Navigation of Deformable Continuum RobotsabstractDeformable continuum robots (DCRs) present unique planning challenges due to nonlinear deformation mechanics and partial state observability, violating the Markov assumptions of conventional reinforcement learning (RL) methods. While Jacobian-Based approaches offer theoretical foundations for rigid manipulators, their direct application to DCRs remains limited by time-varying kinematics and underactuated deformation dynamics. This paper proposes Jacobian Exploratory Dual-Phase RL (JEDP-RL), a framework that decomposes planning into phased Jacobian estimation and policy execution. During each training step, we first perform small-scale local exploratory actions to estimate the deformation Jacobian matrix, then augment the state representation with Jacobian features to restore approximate Markovianity. Extensive SOFA surgical dynamic simulations demonstrate JEDP-RL’s three key advantages over proximal policy optimization (PPO) baselines: 1) Convergence speed: 3.2× faster policy convergence, 2) Navigation efficiency: requires 25% fewer steps to reach the target, and 3) Generalization ability: achieve 92% success rate under material property variations and achieve 83% (33% higher than PPO) success rate in the unseen tissue environment. Chi Kit Ng, Hongliang Ren 0001 |
IROS | 3 |
| 2025 | Body-Temperature-Responsive Balloon Actuator for Adaptive In-Ear Microneedle Electrode DeploymentabstractConventional in-ear electrodes face challenges such as inconsistent skin contact, motion artifacts, and discomfort, limiting their reliability in dynamic conditions. To overcome these limitations, this study presents a body-temperature responsive in-ear balloon actuator (BBA) based on liquid-to-vapor phase change, integrated with silver microneedle electrodes for stable electrophysiological monitoring. The dual-layer balloon structure encapsulates a phase-change fluid core, expanding at 36–37°C to ensure conformal skin contact while minimizing motion artifacts and mechanical pressure. Among 3 tested designs, the dual-material design proved optimal, achieving a peak insertion force of 0.05 N and maintaining impedance fluctuations below 5%. A wearable system was further validated through dynamic tests, demonstrating a 20% reduction in motion artifacts compared to conventional electrodes. These findings highlight the actuator’s potential for stable and comfortable wearable electrophysiological monitoring. Ruizhou Zhao, Wenchao Yue, Entong Li, Chengxi Bai, Hongliang Ren 0001 |
IROS | 5 |
| 2025 | SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
Tong Chen 0011, Shuya Yang, Long Bai 0008, Hongliang Ren 0001, Luping Zhou |
MICCAI (10) | 5 |
| 2025 | Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Yanheng Li 0002, Tong Chen 0011, Jie Wang 0097, Jinlin Wu, Zhen Lei 0001, Hongbin Liu 0001, Hongliang Ren 0001 |
MICCAI (9) | 10 |
| 2025 | SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Kun Yuan 0004, Guankun Wang, Mobarak I. Hoque, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
MICCAI (9) | 9 |
| 2025 | Recognizing Surgical Phases Anywhere: Few-Shot Test-Time Adaptation and Task-Graph Guided Refinement
Kun Yuan 0004, Tingxuan Chen, Joël L. Lavanchy, Christian Heiliger, Ege Özsoy, Yiming Huang 0007, Long Bai 0008, Nassir Navab, Vinkle Srivastav, Hongliang Ren 0001, Nicolas Padoy |
MICCAI (9) | 11 |
| 2025 | CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
Guankun Wang, Han Xiao 0010, Renrui Zhang, Huxin Gao, Long Bai 0008, Xiaoxiao Yang, Zhen Li 0026, Hongsheng Li 0001, Hongliang Ren 0001 |
ACM Multimedia | 9 |
| 2025 | Rethinking data imbalance in class incremental surgical instrument segmentationabstractIn surgical instrument segmentation, the increasing variety of instruments over time poses a significant challenge for existing neural networks, as they are unable to effectively learn such incremental tasks and suffer from catastrophic forgetting. When learning new data, the model experiences a sharp performance drop on previously learned data. Although several continual learning methods have been proposed for incremental understanding tasks in surgical scenarios, the issue of data imbalance often leads to a strong bias in the segmentation head, resulting in poor performance. Data imbalance can occur in two forms: (i) class imbalance between new and old data, and (ii) class imbalance within the same time point of data. Such imbalances often cause the dominant classes to take over the training process of continual semantic segmentation (CSS). To address this issue, we propose SurgCSS, a novel plug-and-play CSS framework for surgical instrument segmentation under data imbalance. Specifically, we generate realistic surgical backgrounds through inpainting and blend instrument foregrounds with the generated backgrounds in a class-aware manner to balance the data distribution in various scenarios. We further propose the Class Desensitization Loss by employing contrastive learning to correct edge biases caused by data imbalance. Moreover, we dynamically fuse the weight parameters of the old and new models to achieve a better trade-off between the biased and unbiased model weights. To investigate the data imbalance problem in surgical scenarios, we construct a new benchmark for surgical instrument CSS by integrating four public datasets: EndoVis 2017, EndoVis 2018, CholecSeg8k, and SAR-RAPR50. Extensive experiments demonstrate the effectiveness of the proposed framework, achieving significant performance improvement against existing baselines. Our method demonstrates excellent potential for clinical applications. The code is publicly available at github.com/Zzsf11/SurgCSS. Shifang Zhao, Long Bai 0008, Kun Yuan 0004, Feng Li 0034, Jieming Yu, Wenzhen Dong, Guankun Wang, Mobarakol Islam, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
Medical Image Anal. | 11 |
| 2025 | V²-SfMLearner: Learning Monocular Depth and Ego-Motion for Multimodal Wireless Capsule EndoscopyabstractDeep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the capsule endoscopies within the gastrointestinal tract cause vibration perturbations in the training data. Existing solutions focus solely on vision-based processing, neglecting other auxiliary signals like vibrations that could reduce noise and improve performance. Therefore, we propose V2-SfMLearner, a multimodal approach integrating vibration signals into vision-based depth and capsule motion estimation for monocular capsule endoscopy. We construct a multimodal capsule endoscopy dataset containing vibration and visual signals, and our artificial intelligence solution develops an unsupervised method using vision-vibration signals, effectively eliminating vibration perturbations through multimodal learning. Specifically, we carefully design a vibration network branch and a Fourier fusion module, to detect and mitigate vibration noises. The fusion framework is compatible with popular vision-only algorithms. Extensive validation on the multimodal dataset demonstrates superior performance and robustness against vision-only algorithms. Without the need for large external equipment, our V2-SfMLearner has the potential for integration into clinical capsule robots, providing real-time and dependable digestive examination tools. The findings show promise for practical implementation in clinical settings, enhancing the diagnostic capabilities of doctors. Note to Practitioners—This paper is motivated by the problem of estimating the depth and ego-motion information for the wireless capsule endoscopy in the human gastrointestinal tract to realize accurate, efficient, robust, and real-time inspection. Our estimation method does not engage any external localization equipment. Instead, inspired by the existing research on integrating capsule endoscopy and inertial measurement units, we introduce vibration signals into vision-based depth and ego-motion estimation approaches, improving the accuracy and robustness of the estimation results based on multimodal learning methods. Research on capsule robots or computer vision can readily be combined with our framework for various clinical and industrial applications. Long Bai 0008, Beilei Cui, Yanheng Li 0002, Shilong Yao, Sishen Yuan, Yanan Wu 0003, Yang Zhang 0053, Max Q.-H. Meng, Zhen Li 0026, Weiping Ding 0001, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 12 |
| 2025 | Learning Anticipatory Decision for Distributed Systems With Robustness GuaranteesabstractThis paper investigates anticipatory decision for unknown distributed systems with robustness concerns. Anticipatory decision focuses on action selection before observations appear at temporal scales. Firstly, anticipatory decision forms sequential feedback with min-max performance guarantees, while causality comes from time series analysis. Next, distribution, robustness and time consistency partition the optimization into spatial and temporal sub-games. The spatial sub-games dispel conflicts on distribution and robustness, while the temporal ones ensure stability and performance through time consistency. Finally, we propose a multi-step reinforcement learning algorithm under causality analysis and game theoretical framework. Numerical results demonstrate the effectiveness of the approach, and practical experiments show potential real-world applications. Note to Practitioners—This framework focuses on anticipatory decision for distributed systems, which suffer from distributed communication, unknown dynamics, environmental disturbances and state observation loss. Our framework has various application scenarios, e.g., internal surgical robots, low-light autonomous driving and non-GPS navigation, and these scenarios mainly involve dynamic environments and weak signal feedback. For example, decision-making in autonomous driving requires not only reacting to current environmental conditions but also anticipating future scenarios and uncertainties due to poor visibility. Most results deal these issues with model-driven approaches, while unknown dynamics render these methods inapplicable. For implementation, we propose a multi-step reinforcement learning algorithm for anticipatory decision framework with stability and robustness guarantees, and details mainly contain three parts: 1) We collect data during offline phase, and form the data structure, namely, current-next observation pair with multi-step decision and accumulated reward; 2) Strategies and value functions are approximated with neural networks through Monte-Carlo methods; 3) The strategy is deployed as sequential feedback in practical systems, and predicts multi-step decisions with single-step state observation. Finally, we select robot consensus with optical sensors as the implementation demo. Peijiang Liu, Xindi Yang, Hongliang Ren 0001, Hao Zhang 0008, Zhuping Wang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Air Shepherd: Trajectory Prediction-Based Target Localization and Circumnavigation in Cluttered EnvironmentsabstractThis paper proposes a trajectory prediction-based target localization and circumnavigation pattern for cluttered three-dimensional environments, which is more realistic and suitable for more complex environments than traditional patterns. The main work of the paper consists of two parts: tracking based on trajectory prediction and circumnavigation based on broadcast information. On the one hand, the tracking Autonomous Aerial vehicle (AAV) obtains target trajectory prediction based on the B-spline curve, and then achieves target localization and tracking through front-end search and back-end optimization. On the other hand, without communicating with each other, a distributed control strategy is presented so that the multiple circumnavigation AAVs can achieve target circumnavigation and reciprocal avoidance by only observing the status of adjacent AAVs. In the simulation, obstacle avoidance vehicles moving freely at different speeds are selected as targets in two scenarios and the simulation results are given to verify the effectiveness of the proposed approach. Furthermore, a hardware-in-the-loop experiment and a overall system validation experiment are designed to verify the feasibility of the algorithm. Kai Rao, Huaicheng Yan 0001, Hongliang Ren 0001, Tan Chen 0004, Youmin Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | RASEC: Rescaling Acquisition Strategy With Energy Constraints Under Fusion Kernel for Active Incision Recommendation in TracheotomyabstractTracheotomy is commonly performed for patients needing prolonged intubation, airway obstruction, and neck injuries. Accurate placement of the incision and the tracheal window is paramount in order to avoid complications. Current surgical technique heavily relies on palpating cartilage landmarks on the neck to place the incision. In order to achieve the accelerated goals of the robot-assisted subtask in a tracheotomy, this paper proposes a novel autonomous palpation-based acquisition strategy - RASEC in the tracheal region, which can interactively predict the next acquisition point to maximize the expected information and minimize the costs of palpation procedure. We employ a Gaussian Process (GP) to model the distribution of hardness and utilize anatomical information as a priori input to guide the point of palpation for medical robots. The dynamic tactile sensor based on the resonant frequency is introduced to measure tissue hardness in the tracheal region by millimeter-scale contact to secure the interaction. We investigate the kernel fusion method to blend the Squared Exponential (SE) kernel with the Ornstein-Uhlenbeck (OU) kernel and optimize the Bayesian optimization search by leveraging the anatomical information of the larynx as a priori knowledge. Moreover, we further regularize the exploration and greed factors. The tactile sensor’s moving distance and the robotic base link’s rotation angle during the incision localization process are considered new factors in the acquisition strategy. Simulation and physical phantom experiments are conducted for comparison with state-of-the-art GP-based exploration approaches. The results show that the sensor’s moving distance was reduced by 53.1% and the rotation angle of the base was reduced by 75.2% of the previous values without sacrificing overall performance capabilities. The satisfying algorithmic index (average precision 0.932, average recall 0.973, average F1 score 0.952) with fewer central estimation distance errors (0.423 mm) and high resolution (1 mm) indicates the performance of the proposed RASEC in terms of exploration efficiency, cost awareness, and localization accuracy for incision localization and recommendation in real robot-assisted subtask in the tracheotomy procedure.Note to Practitioners—This work is well motivated to introduce the Level of Autonomy (LoA) 2 - task-level autonomy, specifically in the context of tracheotomy procedures. The incorporation of robotic palpation techniques aims to provide surgeons with enhanced capabilities for incision recommendations, which directly benefit surgeons to visualize hands-on information and localize the trachea regions more efficiently and further reduce cognitive load. To detect the trachea region for intubation incision without costly ergodic acquisition, this article suggests a highly efficient acquisition strategy utilizing the fusion kernel function and regularized impact factors, eliminating the time consumption for such localization task. The actual clinical value is that our proposed strategy can earn more time for further increasing the probability of patient resuscitation, to facilitate supervised autonomy in the real clinic scene. Wenchao Yue, Fan Bai 0008, Jianbang Liu 0002, Max Q.-H. Meng, Chwee Ming Lim, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Enhancing Anti-Interference of Magnetic Tracking: A MagRobustNet-Based Framework With Self-Supervised Anomaly Detection and Measurements RecoveryabstractMagnetic tracking technology shows great promise for applications in medicine and industry. However, it often suffers from diverse and unpredictable interferences in practical applications, such as hard-/soft-iron interferences and sensor saturation, leading to reduced localization accuracy or even tracking failure. Thus, we propose a two-step framework to mitigate the impact of interferences based on MagRobustNet, a UNet-like autoencoder network. In the first step, disjoint mask sets are used in conjunction with MagRobustNet to detect anomalous measurements subject to disturbances. In the second step, the interfered regions are masked, and MagRobustNet is applied again to recover their expected measurements from neighboring normal data. Experimental results from testing in four interference scenarios showed that the proposed method improved the average position accuracy by 76.2%, enhancing the tracking system's anti-interference capability. In addition, the proposed method can indicate the interfered regions, thereby prompting the adjustment of the magnetometer array to an interference-free location and offering a new potential diagnostic method for localizing ingested foreign bodies in clinical practice. Shijian Su, Huxin Gao, Houde Dai, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Fine-Grained Classification Reveals Angiopathological Heterogeneity of Port Wine Stains Using OCT and OCTA FeaturesabstractAccurate classification of port wine stains (PWS, vascular malformations present at birth), is critical for subsequent treatment planning. However, the current method of classifying PWS based on the external skin appearance rarely reflects the underlying angiopathological heterogeneity of PWS lesions, resulting in inconsistent outcomes with the common vascular-targeted photodynamic therapy (V-PDT) treatments. Conversely, optical coherence tomography angiography (OCTA) is an ideal tool for visualizing the vascular malformations of PWS. Previous studies have shown no significant correlation between OCTA quantitative metrics and the PWS subtypes determined by the current classification approach. In this study, we propose a novel fine-grained classification method for PWS that integrates OCT and OCTA imaging. Utilizing a machine learning-based approach, we subdivided PWS into five distinct subtypes by unearthing the heterogeneity of hypodermic histopathology and vessel structures. Six quantitative metrics, encompassing vascular morphology and depth information of PWS lesions, were designed and statistically analyzed to evaluate angiopathological differences among the subtypes. Our classification reveals significant distinctions across all metrics compared to conventional skin appearance-based subtypes, demonstrating its ability to accurately capture angiopathological heterogeneity. This research marks the first attempt to classify PWS based on angiopathology, potentially guiding more effective subtyping and treatment strategies for PWS. Xiaofeng Deng, Defu Chen, Bowen Liu 0008, Xiwan Zhang, Haixia Qiu, Wu Yuan 0001, Hongliang Ren 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Augment Laminar Jamming Variable Stiffness Through Electroadhesion and Vacuum ActuationabstractVarious variable stiffness mechanisms have been developed to bestow new capabilities for the robotics community by changing the mechanical behaviors of robots. However, variable stiffness is limited in actuation, response speed, stiffness ratio, and, most importantly, modeling. This article proposes hybrid actuated laminar jamming to outperform individual actuated variable stiffness mechanisms. An analytical model for multilayer laminar jamming that accurately characterizes mechanical behaviors in experiments is first built. Comprehensive parametrical analysis based on this model serves as design guidelines for performance improvements of laminar jamming. Feedforward control further proves the validity of the proposed model and exhibits good controllability, showing response speed as fast as 5 ms. The synergy between electroadhesion and vacuum actuation significantly enhances overall performance, resulting in far greater effects than individual contributions. For instance, the proposed device generates a high stiffness that is almost impossible for individual vacuum or electroadhesion. Moreover, vacuuming increases 23% of the breakdown voltage, which leads to a larger electroadhesion force and, hence, a higher stiffness. Cheng Chen 0077, Hongliang Ren 0001, Hongqiang Wang 0003 |
IEEE Trans. Robotics | 2 |
| 2024 | OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted SurgeryabstractIn the realm of automated robotic surgery and computer-assisted interventions, understanding robotic surgical activities stands paramount. Existing algorithms dedicated to surgical activity recognition predominantly cater to pre-defined closed-set paradigms, ignoring the challenges of real-world open-set scenarios. Such algorithms often falter in the presence of test samples originating from classes unseen during training phases. To tackle this problem, we introduce an innovative Open-Set Surgical Activity Recognition (OSSAR) framework. Our solution leverages the hyperspherical reciprocal point strategy to enhance the distinction between known and unknown classes in the feature space. Additionally, we address the issue of over-confidence in the closed set by refining model calibration, avoiding misclassification of unknown classes as known ones. To support our assertions, we establish an open-set surgical activity benchmark utilizing the public JIGSAWS dataset. Besides, we also collect a novel dataset on endoscopic submucosal dissection for surgical activity tasks. Extensive comparisons and ablation experiments on these datasets demonstrate the significant outperformance of our method over existing state-of-the-art approaches. Our proposed solution can effectively address the challenges of real-world surgical scenarios. Our code is publicly accessible at github.com/longbai1006/OSSAR. Long Bai 0008, Guankun Wang, Jie Wang 0097, Xiaoxiao Yang, Huxin Gao, An Wang 0007, Mobarakol Islam, Hongliang Ren 0001 |
ICRA | 9 |
| 2024 | Chained Flexible Capsule Endoscope: Unraveling the Conundrum of Size Limitations and Functional Integration for Gastrointestinal TransitivityabstractCapsule endoscopes, predominantly serving diagnostic functions, provide lucid internal imagery but are devoid of surgical or therapeutic capabilities. Consequently, despite lesion detection, physicians frequently resort to traditional endoscopic or open surgical procedures for treatment, resulting in more complex, potentially risky interventions. To surmount these limitations, this study introduces a chained flexible capsule endoscope (FCE) design concept, specifically conceived to navigate the inherent volume constraints of capsule endoscopes whilst augmenting their therapeutic functionalities. The FCE’s distinctive flexibility originates from a conventional rotating joint design and the incision pattern in the flexible material. In vitro experiments validated the passive navigation ability of the FCE in rugged intestinal tracts. Further, the FCE demonstrates consistent reptile-like peristalsis under the influence of an external magnetic field, and possesses the capability for film expansion and disintegration under high-frequency electromagnetic stimulation. These findings illuminate a promising path toward amplifying the therapeutic capacities of capsule endoscopes without necessitating a size compromise. Sishen Yuan, Baijia Liang, Lailu Li, Qingzhuo Zheng, Shuang Song 0002, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 8 |
| 2024 | Magnetic-Guided Flexible Origami Robot toward Long-Term Phototherapy of H. pylori in the StomachabstractHelicobacter pylori, a pervasive bacterial infection associated with gastrointestinal disorders such as gastritis, peptic ulcer disease, and gastric cancer, impacts approximately 50% of the global population. The efficacy of standard clinical eradication therapies is diminishing due to the rise of antibiotic-resistant strains, necessitating alternative treatment strategies. Photodynamic therapy (PDT) emerges as a promising prospect in this context. This study presents the development and implementation of a magnetically-guided origami robot, incorporating flexible printed circuit units for sustained and stable phototherapy of Helicobacter pylori. Each integrated unit is equipped with wireless charging capabilities, producing an optimal power output that can concurrently illuminate up to 15 LEDs at their maximum intensity. Crucially, these units can be remotely manipulated via a magnetic field, facilitating both translational and rotational movements. We propose an open-loop manual control sequence that allows the formation of a stable, compliant triangular structure through the interaction of internal magnets. This adaptable configuration is uniquely designed to withstand the dynamic squeezing environment prevalent in real-world gastric applications. The research herein represents a significant stride in leveraging technology for innovative medical solutions, particularly in the management of antibiotic-resistant Helicobacter pylori infections. Sishen Yuan, Baijia Liang, Po Wa Wong, Mingjing Xu, Chi Hsuan Li, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 7 |
| 2024 | Inconstant curvature kinematics of parallel continuum robot without static modelabstractIn the study of minimally invasive surgical robots, a mini parallel continuum robot has shown motion advantage after passing through a long and winding working channel. However, due to the interaction force between the elastic wires of the parallel robots during motion generation processes, the constant curvature assumption has shown modeling errors. This causes the current geometric kinematic model to become unreliable. Therefore, there is a need for a more accurate kinematic model in the absence of a complicated static model. This paper aims to solve this issue. The simulation in ANSYS is carried out, and the shape of one of the driving wires, when bending, is fitted by a two-segment polynomial curve. Then, the position of the distal wrist tip can be calculated based on the curve shape. To verify the accuracy of the proposed model, bending simulation and experiment are carried out. The accuracy of the proposed model is compared with that of the kinematic model based on constant curvature assumption. The result shows that the proposed model can get more accurate results, especially when the driving wire displacement increases. For a 10 mm parallel robot, when the displacements of the two pairs of wires are both 3.0 mm, the errors of the two models are 0.42 mm and 5.79 mm (4.2% and 57.9%), respectively. Huxin Gao, Hongliang Ren 0001 |
ICRA | 3 |
| 2024 | ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention UnderstandingabstractSurgical instrument segmentation is crucial in surgical scene understanding, thereby facilitating surgical safety. Existing algorithms directly detected all instruments of predefined categories in the input image, lacking the capability to segment specific instruments according to the surgeon’s intention. During different stages of surgery, surgeons exhibit varying preferences and focus toward different surgical instruments. Therefore, an instrument segmentation algorithm that adheres to the surgeon’s intention can minimize distractions from irrelevant instruments and assist surgeons to a great extent. The recent Segment Anything Model (SAM) reveals the capability to segment objects following prompts, but the manual annotations for prompts are impractical during the surgery. To address these limitations in operating rooms, we propose an audio-driven surgical instrument segmentation framework, named ASI-Seg, to accurately segment the required surgical instruments by parsing the audio commands of surgeons. Specifically, we propose an intention-oriented multimodal fusion to interpret the segmentation intention from audio commands and retrieve relevant instrument details to facilitate segmentation. Moreover, to guide our ASI-Seg segment of the required surgical instruments, we devise a contrastive learning prompt encoder to effectively distinguish the required instruments from the irrelevant ones. Therefore, our ASI-Seg promotes the workflow in the operating rooms, thereby providing targeted support and reducing the cognitive load on surgeons. Extensive experiments are performed to validate the ASI-Seg framework, which reveals remarkable advantages over classical state-of-the-art and medical SAMs in both semantic segmentation and intention-oriented segmentation. The source code is available at https://github.com/Zonmgin-Zhang/ASI-Seg. Zhen Chen 0018, Zongming Zhang, Wenwu Guo, Xingjian Luo, Long Bai 0008, Jinlin Wu, Hongliang Ren 0001, Hongbin Liu 0001 |
IROS | 7 |
| 2024 | Head-Mounted Hydraulic Needle Driver for Targeted Interventions in NeurosurgeryabstractNeedle interventions are crucial in neurosurgery, requiring high precision and stability. This paper presents a 5-DoF head-mounted hydraulic needle robot designed for accurate and targeted needle insertion and neuroimaging in the deep brain. The robot is compact and lightweight by utilizing a hydraulic pipe transmission to connect the needle driver and actuator. The syringe pistons serve as the actuator and executor, enabling synchronized motion, minimal hysteresis, and high-accuracy insertion. The hydraulic transmission system exhibits hysteresis of less than 0.8 mm, with bidirectional insertion accuracy of approximately 0.05 mm. The resulting needle driver features a compact structure measuring 48 mm × 25 mm × 9 mm, accompanied by a 70-mm-long needle guide. The needle driver is mainly 3D printed, while the hydraulic transmission ensures full compatibility with magnetic resonance imaging (MRI) by isolating all electromagnetic parts from the executor. This compact and lightweight robot-assisted needle intervention system significantly enhances the safety, accuracy, and effectiveness of deep-brain neuroimaging. The feasibility of precise positioning and insertion is further demonstrated by deploying an optical coherence tomography (OCT) microneedle in a rat brain. Zhiwei Fang, Chao Xu 0008, Huxin Gao, Danny Tat-Ming Chan, Wu Yuan 0001, Hongliang Ren 0001 |
IROS | 6 |
| 2024 | Towards Electricity-free Pneumatic Miniature Rotation Actuator for Optical Coherence Tomography EndoscopyabstractMiniature rotation actuators have been extensively developed and utilized in optical coherence tomography (OCT) endoscopy, enabling distortion-free OCT imaging in complex and tortuous environments. However, the use of electrical-driven rotation actuators raises safety concerns. Although magnetic-driven rotation actuators have been reported in OCT endoscopy, their use can potentially interfere with other medical devices in clinical settings. Here, we propose a pneumatic miniature rotation actuator that eliminates the electricity and magnetism concerns in circumferential imaging for OCT endoscopy. The rotor of the actuator is designed as a windmill, enabling it to convert air energy into rotation energy. In addition, to maintain the stable rotation, both a sliding bearing with two supporting points and a glass spindle with a half-ball end surface are developed. The rotation speed of our pneumatic actuator can be controlled from 66 to 97 revolutions per second by adjusting the airflow rate from 3.25 to 4.00 liters per minute. By OCT imaging of the human fingers, we demonstrate the feasibility of the pneumatic actuator in electricity-free distal scanning OCT endoscopy. Our pneumatic rotation actuator has wide-ranging potential in various fiber-imaging modalities, including not only OCT but also ultrasound imaging that requires similar rotation capabilities. Tinghua Zhang, Sishen Yuan, Chao Xu 0008, Hongliang Ren 0001, Wu Yuan 0001 |
IROS | 5 |
| 2024 | EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
Long Bai 0008, Tong Chen 0011, Qiaozhi Tan, Wan Jun Nah, Yanheng Li 0002, Zhicheng He 0010, Sishen Yuan, Zhen Chen 0018, Jinlin Wu, Mobarakol Islam, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
MICCAI (7) | 13 |
| 2024 | LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion
Tong Chen 0011, Qingcheng Lyu, Long Bai 0008, Erjian Guo, Huxin Gao, Xiaoxiao Yang, Hongliang Ren 0001, Luping Zhou |
MICCAI (6) | 7 |
| 2024 | EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
Beilei Cui, Mobarakol Islam, Long Bai 0008, An Wang 0007, Hongliang Ren 0001 |
MICCAI (6) | 5 |
| 2024 | Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting
Yiming Huang 0007, Beilei Cui, Long Bai 0008, Mengya Xu, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (6) | 7 |
| 2024 | SimCol3D - 3D reconstruction during colonoscopy challengeabstractColorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps is challenging. A 3D map of the observed surfaces could enhance the identification of unscreened colon tissue and serve as a training platform. However, reconstructing the colon from video footage remains difficult. Learning-based approaches hold promise as robust alternatives, but necessitate extensive datasets. Establishing a benchmark dataset, the 2022 EndoVis sub-challenge SimCol3D aimed to facilitate data-driven depth and pose prediction during colonoscopy. The challenge was hosted as part of MICCAI 2022 in Singapore. Six teams from around the world and representatives from academia and industry participated in the three sub-challenges: synthetic depth prediction, synthetic pose prediction, and real pose prediction. This paper describes the challenge, the submitted methods, and their results. We show that depth prediction from synthetic colonoscopy images is robustly solvable, while pose estimation remains an open research question. Anita Rau, Sophia Bano, Yueming Jin, Pablo Azagra, Javier Morlana, Rawen Kader, Edward Sanderson, Bogdan J. Matuszewski, Erez Posner, Netanel Frank, Varshini Elangovan, Sista Raviteja, Zhengwen Li, Jiquan Liu, Seenivasan Lalithkumar, Mobarakol Islam, Hongliang Ren 0001, Laurence B. Lovat, J. M. M. Montiel, Danail Stoyanov |
Medical Image Anal. | 19 |
| 2024 | Curriculum-Based Augmented Fourier Domain Adaptation for Robust Medical Image SegmentationabstractAccurate and robust medical image segmentation is fundamental and crucial for enhancing the autonomy of computer-aided diagnosis and intervention systems. Medical data collection normally involves different scanners, protocols, and populations, making domain adaptation (DA) a highly demanding research field to alleviate model degradation in the deployment site. To preserve the model performance across multiple testing domains, this work proposes the Curriculum-based Augmented Fourier Domain Adaptation (Curri-AFDA) for robust medical image segmentation. In particular, our curriculum learning strategy is based on the causal relationship of a model under different levels of data shift in the deployment phase, where the higher the shift is, the harder to recognize the variance. Considering this, we progressively introduce more amplitude information from the target domain to the source domain in the frequency space during the curriculum-style training to smoothly schedule the semantic knowledge transfer in an easier-to-harder manner. Besides, we incorporate the training-time chained augmentation mixing to help expand the data distributions while preserving the domain-invariant semantics, which is beneficial for the acquired model to be more robust and generalize better to unseen domains. Extensive experiments on two segmentation tasks of Retina and Nuclei collected from multiple sites and scanners suggest that our proposed method yields superior adaptation and generalization performance. Meanwhile, our approach proves to be more robust under various corruption types and increasing severity levels. In addition, we show our method is also beneficial in the domain-adaptive classification task with skin lesion datasets. The code is available at https://github.com/lofrienger/Curri-AFDA.Note to Practitioners—Medical image segmentation is key to improving computer-assisted diagnosis and intervention autonomy. However, due to domain gaps between different medical sites, deep learning-based segmentation models frequently encounter performance degradation when deployed in a novel domain. Moreover, model robustness is also highly expected to mitigate the effects of data corruption. Considering all these demanding yet practical needs to automate medical applications and benefit healthcare, we propose the Curriculum-based Fourier Domain Adaptation (Curri-AFDA) for medical image segmentation. Extensive experiments on two segmentation tasks with cross-domain datasets show the consistent superiority of our method regarding adaptation and generalization on multiple testing domains and robustness against synthetic corrupted data. Besides, our approach is independent of image modalities because its efficacy does not rely on modality-specific characteristics. In addition, we demonstrate the benefit of our method for image classification besides segmentation in the ablation study. Therefore, our method can potentially be applied in many medical applications and yield improved performance. Future works may be extended by exploring the integration of curriculum learning regime with Fourier domain amplitude fusion in the testing time rather than in the training time like this work and most other existing domain adaptation works. An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Confidence-Aware Paced-Curriculum Learning by Label Smoothing for Surgical Scene UnderstandingabstractCurriculum learning and self-paced learning are the training strategies that gradually feed the samples from easy to more complex. They have captivated increasing attention due to their excellent performance in robotic vision. Most recent works focus on designing curricula based on difficulty levels in input samples or smoothing the feature maps. However, smoothing labels to control the learning utility in a curriculum manner is still unexplored. In this work, we design a paced curriculum by label smoothing (P-CBLS) using paced learning with uniform label smoothing (ULS) for classification tasks and fuse uniform and spatially varying label smoothing (SVLS) for semantic segmentation tasks in a curriculum manner. In ULS and SVLS, a bigger smoothing factor value enforces a heavy smoothing penalty in the true label and limits learning less information. Therefore, we design the curriculum by label smoothing (CBLS). We set a bigger smoothing value at the beginning of training and gradually decreased it to zero to control the model learning utility from lower to higher. We also designed a confidence-aware pacing function and combined it with our CBLS to investigate the benefits of various curricula. The proposed techniques are validated on four robotic surgery datasets of multi-class, multi-label classification, captioning, and segmentation tasks. We also investigate the robustness of our method by corrupting validation data into different severity levels. Our extensive analysis shows that the proposed method improves prediction accuracy and robustness. The code is publicly available at https://github.com/XuMengyaAmy/P-CBLS.Note to Practitioners—The motivation of this article is to improve the performance and robustness of deep neural networks in safety-critical applications such as robotic surgery by controlling the learning ability of the model in a curriculum learning manner and allowing the model to imitate the cognitive process of humans and animals. The designed approaches do not add parameters that require additional computational resources. Mengya Xu, Mobarakol Islam, Ben Glocker, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Sim-to-Real Transfer of Soft Robotic Navigation Strategies That Learns From the Virtual Eye-in-Hand VisionabstractTo steer a soft robot precisely in an unconstructed environment with minimal collision remains an open challenge for soft robots. When the environments are unknown, prior motion planning for navigation may not always be available. This paper presents a novel Sim-to-Real method to guide a cable-driven soft robot in a static environment under the Simulation Open Framework Architecture (SOFA). The scenario aims to resemble one of the steps during a simplified transoral tracheal intubation process where a robotic endotracheal tube is guided to the upper trachea-larynx location by a flexible video-assisted endoscope/stylet. In SOFA, we employ the quadratic programming inverse solver to obtain collision-free motion strategies for the endoscope/stylet manipulation based on the robot model and encode the virtual eye-in-hand vision. Then, we associate the anatomical features recognized by the virtual vision and the joint space motion using a closed-loop nonlinear autoregressive exogenous model (NARX) network. Afterward, we transfer the learned knowledge to the robot prototype, expecting it to navigate to the desired spot in a new phantom environment automatically based on its eye-in-hand vision only. Experiment results indicate that our soft robot can efficaciously navigate through the unstructured phantom to the desired spot with minimal collision motion according to what it has learned from the virtual environment. The results show that the average R-squared coefficient between the closed-loop NARX-forecasted and SOFA-referenced robot's cable and prismatic joint space motion are 0.963 and 0.997, respectively. The eye-in-hand visions also demonstrate good alignment between the robot tip and the glottis. Jiewen Lai, Tian-Ao Ren, Wenchao Yue, Shijian Su, Jason Ying-Kuen Chan, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | A Wearable, Reconfigurable, and Modular Magnetic Tracking System for Wireless Capsule RobotsabstractWearable magnetic tracking systems (MTSs) offer a promising technology for the long-term tracking of wireless-capsule robots within the digestive tract. However, existing wearable MTSs are fixed in size and cannot accommodate patients with diverse abdominal circumferences. To address this limitation, we propose a wearable and reconfigurable MTS. First, we design a reconfigurable sensor array inspired by the structure of bamboo slips, allowing it to conform to the abdominal surface and accommodate individuals with different abdominal circumferences. Next, we formulate a magnetic tracking optimization problem based on the magnetic dipole model and our established kinematic model of the reconfigurable sensor array. Solving the magnetic tracking problem, we achieved outstanding localization accuracy of 1.44$\pm$0.50 mm and 1.07$\pm 0.16^\circ$. Experimental validation demonstrates our proposed system's portability, reconfigurability, and adaptability to varying abdominal circumferences, offering valuable technological means for diagnosing and treating gastrointestinal disorders. Shijian Su, Sishen Yuan, Zhen Li 0026, Miaomiao Ma, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | An Efficient MLP-Based Point-Guided Segmentation Network for Ore Images With Ambiguous BoundaryabstractThe precise segmentation of ore images is critical to the successful execution of the beneficiation process. Due to the homogeneous appearance of the ores, which leads to low contrast and unclear boundaries, accurate segmentation becomes challenging, and recognition becomes problematic. This article proposes a lightweight framework based on multilayer perceptron (MLP), which focuses on solving the problem of edge blurring. Specifically, we introduce a lightweight backbone better suited for efficiently extracting low-level features. Besides, we design a feature pyramid network consisting of two MLP structures that balance local and global information, thus, enhancing detection accuracy. Furthermore, we propose a novel loss function that guides the prediction points to match the instance edge points to achieve clear object boundaries. We have conducted extensive experiments to validate the efficacy of our proposed method. Our approach achieves a remarkable processing speed of over 27 frames per second with a model size of only 73 MB. Moreover, our method delivers a consistently high level of accuracy, with impressive performance scores of 60.4 and 48.9 in$AP_{50}^{\text{box}}$and$AP_{50}^{\text{mask}}$, respectively, as compared with the currently available state-of-the-art techniques, when tested on the ORE image dataset. Guodong Sun 0002, Yuting Peng 0001, Mengya Xu, An Wang 0007, Hongliang Ren 0001, Yang Zhang 0053 |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | TransFSM: Fetal Anatomy Segmentation and Biometric Measurement in Ultrasound Images Using a Hybrid TransformerabstractBiometric parameter measurements are powerful tools for evaluating a fetus's gestational age, growth pattern, and abnormalities in a 2D ultrasound. However, it is still challenging to measure fetal biometric parameters automatically due to the indiscriminate confusing factors, limited foreground-background contrast, variety of fetal anatomy shapes at different gestational ages, and blurry anatomical boundaries in ultrasound images. The performance of a standard CNN architecture is limited for these tasks due to the restricted receptive field. We propose a novel hybrid Transformer framework, TransFSM, to address fetal multi-anatomy segmentation and biometric measurement tasks. Unlike the vanilla Transformer based on a single-scale input, TransFSM has a deformable self-attention mechanism so it can effectively process multi-scale information to segment fetal anatomy with irregular shapes and different sizes. We devised a BAD to capture more intrinsic local details using boundary-wise prior knowledge, which compensates for the defects of the Transformer in extracting local features. In addition, a Transformer auxiliary segment head is designed to improve mask prediction by learning the semantic correspondence of the same pixel categories and feature discriminability among different pixel categories. Extensive experiments were conducted on clinical cases and benchmark datasets for anatomy segmentation and biometric measurement tasks. The experiment results indicate that our method achieves state-of-the-art performance in seven evaluation metrics compared with CNN-based, Transformer-based, and hybrid approaches. By Knowledge distillation, the proposed TransFSM can create a more compact and efficient model with high deploying potential in resource-constrained scenarios. Our study serves as a unified framework for biometric estimation across multiple anatomical regions to monitor fetal growth in clinical practice. Lei Zhao 0013, Guanghua Tan, Bin Pu, Qianghui Wu, Hongliang Ren 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | FARN: Fetal Anatomy Reasoning Network for Detection With Global Context Semantic and Local Topology RelationshipabstractAccurate recognition of fetal anatomical structure is a pivotal task in ultrasound (US) image analysis. Sonographers naturally apply anatomical knowledge and clinical expertise to recognizing key anatomical structures in complex US images. However, mainstream object detection approaches usually treat each structure recognition separately, overlooking anatomical correlations between different structures in fetal US planes. In this work, we propose a Fetal Anatomy Reasoning Network (FARN) that incorporates two kinds of relationship forms: a global context semantic block summarized with visual similarity and a local topology relationship block depicting structural pair constraints. Specifically, by designing the Adaptive Relation Graph Reasoning (ARGR) module, anatomical structures are treated as nodes, with two kinds of relationships between nodes modeled as edges. The flexibility of the model is enhanced by constructing the adaptive relationship graph in a data-driven way, enabling adaptation to various data samples without the need for predefined additional constraints. The feature representation is further refined by aggregating the outputs of the ARGR module. Comprehensive experimental results demonstrate that FARN achieves promising performance in detecting 37 anatomical structures across key US planes in tertiary obstetric screening. FARN effectively utilizes key relationships to improve detection performance, demonstrates robustness to small-scale, similar, and indistinct structures, and avoids some detection errors that deviate from anatomical norms. Overall, our study serves as a resource for developing efficient and concise approaches to model inter-anatomy relationships. Lei Zhao 0013, Guanghua Tan, Qianghui Wu, Bin Pu, Hongliang Ren 0001, Shengli Li 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Needle Trajectory Prediction for Percutaneous Kidney Biopsy in 5G-Powered Teleultrasound Navigation SystemabstractNeedle insertion is a critical component of many remote surgical procedures, including biopsies, injections, neurosurgery, and brachytherapy cancer treatments. However, precise visualization of the biopsy needle trajectory remains challenging due to specular reflection, speckle noise, and needle-like anatomical features. This paper proposes a visual feedback prediction framework for ultrasound-assisted percutaneous kidney biopsy in 5G remote surgery, aiming to enhance operator confidence, reduce procedure time, and minimize the risk of unintended bleeding. Building upon this framework, we design a Lightweight-Accuracy Needle Trajectory Prediction (LA-NTP) model by minimizing the backbone and optimizing the multi-module prediction process, incorporating innovative training strategies (i.e., angle-aware geometric and trajectory augmentation losses). The experimental results demonstrate that it achieves competitive performance with only 20.3% of the model size of the previous best real-time method and a 3.7-fold increase in inference speed. Even in challenging scenarios involving large insertion depths and steep angles, our method provides stable and precise navigation. Lei Zhao 0013, Guanghua Tan, Jiewen Lai, Chwee Ming Lim, Weng Kin Wong, Hongliang Ren 0001, Kenli Li 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic SurgeryabstractDeep Neural Networks (DNNs) based semantic segmentation of the robotic instruments and tissues can enhance the precision of surgical activities in robot-assisted surgery. However, in biological learning, DNNs cannot learn incremental tasks over time and exhibit catastrophic forgetting, which refers to the sharp decline in performance on previously learned tasks after learning a new one. Specifically, when data scarcity is the issue, the model shows a rapid drop in performance on previously learned instruments after learning new data with new instruments. The problem becomes worse when it limits releasing the dataset of the old instruments for the old model due to privacy concerns and the unavailability of the data for the new or updated version of the instruments for the continual learning model. For this purpose, we develop a privacy-preserving synthetic continual semantic segmentation framework by blending and harmonizing (i) open-source old instruments foreground to the synthesized background without revealing real patient data in public and (ii) new instruments foreground to extensively augmented real background. To boost the balanced logit distillation from the old model to the continual learning model, we design overlapping class-aware temperature normalization (CAT) by controlling model learning utility. We also introduce multi-scale shifted-feature distillation (SD) to maintain long and short-range spatial relationships among the semantic objects where conventional short-range spatial features with limited information reduce the power of feature distillation. We demonstrate the effectiveness of our framework on the EndoVis 2017 and 2018 instrument segmentation dataset with a generalized continual learning setting. Code is available at https://github.com/XuMengyaAmy/Synthetic_CAT_SD. Mengya Xu, Mobarakol Islam, Long Bai 0008, Hongliang Ren 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Surgical-VQLA:Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic SurgeryabstractDespite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, expert surgeons are often overloaded with clinical and academic workloads and limit their time in answering. For this purpose, we develop a surgical question-answering system to facilitate robot-assisted surgical scene and activity understanding from recorded videos. Most of the existing visual question answering (VQA) methods require an object detector and regions based feature extractor to extract visual features and fuse them with the embedded text of the question for answer generation. However, (i) surgical object detection model is scarce due to smaller datasets and lack of bounding box annotation; (ii) current fusion strategy of heterogeneous modalities like text and image is naive; (iii) the localized answering is missing, which is crucial in complex surgical scenarios. In this paper, we propose Visual Question Localized-Answering in Robotic Surgery (Surgical-VQLA) to localize the specific surgical area during the answer prediction. To deal with the fusion of the heterogeneous modalities, we design gated vision-language embedding (GVLE) to build input patches for the Language Vision Transformer (LViT) to predict the answer. To get localization, we add the detection head in parallel with the prediction head of the LViT. We also integrate generalized intersection over union (GIoU) loss to boost localization performance by preserving the accuracy of the question-answering model. We annotate two datasets of VQLA by utilizing publicly available surgical videos from EndoVis-17 and 18 of the MICCAI challenges. Our validation results suggest that Surgical-VQLA can better understand the surgical scene and localized the specific area related to the question-answering. GVLE presents an efficient language-vision embedding technique by showing superior performance over the existing benchmarks. Long Bai 0008, Mobarakol Islam, Seenivasan Lalithkumar, Hongliang Ren 0001 |
ICRA | 4 |
| 2023 | Generalizing Surgical Instruments Segmentation to Unseen Domains with One-to-Many SynthesisabstractDespite their impressive performance in various surgical scene understanding tasks, deep learning-based methods are frequently hindered from deploying to real-world surgical applications for various causes. Particularly, data collection, annotation, and domain shift in-between sites and patients are the most common obstacles. In this work, we mitigate data-related issues by efficiently leveraging minimal source images to generate synthetic surgical instrument segmentation datasets and achieve outstanding generalization performance on unseen real domains. Specifically, in our framework, only one background tissue image and at most three images of each foreground instrument are taken as the seed images. These source images are extensively transformed and employed to build up the foreground and background image pools, from which randomly sampled tissue and instrument images are composed with multiple blending techniques to generate new surgical scene images. Besides, we introduce hybrid training-time augmentations to diversify the training data further. Extensive evaluation on three real-world datasets, i.e., Endo2017, Endo2018, and RoboTool, demonstrates that our one-to-many synthetic surgical instruments datasets generation and segmentation framework can achieve encouraging performance compared with training with real data. Notably, on the RoboTool dataset, where a more significant domain gap exists, our framework shows its superiority of generalization by a considerable margin. We expect that our inspiring results will attract research attention to improving model generalization with data synthesizing. An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
IROS | 4 |
| 2023 | LLCaps: Learning to Illuminate Low-Light Capsule Endoscopy with Curved Wavelet Attention and Reverse Diffusion
Long Bai 0008, Tong Chen 0011, Yanan Wu 0003, An Wang 0007, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (10) | 6 |
| 2023 | Revisiting Distillation for Continual Learning on Visual Question Localized-Answering in Robotic Surgery
Long Bai 0008, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (9) | 3 |
| 2023 | CAT-ViL: Co-attention Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery
Long Bai 0008, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (9) | 3 |
| 2023 | Rectifying Noisy Labels with Sequential Prior: Multi-scale Temporal Feature Affinity Learning for Robust Video Segmentation
Beilei Cui, Minqing Zhang, Mengya Xu, An Wang 0007, Wu Yuan 0001, Hongliang Ren 0001 |
MICCAI (9) | 6 |
| 2023 | SurgicalGPT: End-to-End Language-Vision GPT for Visual Question Answering in Surgery
Seenivasan Lalithkumar, Mobarakol Islam, Gokul Kannan, Hongliang Ren 0001 |
MICCAI (9) | 4 |
| 2023 | S2ME: Spatial-Spectral Mutual Teaching and Ensemble Learning for Scribble-Supervised Polyp Segmentation
An Wang 0007, Mengya Xu, Yang Zhang 0053, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (1) | 5 |
| 2023 | Two-stage contextual transformer-based convolutional neural network for airway extraction from CT images
Yanan Wu 0003, Shuiqing Zhao, Shouliang Qi, Jie Feng 0009, Haowen Pang, Runsheng Chang, Long Bai 0008, Shuyue Xia, Wei Qian 0001, Hongliang Ren 0001 |
Artif. Intell. Medicine | 11 |
| 2023 | CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu 0009, Armine Vardazaryan, Fangfang Xia, Tong Xia, Fucang Jia, Yuxuan Yang 0007, Hao Wang 0081, Derong Yu, Guoyan Zheng, Xiaotian Duan, Neil Getty, Ricardo Sanchez-Matilla, Maria Robu, Li Zhang 0040, Huabin Chen, Jiacheng Wang 0002, Liansheng Wang 0002, Beerend G. A. Gerats, Sista Raviteja, Rachana Sathish, Rong Tao, Satoshi Kondo, Winnie Pang, Hongliang Ren 0001, Julian Ronald Abbing, Mohammad Hasan Sarhan, Sebastian Bodenstedt, Nithya Bhasker, Bruno Oliveira 0002, Helena R. Torres, Finn Gaida, Tobias Czempiel, João L. Vilaça, Pedro Morais, Jaime C. Fonseca 0001, Ruby Mae Egging, Inge Nicole Wijma, Chen Qian 0006, Guibin Bian, Zhen Li 0026, Velmurugan Balasubramanian, Debdoot Sheet, Imanol Luengo, Yuanbo Zhu, Shuai Ding 0001, Jakob-Anton Aschenbrenner, Nicolas Elini van der Kar, Mengya Xu, Mobarakol Islam, Seenivasan Lalithkumar, Alexander Jenke, Danail Stoyanov, Didier Mutter, Pietro Mascagni, Barbara Seeliger, Cristians Gonzalez, Nicolas Padoy |
Medical Image Anal. | 28 |
| 2023 | SAVAnet: Surgical Action-Driven Visual Attention Network for Autonomous Endoscope ControlabstractAn endoscope holder must understand the detailed surgical actions and the surgeons’ visual attention to keep important targets in the field of endoscopic view during operations. From an intensive analysis of the surgeons’ attention mechanism, we included that surgical actions, like cutting, suturing, etc., play an important role in determining the positions and weights of visual attention points during a dynamic surgical scene. To perform this process, this work proposes a Surgical Action-driven Visual Attention network (SAVAnet) and applies the network in autonomous endoscope control. Four scenarios are constructed in the da Vinci V-rep simulator: pick&place and needle exercise in a general laparoscopic training environment, needle driving with and without obstacle removal in an abdominal cavity, to create datasets for network training. The results show that the network has an outstanding performance in surgical action prediction with a high average accuracy of over 91%. Additionally, with surgical action guidance, the attention point prediction has higher accuracy and accords with surgeons’ visual attention. Finally, the acquired attention points are utilized to execute visual servoing in simulation. The results verify that the SAVAnet is feasible for autonomous endoscope control in real-time and lays a theoretical foundation for future sim-to-real execution. Note to Practitioners—This paper was motivated by the problem of endowing an endoscope with surgeons’ visual attention mechanism, which is affected by surgical actions, for autonomous endoscope control. An eye-tracking device has been utilized to detect surgeon’s visual attention in real-time and then control the endoscope to follow what the surgeon is looking at. However, this approach is susceptible to the surgical environment. Besides, many instrument detection and segmentation algorithms are developed for automatic surgical instrument tracking. However, surgeons’ visual attention does not always focus on the instruments during operations. In this work, we propose a novel SAVAnet to determine visual attention based on surgical actions. We prove from many qualitative and quantitative experiments that surgical actions play a significant role in determining visual attention. The designed SAVAnet can predict surgical actions correctly and then effectively guide the choice of visual attention. Finally, the simulation results show that the SAVAnet can endow endoscope with surgeons’ visual attention to perform self-control in real time. In future research, we will train the SAVAnet using real datasets and conduct more physical experiments on real surgical robots. Huxin Gao, Weichen Fan, Liang Qiu 0002, Xiaoxiao Yang, Zhen Li 0026, Xiuli Zuo, Max Q.-H. Meng, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2023 | AMagPoseNet: Real-Time Six-DoF Magnet Pose Estimation by Dual-Domain Few-Shot Learning From Prior ModelabstractTraditional magnetic tracking approaches based on mathematical models and optimization algorithms are computationally intensive, depend on initial guesses, and do not guarantee convergence to a global optimum. Although fully supervised data-driven deep learning can solve the above issues, the demand for a comprehensive dataset hampers its applicability in magnetic tracking. Thus, we propose an annular magnet pose estimation network (called AMagPoseNet) based on dual-domain few-shot learning from a prior mathematical model, which consists of two subnetworks: PoseNet and CaliNet. PoseNet learns to estimate the magnet pose from the prior mathematical model, and CaliNet is designed to narrow the gap between the mathematical model domain and the real-world domain. Experimental results reveal that the AMagPoseNet outperforms the optimization-based method regarding localization accuracy (1.87$\pm$1.14 mm, 1.89$\pm \text{0.81}^{\circ }$), robustness (nondependence on initial guesses), and computational latency (2.08$\pm$0.02 ms). In addition, the six-degree-of-freedom pose of the magnet could be estimated when discriminative magnetic field features are provided. With the assistance of the mathematical model, the AMagPoseNet requires only a few real-world samples and has excellent performance, showing great potential for practical biomedical and industrial applications. Shijian Su, Sishen Yuan, Mengya Xu, Huxin Gao, Xiaoxiao Yang, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | Federated Semi-Supervised Learning for Medical Image Segmentation via Pseudo-Label DenoisingabstractDistributed big data and digital healthcare technologies have great potential to promote medical services, but challenges arise when it comes to learning predictive model from diverse and complex e-health datasets. Federated Learning (FL), as a collaborative machine learning technique, aims to address the challenges by learning a joint predictive model across multi-site clients, especially for distributed medical institutions or hospitals. However, most existing FL methods assume that clients possess fully labeled data for training, which is often not the case in e-health datasets due to high labeling costs or expertise requirement. Therefore, this work proposes a novel and feasible approach to learn a Federated Semi-Supervised Learning (FSSL) model from distributed medical image domains, where a federated pseudo-labeling strategy for unlabeled clients is developed based on the embedded knowledge learned from labeled clients. This greatly mitigates the annotation deficiency at unlabeled clients and leads to a cost-effective and efficient medical image analysis tool. We demonstrated the effectiveness of our method by achieving significant improvements compared to the state-of-the-art in both fundus image and prostate MRI segmentation tasks, resulting in the highest Dice scores of 89.23% and 91.95% respectively even with only a few labeled clients participating in model training. This reveals the superiority of our method for practical deployment, ultimately facilitating the wider use of FL in healthcare and leading to better patient outcomes. Liang Qiu 0002, Jierong Cheng, Huxin Gao, Wei Xiong 0001, Hongliang Ren 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Magnetic Tracking With Real-Time Geomagnetic Vector Separation for Robotic Dockable ChargingabstractHigh-precision pose adjustment for the self-charging of mobile robots remains a significant challenge. Permanent magnet (PM)-based magnetic tracking technique is a promising technical solution, with occlusion-free and simultaneous positioning and orientation tracking. However, the superposition of the geomagnetic vector and the magnetic field vector generated by the PM leads to the degrading of magnetic tracking performance. Thus, a magnetic tracking technique with real-time geomagnetic vector separation is investigated in this study. Firstly, the environmental magnetic field is accurately modeled, consisting of the PM field, uniform disturbance field, and non-uniform disturbance field. For the uniform disturbance field, we combine it with the PM pose as unknown parameters to be estimated. For the non-uniform disturbance field, a robust kernel function is employed to diminish its influence on positioning performance. Finally, the PM pose and geomagnetic vector are simultaneously estimated by optimization algorithms. A docking experiment for self-charging mobile robots was carried out based on the proposed tracking technique. The robot can successfully recharge its battery with only one alignment operation, where the repeat parking accuracy at the anchor point is 1.38 mm and ±1.27°, respectively. Shijian Su, Houde Dai, Sishen Yuan, Shuang Song 0002, Hongliang Ren 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | GESRsim: Gastrointestinal Endoscopic Surgical Robot SimulatorabstractRobot-assisted gastrointestinal endoscopic surgery (GES) as a kind of natural orifice transluminal endoscopic surgery (NOTES) is the next-generation minimally invasive surgery (MIS). Besides, rendering certain autonomy to a Gas-trointestinal Endoscopic Surgical Robot (GESR) is promising but highly challenging. Therefore, to accelerate the development and augment the autonomy of GESR, we use CoppeliaSim to develop the first robotic simulator for the GESR system (GESRsim) based on our previous design. The GESRsim provides several 3D models and kinematics of our designed manipulators and endoscopic snake bone. Additionally, we build several scenes for robotic GES training and then utilize different programming interfaces to perform teleoperation. Furthermore, several advanced control algorithms, including visual servoing (VS) and deep reinforcement learning (DRL), are implemented to verify the performance of the GESRsim. Huxin Gao, Zedong Zhang, Xiao Xiao 0006, Liang Qiu 0002, Xiaoxiao Yang, Ruoyi Hao, Xiuli Zuo, Hongliang Ren 0001 |
IROS | 10 |
| 2022 | Model-free and Uncalibrated Visual-feedback Control of Magnetically-Actuated Flexible EndoscopesabstractMagnetically-actuated flexible endoscopes (MAFE) have been well used in minimally-invasive surgery because they can be steered by a magnetic field thus more flexible than traditional endoscopes. Model-free and uncalibrated visual-feedback control makes it possible to manipulate MAFE with a magnetic field without external tracking systems. Because no extra sensor is required to obtain position and posture information, the size of MAFE can be made smaller. However, the traditional control method focuses on 2DoF control, which lacks control over the posture of the end of MAFE. This may result in unnecessary contact between MAFE and tissue and cause injury during the advancement of the endoscope. In this letter, we propose algorithms to enhance the pose control of MAFE to 4DoF and 5DoF based on model-free and uncalibrated visual-feedback control. Experiments in structured environments verify that the control algorithms are able to realize 4DoF manual navigation and 5DoF automatic navigation. Jiewen Tan, Junnan Xue, Xing Yang 0005, Sishen Yuan, Wei Liu 0134, Hongliang Ren 0001, Shuang Song 0002, Jiaole Wang |
IROS | 6 |
| 2022 | Surgical-VQA: Visual Question Answering in Surgical Scenes Using Transformer
Seenivasan Lalithkumar, Mobarakol Islam, Adithya K. Krishna, Hongliang Ren 0001 |
MICCAI (8) | 4 |
| 2022 | Rethinking Surgical Instrument Segmentation: A Background Image Can Be All You Need
An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
MICCAI (8) | 4 |
| 2022 | Rethinking Surgical Captioning: End-to-End Window-Based MLP Transformer Using Patches
Mengya Xu, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (8) | 3 |
| 2022 | RSegNet: A Joint Learning Framework for Deformable Registration and SegmentationabstractMedical image segmentation and registration are two tasks to analyze the anatomical structures in clinical research. Still, deep-learning solutions utilizing the connections between segmentation and registration remain underdiscovered. This article designs a joint learning framework named RSegNet that can realize concurrent deformable registration and segmentation by minimizing an integrated loss function, including three parts: diffeomorphic registration loss, segmentation similarity loss, and dual-consistency supervision loss. The probabilistic diffeomorphic registration branch could benefit from the auxiliary segmentations available from the segmentation branch to achieve anatomical consistency and better deformation regularity by dual-consistency supervision. Simultaneously, the segmentation performance could also be improved by data augmentation based on the registration with well-behaved diffeomorphic guarantees. Experiments on the human brain 3-D magnetic resonance images have been implemented to demonstrate the effectiveness of our approach. We trained and validated RSegNet with 1000 images and tested its performances on four public datasets, which shows that our method successfully yields concurrent improvements of both segmentation and registration compared with separately trained networks. Specifically, our method can increase the accuracy of segmentation and registration by 7.0% and 1.4%, respectively, in terms of Dice scores.Note to Practitioners—Registration and segmentation of medical images are two significant tasks in medical research and clinical application. However, most existing approaches consider these two tasks independently while neglecting the potential association between them. Therefore, we suggest a new approach that combines these two tasks into one joint deep learning framework, boosting registration, and segmentation performance by introducing dual-consistency supervision. Besides, our framework could generate outputs within 1 s by taking an affinely aligned medical image pair as input, which is suitable for time-critical requirements in a clinic. We tested it on four public datasets and achieved state-of-the-art performance to demonstrate the proposed method’s feasibility and robustness. Furthermore, our proposed RSegNet is a general learning framework suitable for various image modalities and anatomical structures. Hence, we expect our framework to serve as a practical clinical tool to speed up medical image analysis procedures and improve diagnostic accuracy. Liang Qiu 0002, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2021 | Kirigami Strain Sensing on Balloon Catheters with Temporary Tattoo PaperabstractThe current state of the art of balloon catheters relies solely on the application of a predetermined quantity of mechanical strain to the balloon during diagnostic and therapeutic procedures. In some cases, the surgeons can use radioactive-contrasting agents and x-ray screening to identify the correct position and size of the inflated balloon. Otherwise, there is little information on the inflated size of the balloon catheter in the occluded lumen. This gap in quantitative feedback of the ballooning behavior needs to be addressed to ensure safe operation. With the advancement in technology and breakthrough in flexible electronics in recent years, kirigami, an ancient cutting, bending and folding technique, is explored in the stretchable sensing field due to its ability to transform 2D planar patterns 3D geometry structures. On top of that, kirigami can increase the mechanical strain by over 300% depending on the different cuts and folds and sensitivity over 80%. This manuscript will address this limitation of conventional balloon catheters by introducing strain sensing using kirigami technology to achieve a better, safer, and more efficient treatment procedure. Experimental results show that the change in normalized resistance of the sensor is directly proportional to the change in the size of the balloon. Bok Seng Yeow, Godwin Ponraj, Kirthika Senthil Kumar, Catherine Jiayi Cai, Hongliang Ren 0001 |
BSN | 6 |
| 2021 | Remote-Center-of-Motion Recommendation toward Brain Needle Intervention Using Deep Reinforcement LearningabstractBrain needle intervention is a specific diagnosis and therapy procedure in brain disorders, such as brain tumors and Parkinson’s disease. Preoperative needle path planning is a vital step to guarantee the patient’s safety and reduce lesions. For positioning accuracy in the CT/MRI environment, we have developed a novel needle intervention robot in our previous work. Because the robot is currently designed for the rigid needle, the task of preoperative path-planning is to search for an optimal Remote Center of Motion (RCM) for needle insertion. Therefore, this work proposes an RCM recommendation system using deep reinforcement learning. Considering the robot kinematics, this system takes the following criteria/constraints into consideration: clinical obstacle (blood vessels, tissues) avoidance (COA), mechanically inverse kinematics (MIK) and mechanically less motion (MLM) for the robot. We design a reward function to combine the above three criteria based on their corresponding importance level and utilize proximal policy optimization (PPO) as the main agent of reinforcement learning (RL). RL methods are proved to be competent in searching the RCM, which satisfies the above criteria simultaneously. On the one hand, the results present that RL agents obtain the success rate of finishing the designed task at 93%, which has reached the human level in the tests. On the other hand, the RL agents have the remarkable capability of combining more complex criteria/constraints in future work. Huxin Gao, Xiao Xiao 0006, Liang Qiu 0002, Max Q.-H. Meng, Nicolas Kon Kam King, Hongliang Ren 0001 |
ICRA | 6 |
| 2021 | Chip-Less Wireless Sensing of Kirigami Structural Morphing Under Various Mechanical Stimuli Using Home-Based Ink-Jet Printable MaterialsabstractThe feasibility of using chip-based RFID designs as wireless sensor tags open a wide range of application possibilities in the field of robotics. However, multi-step lithography manufacturing and/or MEMS techniques are often required for the industrial-grade fabrication of such sensors. In this paper, we present a simple, home-based, two-step fabrication process to produce chipless RF-based wireless sensors. We use an office-based inkjet printer to produce the antenna traces using silver conductive ink. Kirigami-inspired designs are used to produce four sensors responsive to various mechanical stimuli commonly used for DIY robotic projects (contact, compression, extension, and bend). We demonstrate sensing and wireless transmission of the detected mechanical stimuli through the proposed chipless sensor tags with reliable consistency. The tags can be replicated quickly with the inkjet-printing method. This paper also contains analyses on the effects of varying the dimensions and electrical parameters of the tags. The developed antennas in this work can be used as wireless mechano-responsive sensors for robotic applications. Godwin Ponraj, Wei Le Yeo, Kirthika Senthil Kumar, Manivannan Sivaperuman Kalairaj, Catherine Jiayi Cai, Hongliang Ren 0001 |
ICRA | 6 |
| 2021 | Magnetically-Connected Modular Reconfigurable Mini-robotic System with Bilateral Isokinematic Mapping and Fast On-site Assembly towards Minimally Invasive ProceduresabstractThis paper presents a modular and reconfigurable mini-robotic system with 5 degrees of freedom (DoFs) towards minimally invasive surgery (MIS). The mini-robotic system consists of two modules, a 2-DoFs rotational end-effector, and a 3-DoFs positioning platform. The 2-DoFs rotational end-effector is based on a spring-spherical joint mechanism, whose rotation is controlled by Bowden-cable. The 3-DoFs positioning platform is based on the linear Delta parallel mechanism. Magnetic spherical joints are adopted to replace the traditional spherical joint. The magnetic joint connections enable fast assembling and disassemble of the end platform and kinematic chains. Different surgical instruments can be installed without changing the driver and control system. A flexible shaft actuates the 3-DoFs positioning platform to arrange the motors away from the manipulator side. Based on these structure characteristics, the 3-DoFs positioning platform’s size is dramatically reduced. The outer diameter of the current prototype is 32.5 mm. The single-axis positioning accuracy of the 3-DoFs positioning platform is within -1 mm to 0.85 mm. Three axes tracking experiments are also carried out, with the positioning errors of ± 1.2 mm for cylindrical curves and -1.5 mm to 2 mm for spherical helix curves. Static and dynamic load capabilities are also tested. Finally, the feasibility of the proposed system is demonstrated. Xiao Xiao 0006, Shilei Xu, Huxin Gao, Max Q.-H. Meng, Hongliang Ren 0001 |
ICRA | 7 |
| 2021 | Learning Domain Adaptation with Model Calibration for Surgical Report Generation in Robotic SurgeryabstractGenerating a surgical report in robot-assisted surgery, in the form of natural language expression of surgical scene understanding, can play a significant role in document entry tasks, surgical training, and post-operative analysis. Despite the state-of-the-art accuracy of the deep learning algorithm, the deployment performance often drops when applied to the Target Domain (TD) data. For this purpose, we develop a multi-layer transformer-based model with the gradient reversal adversarial learning to generate a caption for the multi-domain surgical images that can describe the semantic relationship between instruments and surgical Region of Interest (ROI). In the gradient reversal adversarial learning scheme, the gradient multiplies with a negative constant and updates adversarially in backward propagation, discriminating between the source and target domains and emerging domain-invariant features. We also investigate model calibration with label smoothing technique and the effect of a well-calibrated model for the penultimate layer’s feature representation and Domain Adaptation (DA). We annotate two robotic surgery datasets of MICCAI robotic scene segmentation and Transoral Robotic Surgery (TORS) with the captions of procedures and empirically show that our proposed method improves the performance in both source and target domain surgical reports generation in the manners of unsupervised, zero-shot, one-shot, and few-shot learning. Mengya Xu, Mobarakol Islam, Chwee Ming Lim, Hongliang Ren 0001 |
ICRA | 4 |
| 2021 | Origami-Inspired Snap-through Bistability in Parallel and Curved Mechanisms Through the Inflection of Degree Four VertexesabstractOrigami, the art of folding paper, can impart useful design inspirations to the creation of mechanical structures and mechanisms. Bistability is a useful property for origami designs, which can help compartmentalize different actuations and stiffness tuning regimes. Given the benefits of bistability, we investigated origami designs used to build robots and deployable structures. We show snap-through bistable designs applied to parallel and curved mechanisms, which are of value to robotic mechanical design. The designs proposed were investigated through geometry analysis and stress-strain experiments. The origami designs were modified to show that the mechanical properties of our bistable designs can be modulated. Initial actuation utilized magnetic and tendon-driven mechanisms for the parallel and curved structures, respectively. We anticipate that these bistable snap-through designs can contribute to deployable mechanisms and give such devices additional capabilities in mechanical response tuning. Bok Seng Yeow, Catherine Jiayi Cai, Manivannan Sivaperuman Kalairaj, Feng Wen Hoo, Zu Xuan Lee, Janice Chui Shien Tan, Jian Rong Ho, Vienna Minhui Ma, Hongliang Ren 0001 |
ICRA | 10 |
| 2021 | Class-Incremental Domain Adaptation with Smoothing and Calibration for Surgical Report Generation
Mengya Xu, Mobarakol Islam, Chwee Ming Lim, Hongliang Ren 0001 |
MICCAI (4) | 4 |
| 2021 | U-RSNet: An unsupervised probabilistic model for joint registration and segmentation
Liang Qiu 0002, Hongliang Ren 0001 |
Neurocomputing | 2 |
| 2021 | ST-MTL: Spatio-Temporal multitask learning model to predict scanpath while tracking instruments in robotic surgery
Mobarakol Islam, Vibashan VS, Chwee Ming Lim, Hongliang Ren 0001 |
Medical Image Anal. | 4 |
| 2021 | Soft Robotic Gripper Driven by Flexible Shafts for Simultaneous Grasping and In-Hand Cap ManipulationabstractPerforming a successful robotic grasping to uncertain objects in unstructured environments is challenging. This study presents a new compliant soft robotic gripper for objects handling and cap manipulation through the coordination of three soft fingers and in-hand manipulation. The experiments are conducted to validate that the soft robotic gripper can successfully realize simultaneous grasping and capping manipulations with only one flexible shaft actuation for every single soft finger.Note to Practitioners—Uncertain object manipulation tasks pose significant challenges to a robotic gripper while grasping and capping unknown objects without damaging them. The existing rigid grippers have experienced flexible manipulation through multiple degrees of freedom (DoFs) by complex mechanical structures, and the soft gripper can realize stiffness-compliant manipulation differently. The proposed novel robotic in-hand manipulation can execute grasping and cap manipulation by a single flexible shaft to simultaneously achieve bending and rotational movements. The relationship between stretching force and finger’s curvature can enable a custom design for user-specific applications. Quanquan Liu 0001, Ning Tan 0003, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | Feature-Guided Nonrigid 3-D Point Set Registration Framework for Image-Guided Liver Surgery: From Isotropic Positional Noise to Anisotropic Positional NoiseabstractRegistration is an essential problem in image-guided surgery (IGS) since it brings different involved coordinate frames together. Nonrigid or deformable registration still faces many challenges, such as two point sets (PSs) are partially overlapped. To tackle the challenges in the nonrigid registration, we introduce a new two-step point-based registration pipeline that includes two steps. In the first step, the rigid transformation between the two spaces is recovered where the orientation vectors are adopted. In the second step, built on the nonrigid coherent point drift (CPD) approach, the anisotropic positional noise is also assumed. Registration results on the human liver verify the proposed approach' great improvements over the other methods. First, the rotation and translation are recovered with smaller error values than the existing methods. Second, our registration method's performance is much more robust to the partial overlapping between two PSs. Third, the two-step registration framework achieves the best performances in most test cases when there is a localization error in acquiring the intraoperative data. Note to Practitioners-A novel registration approach is presented for image-guided liver surgery (LGLS). Compared with existing nonrigid registration methods, two significant changes (or improvements) exist in the proposed registration framework: 1) the normal vectors are extracted and utilized in the rigid registration step and 2) the anisotropic positional uncertainties are considered. In both steps, the registration problems are formulated as a maximum likelihood (ML) problems and dealt with the expectation-maximization (EM) technique. In both steps, the matrix form of the updated positional covariance is provided and can speed up the computational process. The readers are reminded that with extra information and a more general positional error assumption, our approach demonstrates improved performances in the case of partial-to-full alignment. Zhe Min, Delong Zhu 0001, Hongliang Ren 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2020 | AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted SurgeryabstractSurgical scene understanding and multi-tasking learning are crucial for image-guided robotic surgery. Training a real-time robotic system for the detection and segmentation of high-resolution images provides a challenging problem with the limited computational resource. The perception drawn can be applied in effective real-time feedback, surgical skill assessment, and human-robot collaborative surgeries to enhance surgical outcomes. For this purpose, we develop a novel end-to-end trainable real-time Multi-Task Learning (MTL) model with weight-shared encoder and task-aware detection and segmentation decoders. Optimization of multiple tasks at the same convergence point is vital and presents a complex problem. Thus, we propose an asynchronous task-aware optimization (ATO) technique to calculate task-oriented gradients and train the decoders independently. Moreover, MTL models are always computationally expensive, which hinder real-time applications. To address this challenge, we introduce a global attention dynamic pruning (GADP) by removing less significant and sparse parameters. We further design a skip squeeze and excitation (SE) module, which suppresses weak features, excites significant features and performs dynamic spatial and channel-wise feature re-calibration. Validating on the robotic instrument segmentation dataset of MICCAI endoscopic vision challenge, our model significantly outperforms state-of-the-art segmentation and detection models, including best-performed models in the challenge. Mobarakol Islam, Vibashan VS, Hongliang Ren 0001 |
ICRA | 3 |
| 2020 | Learning and Reasoning with the Graph Structure Representation in Robotic Surgery
Mobarakol Islam, Seenivasan Lalithkumar, Chwee Ming Lim, Hongliang Ren 0001 |
MICCAI (3) | 4 |
| 2020 | Additional planning with multiple objectives for reinforcement learning
Anqi Pan, Wenjun Xu 0005, Lei Wang 0006, Hongliang Ren 0001 |
Knowl. Based Syst. | 4 |
| 2020 | Heuristic orientation adjustment for better exploration in multi-objective optimization
Anqi Pan, Lei Wang 0006, Weian Guo, Hongliang Ren 0001, Qidi Wu |
Neural Comput. Appl. | 4 |
| 2020 | Statistical Model of Total Target Registration Error in Image-Guided SurgeryabstractIn a paired-point rigid registration, target registration error (TRE) is deemed to be the most important quality metric. TRE usually cannot be directly measured, and thus many TRE estimation algorithms have been proposed. However, target localization errors (TLEs) in two spaces are not considered in the definition of TRE. In this paper, we propose a new type of evaluation metric that is referred to as total TRE (TTRE) at a given target point with TLE incorporated. Statistics including mean, root mean square (rms), and covariance matrix of TTRE are derived without making any assumption of the TLE magnitude. TTRE and fiducial registration error (FRE) are proved to be uncorrelated when an ideal weighting scheme is adopted in solving the registration problem. The proposed error model is validated through extensive experiments. In the first experiment with random fiducials and targets, in 90% of the test cases, there shows no difference between the predicted and simulated TTRE statistics when six fiducials are used. In the second experiment of deep-brain stimulation surgery, the mean value of CC(TTRE,FRE) being 8.9246 × 10-4± 0.0389 was observed, which indicates that TTRE and FRE are uncorrelated. In the third experiment of surgical tool-tip tracking, the mean and standard deviation of percentage differences between predicted and simulated TTRE rms values are 2.22% ± 0.77% for the planar tool and 2.62% ± 0.59% for the textral tool. In summary, our proposed algorithm can well model the TTRE metric. Zhe Min, Hongliang Ren 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Depth Estimation of Hard Inclusions in Soft Tissue by Autonomous Robotic Palpation Using Deep Recurrent Neural NetworkabstractAccurately detecting tumors and estimating the depth of tumors is essential in the surgical removal of tumors. In robotic-assisted surgery, autonomous robotic palpation has the potential to provide more precise detection, tumors' depth estimation, and less intrusion when normal tissues surround tumors. In this article, by mimicking the human finger touch, we propose a tactile sensing-based deep recurrent neural network (DRNN) with long short-term memory (LSTM) architecture to improve the accuracy of the detection and depth estimation of tumors embedded in soft tissue. In the experimental setup, the hard inclusions simulate the tumors, while the phantom tissue is fabricated by silicon to simulate the soft tissue. During the experiment, the data from the force sensor and displacement of the robot palpation probe are for detection and depth estimation purposes. The collected sequential data set of the force and the displacement of the probe during one completed palpation process will go through the proposed DRNN network with deep LSTM architecture, in which the temporal dependencies of the sequential data will be captured in the cell states in the deep LSTM layers. Subsequently, the softmax classifier is adopted to determine if there is any hard inclusion exists and offer the depth estimation of the hard inclusions. Experiments based on 396 real data sets demonstrate that the detection accuracy for the testing data set is 99.2% and the depth estimation accuracy for the testing data set is 95.8%. The accuracy of the proposed method is best when comparing with other widely used methods. Note to Practitioners-The palpation of tumors motivated this article in the robot-assisted surgical systems through tactile feedback. In order to mimic the human touch on the soft tissue, this article presents a deep-learning-based approach to estimate the depth of the hard inclusions in the phantom tissue through force information. The displacement of the palpation probe and the touch force during one palpation are recorded as data sequences to train the deep model, which aims to capture dynamics and long-term dependence of the palpation process. In this article, we made the first successful attempt to accurately estimate the depth of the hard inclusions buried at different locations of the phantom tissue using only force information. The proposed approach can work in different robot-assisted scenarios, such as master-slave robotic surgery. In the clinic applications, the force sensor will be integrated at the end-effector of the robotic manipulator. According to the specific requirements, the force sensor and the robotic manipulator might be different from those used in this article. For some applications, such as the laparoscopic interventions, the complete vertical contact tends to be difficult to obtain due to the laparoscopic port effects. The projection of the recorded force data and displacement can obtain the information in the normal direction. The future work is going to be extended to tissue environments with arbitrary surface and tumors with various shapes/depths for more complex and prospective clinical applications. Bo Xiao 0002, Wenjun Xu 0005, Jing Guo 0007, Hak-Keung Lam, Guangyu Jia, Wuzhou Hong, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2019 | Learning Where to Look While Tracking Instruments in Robot-Assisted Surgery
Mobarakol Islam, Yueyuan Li, Hongliang Ren 0001 |
MICCAI (5) | 3 |
| 2019 | Ultrasound-Assisted Guidance With Force Cues for Intravascular InterventionsabstractImage guidance during minimally invasive intravascular interventions is primarily achieved based on X-ray fluoroscopy, which has several limitations including limited 3-D imaging capability, significant doses of radiation to operators, and lack of contact force measurement between the cardiovascular tissue and interventional tools. Ultrasound imaging can be adopted to complement or possibly replace 2-D fluoroscopy for intravascular interventions due to its portability, safety to use, and the ability of providing depth information. However, it is challenging to precisely visualize catheters and guidewires in the ultrasound images. In this paper, we propose a novel method to figure out both the position and orientation of the catheter tip in 2-D ultrasound images in real time by detecting and tracking a passive marker attached to the catheter tip. Moreover, the contact force can be estimated simultaneously as well via measuring the length variation of the marker. A geometrical model-based method is introduced to detect the initial position of the marker, and a Kanade-Lucas-Tomasi-based algorithm is developed to track the position, orientation, and length of the marker. The ex vivo experiment results validate the effectiveness of the proposed approach in automatically locating the catheter tip in ultrasound images and its capability of sensing the contact force. Therefore, it can be concluded that the presented method can be utilized to better facilitate operators during intravascular interventions. Jin Guo 0006, Chaoyang Shi, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2019 | Pose Characterization and Analysis of Soft Continuum Robots With Modeling Uncertainties Based on Interval ArithmeticabstractThis paper introduces a systematical interval-based framework of inherent uncertainties and pose evaluation for a class of soft continuum robots driven by flexible shafts. A more general model of continuum robots driven by shaft tendons is extended from prior kinematic models. On top of the proposed model, the interval-based analysis is presented to analyze and characterize the pose of continuum robots considering uncertainties in kinematic parameters and joint inputs. A 3-D printed bending actuator driven by a flexible shaft is evaluated for case study based on the proposed interval-valued framework. This paper investigates and compares a couple of refinement methods and proposes a new way of sensitivity analysis of model parameters based on interval arithmetic. The kinematic and mechanics parameters are measured and identified experimentally with a representation of intervals. The in-plane motion experiment validates that the computed bounds can enclose all the measured tip positions with consideration of the measurement uncertainty. The method is also validated when external loading is exerted. Ning Tan 0003, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2019 | Surgical Instrument Tracking By Multiple Monocular Modules and a Sensor Fusion ApproachabstractThis paper presents a sensor fusion-based surgical instrument tracking system which uses multiple monocular modules. The system is an optical tracking system, which has been widely utilized in the image-guide surgery because of its high accuracy and precision. However, the line-of-sight occlusion problem which remains unresolved in current systems frustrates surgeons during the operation. To address this challenge, we propose a surgical instrument tracking system based on multiple monocular modules. The rationale is to enable the system to track the surgical instruments inside the surgical site from different views. Three sensor fusion algorithms are proposed to integrate all sensor data from the multimodule system. In order to show the feasibility of the tracking system, simulations and comparison experiments have been carried out. The intensive investigation results give a practical instruction to the real implementation of the proposed system in image-guided interventions. Moreover, an image-guided surgical trial by using a cadaver head has been carried out to validate the feasibility of the proposed system and the tracking algorithms. The results from both the simulation and the cadaver trial have shown the effectiveness of the proposed robust fusion algorithm. Jiaole Wang, Shuang Song 0002, Hongliang Ren 0001, Chwee Ming Lim, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2019 | Multilateral Teleoperation With New Cooperative Structure Based on Reconfigurable Robots and Type-2 Fuzzy LogicabstractThis paper develops an innovative multilateral teleoperation system with two haptic devices on the master side and a newly designed reconfigurable multi-fingered robot on the slave side. A novel nonsingular fast terminal sliding-mode algorithm, together with varying dominance factors for cooperation, is proposed to offer this system's fast position and force tracking, as well as an integrated perception for the operator on the reconfigurable slave robot (manipulator). The Type-2 fuzzy model is used to describe the overall system dynamics, and accordingly a new fuzzy-model-based state observer is proposed to compensate for system uncertainties. A sliding-mode adaptive controller is designed to deal with the varying zero drift of the force sensors and force observers. The stability of the closed-loop system under time-varying delays is proved using Lyapunov-Krasovskii functions. Finally, experiments to grasp different objects are performed to verify the effectiveness of this multilateral teleoperation system. Da Sun, Qianfang Liao, Hongliang Ren 0001 |
IEEE Trans. Cybern. | 5 |
| 2019 | A Robotic System With Multichannel Flexible Parallel Manipulators for Single Port Access SurgeryabstractRobot-assisted single port access surgery (SPAS) brings significant benefits to the patients. However, it is challenging due to the limited incision size and requirements in precision, load capacity, and dexterity. To address these challenges, this paper presents a multichannel SPAS robotic system consisting of two channels of 5-degree-of-freedom (DOF) flexible manipulators and an endoscope channel. Each channel of the manipulators is designed based on a 3-DOF parallel mechanism with three sets of super-elastic Ni-Ti rods and universal joints. This design optimizes the structure for flexibility and safety considerations with a simplified structure compared with conventional parallel mechanisms. Flexible shafts are used for torque transmission from the actuation unit, allowing the motors to be farther away from the patient-side robotic system. The kinematics of the manipulator is derived, and then the reachable workspace and dexterity are analyzed. Furthermore, a prototype of the proposed robotic system is presented and evaluated through adequate experiments. The flexibility of the manipulator is verified via stiffness characterization test. The results of the accuracy tests confirmed that the robotic system can implement manipulation arm with acceptable accuracy. The feasibility and effectiveness of applying the robotic system to the practical operation are also demonstrated through experimentations. Xiao Xiao 0006, Chwee Ming Lim, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | Estimation of Object Orientation Using Conductive Ink and Fabric Based Multilayered Tactile SensorabstractRobots, either made from hard or soft materials, are being used increasingly for unknown object manipulation tasks in unstructured environments. To hold an object firmly and do the required task, the orientation of the object with respect to the robotic hand gripper is one of the necessary key information. Humans can sense the orientation of an object based on vision, kinesthetic sensation, or touch sensation. Similarly, robots shall also be able to estimate grasped object orientation just based on tactile sensing without knowing manipulator kinesthetic or kinematics. Existing sensors are mostly based on vision or rigid inertial measurement units (such as accelerometers, gyroscopes, magnetometers), which require additional setup and are typically not compliant with the arbitrary surfaces of the unknown obj ects. This paper discusses a new approach for object orientation estimation from a multilayered tactile sensor based on conductive silver ink and conductive fabric. Fabric based sensors are flexible and can confer to both hard and soft surfaces. Experimental results show that the tactile sensor developed is able to estimate the tilt angle without any information about manipulator kinematics. Godwin Ponraj, Hongliang Ren 0001 |
ICRA | 2 |
| 2018 | Tracking control design of interval type-2 polynomial-fuzzy-model-based systems with time-varying delay
Bo Xiao 0002, Hak-Keung Lam, Xiaozhan Yang, Yan Yu 0001, Hongliang Ren 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2018 | Data-Defect Inspection With Kernel-Neighbor-Density-Change Outlier FactorabstractData-defect would affect the data quality and the analysis results of data mining. This paper presents a data-defect inspection method with kernel-neighbor-density-change outlier factor (KNDCOF). The definition of kernel neighbor density is proposed to represent the density of each object in database, and the ascending distance series (ADS) of each object is calculated based on the kernel distance between the object and its neighbors. Then, the average density fluctuation (ADF) of the object is established according to the weighted sum of the square of density difference between the object and others in ADS. Finally, the KNDCOF of the object is equal to the ratios of the ADF of the object and the average ADF of neighbors of the object. The degree of the object being an outlier is indicated by the KNDCOF value. The experiments are performed on three real data sets to evaluate the effectiveness of the proposed method. The experimental results verify that the proposed method has higher quality of data-defect inspection and does not increase the time complexity. Hui Cao 0003, Hongliang Ren 0001, Shuzhi Sam Ge |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | Three-Dimensional Intravascular Reconstruction Techniques Based on Intravascular Ultrasound: A Technical ReviewabstractIntravascular ultrasound (IVUS) imaging provides two-dimensional (2-D) real-time luminal and transmural cross-sectional images of intravascular vessels with detailed pathological information. It has offered significant advantages in terms of diagnosis and guidance and has been increasingly introduced from coronary interventions into more generalized endovascular surgery. However, IVUS itself does not provide spatial pose information for its generated images, making it difficult to construct a 3-D intravascular visualization. To address this limitation, IVUS imaging-driven 3-D intravascular reconstruction techniques have been developed. These techniques enable accurate diagnosis and quantitative measurements of intravascular diseases to facilitate optimal treatment determination. Such reconstruction extends the IVUS imaging modality from pure diagnostic assistance to intraoperative navigation and guidance and supports both therapeutic options and interventional operations. This paper presents a comprehensive survey of technological advances and recent progress on IVUS imaging-based 3-D intravascular reconstruction and its state-of-the-art applications. Limitations of existing technologies and prospects of new technologies are also discussed. Chaoyang Shi, Xióngbiao Luó, Jin Guo 0006, Zoran Najdovski, Toshio Fukuda, Hongliang Ren 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | Low-Cost Pyrometry System With Nonlinear Multisense Partial Least SquaresabstractAccurate high-temperature measurement is very important for process monitoring of an industrial system. Infrared thermometers usually can handle no more than 1000 °C and should use some expensive accessories for higher temperature measurements. This paper proposes a low-cost pyrometry system with nonlinear multisense partial least squares (NMSPLS). The ordinary camera with different filters is designed to collect the images of hot object at different wavelengths, and the NMSPLS is presented for predicting the temperature of the hot object from the obtained images. For the proposed method, the obtained images are represented by the multisense tensor, where red, green, and blue are regarded as three different dimensions in a sense of the tensor, respectively. The proposed method integrates an outer model and a nonlinear inner model. For the outer model, the independent variables and the dependent variables are projected into a low-dimensional common latent subspace. The weight matrices are calculated from the independent variables by the tucker decomposition, and the single value decomposition is adopted for extracting the latent variables (Lvs) based on the covariance between the independent variables and the dependent variables. For the nonlinear inner model, the neural network is adopted and the extracted Lvs are used as the input and the output of the neural network, respectively. Two real experiments are performed for estimating the proposed method. The experimental results verify that the proposed method can be applied for pyrometry and have higher effectiveness. Hui Cao 0003, Hongliang Ren 0001, Shuzhi Sam Ge |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | Semi-Automated Segmentation of Glioblastomas in Brain MRI Using Machine Learning TechniquesabstractGlioblastomas (GBMs) are cancerous brain tumors that require careful and intricate analysis for surgical planning. Physicians employ Magnetic Resonance Imaging (MRI) in order to diagnose glioblastomas. The segmentation of the tumor is a crucial step in surgical planning. Clinicians manually segment the tumor voxel-by-voxel; however, this is very time consuming. Hence, extensive research has been conducted to semi-automate and fully-automate this segmentation process. This project explores manual segmentation and utilizes k-means clustering technique for semi-automated segmentation. The accuracy of the k-means clustering segmentation was measured using the Dice Coefficient (DC). The results show that k-means clustering provides high accuracy for the segmentation of the enhanced region of tumor (which appears bright in the T1 post contrast MR image) and hence, it can be efficiently used to speed up manual segmentation. Naomi Joseph, Parita Sanghani, Hongliang Ren 0001 |
ICMLA | 3 |
| 2017 | Open-Source Development of a Low-Cost Stereo-Endoscopy System for Natural Orifice Transluminal Endoscopic Surgery
Jia Xin Koh, Hongliang Ren 0001 |
ICVS | 2 |
| 2017 | TTRE: A new type of error to evaluate the accuracy of a paired-point rigid registrationabstractTarget registration error (TRE) is widely adopted to evaluate the accuracy of a paired-point rigid registration (PPRR). However, TRE is defined in such a way that target localization error (TLE) is not considered. In this paper, we first propose a new type of error that is referred to as total target registration error (TTRE). The statistical model of TTRE is derived that we take the TLE in two spaces to be registered into consideration. Results in the first simulation show that the developed model can accurately estimate the simulated TTRE root-mean-square (RMS) (RMS percent differences|| <; 1.5% ± 2%) in all test cases. When all elements of diagonal FLE and TLE covariance matrices are independently generated from a uniform distribution that spans from 0 to 1mm and the number of fiducials N ≥ 6, the mean and covariance matrix of TTRE are well modelled. We have also theoretically proved and validated through the second simulation that TTRE and fiducial registration error (FRE) are uncorrelated (correlation coefficient (CC) <; 0.1). Finally, TTRE and TRE were found to exhibit a low correlation (0.37 <; CC <; 0.46). Zhe Min, Hongliang Ren 0001, Max Q.-H. Meng |
IROS | 2 |
| 2017 | Finding the Kinematic Base Frame of a Robot by Hand-Eye Calibration Using 3D Position DataabstractWhen a robot is required to perform specific tasks defined in the world frame, there is a need for finding the coordinate transformation between the kinematic base frame of the robot and the world frame. The kinematic base frame used by the robot controller to define and evaluate the kinematics may deviate from the mechanical base frame constructed based on structural features. Besides, by using kinematic modeling rules such as the product of exponentials (POE) formula, the base frame can be arbitrarily located, and does not have to be related to any feature of the mechanical structure. As a result, the kinematic base frame cannot be measured directly. This paper proposes to find the kinematic base frame by solving a hand-eye calibration problem using 3D position measurements only, which avoids the inconvenience and inaccuracy of measuring orientations and thus significantly facilitates practical operations. A closed-form solution and an iterative solution are explicitly formulated and proved effective by simulations. Comprehensive analyses of the impact of key parameters to the accuracy of the solution are also carried out, providing four guidelines to better conduct practical operations. Finally, experiments on a 7-DOF industrial robot are performed with an optical tracking system to demonstrate the superiority of the proposed method using position data only over the method using full pose data. Liao Wu, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2017 | Finite Time Fault Tolerant Control for Robot Manipulators Using Time Delay Estimation and Continuous Nonsingular Fast Terminal Sliding Mode ControlabstractIn this paper, a novel finite time fault tolerant control (FTC) is proposed for uncertain robot manipulators with actuator faults. First, a finite time passive FTC (PFTC) based on a robust nonsingular fast terminal sliding mode control (NFTSMC) is investigated. Be analyzed for addressing the disadvantages of the PFTC, an AFTC are then investigated by combining NFTSMC with a simple fault diagnosis scheme. In this scheme, an online fault estimation algorithm based on time delay estimation (TDE) is proposed to approximate actuator faults. The estimated fault information is used to detect, isolate, and accommodate the effect of the faults in the system. Then, a robust AFTC law is established by combining the obtained fault information and a robust NFTSMC. Finally, a high-order sliding mode (HOSM) control based on super-twisting algorithm is employed to eliminate the chattering. In comparison to the PFTC and other state-of-the-art approaches, the proposed AFTC scheme possess several advantages such as high precision, strong robustness, no singularity, less chattering, and fast finite-time convergence due to the combined NFTSMC and HOSM control, and requires no prior knowledge of the fault due to TDE-based fault estimation. Finally, simulation results are obtained to verify the effectiveness of the proposed strategy. Mien Van, Shuzhi Sam Ge, Hongliang Ren 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Robust Fault-Tolerant Control for a Class of Second-Order Nonlinear Systems Using an Adaptive Third-Order Sliding Mode ControlabstractDue to the robustness against the uncertainties, conventional sliding mode control (SMC) has been extensively developed for fault-tolerant control (FTC) system. However, the FTCs based on conventional SMC provide several disadvantages such as large transient state error, less robustness, and large chattering, that limit its application for real application. In order to enhance the performance, a novel adaptive third-order SMC, which combines a novel third-order sliding mode surface, a continuous strategy and an adaptation law, is proposed. Compared with other innovation approaches, the proposed controller has an excellent capability to tackle several types of actuator faults with an enhancing on robustness, precision, chattering reduction, and time of convergence. The proposed method is then applied for an attitude control of a spacecraft and the results demonstrate the superior performance. Mien Van, Shuzhi Sam Ge, Hongliang Ren 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Human-compliant body-attached soft robots towards automatic cooperative ultrasound imagingabstractUltrasound imaging procedures are deemed as one of the most convenient and least invasive medical diagnostic imaging modalities and have been widely utilized in health care providers, which are expecting semiautomatic or fully-automatic imaging systems to reduce the current clinical workloads. This paper presents a portable and wearable soft robotic system which has been designed with the purpose of replacing the manual operation to cooperatively steer the ultrasound probe. This human-compliant soft robotic system, which is equipped with four separated parallel soft pneumatic actuators and is able to achieve movements in three directions. Vacuum suction force is introduced to attach the robot onto the intended body location. The design and fabrication of this soft robotic system are illustrated. To our knowledge, this is the first body-attached soft robot for compliant ultrasound imaging. The feasibility of the system is demonstrated through proof-of-concept experiments. Hongliang Ren 0001, Koon Lin Tan |
CSCWD | 1 |
| 2016 | Automate surgical tasks for a flexible Serpentine Manipulator via learning actuation space trajectory from demonstrationabstractSurgical robotic systems with miniaturized flexible Tendon-driven Serpentine Manipulators (TSM) have enjoyed increasing popularities among surgeons and researchers for their advantages of working in constrained and torturous human lumen such as oral cavity and upper GI tract. However, they suffer from sufficient nonlinearities and model uncertainties due to friction, tension varying, tendon slacking, etc. Model based control is insufficient to overcome such uncertainties and automate challenging surgical related tasks. The objective of this work is to automate certain clinical tasks to alleviate surgeon fatigue and promote task efficiency in kinematics free and sensor free circumstances. We present a data-driven approach based on Learning from Demonstration (LfD), which utilizes statistical machine learning models to encode system underlying dynamics and generalize smooth motor trajectories by direct actuation space learning. Motion segmentation is enabled with soft margin Support Vector Machine (soft-SVM) in complicated tasks. We also make attempts to retrieve task-specific properties by Locally Weighted Regression (LWR). We evaluated the approach on two surgical related tasks: compliant insertion and simplified Endoscopic Submucosal Dissection (ESD). The flexible TSM successfully reproduced both tasks and demonstrated superior trajectory performance. A video is available at: https://youtu.be/rLQo6xKtyMI. Wenjun Xu 0005, Jie Chen 0028, Henry Y. K. Lau, Hongliang Ren 0001 |
ICRA | 4 |
| 2016 | Towards hybrid control of a flexible curvilinear surgical robot with visual/haptic guidanceabstractComprised of multiple telescoptic precurved tubes that can independently rotate and translate, concentric tube robots (CTRs) are favorable in minimally invasive surgeries thanks to their small size and considerable dexterity along with curvilinear accessibility. However, there is a lack of investigation on improvement of the surgeons' perception which in turn can be used to guide the telemanipulation. In this work, we proposed an eye-in-hand configuration for the concentric tube robot by adding an endoscope to the tip of the inner tube, which provides direct and intuitive visual sensing ability for the operator. Based on this visual feedback, we further developed two frameworks for the hybrid control of CTR, namely Teleoperation Before Visual Servoing (TBVS) and Teleoperation During Visual Servoing (TDVS). The structures of these two frameworks were elaborated with key algorithms derived. The effectiveness of the proposed methods were demonstrated through a series of experiments both in free space and in a confined environment (inside a skull model). The results manifested that the visual guidance had the potential of assisting the operator to control the CTR more efficiently. Liao Wu, Keyu Wu 0001, Hongliang Ren 0001 |
IROS | 3 |
| 2016 | Self-triggered output feedback control for consensus of multi-agent systems
Miaomiao Wu, Hao Zhang 0008, Huaicheng Yan 0001, Hongliang Ren 0001 |
Neurocomputing | 4 |
| 2016 | Kinematic Analysis and Motion Control of Wheeled Mobile Robots in Cylindrical WorkspacesabstractWheeled mobile robots (WMRs) are often used for maintenance of round pipes or ducts, which can typically be represented as a cylindrical workspace. Working in round pipes or ducts, kinematic models of WMRs are different from those applying on a plane and thus pose significant challenges in terms of kinematic analysis and motion control. To address these challenges, the kinematic properties of WMRs in a cylindrical workspace are analyzed in this paper. First, we discuss the kinematic properties of a single wheel in a cylindrical workspace. Then, we analyze the geometric constraints of WMRs in round pipes or ducts with analytical geometry. Based on these analyses, kinematic properties of WMRs in cylindrical workspaces are discussed with screw theory. A control law based on biaxial clinometer information is proposed, and it enables the robot to move horizontally in round pipes or ducts. Finally, the motion of a single wheel purely rolling in a cylindrical workspace is simulated. Experiments using a car-like mobile robot moving in round ducts are carried out to show the feasibility of the proposed algorithm. Zhangjun Song, Hongliang Ren 0001, Jianwei Zhang 0001, Shuzhi Sam Ge |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2016 | Parameterized Distortion-Invariant Feature for Robust Tracking in Omnidirectional VisionabstractCentral catadioptric omnidirectional images exhibit serious nonlinear distortions due to the involved quadratic mirrors. Therefore, features based on the conventional pin-hole model are hard to achieve satisfactory performances when directly applied to the distorted omnidirectional images. This paper analyzes the catadioptric geometry to facilitate modeling the nonlinear distortions of omnidirectional images. Different to the conventional imaging model, the prior information is considered in catadioptric system. A parameterized neighborhood mapping model is proposed to efficiently calculate the neighborhood of an object based on its measurable radial distance in the image plane. On the basis of the parameterized nonlinear model, a distortion-invariant fragment-based joint-feature mixture model of Gaussian is presented for human target tracking in omnidirectional vision. Under the framework of Gaussian Mixture Model, the problem of feature matching is converted into the feature clustering. The joint probability distribution of a joint-feature class is modeled by a mixture of Gaussian. A weight contribution mechanism is designed to flexibly weight the fragments contribution based on their responses, which leads to a robust tracking even under serious partial occlusion. Finally, experiments validate the advantage of the proposed algorithm over other conventional approaches. Catadioptric omnidirectional cameras have been widely used in robotics and surveillance fields for visual sensing due to its big field-of-view. However, conventional visual models use large-scale statistical sampling for feature extraction in catadioptric sensor, which may consume lot of computational cost. For practical applications, a parameterized model that can accurately and efficiently formulate distortion of catadioptric image is desirable. Integrating of the priori of system, a parameterized neighborhood model is presented to directly extract distorted image content in image, which can significantly improve the efficiency of algorithm. To robustly handle challenging occlusion in the distorted image, a flexible fragment-based joint-feature framework is presented for robust non-rigid human target tracking. Compared with the conventional tracking methods applied to catadioptric vision, the proposed tracking approaches leads to much better performance from the perspective of efficiency and robustness. Yazhe Tang, Youfu Li 0001, Shuzhi Sam Ge, Jun Luo 0006, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2016 | Fault Diagnosis in Image-Based Visual Servoing With Eye-in-Hand Configurations Using Kalman FilterabstractIn this paper, the fault diagnosis (FD) problem in image-based visual servoing with eye-in-hand configurations is investigated. The potential failures are detected and isolated based on approximating parameters related. First, the failure scenarios of the visual servoing systems are reviewed and classified into the actuator and sensor faults. Second, a residual generator is proposed to detect the failure occurrences, based on the Kalman filter. Third, a decision table is proposed to isolate the fault type. Finally, simulation and experimental results are given to validate the efficacy and the efficiency of the proposed FD strategies. Mien Van, Denglu Wu, Shuzhi Sam Ge, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2016 | Simultaneous Hand-Eye, Tool-Flange, and Robot-Robot Calibration for Comanipulation by Solving the AXB=YCZ ProblemabstractMultirobot comanipulation shows great potential in surpassing the limitations of single-robot manipulation in complicated tasks such as robotic surgeries. However, a dynamic multirobot setup in unstructured environments poses great uncertainties in robot configurations. Therefore, the coordination relationships between the end-effectors and other devices, such as cameras (hand–eye calibration) and tools (tool–flange calibration), as well as the relationships among the base frames (robot–robot calibration) have to be determined timely to enable accurate robotic cooperation for the constantly changing configuration of the systems. We formulated the problem of hand–eye, tool–flange, and robot–robot calibration to a matrix equation$\mathbf{AXB=YCZ}$. A series of generic geometric properties and lemmas were presented, leading to the derivation of the final simultaneous algorithm. In addition to the accurate iterative solution, a closed-form solution was also introduced based on quaternions to give an initial value. To show the feasibility and superiority of the simultaneous method, two nonsimultaneous methods were compared through thorough simulations under various robot movements and noise levels. Comprehensive experiments on real robots were also performed to further validate the proposed methods. The comparison results from both simulations and experiments demonstrated the superior accuracy and efficiency of the proposed simultaneous calibration method. Liao Wu, Jiaole Wang, Lin Qi 0002, Keyu Wu 0001, Hongliang Ren 0001, Max Q.-H. Meng |
IEEE Trans. Robotics | 5 |
| 2015 | Soft oral interventional rehabilitation robot based on low-profile soft pneumatic actuatorabstractMandibular mobility plays a significant role in human daily lives, enabling food intake, respiration, speaking, and other oral activities. However, as for the patients suffering from mandibular mobility disorders, their mandible functions are deteriorated, which severely affects their quality of life. In this paper, we present a new solution to recover mandibular mobility: a soft oral rehabilitation robot (SORR), which is actuated by a novel type of soft pneumatic actuator (SPA). After identifying the biometrics of the human mandible, we illustrate the application-oriented design of the robot and the SPA. A static model of the SPA was established to predict its behavior and eligibility for the application. In the experimental characterization, we measured the elongation and force output of the SPA, and the detailed results are presented. The comparison and analysis of the results provide physical insight into the mechanisms of the SPA. As the first application of soft robot inside human body, this work enlightens profound application potentials of the silicone-based SPAs. Yi Sun 0008, Chwee Ming Lim, Hee Hon Tan, Hongliang Ren 0001 |
ICRA | 4 |
| 2015 | Distortion invariant joint-feature for visual tracking in catadioptric omnidirectional visionabstractCentral catadioptric omnidirectional images exhibit serious nonlinear distortions due to quadratic mirrors involved. Conventional visual features developed based on the perspective model are hard to achieve a satisfactory performance when directly applied to the distorted omnidirectional image. This paper presents a parameterized neighborhood model to efficiently calculate the adaptive neighborhood of an object based on the measurable radial distance in image plane. On the basis of the parameterized neighborhood model, a distortion invariant joint-feature framework implemented with contour-color fragment mixture model of Gaussian is proposed for visual tracking in catadioptric omnidirectional camera system. Under the framework of Gaussian Mixture Model, the problem of feature matching is converted into feature clustering. A weight contribution mechanism is presented to flexibly weight the fragments based on their responses, which makes the system robustly guided by limited visible fragments even when serious partial occlusion happens. The experiments validate the performance of the proposed algorithm. Yazhe Tang, Youfu Li 0001, Shuzhi Sam Ge, Jun Luo 0006, Hongliang Ren 0001 |
ICRA | 5 |
| 2015 | Minimum sweeping area motion planning for flexible serpentine surgical manipulator with kinematic constraintsabstractFlexible serpentine manipulators are widely used in surgical robots as it can be operated inside the patient's body cavity by backbone bending. However, during the bending the manipulator sweeps over a region, where sensitive organs may locate. This raises the safety concern. In this paper, a motion planning algorithm focusing on minimize the sweeping area for flexible serpentine manipulators is presented. Particularly, a three dimensional backward average neural dynamic model (BANDM) is proposed to build minimum sweeping area planning field in the configuration space of the serpentine manipulator. Given a target position, the motion sequence is generated automatically based on the established planning field. The simulations and experimental results validate the effectiveness and superiority of the proposed planning approach over conventional planning algorithms in terms of sweeping area with keeping target reach and obstacle avoidance. Zheng Li 0012, Wenjun Xu 0005, Yaonan Wang 0001, Hongliang Ren 0001 |
IROS | 5 |
| 2015 | A novel constrained tendon-driven serpentine manipulatorabstractIn this paper, a novel constrained tendon-driven serpentine manipulator (CTSM) suited for minimally invasive surgery is presented. It comprises of a flexible backbone, a set of controlling tendons and a constraint. In the CTSM not only the curvature of the bending section can be controlled but also the length. Specifically, the curvature is controlled by the tendons, and the length is controlled by a constraint tube, which is translational and is concentric with the flexible backbone. The kinematic model of the CTSM is developed based on the piecewise constant curvature assumption. Analysis shows that by introducing the translational constraint both the workspace and dexterity of the manipulator are improved. The stiffer the constraint the larger the workspace expansion and the smaller the dexterity enhancement. A prototype is developed and the experimental results validate the design idea and analysis. Zheng Li 0012, Haoyong Yu, Hongliang Ren 0001, Philip W. Y. Chiu, Ruxu Du |
IROS | 3 |
| 2015 | Motion planning of continuum tubular robots based on centerlines extracted from statistical atlasabstractContinuum tubular robots, which are constructed by telescoping pre-curved elastic tubes, are capable of balancing the force application and steerability during minimally invasive surgeries. These devices are able to reach the desired surgical sites in body cavities without colliding with critical blood vessels, nerves and tissues. However, the motion planning of continuum tubular robots is quite challenging because of their complicated kinematics as well as the high dimensional configuration space. In this paper, a sampling-based motion planning method is proposed based on the Rapidly-exploring Random Tree (RRT) algorithm for continuum tubular robots in 3D environments, such as medullary cavities. The proposed motion planner enables a continuum tubular robot to maneuver roughly along the central axis of the statistical humerus atlas in an approximate follow-the-leader manner. The experiment results have demonstrated the effectiveness and superiority of the proposed motion planning algorithm. Keyu Wu 0001, Liao Wu, Hongliang Ren 0001 |
IROS | 3 |
| 2015 | No-reference blur assessment based on edge modeling
Jingwei Guan, Wei Zhang 0021, Jason Gu, Hongliang Ren 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | Drift analysis of mutation operations for biogeography-based optimization
Weian Guo, Lei Wang 0006, Shuzhi Sam Ge, Hongliang Ren 0001, Yanfen Mao |
Soft Comput. | 4 |
| 2015 | Intensity-Based Visual Servoing for Instrument and Tissue Tracking in 3D Ultrasound VolumesabstractThis paper presents a three dimensional ultrasound (3DUS)-based visual servoing technique for intraoperative tracking of the motion of both surgical instruments and tissue targets. In the proposed approach, visual servoing techniques are used to control the position of a virtual ultrasound probe so as to keep a target centered within the virtual probe's field-of-view. Multiple virtual probes can be servoed in parallel to provide simultaneous tracking of instruments and tissue. The technique is developed in the context of robotic beating-heart intracardiac surgery in which the goal of tracking is to both provide guidance to the operator as well as to provide the means to automate the surgical procedure. To deal with the low signal-to-noise ratio (SNR) of the 3DUS volumes, an intensity-based method is proposed that requires no primitive extraction or image segmentation since it directly utilizes the image intensity information as a visual feature. This approach is computationally efficient and can be applied to a wide range of tissue types and medical instruments. This paper presents the first validation of these techniques through offline robot and tissue tracking using actual in vivo cardiac volume sequences from a robotic beating-heart surgery. Caroline Vienne, Hongliang Ren 0001, Alexandre Krupa, Pierre E. Dupont |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2015 | Towards Occlusion-Free Surgical Instrument Tracking: A Modular Monocular Approach and an Agile Calibration MethodabstractOptical means of instrument tracking has been widely used in image-guided interventions and considered the de facto standard for tracking rigid bodies with a direct line-of-sight. However, the occlusion problem which remains unresolved in current systems frustrates surgeons during the operation. To address this challenge, we propose a surgical instrument tracking system based on multiple reconfigurable monocular modules. The main approach is to enable the system to dynamically reconfigure the multiple monocular modules when occlusion occurs partially within the workspace. In this paper, we focus on the system architecture and an agile multicamera calibration method which only uses the customized tool for the surgical instrument tracking scenario. Additionally, two fast non-iterative algorithms are proposed and studied. In order to show the feasibility and superiority of the corresponding multicamera calibration algorithm, comparison experiments have carried out. The intensive investigation results give a practical instruction to the real implementation of the proposed system in image-guided interventions. Jiaole Wang, Max Q.-H. Meng, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2015 | A Minimal POE-Based Model for Robotic Kinematic Calibration With Only Position MeasurementsabstractThis paper proposes an algorithm for robotic kinematic calibration based on a minimal product of exponentials (POE)-based model for the applications where only position measurements are required. Both joint zero-offset errors and initial frame twist error can be involved in this model. Analysis of the identifiability of these errors shows that at most six elements of these parameters can be identified. It also suggests that at least three noncollinear points on the end-effector should be measured to maximize the identifiability. Compared with the traditional POE-based model with full pose (position and orientation) measurements, the minimal model with only position measurements outperforms in terms of convenience, efficiency, and accuracy. Note to Practitioners-Kinematic calibration is pivotal to improve the position accuracy of a robot. To avoid the disadvantages of measuring the orientation of the end-effector during calibration, an algorithm using only position measurements is presented, with which one needs only position measurements of several points fixed on the end-effector without orientation information during the whole calibration process. This will greatly facilitate the scheme design as well as the practical operations. The identifiability of the parameters is then analyzed with two conclusions: 1) at most six elements of the joint zero-offsets and the initial frame twist in total can be identified simultaneously and 2) at least three points on the end-effector which are not collinear should be measured so as to make the identifiability maximum. According to these conclusions, one should carefully select parameters to formulate the error model and measure sufficient points on the end-effector during the calibration procedure. Liao Wu, Xiangdong Yang, Ken Chen 0002, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2014 | Towards simultaneous coordinate calibrations for cooperative multiple robotsabstractTasks that are too hard for single robot can be easily carried out by multiple robots in a cooperative manner. If some/all robots have mobile bases, the cooperation is subjected to great uncertainties in both the robotic system and environment. Therefore, the relationships among all the base frames (robot-robot calibration) and the relationships between the end-effectors and the other devices such as cameras and tools (hand-eye and tool-flange calibrations) have to be calculated to enable the robots to cooperate. To address these challenges, in this paper, we propose a simultaneous hand-eye, tool-flange and robot-robot calibration method. Thorough simulations are conducted to show the superiority of the proposed simultaneous method under different noise levels and various numbers of robot movements. Furthermore, the comparison to two non-simultaneous calibration methods has also been carried out to show the efficiency and robustness of the proposed simultaneous method. Jiaole Wang, Liao Wu, Max Q.-H. Meng, Hongliang Ren 0001 |
IROS | 4 |
| 2014 | Marker-Based Surgical Instrument Tracking Using Dual Kinect SensorsabstractThis paper presents a passive-marker-based optical tracking system utilizing dual Kinect sensors and additional custom optical tracking components. To obtain sub-millimeter tracking accuracy, we introduce robust calibration of dual infrared sensors and point correspondence establishment in a stereo configuration. The 3D localization is subsequently accomplished using multiple back projection lines. The proposed system extends existing inexpensive consumer electronic devices, implements tracking algorithms, and shows the feasibility of applying the proposed low-cost system to surgical training for computer assisted surgeries. Hongliang Ren 0001, Andy Lim |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2014 | Treatment Planning and Image Guidance for Radiofrequency Ablation of Large TumorsabstractThis article addresses the two key challenges in computer-assisted percutaneous tumor ablation: planning multiple overlapping ablations for large tumors while avoiding critical structures, and executing the prescribed plan. Toward semiautomatic treatment planning for image-guided surgical interventions, we develop a systematic approach to the needle-based ablation placement task, ranging from preoperative planning algorithms to an intraoperative execution platform. The planning system incorporates clinical constraints on ablations and trajectories using a multiple objective optimization formulation, which consists of optimal path selection and ablation coverage optimization based on integer programming. The system implementation is presented and validated in both phantom and animal studies. The presented system can potentially be further extended for other ablation techniques such as cryotherapy. Hongliang Ren 0001, Enrique Campos-Náñez, Ziv Yaniv, Filip Banovac, Hernán G. Abeledo, Nobuhiko Hata, Kevin Cleary |
IEEE J. Biomed. Health Informatics | 1 |
| 2013 | Force Efficient Analysis of a Hybrid Magnetic Actuation System for Minimally Invasive Diagnostics and Interventions
Jinji Sun, Hongliang Ren 0001, Keyu Wu 0001 |
ICOST | 2 |
| 2012 | Tubular Enhanced Geodesic Active Contours for continuum robot detection using 3D ultrasoundabstractThree dimensional ultrasound is a promising imaging modality for minimally invasive robotic surgery. As the robots are typically metallic, they interact strongly with the sound waves in ways that are not modeled by the ultrasound system's signal processing algorithms. Consequently, they produce substantial imaging artifacts that can make image guidance difficult, even for experienced surgeons. This paper introduces a new approach for detecting curved continuum robots in 3D ultrasound images. The proposed approach combines geodesic active contours with a speed function that is based on enhancing the "tubularity" of the continuum robot. In particular, it takes advantage of the known robot diameter along its length. It also takes advantage of the fact that the robot surface facing the ultrasound probe provides the most accurate image. This method, termed Tubular Enhanced Geodesic Active Contours (TEGAC), is demonstrated through ex vivo intracardiac experiments to offer superior performance compared to conventional active contours. Hongliang Ren 0001, Pierre E. Dupont |
ICRA | 1 |
| 2012 | Multisensor Data Fusion in an Integrated Tracking System for Endoscopic SurgeryabstractSurgical planning and navigation systems are vital for minimally invasive endoscopic surgeries but it is challenging to track the position and orientation of intrabody surgical instruments in these procedures. In order to address this problem, we propose a tracking system including multiple-sensor integration and data fusion. The proposed tracking approach is free of the constraints of line-of-sight, less subject to environmental distortion, and with higher update rate. By incorporating electromagnetic and inertial sensors, the system yields continuous 6-DOF information. Based on a system dynamic model and estimation theories, a new multisensor fusion algorithm, cascade orientation and position-estimation algorithm, is proposed for the integrated tracking device. The experimental results show that the proposed algorithms achieve accurate orientation and position tracking with robustness. Hongliang Ren 0001, Denis Rank, Martin Merdes, Jan Stallkamp, Peter Kazanzides |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2012 | Passive Markers for Tracking Surgical Instruments in Real-Time 3-D Ultrasound ImagingabstractA family of passive echogenic markers is presented by which the position and orientation of a surgical instrument can be determined in a 3-D ultrasound volume, using simple image processing. Markers are attached near the distal end of the instrument so that they appear in the ultrasound volume along with the instrument tip. They are detected and measured within the ultrasound image, thus requiring no external tracking device. This approach facilitates imaging instruments and tissue simultaneously in ultrasound-guided interventions. Marker-based estimates of instrument pose can be used in augmented reality displays or for image-based servoing. Design principles for marker shapes are presented that ensure imaging system and measurement uniqueness constraints are met. An error analysis is included that can be used to guide marker design and which also establishes a lower bound on measurement uncertainty. Finally, examples of marker measurement and tracking algorithms are presented along with experimental validation of the concepts. Jeffrey A. Stoll, Hongliang Ren 0001, Pierre E. Dupont |
IEEE Trans. Medical Imaging | 2 |
| 2011 | Detection of curved robots using 3D ultrasoundabstractThree-dimensional ultrasound can be an effective imaging modality for image-guided interventions since it enables visualization of both the instruments and the tissue. For robotic applications, its realtime frame rates create the potential for image-based instrument tracking and servoing. These capabilities can enable improved instrument visualization, compensation for tissue motion as well as surgical task automation. Continuum robots, whose shape comprises a smooth curve along their length, are well suited for minimally invasive procedures. Existing techniques for ultrasound tracking, however, are limited to straight, laparoscopic-type instruments and thus are not applicable to continuum robot tracking. Toward the goal of developing tracking algorithms for continuum robots, this paper presents a method for detecting a robot comprised of a single constant curvature in a 3D ultrasound volume. Computational efficiency is achieved by decomposing the six-dimensional circle estimation problem into two sequential three-dimensional estimation problems. Simulation and experiment are used to evaluate the proposed method. Hongliang Ren 0001, Nikolay V. Vasilyev, Pierre E. Dupont |
IROS | 1 |
| 2009 | Node localization during power adjustment in wireless sensor networksabstractNode localization is a challenging problem in wireless sensor networks, especially in the scenarios of tuning multiple transmit-powers. In this article, we utilized particle filter to infer static node position from the correlations between radio frequency (RF) received signal strength indication (RSSI) and distance under multiple power settings. The RSSI based stochastic measurement model was analyzed and followed by the particle filter design. The simulation results verified the performance of proposed algorithm for localization. The proposed method is contributive in terms of making advantages of multiple transmit power for localization. Hongliang Ren 0001, Max Q.-H. Meng |
ICRA | 1 |
| 2009 | Game-Theoretic Modeling of Joint Topology Control and Power Scheduling for Wireless Heterogeneous Sensor NetworksabstractWireless Heterogeneous Sensor Network (WHSN) facilitates ubiquitous information acquisition for Ambient Intelligence (AmI) systems. It is of great importance of power management and topology control for WHSN to achieve desirable network performances, such as clustering properties, connectivity and power efficiency. This paper proposes a game theoretic model of topology control to analyze the decentralized interactions among heterogeneous sensors. We study the utility function for nodes to achieve desirable frame success rate and node degree, while minimizing the power consumption. Specifically, we propose a static complete-information game formulation for power scheduling and then prove the existence of the Nash equilibrium with simultaneous move. Because the heterogeneous sensors typically react to neighboring environment based on local information and the states of sensors are evolving over time, the power-scheduling problem in WHSN is further formulated into a more realistic incomplete-information dynamic game model with sequential move. We then analyze the separating equilibrium, one of the perfect Bayesian equilibriums resulted from the dynamic game, with the sensors revealing their operational states from their actions. The sufficient and necessary conditions for the existence of separating equilibrium are derived for the dynamic Bayesian game, which provide theoretical basis to the proposed power scheduling algorithms, NEPow and BEPow. The primary contributions of this paper include applying game theory to analyze the distributed decision-making process of individual sensor nodes and to analyze the desirable utilities of heterogeneous sensor nodes. Simulations are presented to validate the proposed algorithms and the results show their ability of maintaining reliable connectivity, reducing power consumption, while achieving desirable network performances. Hongliang Ren 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2006 | Investigating Network Optimization Approaches in Wireless Sensor NetworksabstractMore and more wireless mobile sensor networks are employed by robotics to perform harsh tasks such as disaster rescue, emission sources localization or hazardous contaminants localization. There are a lot of network optimization problems to be solved in the protocol design of wireless mobile sensor networks (WMSN), such as rate control, flow control, congestion control, medium access control, queue management, power control and topology control etc. These issues involve several layers of the network protocol stack so that it's quite difficult to consider every single optimization problem of them in a holistic view. The majority of contemporary research works mainly deal with one or some of them in terms of certain applications or objectives. However, most of the proposed protocols are based on simulations or experiments which lack of sufficient mathematical or theoretical analysis to fully understand the convergence or stability. In order to study the theoretical basis of network algorithms, this paper briefly outlines the current methodologies exploited to design and optimize the performance of sensor networks. In addition, the paper investigates the theoretical aspects to make sense of the network optimization algorithms and give a survey mainly in terms of convex optimization, game theory and artificial intelligence Hongliang Ren 0001, Max Q.-H. Meng, Xijun Chen |
IROS | 1 |
| 2006 | Wireless Assistive Sensor Networks for the DeafabstractIn this paper, a wireless sensor network based assistive system (WASN) is developed to help the deaf or the hearing-impaired people, particularly, help them to be aware of their environments. A brief review on the current assistive devices is also addressed first. The system architecture, components and specifications are presented for two typical application scenarios: smart home and smart school playground. A node platform is developed to implement the system. Finally, the system performance is evaluated by simulations, followed by the analysis of the feasibility and availability Hongliang Ren 0001, Max Q.-H. Meng, Xijun Chen |
IROS | 1 |
| 2006 | Rate Control to Reduce Bioeffects in Wireless Biomedical Sensor NetworksabstractDuring the course of physiological information monitoring by wireless biomedical sensor networks, adverse biological effects will be caused by wireless radio frequency radiation, especially for long term, intensive and close inspection. This paper is concerned with bioeffect metric to evaluate the performance of wireless biosensor networks in terms of health effect consideration, and a price-based rate control algorithm to reduce the bioeffect. This paper first investigates the bioeffects caused by radiofrequency transmission of sensor node, including thermal effects and athermal effects. The bioeffects model is studied in both near-field and far-field, in relation to specific absorption rate (SAR). The main contribution of this work is that a normalized bioeffect metric, equivalent coefficient-of-absorption-and-bioeffects (CAB), is derived to evaluate and design the communication protocols for wireless biosensor networks. After identifying the factors that can reduce the adverse health effects in the communication system we present a bioeffect aware rate control algorithm for the system Hongliang Ren 0001, Max Q.-H. Meng |
MobiQuitous | 1 |
| 2006 | Bioeffects Control in Wireless Biomedical Sensor NetworksabstractWireless sensor networks have been employed to enhance the healthcare system recently. During the course of physiological information monitoring by wireless biomedical sensor networks, adverse biological effects will be caused by wireless radio frequency radiation, especially for long term, intensive and close inspection. This paper is concerned with bioeffects metric to evaluate the performance of wireless biosensor networks in terms of health effect consideration, and a price-based rate control algorithm to reduce the bioeffects. This paper first investigates the bioeffects caused by radiofrequency transmission of sensor node, including thermal effects and athermal effects. The bioeffects model is studied in both near-field and far-field, in relation to specific absorption rate (SAR). The main contribution of this work is that a normalized bioeffects metric, equivalent coefficient-of-absorption-and-bioeffects (CAB), is derived to evaluate and design the communication protocols for wireless biosensor networks. After identifying the factors that can reduce the adverse health effects in the communication system, we use the proposed bioeffects metric to evaluate the performance of power control algorithm compared with the one without power control. Finally, we present a bioeffects aware rate control algorithm for the system Hongliang Ren 0001, Max Q.-H. Meng |
SECON | 1 |