VLDB 2026 Research / reviewers in the wild / expert
Zhen Li 0049
dblp:74/2397-49
· DBLP profile ↗
30ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0002-4033-8650ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 16 since 2021Systems, architecture and hardware · 10 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-supervised exceptional prototypical network for few-shot grading of gastric intestinal metaplasia
Xuanchi Chen, Zhen Li 0049, Mingzhe Zhang 0001, Han Yu 0001, Li-Zhen Cui 0001, Xiangwei Zheng 0001 |
Neural Networks | 3 |
| 2025 | Spatiotemporal Motion Prediction of Intraocular Microsurgical Robot in Non-Visible RegionsabstractIn intraocular microsurgery with minute operational scales, instruments pass through non-visible regions of the anterior segment, where robot-assisted surgery, which heavily relies on visual perception, fails to determine the instrument’s attitude relative to the eyeball. This compromises surgical flexibility, increases risks, and hinders autonomous surgery development. Therefore, a framework for predicting instrument trajectories in non-visible regions during robot-assisted microsurgery has been proposed to mitigate the risks of retinal and lens injuries caused by blind operations and enhance surgical procedures’ intelligence and autonomy. First, a lightweight reconstruction of the anterior segment environment is performed under controlled knowledge guidance to construct a global map. Second, the tip position of the surgical instrument is detected through multi-sensor fusion, enabling the perception of instrument-environment interactions under visual constraints. Based on this, a long short-term spatiotemporal aggregation algorithm for instrument trajectory prediction is proposed, which enhances surgical safety by providing high-precision predictions of the instrument tip’s motion trajectory. Experiments show that the framework achieved a 0.0435 mm average prediction error in non-visible regions, corresponding to 0.03% of the region in a single dimension and 7.25% of the surgical instrument’s diameter. This significantly enhances the precision of robot-assisted surgery under visual constraints and provides robust technical support for safe, intelligent, and autonomous intraocular robotic surgery. Ya-Wen Deng, Zhen Li 0049, Yu-Peng Zhai, Weihong Yu, Zhangguo Yu, Guibin Bian |
IROS | 2 |
| 2025 | Implicit Disparity-Blur Alignment for Fast and Precise Autofocus in Robotic Microsurgical ImagingabstractCreating an intelligent surgical environment requires not only advanced robotic systems but also optimized microscopic imaging. However, autofocus remains a fundamental challenge, with current methods suffering from slow iterative processes or directional ambiguity, which compromises real-time performance. This paper presents an implicit disparity-blur alignment approach for robotic microsurgical autofocus, integrating stereo geometry’s monotonic depth cues with de-focus characteristics for rapid convergence. A novel physics-guided dual-stream network is developed to encode implicit depth representations through hierarchical cross-pathway feature fusion, enabling reliable focus prediction without explicit stereo matching in blur-degraded regions. An ROI-aware attention module is proposed to dynamically optimize focus-critical regions, coupled with learnable physics-guided kernel learning for precise Z-offset estimation. The approach achieves a top directional accuracy of 94.85% and a single-pass focus error of 0.20 mm with an inference time of 53 ms on a surgical dataset, which outperforms state-of-the-art methods in reducing iteration count by 22.8% and inference time by 51.8%. An intelligent robotic microscope prototype is developed, with validation through ex vivo tests demonstrating its ability to enable fast and precise multi-region focusing for microsurgeries. Pan Fu, Zhen Li 0049, Ming-Yang Zhang, Yu-Peng Zhai, Wen-Hao He, Guibin Bian |
IROS | 2 |
| 2025 | Dynamic Action Localization and Recognition for Intelligent Perception of Surgical RobotsabstractRobot-assisted surgery has significantly advanced surgical precision, yet the development of autonomous surgical robots remains hindered by their limited understanding of complex surgical actions. Current systems lack the ability to effectively perceive and interpret intricate surgical relationships, which restricts their capability to assist surgeons in dynamic surgical environments. To overcome these challenges, a novel self-supervised learning method for surgical action recognition has been proposed, aimed at enhancing the understanding of surgical actions. The method has introduced a dynamic masking with attention-based action localization module to focus the model on critical spatial regions where actions occur, enabling surgical view guidance for intelligent surgical robot while extracting key features. Moreover, a graph-enhanced adaptive feature selection module is employed to assign relevance to features and capture the temporal relationships between adjacent frames. Long Short-Term Memory has been utilized to model long-term dependencies across video sequences, while multi-view contrastive learning facilitates the extraction of discriminative features from both masked and unmasked sequences. Experimental results demonstrate a 3.4% improvement in Average Precision and an Area Under Receiver Operating Characteristic Curve of 92.9% on Neuro67 dataset for surgical action recognition. The method enables dynamic adjustments to the surgical view, achieving surgical visual navigation. These advancements contribute to the development of intelligent and autonomous surgical robots capable of assisting surgeons in complex and dynamic surgical settings. Yaqin Peng, Guibin Bian, Zhen Li 0049 |
IROS | 3 |
| 2025 | High-Precision Tracking of Time-Varying Trajectories for Microsurgical Robots in Constrained EnvironmentsabstractThis research addresses the challenge of achieving high-precision tracking of time-varying trajectories under nonlinear disturbances and motion constraints in microsurgical robots. A hybrid control framework integrating fuzzy adaptive sliding mode control with radial basis function neural networks is proposed. This framework dynamically adjusts the sliding mode gain to suppress high-frequency jitter and compensate for unmodeled disturbances such as joint friction and tissue contact forces. Experiments conducted on a self-developed microscopic ophthalmic robot platform demonstrated that the trajectory tracking error was reduced to 1.1 μm, representing improvements of 85.9%, 76.1%, and 66.7% compared to PID control, sliding mode control and non-singular fast terminal sliding mode control respectively. The tracking delay was 19 milliseconds. In experiments on living pigs with central retinal artery occlusion, the system successfully performed intravascular injection, with a maximum error of 3.97 μm. This solution, through optimization via fuzzy logic and neural networks, achieves micron-level precision and robustness, effectively solving high-frequency control noise and low-frequency environmental disturbances, ensuring both the accuracy and safety of the microsurgical robot. Yu-Peng Zhai, Guibin Bian, Zhen Li 0049, Tian-Qi Deng, Ming-Yang Zhang, Pan Fu, Wen-Hao He, Ya-Wen Deng |
IROS | 3 |
| 2025 | Incomplete Multi-View Drug Recommendation via Multi-Level Representation Learning and Curriculum LearningabstractThe drug recommendation task aims to provide effective and safe prescription decision support for clinical treatment based on patients' past Electronic Health Records (EHR). However, the prevalent phenomenon of missing views in multi-source heterogeneous EHR data may cause performance degradation. This is due to the lack of sufficient information and increased learning difficulties, which limit the practical effectiveness of drug recommendation models in medical applications. In this paper, we emphasize the problems of incompleteness in practical drug recommendation and propose the Incomplete Multi-View Drug Recommendation model via Multi-Level Representation Learning and Curriculum Learning named IMDR. In particular, IMDR employs a Multi-Level Representation Learning architecture equipped with a Medical Code-Level Drug Knowledge Infusion Module and a Visit-Level Cross-View Information Module for patient representation learning to overcome the information loss caused by incomplete data. And then, a Gaussian-guided curriculum learning strategy is proposed to assist the learning process of IMDR with a novel difficulty measure to achieve effective progressive learning under missing medical views. Systematic evaluation on two large-scale real-world medical datasets, MIMIC-III and MIMIC-IV, demonstrates that IMDR reduces the Drug-Drug Interaction (DDI) rate by 2.97% compared to existing state-of-the-art drug recommendation baselines, while achieving significant improvements of 3.29% and 1.97% in Jaccard similarity scores and F1 score, respectively. Furthermore, compared to advanced incomplete multi-view learning (IML) models, IMDR's advantages in Jaccard similarity scores and F1 score further expand to 4.03% and 2.41%. Ning Liu 0014, Yunsen Tang, Haitao Yuan 0002, Hongtao Lv, Lili Jiang 0002, Zhen Li 0049, Wei Zhang 0056, Jianyong Wang 0001 |
KDD (2) | 6 |
| 2025 | A spatiotemporal dynamic fusion network for surgical action recognition
Guibin Bian, Yaqin Peng, Zhen Li 0049 |
Neurocomputing | 3 |
| 2025 | Automatic Robotic Cranium-Milling: A Motion Control Study of In Vitro Animal ExperimentsabstractAutonomous robotic surgery offers enhanced effectiveness, precision, and reliability, regardless of the surgeons’ expertise. Prior neurosurgery robot studies involved surgeons manually assisting the robot in skull-milling tasks by holding the milling cutter shank, constraining the robot’s autonomy. A model-free adaptive nonlinear force control algorithm is designed to accomplish automatic cranial-milling tasks. Furthermore, a skull-milling breakthrough detection algorithm by monitoring the change of feed force is proposed to determine the completion of the milling task autonomously. A robotic system is developed for automatic cranium-milling and 72 in vitro skull-milling experiments indicate that when using the proposed control algorithm, the maximum root mean square error percentage of the vertical force is 0.99$\%$, while the control error percentages of other mainstream methods are all above 5.5$\%$. Moreover, the success rate of breakthrough detection is 98.61$\%$and the robot autonomously performs the skull milling task with minimal human intervention during the whole experiment. The results demonstrate that the proposed method provides the potential to improve the intelligence of neurosurgery.Note to Practitioners— The purpose is to propose a model-free adaptive nonlinear force control method for automatic skull-milling tasks. In previous studies, the involvement of surgeons manually holding the milling cutter shank to assist neurosurgery robots in skull-milling tasks has been observed. However, the autonomy of the robot is restricted and its potential for precise control tasks is failed to leverage. Therefore, a model-free adaptive nonlinear force control algorithm is proposed and a robotic system is built to enable the robot to autonomously perform cranial-milling tasks in this work. This application aims to enhance the autonomy of robot-assisted neurosurgery, making it a potential solution for remote surgery and addressing the shortage of medical resources in rural areas. Guibin Bian, Chen Qian 0006, Zhen Li 0049, Pei-Cong Ge, Jizong Zhao |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Few-Human-Interaction Reinforcement Learning for Autonomous Transbronchial InterventionabstractThe transbronchial interventional surgery presents challenges with winding and convoluted pathways, prone to compression and friction. Current autonomous planning struggles to reach deeper bronchial positions, and hard to consider multiple conflicting goals simultaneously. This article introduces an innovative planning scheme with preference weights to achieve smooth, frictionless, and collision-free autonomous transbronchial intervention with continuum robot (CR). A few-human-interaction twin-delayed deep deterministic policy gradient (FHITD3) generated from surgeon preference guidance is proposed, which determines the optimal strategy for the motion of CR. Preference knowledge is generated through interaction between human and few diversity samples. An abstract actuator space description is proposed for the posture and position representation of CR during movement within bronchus. A contact motion analysis strategy is proposed to calculate real-time attitude of CR in contact with bronchus. In addition, an oscillation suppression approach to address CR's unsmooth distal end trajectory is proposed. Simulated experiments show that the CR autonomously completes intervention tasks with a smooth and stable trajectory, reducing distal end oscillation by over 45%. It achieves a target endpoint within the fourth level bronchus (approximately 5 mm diameter) with over 90% probability. Guibin Bian, Xiang-Rong Tang, Zhen Li 0049, Ming-Yang Zhang, Yu-Peng Zhai |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Procedure Recognition by Knowledge-Driven Segmentation in Robotic-Assisted Vitreoretinal SurgeryabstractInternal limiting membrane (ILM) peeling is a vital vitreoretinal surgery procedure. However, due to the thickness of just 1-2 micrometers and the intricacies associated with its varying density and adhesion, the difficulty of manipulation exceeds the physiological limits of human perception and operation. Surgical robot is characterized by high precision and stability. However, navigating intricate intraocular environments and handling minuscule high-precision areas remain enormous challenges. These include issues of uneven lighting, field-of-view loss, and motion blur. This paper proposed a perception method named ‘Multimodal Surgical Process Recognition based on Domain Knowledge and Segmentation (MSPR-DKS),’ designed to address these challenges and provide input for the precise control of robots. Moreover, a comprehensive dataset focused on ILM peeling during macular hole surgeries was established. Experimental results underscore the efficacy of this approach, with segmentation accuracies exceeding 99.27% for instruments and macular holes and an average accuracy of 98.97% in recognizing surgical processes. This study paves the way for leveraging domain knowledge and image segmentation to improve robot-assisted manipulation of soft tissues in ophthalmology. Zhen Li 0049, Ya-Wen Deng, Weihong Yu, Haoxiang Qi, Yaliang Liu, Zhangguo Yu, Guibin Bian |
ICRA | 1 |
| 2024 | A Hybrid Admittance Control Algorithm for Automatic Robotic Cranium-MillingabstractPrior robot-assisted cranium-milling studies only considered controlling the force in the skull’s vertical direction and neglected the milling cutter’s feed force. Additionally, achieving stable force control in multiple directions is challenging for robots due to the uneven skull surface. Here a hybrid admittance control algorithm incorporating a model-free adaptive nonlinear force control and fuzzy control algorithms is proposed to accomplish effective automatic cranial-milling tasks. First, a pure data-driven model-free adaptive control method based on partial form dynamic linearization is used to control the feed force. Second, fuzzy control minimizes the total error of both the vertical and feed force by adaptively adjusting the milling cutter’s velocity and position. 42 ex vivo animal skull-milling experiments conducted by the automatic robotic cranium-milling system indicate that when using the proposed control algorithm, the force error percentage can be maintained below 5.0% within 3 s and the maximal root mean square error percentages for vertical and feed force are 1.85% and 1.94%, respectively. Moreover, no instances of dura mater damage are observed and the robotic system exhibits a high level of autonomy as it performs the skull milling task with minimal human involvement throughout the entire experiment. The results suggest the potential for advancing the intelligence level of neurosurgery in the future. Chen Qian 0006, Zhen Li 0049, Pei-Cong Ge, Jizong Zhao, Guibin Bian |
ICRA | 2 |
| 2024 | Design and Modeling of a Thin-walled Multi-segment Continuum Robotic BronchoscopeabstractCable-driven continuum robots in bronchoscopic procedures hold immense potential to revolutionize the diagnosis and treatment of lung cancer. However, robotic bronchoscopes in current studies are typically large in size and inflexible. Therefore, this article introduces a novel cable-driven continuum robot bronchoscopy system that achieves modular design between the actuation and operation ends. A continuum structure with a dual-segment notched flexible skeleton, featuring a wall thickness of 0.45 mm, has been designed to perform bending movements exceeding 190°. This enhances flexibility and increases the spatial capacity of the working channels. A kinematic model was developed, integrating the actuation force and the mechanical characteristics of the driving cables for error compensation, estimating the correlation between the displacement of the driving cables and the position of the continuum robot’s end-effector. The verification showed that the root mean square error (RMSE) of the end-effector position is 2.57 mm, which accounts for 4.8% of the continuum’s length. A prototype of the robotic bronchoscopy system was created, and its performance and potential applications in bronchoscopic intervention surgeries were validated through vivo pig intervention experiments. Guibin Bian, Ming-Yang Zhang, Yu-Peng Zhai, Zhen Li 0049 |
IROS | 7 |
| 2024 | Multilevel Causality Learning for Multi-label Gastric Atrophy Diagnosis
Xiaoxiao Cui, Shanzhi Jiang, Baolin Sun, Yankun Cao, Zhen Li 0049, Chaoyang Lv, Zhi Liu 0004, Li-Zhen Cui 0001, Shuo Li 0001 |
MICCAI (3) | 6 |
| 2024 | BiMNet: A Multimodal Data Fusion Network for continuous circular capsulorhexis Action Segmentation
Guibin Bian, Zhen Li 0049, Pan Fu, Chen Xin 0003, Daniel Santos da Silva, Victor Hugo C. de Albuquerque |
Expert Syst. Appl. | 3 |
| 2024 | Self-supervised visual-textual prompt learning for few-shot grading of gastric intestinal metaplasia
Xuanchi Chen, Xiangwei Zheng 0001, Zhen Li 0049, Mingzhe Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2023 | MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report GenerationabstractAutomatic medical report generation is an essential task in applying artificial intelligence to the medical domain, which can lighten the workloads of doctors and promote clinical automation. The state-of-the-art approaches employ Transformer-based encoder-decoder architectures to generate reports for medical images. However, they do not fully explore the relationships between multi-modal medical data, and generate inaccurate and inconsistent reports. To address these issues, this paper proposes a Multi-modal Memory Transformer Network (MMTN) to cope with multi-modal medical data for generating image-report consistent medical reports. On the one hand, MMTN reduces the occurrence of image-report inconsistencies by designing a unique encoder to associate and memorize the relationship between medical images and medical terminologies. On the other hand, MMTN utilizes the cross-modal complementarity of the medical vision and language for the word prediction, which further enhances the accuracy of generating medical reports. Extensive experiments on three real datasets show that MMTN achieves significant effectiveness over state-of-the-art approaches on both automatic metrics and human evaluation. Li-Zhen Cui 0001, Lei Zhang 0199, Fuqiang Yu, Zhen Li 0049 |
AAAI | 5 |
| 2023 | CMT: Cross-modal Memory Transformer for Medical Image Report Generation
Li-Zhen Cui 0001, Lei Zhang 0199, Fuqiang Yu, Zhen Li 0049, Chunyan Miao |
DASFAA (3) | 6 |
| 2023 | Automated Key Action Detection for Closed Reduction of Pelvic Fractures by Expert Surgeons in Robot-Assisted SurgeryabstractPelvic fractures are one of the most serious traumas in orthopedics, and the technical proficiency and expertise of the surgical team strongly influence the quality of reduction results. With the advancement of information technology and robotics, robot-assisted pelvic fracture reduction surgery is expected to reduce the impact caused by inexperienced doctors and improve the accuracy and stability of pelvic reduction. However, this requires the robot to detect key surgeon actions from time-series data, enabling the robot to independently perceive the surgical status, predict the surgeon's intentions, assess the demonstrated level of professional competence, and assess the progress of the surgery. Therefore, a multi-task deep learning neural network architecture is proposed, which incorporates Convolutional Neural Network-Bidirectional Long Short-Term Memory (CNN-BiLSTM) along with tri-modality fusion and feature extraction techniques. The proposed framework aims to achieve key action detection in closed reduction operations for pelvic fractures. Subsequently, a trimodal fine-grained dataset was constructed, wherein 29, 32, and 14 labels were marked on flexion, position, and pressure data for 14 key closed reduction actions. The experimental results show that the correct detection rate of closed reduction actions is 92.3 %, significantly higher than the commonly used recognition algorithms. This work provides a method for the robot to learn the surgeon's professional knowledge, provides the basis for the operation's motion perception, and contributes to the autonomy of the robot-assisted closed reduction surgery of pelvic fractures. Mingzhang Pan, Ya-Wen Deng, Zhen Li 0049, Xiao-Lan Liao, Guibin Bian |
IROS | 3 |
| 2023 | Learning surgical skills under the RCM constraint from demonstrations in robot-assisted minimally invasive surgery
Guibin Bian, Zhen Li 0049, Bing-Ting Wei, Wei-Peng Liu, Daniel Santos da Silva, Victor Hugo C. de Albuquerque |
Expert Syst. Appl. | 3 |
| 2023 | Dynamic Multiaction Recognition and Expert Movement Mapping for Closed Pelvic ReductionabstractPelvic fractures are one of the most serious traumas in orthopedic care, and reduction during routine surgery is a significant challenge. Because there are so many vital organs, blood vessels, and nerves around the pelvis, and the reduction force is large, the operational requirements for the surgeon are extremely strict and require extensive experience and surgical skills. This article proposes a method for collecting and digitizing doctors’ reduction movements, which aims to help intelligent devices recognize surgeons’ reduction actions and provides a means to learn from expert experience to improve the accuracy of surgery. First, the convolutional bidirectional long short-term memory algorithm with multilayer cross-fused features is proposed. It extracts time and spatial correlations between multimodal data in a hierarchical manner. Second, discrete dynamic motion primitives are adopted for mapping the surgeon's palm movement trajectory. Finally, this article constructs a data acquisition platform and collects data from surgeons with varying proficiency in closed reduction. Experiment results show that the closed reduction action recognition accuracy is 99% and posture recognition accuracy is 95.5%. The recognition algorithm proposed by this article is significantly higher than the commonly used algorithms in terms of Accuracy, Precision, Recall, and F1-Score. This article provides methods and means for the digitization of surgical expertise and transfers learning for robot-assisted surgery. Mingzhang Pan, Ya-Wen Deng, Zhen Li 0049, Xiao-Lan Liao, Guibin Bian |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Motion Decoupling Network for Intra-Operative Motion Estimation Under OcclusionabstractIn recent intelligent-robot-assisted surgery studies, an urgent issue is how to detect the motion of instruments and soft tissue accurately from intra-operative images. Although optical flow technology from computer vision is a powerful solution to the motion-tracking problem, it has difficulty obtaining the pixel-wise optical flow ground truth of real surgery videos for supervised learning. Thus, unsupervised learning methods are critical. However, current unsupervised methods face the challenge of heavy occlusion in the surgical scene. This paper proposes a novel unsupervised learning framework to estimate the motion from surgical images under occlusion. The framework consists of a Motion Decoupling Network to estimate the tissue and the instrument motion with different constraints. Notably, the network integrates a segmentation subnet that estimates the segmentation map of instruments in an unsupervised manner to obtain the occlusion region and improve the dual motion estimation. Additionally, a hybrid self-supervised strategy with occlusion completion is introduced to recover realistic vision clues. Extensive experiments on two surgical datasets show that the proposed method achieves accurate motion estimation for intra-operative scenes and outperforms other unsupervised methods, with a margin of 15% in accuracy. The average estimation error for tissue is less than 2.2 pixels on average for both surgical datasets. Guibin Bian, Li Zhang 0040, He Chen 0003, Zhen Li 0049, Pan Fu, Wen-Qian Yue, Yu-Wen Luo, Pei-Cong Ge, Weipeng Liu |
IEEE Trans. Medical Imaging | 4 |
| 2022 | KdINet: Knowledge-driven Interpretable Network for Medical Imaging DiagnosisabstractAutomatic diagnosis for medical images is a significant research problem of computer-aided diagnosis to reduce the workload of doctors in recent years. However, existing deep learning approaches for diagnoses are usually black-box models with implicit decision-making processes that make them inexplainable. To alleviate the issue, in this paper, we propose a Knowledge-driven Interpretable Network (KdINet) for interpretable disease classification of medical images. KdINet first exploits the pretrained CNN module and the hierarchical representation module to learn two different disease feature representations (i.e., visual disease features and hierarchical disease features). Subsequently, the joint training of two disease features by KdINet’s disease classifier learns a hierarchical classification criterion to infer the diagnosis and generate corresponding interpretable justifications (i.e., ancestral disease paths of the diagnosis). Extensive experiments on three datasets demonstrate that our KdINet achieves significantly higher effectiveness than state-of-the-art approaches on disease classification metrics. Li-Zhen Cui 0001, Lei Zhang 0199, Fuqiang Yu, Zhen Li 0049, Chunyan Miao |
BIBM | 5 |
| 2022 | KdTNet: Medical Image Report Generation via Knowledge-Driven Transformer
Li-Zhen Cui 0001, Fuqiang Yu, Lei Zhang 0199, Zhen Li 0049, Ning Liu 0014 |
DASFAA (3) | 5 |
| 2022 | SurgiNet: Pyramid Attention Aggregation and Class-wise Self-Distillation for Surgical Instrument Segmentation
Zhen-Liang Ni, Xiao-Hu Zhou, Guan'an Wang, Wen-Qian Yue, Zhen Li 0049, Guibin Bian, Zeng-Guang Hou |
Medical Image Anal. | 5 |
| 2022 | Space Squeeze Reasoning and Low-Rank Bilinear Feature Fusion for Surgical Image SegmentationabstractSurgical image segmentation is critical for surgical robot control and computer-assisted surgery. In the surgical scene, the local features of objects are highly similar, and the illumination interference is strong, which makes surgical image segmentation challenging. To address the above issues, a bilinear squeeze reasoning network is proposed for surgical image segmentation. In it, the space squeeze reasoning module is proposed, which adopts height pooling and width pooling to squeeze global contexts in the vertical and horizontal directions, respectively. The similarity between each horizontal position and each vertical position is calculated to encode long-range semantic dependencies and establish the affinity matrix. The feature maps are also squeezed from both the vertical and horizontal directions to model channel relations. Guided by channel relations, the affinity matrix is expanded to the same size as the input features. It captures long-range semantic dependencies from different directions, helping address the local similarity issue. Besides, a low-rank bilinear fusion module is proposed to enhance the model's ability to recognize similar features. This module is based on the low-rank bilinear model to capture the inter-layer feature relations. It integrates the location details from low-level features and semantic information from high-level features. Various semantics can be represented more accurately, which effectively improves feature representation. The proposed network achieves state-of-the-art performance on cataract image segmentation dataset CataSeg and robotic image segmentation dataset EndoVis 2018. Zhen-Liang Ni, Guibin Bian, Zhen Li 0049, Xiao-Hu Zhou, Rui-Qi Li, Zeng-Guang Hou |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Attention-Guided Lightweight Network for Real-Time Segmentation of Robotic Surgical InstrumentsabstractThe real-time segmentation of surgical instruments plays a crucial role in robot-assisted surgery. However, it is still a challenging task to implement deep learning models to do real-time segmentation for surgical instruments due to their high computational costs and slow inference speed. In this paper, we propose an attention-guided lightweight network (LWANet), which can segment surgical instruments in real-time. LWANet adopts encoder-decoder architecture, where the encoder is the lightweight network MobileNetV2, and the decoder consists of depthwise separable convolution, attention fusion block, and transposed convolution. Depthwise separable convolution is used as the basic unit to construct the decoder, which can reduce the model size and computational costs. Attention fusion block captures global contexts and encodes semantic dependencies between channels to emphasize target regions, contributing to locating the surgical instrument. Transposed convolution is performed to upsample feature maps for acquiring refined edges. LWANet can segment surgical instruments in real-time while takes little computational costs. Based on 960x544 inputs, its inference speed can reach 39 fps with only 3.39 GFLOPs. Also, it has a small model size and the number of parameters is only 2.06 M. The proposed network is evaluated on two datasets. It achieves state-of-the- art performance 94.10% mean IOU on Cata7 and obtains a new record on EndoVis 2017 with a 4.10% increase on mean IOU. Zhen-Liang Ni, Guibin Bian, Zeng-Guang Hou, Xiao-Hu Zhou, Xiaoliang Xie, Zhen Li 0049 |
ICRA | 6 |
| 2020 | BARNet: Bilinear Attention Network with Adaptive Receptive Fields for Surgical Instrument SegmentationabstractSurgical instrument segmentation is crucial for computer-assisted surgery. Different from common object segmentation, it is more challenging due to the large illumination variation and scale variation in the surgical scenes. In this paper, we propose a bilinear attention network with adaptive receptive fields to address these two issues. To deal with the illumination variation, the bilinear attention module models global contexts and semantic dependencies between pixels by capturing second-order statistics. With them, semantic features in challenging areas can be inferred from their neighbors, and the distinction of various semantics can be boosted. To adapt to the scale variation, our adaptive receptive field module aggregates multi-scale features and selects receptive fields adaptively. Specifically, it models the semantic relationships between channels to choose feature maps with appropriate scales, changing the receptive field of subsequent convolutions. The proposed network achieves the best performance 97.47% mean IoU on Cata7. It also takes the first place on EndoVis 2017, exceeding the second place by 10.10% mean IoU. Zhen-Liang Ni, Guibin Bian, Guan'an Wang, Xiao-Hu Zhou, Zeng-Guang Hou, Xiaoliang Xie, Zhen Li 0049, Yuhan Wang 0017 |
IJCAI | 7 |
| 2019 | RAUNet: Residual Attention U-Net for Semantic Segmentation of Cataract Surgical Instruments
Zhen-Liang Ni, Guibin Bian, Xiao-Hu Zhou, Zeng-Guang Hou, Xiaoliang Xie, Chen Wang 0122, Yan-Jie Zhou, Rui-Qi Li, Zhen Li 0049 |
ICONIP (2) | 9 |
| 2019 | Path Planning for Surgery Robot with Bidirectional Continuous Tree Search and Neural NetworkabstractSolving a thorny issue of real-time path planning for surgery robot in uncertain environments, a novel algorithm named bidirectional continuous tree search (BCTS) is proposed. Most partially observable markov decision process (POMDP) planners address challenges of unknown environments with discrete states, observations and actions, which are fail to automate the operative procedure. However, the BCTS method addresses the issue by handling POMDPs in continuous state, observation and action spaces. The proposed approach has a bidirectional search structure with the intent of greatly improving the calculation efficiency. Meanwhile, Bayesian optimization (BO) algorithm is considered to dynamically sample promising actions while we construct a belief tree. In view of the speed of BO process, the upper and lower bounds of the optimal action values given by fast informed bound (FIB) and point-based value iteration (PBVI) limit the search scope, so we can improve the speed of BO. In addition, we apply an optimal path planning generator, radial basis function neural network (RBFNN), to obtain a smoother trajectory. Finally, simulation of glaucoma surgery has been carried out to explore the best surgical approach. The results show that the introduced structure can effectively guide the surgery robot to perform surgical procedures and receive a real-time as well as smooth path. Rui-Jian Huang, Guibin Bian, Chen Xin 0003, Zhen Li 0049, Zeng-Guang Hou |
IROS | 4 |
| 2019 | An Extremely Fast and Precise Convolutional Neural Network for Recognition and Localization of Cataract Surgical Tools
Dongqing Zang, Guibin Bian, Yunlai Wang, Zhen Li 0049 |
MICCAI (5) | 4 |