Zhen-Liang Ni

dblp:241/7013 · also Zhenliang Ni · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-3358-1994ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation
abstract
Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic range conditions due to limitations of frame cameras. Event cameras offer complementary advantages such as high temporal resolution and low latency, yet lack color and texture, making them insufficient on their own. To address this, recent research has explored multimodal fusion of RGB and event data; however, many existing approaches are computationally expensive and focus primarily on spatial fusion, neglecting the temporal dynamics inherent in event streams. In this work, we propose MambaSeg, a novel dual-branch semantic segmentation framework that employs parallel Mamba encoders to efficiently model RGB images and event streams. To reduce cross-modal ambiguity, we introduce the Dual-Dimensional Interaction Module (DDIM), comprising a Cross-Spatial Interaction Module (CSIM) and a Cross-Temporal Interaction Module (CTIM), which jointly perform fine-grained fusion along both spatial and temporal dimensions. This design improves cross-modal alignment, reduces ambiguity, and leverages the complementary properties of each modality. Extensive experiments on the DDD17 and DSEC datasets demonstrate that MambaSeg achieves state-of-the-art segmentation performance while significantly reducing computational cost, showcasing its promise for efficient, scalable, and robust multimodal perception.
Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji, Chao Chen 0004, Qingyi Gu, Zhen-Liang Ni
AAAI7
2026 GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
abstract
The rapid advancement of video generation models has made it increasingly challenging to distinguish AI-generated videos from real ones. This issue underscores the urgent need for effective AI-generated video detectors to prevent the dissemination of false information via such videos. However, the development of high-performance AI-generated video detectors is currently impeded by the lack of large-scale, high-quality datasets specifically designed for generative video detection. To this end, we introduce GenVidBench, a challenging AI-generated video detection dataset with several key advantages: 1) Large-scale video collection: The dataset contains 6.78 million videos and is currently the largest dataset for AI-generated video detection. 2) Cross-Source and Cross-Generator: The cross-source generation reduces the interference of video content on the detection. The cross-generator ensures diversity in video attributes between the training and test sets, preventing them from being overly similar. 3) State-of-the-Art Video Generators: The dataset includes videos from 11 state-of-the-art AI video generators, ensuring that it covers the latest advancements in the field of video generation. These generators ensure that the datasets are not only large in scale but also diverse, aiding in the development of generalized and effective detection models. Additionally, we present extensive experimental results with advanced video classification models. With GenVidBench, researchers can efficiently develop and evaluate AI-generated video detection models.
Zhen-Liang Ni, Qiangyu Yan, Mouxiao Huang, Tianning Yuan, Yehui Tang 0001, Hailin Hu 0002, Xinghao Chen 0001, Yunhe Wang 0001
AAAI1
2025 TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
abstract
Mamba has shown great potential for computer vision due to its linear complexity in modeling the global context with respect to the input length. However, existing lightweight Mamba-based backbones cannot demonstrate performance that matches Convolution or Transformer-based methods. By observing, we find that simply modifying the scanning path in the image domain is not conducive to fully exploiting the potential of vision Mamba. In this paper, we first perform comprehensive spectral and quantitative analyses, and verify that the Mamba block mainly models low-frequency information under Convolution-Mamba hybrid architecture. Based on the analyses, we introduce a novel Laplace mixer to decouple the features in terms of frequency and input only the low-frequency components into the Mamba block. In addition, considering the redundancy of the features and the different requirements for high-frequency details and low-frequency global information at different stages, we introduce a frequency ramp inception, i.e., gradually reduce the input dimensions of the high-frequency branches, so as to efficiently trade-off the high-frequency and low-frequency components at different layers. By integrating mobile-friendly convolution and efficient Laplace mixer, we build a series of tiny hybrid vision Mamba called TinyViM. The proposed TinyViM achieves impressive performance on several downstream tasks including image classification, semantic segmentation, object detection and instance segmentation. In particular, TinyViM outperforms Convolution, Transformer and Mamba-based models with similar scales, and the throughput is about 2-3 times higher than that of other Mamba-based models. Code is available at https://github.com/xwmaxwma/TinyViM.
Zhen-Liang Ni, Xinghao Chen 0001
ICCV2
2025 TimePro: Efficient Multivariate Long-term Time Series Forecasting with Variable- and Time-Aware Hyper-state
abstract
In long-term time series forecasting, different variables often influence the target variable over distinct time intervals, a challenge known as the multi-delay issue. Traditional models typically process all variables or time points uniformly, which limits their ability to capture complex variable relationships and obtain non-trivial time representations. To address this issue, we propose TimePro, an innovative Mamba-based model that constructs variate- and time-aware hyper-states. Unlike conventional approaches that merely transfer plain states across variable or time dimensions, TimePro preserves the fine-grained temporal features of each variate token and adaptively selects the focused time points to tune the plain state. The reconstructed hyper-state can perceive both variable relationships and salient temporal information, which helps the model make accurate forecasting. In experiments, TimePro performs competitively on eight real-world long-term forecasting benchmarks with satisfactory linear complexity. Code is available at https://github.com/xwmaxwma/TimePro.
Zhen-Liang Ni, Xinghao Chen 0001
ICML2
2024 Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
Zhen-Liang Ni, Xinghao Chen 0001, Yingjie Zhai, Yehui Tang 0001, Yunhe Wang 0001
ECCV (52)1
2024 SSA-Seg: Semantic and Spatial Adaptive Pixel-level Classifier for Semantic Segmentation
abstract
Vanilla pixel-level classifiers for semantic segmentation are based on a certain paradigm, involving the inner product of fixed prototypes obtained from the training set and pixel features in the test image. This approach, however, encounters significant limitations, i.e., feature deviation in the semantic domain and information loss in the spatial domain. The former struggles with large intra-class variance among pixel features from different images, while the latter fails to utilize the structured information of semantic objects effectively. This leads to blurred mask boundaries as well as a deficiency of fine-grained recognition capability. In this paper, we propose a novel Semantic and Spatial Adaptive Classifier (SSA-Seg) to address the above challenges. Specifically, we employ the coarse masks obtained from the fixed prototypes as a guide to adjust the fixed prototype towards the center of the semantic and spatial domains in the test image. The adapted prototypes in semantic and spatial domains are then simultaneously considered to accomplish classification decisions. In addition, we propose an online multi-domain distillation learning strategy to improve the adaption process. Experimental results on three publicly available benchmarks show that the proposed SSA-Seg significantly improves the segmentation performance of the baseline models with only a minimal increase in computational cost.
Zhen-Liang Ni, Xinghao Chen 0001
NeurIPS2
2023 Dual Relation Knowledge Distillation for Object Detection
abstract
Knowledge distillation is an effective method for model compression. However, it is still a challenging topic to apply knowledge distillation to detection tasks. There are two key points resulting in poor distillation performance for detection tasks. One is the serious imbalance between foreground and background features, another one is that small object lacks enough feature representation. To solve the above issues, we propose a new distillation method named dual relation knowledge distillation (DRKD), including pixel-wise relation distillation and instance-wise relation distillation. The pixel-wise relation distillation embeds pixel-wise features in the graph space and applies graph convolution to capture the global pixel relation. By distilling the global pixel relation, the student detector can learn the relation between foreground and background features, and avoid the difficulty of distilling features directly for the feature imbalance issue. Besides, we find that instance-wise relation supplements valuable knowledge beyond independent features for small objects. Thus, the instance-wise relation distillation is designed, which calculates the similarity of different instances to obtain a relation matrix. More importantly, a relation filter module is designed to highlight valuable instance relations. The proposed dual relation knowledge distillation is general and can be easily applied for both one-stage and two-stage detectors. Our method achieves state-of-the-art performance, which improves Faster R-CNN based on ResNet50 from 38.4% to 41.6% mAP and improves RetinaNet based on ResNet50 from 37.4% to 40.3% mAP on COCO 2017.
Zhen-Liang Ni, Fukui Yang, Shengzhao Wen
IJCAI1
2023 Learning Skill Characteristics From Manipulations
abstract
Percutaneous coronary intervention (PCI) has increasingly become the main treatment for coronary artery disease. The procedure requires high experienced skills and dexterous manipulations. However, there are few techniques to model PCI skill so far. In this study, a learning framework with local and ensemble learning is proposed to learn skill characteristics of different skill-level subjects from their PCI manipulations. Ten interventional cardiologists (four experts and six novices) were recruited to deliver a medical guidewire to two target arteries on a porcine model for in vivo studies. Simultaneously, translation and twist manipulations of thumb, forefinger, and wrist are acquired with electromagnetic (EM) and fiber-optic bend (FOB) sensors, respectively. These behavior data are then processed with wavelet packet decomposition (WPD) under 1-10 levels for feature extraction. The feature vectors are further fed into three candidate individual classifiers in the local learning layer. Furthermore, the local learning results from different manipulation behaviors are fused in the ensemble learning layer with three rule-based ensemble learning algorithms. In subject-dependent skill characteristics learning, the ensemble learning can achieve 100% accuracy, significantly outperforming the best local result (90%). Furthermore, ensemble learning can also maintain 73% accuracy in subject-independent schemes. These promising results demonstrate the great potential of the proposed method to facilitate skill learning in surgical robotics and skill assessment in clinical practice.
Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Zhen-Liang Ni, Yan-Jie Zhou, Rui-Qi Li, Mei-Jiang Gui, Chen-Chen Fan, Zhen-Qiu Feng, Guibin Bian, Zeng-Guang Hou
IEEE Trans. Neural Networks Learn. Syst.4
2022 A Dual-Stream Architecture for Real-Time Morphological Analysis of Aneurysm in Robot-Assisted Minimally Invasive Surgery
abstract
Real-time and precise morphological analysis of intraoperative AAA is a significant pre-imperative for robot-assisted minimally invasive surgery (RMIS). However, this task is frequently accompanied by the difficulties of ambiguous boundaries and obscured surfaces of aneurysms. To remedy these problems, we propose a Light-Weight Dual-Stream Boundary-Aware Network (DSB-Net) and a novel diagnosis algorithm for real-time morphological analysis of AAA. In the network, the features at the boundaries are preserved by incorporating a boundary localization stream, while the interior segmentation accuracy is guaranteed with a mask prediction stream. Moreover, the diagnosis algorithm is developed to measure the exact size of AAA. Quantitative and qualitative assessments on two different types of datasets illustrate that (1) The presented DSB-Net remarkably outperforms the other previously proposed medical networks with the inference rate of 10.8 FPS, which meets the real-time clinical necessities. (2) The developed algorithm provides accurate size measurements for AAA, which indicates the proposed approach can be integrated into the robotic navigation framework for RMIS.
Yan-Jie Zhou, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Rui-Qi Li, Zhen-Liang Ni, Chen-Chen Fan
ICRA7
2022 SurgiNet: Pyramid Attention Aggregation and Class-wise Self-Distillation for Surgical Instrument Segmentation
Zhen-Liang Ni, Xiao-Hu Zhou, Guan'an Wang, Wen-Qian Yue, Zhen Li 0049, Guibin Bian, Zeng-Guang Hou
Medical Image Anal.1
2022 A Multilayer and Multimodal-Fusion Architecture for Simultaneous Recognition of Endovascular Manipulations and Assessment of Technical Skills
abstract
The clinical success of the percutaneous coronary intervention (PCI) is highly dependent on endovascular manipulation skills and dexterous manipulation strategies of interventionalists. However, the analysis of endovascular manipulations and related discussion for technical skill assessment are limited. In this study, a multilayer and multimodal-fusion architecture is proposed to recognize six typical endovascular manipulations. The synchronously acquired multimodal motion signals from ten subjects are used as the inputs of the architecture independently. Six classification-based and two rule-based fusion algorithms are evaluated for performance comparisons. The recognition metrics under the determined architecture are further used to assess technical skills. The experimental results indicate that the proposed architecture can achieve the overall accuracy of 96.41%, much higher than that of a single-layer recognition architecture (92.85%). In addition, the multimodal fusion brings significant performance improvement in comparison with single-modal schemes. Furthermore, the K -means-based skill assessment can obtain an accuracy of 95% to cluster the attempts made by different skill-level groups. These hopeful results indicate the great possibility of the architecture to facilitate clinical skill assessment and skill learning.
Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Qiu Feng, Zeng-Guang Hou, Guibin Bian, Rui-Qi Li, Zhen-Liang Ni, Shiqi Liu 0004, Yan-Jie Zhou
IEEE Trans. Cybern.7
2022 Space Squeeze Reasoning and Low-Rank Bilinear Feature Fusion for Surgical Image Segmentation
abstract
Surgical image segmentation is critical for surgical robot control and computer-assisted surgery. In the surgical scene, the local features of objects are highly similar, and the illumination interference is strong, which makes surgical image segmentation challenging. To address the above issues, a bilinear squeeze reasoning network is proposed for surgical image segmentation. In it, the space squeeze reasoning module is proposed, which adopts height pooling and width pooling to squeeze global contexts in the vertical and horizontal directions, respectively. The similarity between each horizontal position and each vertical position is calculated to encode long-range semantic dependencies and establish the affinity matrix. The feature maps are also squeezed from both the vertical and horizontal directions to model channel relations. Guided by channel relations, the affinity matrix is expanded to the same size as the input features. It captures long-range semantic dependencies from different directions, helping address the local similarity issue. Besides, a low-rank bilinear fusion module is proposed to enhance the model's ability to recognize similar features. This module is based on the low-rank bilinear model to capture the inter-layer feature relations. It integrates the location details from low-level features and semantic information from high-level features. Various semantics can be represented more accurately, which effectively improves feature representation. The proposed network achieves state-of-the-art performance on cataract image segmentation dataset CataSeg and robotic image segmentation dataset EndoVis 2018.
Zhen-Liang Ni, Guibin Bian, Zhen Li 0049, Xiao-Hu Zhou, Rui-Qi Li, Zeng-Guang Hou
IEEE J. Biomed. Health Informatics1
2022 TR-GAN: Multi-Session Future MRI Prediction With Temporal Recurrent Generative Adversarial Network
abstract
Magnetic Resonance Imaging (MRI) has been proven to be an efficient way to diagnose Alzheimer's disease (AD). Recent dramatic progress on deep learning greatly promotes the MRI analysis based on data-driven CNN methods using a large-scale longitudinal MRI dataset. However, most of the existing MRI datasets are fragmented due to unexpected quits of volunteers. To tackle this problem, we propose a novel Temporal Recurrent Generative Adversarial Network (TR-GAN) to complete missing sessions of MRI datasets. Unlike existing GAN-based methods, which either fail to generate future sessions or only generate fixed-length sessions, TR-GAN takes all past sessions to recurrently and smoothly generate future ones with variant length. Specifically, TR-GAN adopts recurrent connection to deal with variant input sequence length and flexibly generate future variant sessions. Besides, we also design a multiple scale & location (MSL) module and a SWAP module to encourage the model to better focus on detailed information, which helps to generate high-quality MRI data. Compared with other popular GAN architectures, TR-GAN achieved the best performance in all evaluation metrics of two datasets. After expanding the Whole MRI dataset, the balanced accuracy of AD vs. cognitively normal (CN) vs. mild cognitive impairment (MCI) and stable MCI vs. progressive MCI classification can be increased by 3.61% and 4.00%, respectively.
Chen-Chen Fan, Hongjun Yang, Xiao-Hu Zhou, Zhen-Liang Ni, Guan'an Wang, Yan-Jie Zhou, Zeng-Guang Hou
IEEE Trans. Medical Imaging6
2021 Group Feature Learning and Domain Adversarial Neural Network for aMCI Diagnosis System Based on EEG
abstract
Medical diagnostic robot systems have been paid more and more attention due to its objectivity and accuracy. The diagnosis of mild cognitive impairment (MCI) is considered an effective means to prevent Alzheimer's disease (AD). Doctors diagnose MCI based on various clinical examinations, which are expensive and the diagnosis results rely on the knowledge of doctors. Therefore, it is necessary to develop a robot diagnostic system to eliminate the influence of human factors and obtain a higher accuracy rate. In this paper, we propose a novel Group Feature Domain Adversarial Neural Network (GF- DANN) for amnestic MCI (aMCI) diagnosis, which involves two important modules. A Group Feature Extraction (GFE) module is proposed to reduce individual differences by learning group- level features through adversarial learning. A Dual Branch Domain Adaptation (DBDA) module is carefully designed to reduce the distribution difference between the source and target domain in a domain adaption way. On three types of data set, GF-DANN achieves the best accuracy compared with classic machine learning and deep learning methods. On the DMS data set, GF-DANN has obtained an accuracy rate of 89.47%, and the sensitivity and specificity are 90% and 89%. In addition, by comparing three EEG data collection paradigms, our results demonstrate that the DMS paradigm has the potential to build an aMCI diagnose robot system.
Chen-Chen Fan, Haiqun Xie, Hongjun Yang, Zhen-Liang Ni, Guan'an Wang, Yan-Jie Zhou, Zhijie Fang, Shuyun Huang, Zeng-Guang Hou
ICRA5
2021 A Real-Time Multi-Task Framework for Guidewire Segmentation and Endpoint Localization in Endovascular Interventions
abstract
Real-time guidewire segmentation and endpoint localization play a pivotal role in robot-assisted minimally invasive surgery, which is helpful to reduce radiation dose and procedure time. Nevertheless, the tasks often come with the challenge of limited computational resources. For this purpose, a real-time multi-task framework with two stages is developed. In the first stage, a Fast Attention-fused Network (FAD-Net) is proposed to obtain accurate guidewire segmentation masks. In the second stage, a lightweight localization network and a post-processing algorithm are designed to robustly predict the guidewire endpoint position. Quantitative and qualitative evaluations on intraoperative X-ray sequences from 30 patients demonstrate that the developed framework outperforms the previously-published results for the tasks, achieving state-of-the-art performance. Moreover, the inference rate of the developed framework is approximately 10.6 FPS, which meets the real-time requirement of X-ray fluoroscopy. These results indicate the proposed approach has the potential to be integrated into the robotic navigation framework for endovascular interventions, enabling robotic-assisted minimally invasive surgery.
Yan-Jie Zhou, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Guan'an Wang, Zeng-Guang Hou, Rui-Qi Li, Zhen-Liang Ni, Chen-Chen Fan
ICRA8
2021 Vessel Width Estimation via Convolutional Regression
Rui-Qi Li, Guibin Bian, Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Liang Ni, Yan-Jie Zhou, Yuhan Wang 0017, Zeng-Guang Hou
MICCAI (6)5
2021 Comparative validation of multi-instance instrument segmentation in endoscopy: Results of the ROBUST-MIS 2019 challenge
abstract
Intraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts).
Tobias Roß, Annika Reinke, Peter M. Full, Martin Wagner 0001, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mîndroc-Filimon, Patrick Godau, Thuy Nuong Tran, Pierangela Bruno, Pablo Andrés Arbeláez, Guibin Bian, Sebastian Bodenstedt, Jon Lindström Bolmgren, Laura Bravo-Sánchez, Hua-Bin Chen, Cristina González, Pål Halvorsen, Pheng-Ann Heng, Enes Hosgor, Zeng-Guang Hou, Fabian Isensee, Debesh Jha, Tingting Jiang 0001, Yueming Jin, Kadir Kirtaç, Sabrina Kletz, Stefan Leger, Klaus H. Maier-Hein, Zhen-Liang Ni, Michael Riegler 0001, Klaus Schöffmann, Ruohua Shi, Stefanie Speidel, Michael Stenzel, Isabell Twick, Guotai Wang, Jiacheng Wang 0002, Liansheng Wang 0002, Lu Wang 0002, Yan-Jie Zhou, Lei Zhu 0003, Manuel Wiesenfarth, Annette Kopp-Schneider, Beat P. Müller-Stich, Lena Maier-Hein
Medical Image Anal.33
2021 Real-Time Multi-Guidewire Endpoint Localization in Fluoroscopy Images
abstract
The real-time localization of the guidewire endpoints is a stepping stone to computer-assisted percutaneous coronary intervention (PCI). However, methods for multi-guidewire endpoint localization in fluoroscopy images are still scarce. In this paper, we introduce a framework for real-time multi-guidewire endpoint localization in fluoroscopy images. The framework consists of two stages, first detecting all guidewire instances in the fluoroscopy image, and then locating the endpoints of each single guidewire instance. In the first stage, a YOLOv3 detector is used for guidewire detection, and a post-processing algorithm is proposed to refine the guidewire detection results. In the second stage, a Segmentation Attention-hourglass (SA-hourglass) network is proposed to predict the endpoint locations of each single guidewire instance. The SA-hourglass network can be generalized to the keypoint localization of other surgical instruments. In our experiments, the SA-hourglass network is applied not only on a guidewire dataset but also on a retinal microsurgery dataset, reaching the mean pixel error (MPE) of 2.20 pixels on the guidewire dataset and the MPE of 5.30 pixels on the retinal microsurgery dataset, both achieving the state-of-the-art localization results. Besides, the inference rate of our framework is at least 20FPS, which meets the real-time requirement of fluoroscopy images (6-12FPS).
Rui-Qi Li, Xiaoliang Xie, Xiao-Hu Zhou, Shiqi Liu 0004, Zhen-Liang Ni, Yan-Jie Zhou, Guibin Bian, Zeng-Guang Hou
IEEE Trans. Medical Imaging5
2020 Pyramid Attention Aggregation Network for Semantic Segmentation of Surgical Instruments
abstract
Semantic segmentation of surgical instruments plays a critical role in computer-assisted surgery. However, specular reflection and scale variation of instruments are likely to occur in the surgical environment, undesirably altering visual features of instruments, such as color and shape. These issues make semantic segmentation of surgical instruments more challenging. In this paper, a novel network, Pyramid Attention Aggregation Network, is proposed to aggregate multi-scale attentive features for surgical instruments. It contains two critical modules: Double Attention Module and Pyramid Upsampling Module. Specifically, the Double Attention Module includes two attention blocks (i.e., position attention block and channel attention block), which model semantic dependencies between positions and channels by capturing joint semantic information and global contexts, respectively. The attentive features generated by the Double Attention Module can distinguish target regions, contributing to solving the specular reflection issue. Moreover, the Pyramid Upsampling Module extracts local details and global contexts by aggregating multi-scale attentive features. It learns the shape and size features of surgical instruments in different receptive fields and thus addresses the scale variation issue. The proposed network achieves state-of-the-art performance on various datasets. It achieves a new record of 97.10% mean IOU on Cata7. Besides, it comes first in the MICCAI EndoVis Challenge 2017 with 9.90% increase on mean IOU.
Zhen-Liang Ni, Guibin Bian, Guan'an Wang, Xiao-Hu Zhou, Zeng-Guang Hou, Hua-Bin Chen, Xiaoliang Xie
AAAI1
2020 CAU-net: A Novel Convolutional Neural Network for Coronary Artery Segmentation in Digital Substraction Angiography
Rui-Qi Li, Guibin Bian, Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Liang Ni, Zeng-Guang Hou
ICONIP (1)5
2020 Attention-Guided Lightweight Network for Real-Time Segmentation of Robotic Surgical Instruments
abstract
The real-time segmentation of surgical instruments plays a crucial role in robot-assisted surgery. However, it is still a challenging task to implement deep learning models to do real-time segmentation for surgical instruments due to their high computational costs and slow inference speed. In this paper, we propose an attention-guided lightweight network (LWANet), which can segment surgical instruments in real-time. LWANet adopts encoder-decoder architecture, where the encoder is the lightweight network MobileNetV2, and the decoder consists of depthwise separable convolution, attention fusion block, and transposed convolution. Depthwise separable convolution is used as the basic unit to construct the decoder, which can reduce the model size and computational costs. Attention fusion block captures global contexts and encodes semantic dependencies between channels to emphasize target regions, contributing to locating the surgical instrument. Transposed convolution is performed to upsample feature maps for acquiring refined edges. LWANet can segment surgical instruments in real-time while takes little computational costs. Based on 960x544 inputs, its inference speed can reach 39 fps with only 3.39 GFLOPs. Also, it has a small model size and the number of parameters is only 2.06 M. The proposed network is evaluated on two datasets. It achieves state-of-the- art performance 94.10% mean IOU on Cata7 and obtains a new record on EndoVis 2017 with a 4.10% increase on mean IOU.
Zhen-Liang Ni, Guibin Bian, Zeng-Guang Hou, Xiao-Hu Zhou, Xiaoliang Xie, Zhen Li 0049
ICRA1
2020 A Multilayer-Multimodal Fusion Architecture for Pattern Recognition of Natural Manipulations in Percutaneous Coronary Interventions
abstract
The increasingly-used robotic systems can provide precise delivery and reduce X-ray radiation to medical staff in percutaneous coronary interventions (PCI), but natural manipulations of interventionalists are forgone in most robot-assisted procedures. Therefore, it is necessary to explore natural manipulations to design more advanced human-robot interfaces (HRI). In this study, a multilayer-multimodal fusion architecture is proposed to recognize six typical subpatterns of guidewire manipulations in conventional PCI. The synchronously acquired multimodal behaviors from ten subjects are used as the inputs of the fusion architecture. Six classification-based and two rule-based fusion algorithms are evaluated for performance comparisons. Experimental results indicate that the multimodal fusion brings significant accuracy improvement in comparison with single-modal schemes. Furthermore, the proposed architecture can achieve the overall accuracy of 96.90%, much higher than that of a singlelayer recognition architecture (92.56%). These results have indicated the potential of the proposed method for facilitating the development of HRI for robot-assisted PCI.
Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Qiu Feng, Zeng-Guang Hou, Guibin Bian, Rui-Qi Li, Zhen-Liang Ni, Shiqi Liu 0004, Yan-Jie Zhou
ICRA7
2020 BARNet: Bilinear Attention Network with Adaptive Receptive Fields for Surgical Instrument Segmentation
abstract
Surgical instrument segmentation is crucial for computer-assisted surgery. Different from common object segmentation, it is more challenging due to the large illumination variation and scale variation in the surgical scenes. In this paper, we propose a bilinear attention network with adaptive receptive fields to address these two issues. To deal with the illumination variation, the bilinear attention module models global contexts and semantic dependencies between pixels by capturing second-order statistics. With them, semantic features in challenging areas can be inferred from their neighbors, and the distinction of various semantics can be boosted. To adapt to the scale variation, our adaptive receptive field module aggregates multi-scale features and selects receptive fields adaptively. Specifically, it models the semantic relationships between channels to choose feature maps with appropriate scales, changing the receptive field of subsequent convolutions. The proposed network achieves the best performance 97.47% mean IoU on Cata7. It also takes the first place on EndoVis 2017, exceeding the second place by 10.10% mean IoU.
Zhen-Liang Ni, Guibin Bian, Guan'an Wang, Xiao-Hu Zhou, Zeng-Guang Hou, Xiaoliang Xie, Zhen Li 0049, Yuhan Wang 0017
IJCAI1
2019 RAUNet: Residual Attention U-Net for Semantic Segmentation of Cataract Surgical Instruments
Zhen-Liang Ni, Guibin Bian, Xiao-Hu Zhou, Zeng-Guang Hou, Xiaoliang Xie, Chen Wang 0122, Yan-Jie Zhou, Rui-Qi Li, Zhen Li 0049
ICONIP (2)1
2019 A Two-Stage Framework for Real-Time Guidewire Endpoint Localization
Rui-Qi Li, Guibin Bian, Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Liang Ni, Zeng-Guang Hou
MICCAI (5)5