Xiao-Hu Zhou

dblp:195/8928 · DBLP profile ↗
← Back
51ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0002-7602-4848ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 4 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 VasoMIM: Vascular Anatomy-Aware Masked Image Modeling for Vessel Segmentation
abstract
Accurate vessel segmentation in X-ray angiograms is crucial for numerous clinical applications. However, the scarcity of annotated data presents a significant challenge, which has driven the adoption of self-supervised learning (SSL) methods such as masked image modeling (MIM) to leverage large-scale unlabeled data for learning transferable representations. Unfortunately, conventional MIM often fails to capture vascular anatomy because of the severe class imbalance between vessel and background pixels, leading to weak vascular representations. To address this, we introduce Vascular anatomy-aware Masked Image Modeling (VasoMIM), a novel MIM framework tailored for X-ray angiograms that explicitly integrates anatomical knowledge into the pre-training process. Specifically, it comprises two complementary components: anatomy-guided masking strategy and anatomical consistency loss. The former preferentially masks vessel-containing patches to focus the model on reconstructing vessel-relevant regions. The latter enforces consistency in vascular semantics between the original and reconstructed images, thereby improving the discriminability of vascular representations. Empirically, VasoMIM achieves state-of-the-art performance across three datasets. These findings highlight its potential to facilitate X-ray angiogram analysis.
De-Xing Huang, Xiao-Hu Zhou, Mei-Jiang Gui, Xiaoliang Xie, Shuang-Yi Wang, Tian-Yu Xiang, Rui-Ze Ma, Nu-Fang Xiao, Zeng-Guang Hou
AAAI2
2026 HGATSolver: A Heterogeneous Graph Attention Solver for Fluid-Structure Interaction
abstract
Fluid–structure interaction (FSI) systems involve distinct physical domains, fluid and solid, governed by different partial differential equations and coupled at a dynamic interface. While learning-based solvers offer a promising alternative to costly numerical simulations, existing methods struggle to capture the heterogeneous dynamics of FSI within a unified framework. This challenge is further exacerbated by inconsistencies in response across domains due to interface coupling and by disparities in learning difficulty across fluid and solid regions, leading to instability during prediction. To address these challenges, we propose the Heterogeneous Graph Attention Solver (HGATSolver). HGATSolver encodes the system as a heterogeneous graph, embedding physical structure directly into the model via distinct node and edge types for fluid, solid, and interface regions. This enables specialized message-passing mechanisms tailored to each physical domain. To stabilize explicit time stepping, we introduce a novel physics-conditioned gating mechanism that serves as a learnable, adaptive relaxation factor. Furthermore, an Inter-domain Gradient-Balancing Loss dynamically balances the optimization objectives across domains based on predictive uncertainty. Extensive experiments on two constructed FSI benchmarks and a public dataset demonstrate that HGATSolver achieves state-of-the-art performance, establishing an effective framework for surrogate modeling of coupled multi-physics systems.
Haichuan Lin, Linying Cao, Xiao-Hu Zhou, Chen Chen 0036, Shuang-Yi Wang, Zeng-Guang Hou
AAAI6
2026 Pseudo-Label Guided Multi-Task Learning for Abdominal Multi-Branch Vascular Segmentation From Partially Labeled DSA Datasets
abstract
Accurate segmentation of multi-branched blood vessels from Digital Subtraction Angiography (DSA) images is essential to improve efficiency and safety of vascular interventional procedures. However, the high-speed flow of contrast agents may lead to incomplete visualization of multiple blood vessel branches and unclear boundary contours, resulting in the so-called partial labeling issue. This significantly undermines the network’s ability to extract and understand the features of multi-branched vascular structures with uncertain region, thereby severely impairing the accuracy of the recognition results. In this paper, we introduce a novel pseudo-label guided multi-task learning strategy, capable of effectively learning feature representation completion under partial label supervision. Specifically, a pretext task branch that generates boundary pseudo-label signals extracts absence structural information and transfers it to the target task for multi-branch vascular segmentation. To achieve more precise semantic-supplementing between tasks, an affinity-based criss-cross feature propagation (CCFP) module is designed to dynamically fill semantic and structural voids caused by missing categories. Furthermore, to mitigate performance degradation caused by unreliable pseudo-labels, a unique loss function is proposed to constrain redundant information at both the pixel and structural levels. We validate our approach through the creation of an in-house DSA dataset composed of six sub-datasets, each containing different vascular branches. Extensive experimental results demonstrate that our method not only addresses the challenge of partial labeling but also strikes a balance between pixel-wise accuracy and the preservation of structural integrity, offering potential value in the field of clinical applications.
Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Zhi-Chao Lai
IEEE Trans Autom. Sci. Eng.5
2026 Toward Precise Guidance: A Novel Cross-Dimensional Mapping Framework for 3-D Cerebrovascular Surgical Navigation
Haining Zhao 0002, Shiqi Liu 0004, Ji-Chang Luo, Xiao-Hu Zhou, Zeng-Guang Hou, Li-Qun Jiao, Xiyao Ma, Lin-Sen Zhang, Xiaoliang Xie
IEEE Trans Autom. Sci. Eng.4
2026 LEASE: Offline Preference-Based Reinforcement Learning With High Sample Efficiency
abstract
Offline preference-based reinforcement learning (PbRL) offers an effective approach to addressing the challenges of designing rewards and mitigating the high costs associated with online interaction. However, since labeling preference needs real-time human feedback, acquiring sufficient preference labels is challenging. To solve this, this article proposes an offline PbRL with a high sample efficiency (LEASE) algorithm, where a learned transition model is leveraged to generate unlabeled preference data. Considering the pretrained reward model may generate incorrect labels for unlabeled data, we design an uncertainty-aware mechanism to ensure the performance of the reward model, where only high-confidence and low-variance data are selected. Moreover, the generalization bound of the reward model is provided to analyze the factors influencing reward accuracy, and the policy learned byLEASEhas a theoretical improvement guarantee. The above developed theory is based on a state-action pair, which can be easily combined with other offline algorithms. The experimental results show thatLEASEcan achieve comparable performance to the baseline under fewer preference data without online interaction.
Xiao-Yin Liu, Xiao-Hu Zhou, Zeng-Guang Hou
IEEE Trans. Syst. Man Cybern. Syst.3
2025 CAS-IQA: Teaching Vision-Language Models for Synthetic Angiography Quality Assessment
De-Xing Huang, Xiao-Hu Zhou, Mei-Jiang Gui, Nu-Fang Xiao, Jian-Long Hao, Ming-Yuan Liu, Zeng-Guang Hou
ICONIP (2)3
2025 Model-Free Catheter Delivery Strategy for Robotic Transcatheter Tricuspid Valve Replacement
abstract
Transcatheter tricuspid valve replacement (TTVR) has emerged as a promising minimally invasive procedure for treating severe tricuspid regurgitation (TR). However, accurate catheter delivery remains a significant challenge, primarily due to the reliance on 2D vision feedback, complex catheter kinematics, camera-to-robot pose calibration, which are difficult to generalize across patients. To address these issues, this paper presents a model-free robotic catheter delivery strategy for TTVR using Data-Enabled Predictive Control (DeePC). This approach leverages data-driven control to optimize catheter positioning without the need for prior knowledge of the system’s dynamics, eliminating the need for complex kinematic models or camera calibration. The proposed method incorporates environmental constraints to ensure the safety of the procedure, delivering the catheter to the desired location with high accuracy across varying catheters and camera poses. Experimental results demonstrate the effectiveness and versatility of the approach, suggesting its potential for broader applications in robotic-assisted surgeries. This work presents a new perspective for vision based robotic TTVR, as well as other clinical interventions involving robotic catheter control.
Haichuan Lin, Longyue Tan, Weizhao Wang, Yuen Chiu Ng, Xilong Hou 0001, Chen Chen 0036, Xiao-Hu Zhou, Zeng-Guang Hou, Shuangyi Wang
IROS10
2025 Environment-Driven Online LiDAR-Camera Extrinsic Calibration
abstract
LiDAR-camera extrinsic calibration (LCEC) is crucial for multi-modal data fusion in autonomous robotic systems. Existing methods, whether target-based or target-free, typically rely on customized calibration targets or fixed scene types, which limit their applicability in real-world scenarios. To address these challenges, we present EdO-LCEC, the first environment-driven online calibration approach. Unlike traditional target-free methods, EdO-LCEC employs a generalizable scene discriminator to estimate the feature density of the application environment. Guided by this feature density, EdO-LCEC extracts LiDAR intensity and depth features from varying perspectives to achieve higher calibration accuracy. To overcome the challenges of cross-modal feature matching between LiDAR and camera, we introduce dual-path correspondence matching (DPCM), which leverages both structural and textural consistency for reliable 3D-2D correspondences. Furthermore, we formulate the calibration process as a joint optimization problem that integrates global constraints across multiple views and scenes, thereby enhancing overall accuracy. Extensive experiments on real-world datasets demonstrate that EdO-LCEC outperforms state-of-the-art methods, particularly in scenarios involving sparse point clouds or partially overlapping sensor views.
Hongbo Zhao 0009, Ping Zhong 0002, Xiao-Hu Zhou, Wei Ye 0001, Rui Fan 0001
IEEE Trans Autom. Sci. Eng.6
2025 Real-Time 2D/3D Registration via CNN Regression and Centroid Alignment
abstract
Registration of pre-operative 3D volumes and intra-operative 2D images is critical for neurological interventions. In various 2D/3D registration tasks, deep learning-based approaches have become popular and achieved tremendous success. However, due to vast space of transformation parameters, estimation errors are significant in these approaches. To tackle above issues, a novel learning-based framework for 2D/3D registration is proposed, consisting of CNN regression and centroid alignment. The former introduces a residual regression network (Res-RegNet) to preliminarily estimate transformation parameters. To further reduce estimation errors, the latter utilizes target vessel centroids to refine projected images. The proposed framework is individually trained and evaluated on three patients, reaching mean Dice of 76.69%, 78.51%, and 85.39%, respectively, all outperforming baseline methods. Extensive ablation studies demonstrate centroid alignment can significantly improve registration performance. As a normalization layer in Res-RegNet, SPADE can modulate activations using binarized inputs through a spatially-adaptive, learned transformation. Semantic information of inputs is preserved to learn better representations for parameter estimation. Moreover, the inference rate of our framework is about 21 FPS combined with the state-of-the-art segmentation model, significantly surpassing real-time requirements (6$\sim$12 FPS) in clinical practice. These promising results indicate the potential of the framework to facilitate various 2D/3D registration tasks.Note to Practitioners—This paper was motivated by the problem of image-guided neurological interventions. Existing 2D/3D registration methods suffer from 1) long iteration times, which are difficult to meet real-time clinical necessities, or 2) significant parameter estimation errors, leading to poor registration accuracies. Therefore, this paper suggests a new registration framework, combining with CNN regression to give predictions of transformation parameters via a single forward propagation, and centroid alignment to reduce estimation errors by translation transformation. The framework is trained and tested on three patients separately and achieves state-of-the-art performance, demonstrating its superiority. Furthermore, the proposed framework is a learning-based method that is adaptable to various image modalities. Therefore, it has latent capacities to be integrated into surgical navigation systems.
De-Xing Huang, Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Zhen-Qiu Feng, Zeng-Guang Hou
IEEE Trans Autom. Sci. Eng.2
2025 Particle Restoration: A Novel Image Processing Framework for Improving Real Cryo-EM Image Quality in Single Particle Analysis
abstract
Cryo-electron microscopy single particle analysis (cryo-EM SPA) is the most powerful technique for biomacromolecule structure determination. However, many factors such as complicated noise and radiation damage make the quality of cryo-EM images extremely poor, where high-frequency structure details are submerged, limiting the application of deep learning and suppressing the resolution of reconstruction. Thus, image restoration is of vital importance. Some related works explore micrograph restoration, but the particles in restored micrographs are still of poor quality. Moreover, the training approach of existing methods uses noisy observations or simulated data as supervision, leading to reduced performance on real cryo-EM data. In this paper, we define the task of particle restoration and propose a novel 4-step framework to this end. Labels are created for each particle image and paired data is collected within our framework, compensating for the absence of ground truth. A deep neural network with encoder-decoder architecture is designed to learn the mapping from degraded particles to high-quality ones, while other networks can also be employed as a plug-and-play module. Three datasets are constructed from real cryo-EM data and extensive experiments are carried out. Both quantitative metrics and qualitative visualization indicate that our framework is effective for cryo-EM particle restoration. It becomes easier to extract particle features after restoration, aiding in SPA and the effective application of deep learning on cryo-EM images. The downstream task experiments of cryo-EM SPA are also conducted, showing that the proposed framework has the potential to improve cryo-EM SPA performance.
Bin Hu 0001, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Hong-Jia Li, Qing-Bing Zheng, Fa Zhang 0001, Zeng-Guang Hou, Ning-Shao Xia
IEEE Trans. Comput. Biol. Bioinform.5
2025 A Weight-Aware-Based Multisource Unsupervised Domain Adaptation Method for Human Motion Intention Recognition
abstract
Accurate recognition of human motion intention (HMI) is beneficial for exoskeleton robots to improve the wearing comfort level and achieve natural human-robot interaction. A classifier trained on labeled source subjects (domains) performs poorly on unlabeled target subject since the difference in individual motor characteristics. The unsupervised domain adaptation (UDA) method has become an effective way to this problem. However, the labeled data are collected from multiple source subjects that might be different not only from the target subject but also from each other. The current UDA methods for HMI recognition ignore the difference between each source subject, which reduces the classification accuracy. Therefore, this article considers the differences between source subjects and develops a novel theory and algorithm for UDA to recognize HMI, where the margin disparity discrepancy (MDD) is extended to multisource UDA theory and a novel weight-aware-based multisource UDA algorithm (WMDD) is proposed. The source domain weight, which can be adjusted adaptively by the MDD between each source subject and target subject, is incorporated into UDA to measure the differences between source subjects. The developed multisource UDA theory is theoretical and the generalization error on target subject is guaranteed. The theory can be transformed into an optimization problem for UDA, successfully bridging the gap between theory and algorithm. Moreover, a lightweight network is employed to guarantee the real-time of classification and the adversarial learning between feature generator and ensemble classifiers is utilized to further improve the generalization ability. The extensive experiments verify theoretical analysis and show that WMDD outperforms previous UDA methods on HMI recognition tasks.
Xiao-Yin Liu, Xiao-Hu Zhou, Zeng-Guang Hou
IEEE Trans. Cybern.3
2025 Learning Motor Cues in Brain-Muscle Modulation
abstract
Current studies for brain-muscle modulation often analyze selected properties in electrophysiological signals, leading to a partial understanding. This article proposes a cross-modal generative model that converts brain activities measured by electroencephalography (EEG) to corresponding muscular responses recorded by electromyography (EMG). Examining the generation process in the model highlights how the motor cue, representing implicit motor information hidden within brain activities, modulates the interaction between brain and muscle systems. The proposed model employs a two-stage generation process to bridge the semantic gap in cross-modal signals. Initially, the shared movement-related information between EEG and EMG signals is extracted using a contrastive learning framework. These shared representations act as conditional vectors in the subsequent EMG generation stage based on generative adversarial networks (GANs). Experiments on a self-collected multimodal electrophysiological signal data set show the algorithm's superiority over existing time series generative methods in cross-modal EMG generation. Further insights derived from the model's inference process underscore the brain's strategy for muscle control during movements. This research provides a data-driven approach for the neuroscience community, offering a comprehensive perspective of brain-muscular modulation.
Tian-Yu Xiang, Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Mei-Jiang Gui, Hao Li 0077, De-Xing Huang, Zeng-Guang Hou
IEEE Trans. Cybern.2
2025 Upper Limb Motor Sequence Analysis: From Isolated to Sequential
abstract
Motor skills are performed through sequential movements rather than isolated actions. Yet, decoding these sequences from biosignals poses a significant challenge. To address this gap, this study transitions motor decoding from classifying movements in isolated time windows to segmenting sequential movements. The proposed algorithm segments the electromyography (EMG) sequence in a coarse-to-fine manner. It begins with frame-level segmentation and locating the approximate boundaries at the movement-level. A region-growing-inspired fusion strategy is then designed to incorporate the coarse segmentation and localization results for the fined output. Experiments on a self-collected EMG dataset demonstrate impressive results in segmenting movements for participant-dependent/independent setups (accuracy:$94.2\hbox{\%}/74.7\hbox{\%}$; dice coefficient:$92.5\hbox{\%}/61.7\hbox{\%}$; mean Intersection over Union:$80.9\hbox{\%}/51.9\hbox{\%}$). Further analysis shows the algorithm's ability to capture the natural rhythm in participants' movement sequences. This research paves the way for a deep understanding of motor sequences, which benefits various applications, such as rehabilitation engineering.
Tian-Yu Xiang, Xiao-Hu Zhou, Mei-Jiang Gui, Xiaoliang Xie, Shiqi Liu 0004, Hao Li 0077, De-Xing Huang, Jiaxing Wang 0001, Yongqiang Tang, Jiamou Liu, Zeng-Guang Hou
IEEE Trans. Ind. Informatics2
2025 DOMAIN: Mildly Conservative Model-Based Offline Reinforcement Learning
abstract
Model-based reinforcement learning (RL), which learns an environment model from the offline dataset and generates more out-of-distribution model data, has become an effective approach to the problem of distribution shift in offline RL. Due to the gap between the learned and actual environment, conservatism should be incorporated into the algorithm to balance accurate offline data and imprecise model data. The conservatism of current algorithms mostly relies on model uncertainty estimation. However, uncertainty estimation is unreliable and leads to poor performance in certain scenarios, and the previous methods ignore differences between the model data, which brings great conservatism. To address the above issues, this article proposes a mildly conservative model-based offline RL algorithm (DOMAIN) without estimating model uncertainty, and designs the adaptive sampling distribution of model samples, which can adaptively adjust the model data penalty. In this article, we theoretically demonstrate that theQvalue learned by the DOMAIN outside the region is a lower bound of the trueQvalue, the DOMAIN is less conservative than previous model-based offline RL algorithms, and has the guarantee of safety policy improvement. The results of extensive experiments show that DOMAIN outperforms prior RL algorithms and the average performance has improved by 1.8% on the D4RL benchmark.
Xiao-Yin Liu, Xiao-Hu Zhou, Mei-Jiang Gui, Xiaoliang Xie, Shiqi Liu 0004, Shuangyi Wang, Qi-Chao Zhang, Biao Luo 0001, Zeng-Guang Hou
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Transformer-Based Long Time Series Forecasting with Decoupled Information Extraction and Information Complementarity
Jian-Long Hao, Sheng-Bin Duan, Yali Lv, Xiao-Hu Zhou
ICONIP (3)4
2024 A Two-Stage Network for Enhanced Intracranial Artery 3D Segmentation in TOF-MRA Volume
Xiaoliang Xie, Xiao-Hu Zhou, Ji-Chang Luo, De-Lin Liu, Zeng-Guang Hou, Jia-Xing Wang
ICONIP (1)4
2024 MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator
Xiao-Yin Liu, Xiao-Hu Zhou, Hao Li 0077, Mei-Jiang Gui, Tian-Yu Xiang, De-Xing Huang, Zeng-Guang Hou
IJCAI2
2024 Cross-Modal Motor Representation Learning
abstract
Learning motor representations in brains presents a challenge due to the entanglement of motor-related and unrelated information within neural imaging data. This study introduces a cross-modal learning algorithm that utilizes electromyogram (EMG) muscle cues to refine the learning of electroencephalogram (EEG) motor representations. The algorithm begins with original EEG representations from a baseline motor classification model. Subsequently, EMG muscle cues are learned to decompose the original EEG representations into motor-related and unrelated components. The decomposition process is achieved by aligning the EMG representations more closely with motor-related components and less with unrelated ones. Experimental results on a self-collected multi-modal dataset show the proposed algorithm leads to a performance enhancement of approximately 4% across various algorithms compared with the original EEG representations in motor classification. This advancement demonstrates the algorithm’s effectiveness in isolating motor-related information from complex brain activities. The innovative use of muscle cues for EEG motor characteristic learning opens new possibilities for incorporating cross-modal learning in creating more accurate brain-computer interfaces.
Tian-Yu Xiang, Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Mei-Jiang Gui, Hao Li 0077, De-Xing Huang, Zeng-Guang Hou
IJCNN2
2024 3D Ultrasound Image Acquisition and Diagnostic Analysis of the Common Carotid Artery with a Portable Robotic Device
abstract
Ultrasound (US) imaging of the carotid artery (CA) is a non-invasive diagnostic tool widely used in the medical field to assess the condition of the carotid artery, thereby predicting the risk of cardiovascular and cerebrovascular diseases. However, implementing this method in primary healthcare can be challenging due to the requirement for professionally trained sonographers. With the adoption of US robotic devices, the probe pose can be acquired while scanning, offering the possibility for 3D reconstruction and providing analyses that are not dependent on operator experience. This article introduces a method to semi-automatically acquire serialized US images of the common carotid artery (CCA). The method involves a specially designed robotic device built with a 6-RSU parallel mechanism, which is controlled according to robot pose, force sensor data and synchronous US images. To validate the images acquired, a method is proposed to segment the intima-media of CCA and calculate the intima-media thickness (IMT), which is a key indicator for cerebrovascular events prediction. After that, we propose an algorithm to reconstruct CCA into 3D voxel data with patient movement and cardiac cycle compensated, and a longitudinal view US image of CCA can be resliced from the voxel. The methods are tested on human subjects and the results indicate that the system and workflow can provide both quantitative and qualitative information of CCA for further diagnosis.
Longyue Tan, Zhaokun Deng, Mingrui Hao, Xilong Hou 0001, Chen Chen 0036, Xiaolin Gu, Xiao-Hu Zhou, Zeng-Guang Hou, Shuangyi Wang
IROS8
2023 Effective Skill Learning on Vascular Robotic Systems: Combining Offline and Online Reinforcement Learning
Hao Li 0077, Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Mei-Jiang Gui, Tian-Yu Xiang, De-Xing Huang, Zeng-Guang Hou
ICONIP (15)2
2023 A DNN-Based Learning Framework for Continuous Movements Segmentation
Tian-Yu Xiang, Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Zhen-Qiu Feng, Mei-Jiang Gui, Hao Li 0077, Zeng-Guang Hou
ICONIP (3)2
2023 Feature-Fusion-Based Haze Recognition in Endoscopic Images
Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Zhen-Qiu Feng, Zeng-Guang Hou
ICONIP (12)2
2023 An Effective Morphological Analysis Framework of Intracranial Artery in 3D Digital Subtraction Angiography
Haining Zhao 0002, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Liqun Jiao, Jichang Luo, Jia Dong, Bairu Zhang
ICONIP (10)5
2023 Towards Flexible and Universal: A Novel Endpoint-based Framework for Vessel Structural Information Extraction
abstract
In computer-assisted intravascular interventional surgery, extracting detailed information of target vessels from X-ray angiographic images can be meaningful in improving safety and effectiveness. However, large amounts of effort have been dedicated to segmenting the whole blood vessels from the background while ignoring the internal structure, which is limited in clinical application. In this paper, we propose a flexible and universal endpoint-based framework for vessel structural information extraction. The framework first localizes all the endpoints of target vessel segments through a Coarse-to-Fine Keypoint Detection Network (CFKD-Net), in which the designed Multi-branch Feature Aggregation (MFA) module captures both in-patch and cross-patch information to help recognize the points of interest based on global structure. A novel MaskMSELoss is also proposed to disambiguate those irrelevant responses. Then a designed VEssel Segmentation and Analysis (VESA) algorithm will generate the segmentation mask and morphological analysis for each vessel segment simply based on the endpoints. It can also be flexibly applied to analyze variant blood vessels which are not pre-defined before. Extensive experiments on two different coronary artery datasets consistently demonstrate that this framework can achieve state-of-the-art detection performance and successfully extract and analyze target vessel segments. Since the framework shows excellent performance on the coronary arteries with severe deformation and strong noise, it is highly promising for analyzing other vascular images.
Xiyao Ma, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Xinkai Qu, Wenzheng Han, Ming Wang 0001, Lin-Sen Zhang
ACM Multimedia4
2023 Learning Shared Semantic Information from Multimodal Bio-signals for Brain-Muscle Modulation Analysis
abstract
This paper presents a novel learning-based algorithm to investigate the high-level shared semantic information between electroencephalography (EEG) and electromyography (EMG) signals, for understanding brain-muscle modulation during movement execution. The proposed algorithm incorporates a spatial encoder that condenses spatial information obtained from EEG/EMG signals into unified temporal tokens using a learnable correlation matrix. These tokens are then encoded and decoded via a siamese temporal encoder and classification head to extract joint semantic information presented in cross-modal signals. Additionally, an analysis pipeline is designed to examine brain-muscle modulation based on the proposed algorithm. Experimental results from a self-collected multimodal bio-signals dataset validate the efficacy of the proposed algorithm in extracting and analyzing high-level latent semantic information shared in EEG and EMG signals, outperforming the state-of-the-art model by 5.35% in accuracy, 4.69% in precision, and 8.65% in recall. Notably, the designed analysis pipeline can also reveal low-level relationships, such as those related to time and space, between multimodal bio-signals. This research provides neuroscientists with a valuable tool for obtaining enhanced insights into brain-muscle modulation.
Tian-Yu Xiang, Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Hong-Jun Yang, Zhen-Qiu Feng, Mei-Jiang Gui, Hao Li 0077, De-Xing Huang, Zeng-Guang Hou
ACM Multimedia2
2023 A Novel Spatial Position Prediction Navigation System Makes Surgery More Accurate
abstract
During intravascular interventional surgery, the 3D surgical navigation system can provide doctors with 3D spatial information of the vascular lumen, reducing the impact of missing dimension caused by digital subtraction angiography (DSA) guidance and further improving the success rate of surgeries. Nevertheless, this task often comes with the challenge of complex registration problems due to vessel deformation caused by respiratory motion and high requirements for the surgical environment because of the dependence on external electromagnetic sensors. This article proposes a novel 3D spatial predictive positioning navigation (SPPN) technique to predict the real-time tip position of surgical instruments. In the first stage, we propose a trajectory prediction algorithm integrated with instrumental morphological constraints to generate the initial trajectory. Then, a novel hybrid physical model is designed to estimate the trajectory's energy and mechanics. In the second stage, a point cloud clustering algorithm applies multi-information fusion to generate the maximum probability endpoint cloud. Then, an energy-weighted probability density function is introduced using statistical analysis to achieve the prediction of the 3D spatial location of instrument endpoints. Extensive experiments are conducted on 3D-printed human artery and vein models based on a high-precision electromagnetic tracking system. Experimental results demonstrate the outstanding performance of our method, reaching 98.2% of the achievement ratio and less than 3 mm of the average positioning accuracy. This work is the first 3D surgical navigation algorithm that entirely relies on vascular interventional robot sensors, effectively improving the accuracy of interventional surgery and making it more accessible for primary surgeons.
Lin-Sen Zhang, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Chao-Nan Wang, Xinkai Qu, Wenzheng Han, Xiyao Ma
IEEE Trans. Medical Imaging4
2023 Learning Skill Characteristics From Manipulations
abstract
Percutaneous coronary intervention (PCI) has increasingly become the main treatment for coronary artery disease. The procedure requires high experienced skills and dexterous manipulations. However, there are few techniques to model PCI skill so far. In this study, a learning framework with local and ensemble learning is proposed to learn skill characteristics of different skill-level subjects from their PCI manipulations. Ten interventional cardiologists (four experts and six novices) were recruited to deliver a medical guidewire to two target arteries on a porcine model for in vivo studies. Simultaneously, translation and twist manipulations of thumb, forefinger, and wrist are acquired with electromagnetic (EM) and fiber-optic bend (FOB) sensors, respectively. These behavior data are then processed with wavelet packet decomposition (WPD) under 1-10 levels for feature extraction. The feature vectors are further fed into three candidate individual classifiers in the local learning layer. Furthermore, the local learning results from different manipulation behaviors are fused in the ensemble learning layer with three rule-based ensemble learning algorithms. In subject-dependent skill characteristics learning, the ensemble learning can achieve 100% accuracy, significantly outperforming the best local result (90%). Furthermore, ensemble learning can also maintain 73% accuracy in subject-independent schemes. These promising results demonstrate the great potential of the proposed method to facilitate skill learning in surgical robotics and skill assessment in clinical practice.
Xiao-Hu Zhou, Xiaoliang Xie, Shiqi Liu 0004, Zhen-Liang Ni, Yan-Jie Zhou, Rui-Qi Li, Mei-Jiang Gui, Chen-Chen Fan, Zhen-Qiu Feng, Guibin Bian, Zeng-Guang Hou
IEEE Trans. Neural Networks Learn. Syst.1
2022 Towards Automated Segmentation of Human Abdominal Aorta and Its Branches Using a Hybrid Feature Extraction Module with LSTM
Bo Zhang 0104, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Xiyao Ma, Lin-Sen Zhang
ICONIP (7)4
2022 A Dual-Stream Architecture for Real-Time Morphological Analysis of Aneurysm in Robot-Assisted Minimally Invasive Surgery
abstract
Real-time and precise morphological analysis of intraoperative AAA is a significant pre-imperative for robot-assisted minimally invasive surgery (RMIS). However, this task is frequently accompanied by the difficulties of ambiguous boundaries and obscured surfaces of aneurysms. To remedy these problems, we propose a Light-Weight Dual-Stream Boundary-Aware Network (DSB-Net) and a novel diagnosis algorithm for real-time morphological analysis of AAA. In the network, the features at the boundaries are preserved by incorporating a boundary localization stream, while the interior segmentation accuracy is guaranteed with a mask prediction stream. Moreover, the diagnosis algorithm is developed to measure the exact size of AAA. Quantitative and qualitative assessments on two different types of datasets illustrate that (1) The presented DSB-Net remarkably outperforms the other previously proposed medical networks with the inference rate of 10.8 FPS, which meets the real-time clinical necessities. (2) The developed algorithm provides accurate size measurements for AAA, which indicates the proposed approach can be integrated into the robotic navigation framework for RMIS.
Yan-Jie Zhou, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Rui-Qi Li, Zhen-Liang Ni, Chen-Chen Fan
ICRA4
2022 DSP-Net: Deeply-Supervised Pseudo-Siamese Network for Dynamic Angiographic Image Matching
Xiyao Ma, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Yan-Jie Zhou, Lin-Sen Zhang, Chao-Nan Wang
MICCAI (8)4
2022 A Novel Fusion Network for Morphological Analysis of Common Iliac Artery
Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Yan-Jie Zhou, Xiyao Ma
MICCAI (8)4
2022 SurgiNet: Pyramid Attention Aggregation and Class-wise Self-Distillation for Surgical Instrument Segmentation
Zhen-Liang Ni, Xiao-Hu Zhou, Guan'an Wang, Wen-Qian Yue, Zhen Li 0049, Guibin Bian, Zeng-Guang Hou
Medical Image Anal.2
2022 A Multilayer and Multimodal-Fusion Architecture for Simultaneous Recognition of Endovascular Manipulations and Assessment of Technical Skills
abstract
The clinical success of the percutaneous coronary intervention (PCI) is highly dependent on endovascular manipulation skills and dexterous manipulation strategies of interventionalists. However, the analysis of endovascular manipulations and related discussion for technical skill assessment are limited. In this study, a multilayer and multimodal-fusion architecture is proposed to recognize six typical endovascular manipulations. The synchronously acquired multimodal motion signals from ten subjects are used as the inputs of the architecture independently. Six classification-based and two rule-based fusion algorithms are evaluated for performance comparisons. The recognition metrics under the determined architecture are further used to assess technical skills. The experimental results indicate that the proposed architecture can achieve the overall accuracy of 96.41%, much higher than that of a single-layer recognition architecture (92.85%). In addition, the multimodal fusion brings significant performance improvement in comparison with single-modal schemes. Furthermore, the K -means-based skill assessment can obtain an accuracy of 95% to cluster the attempts made by different skill-level groups. These hopeful results indicate the great possibility of the architecture to facilitate clinical skill assessment and skill learning.
Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Qiu Feng, Zeng-Guang Hou, Guibin Bian, Rui-Qi Li, Zhen-Liang Ni, Shiqi Liu 0004, Yan-Jie Zhou
IEEE Trans. Cybern.1
2022 Space Squeeze Reasoning and Low-Rank Bilinear Feature Fusion for Surgical Image Segmentation
abstract
Surgical image segmentation is critical for surgical robot control and computer-assisted surgery. In the surgical scene, the local features of objects are highly similar, and the illumination interference is strong, which makes surgical image segmentation challenging. To address the above issues, a bilinear squeeze reasoning network is proposed for surgical image segmentation. In it, the space squeeze reasoning module is proposed, which adopts height pooling and width pooling to squeeze global contexts in the vertical and horizontal directions, respectively. The similarity between each horizontal position and each vertical position is calculated to encode long-range semantic dependencies and establish the affinity matrix. The feature maps are also squeezed from both the vertical and horizontal directions to model channel relations. Guided by channel relations, the affinity matrix is expanded to the same size as the input features. It captures long-range semantic dependencies from different directions, helping address the local similarity issue. Besides, a low-rank bilinear fusion module is proposed to enhance the model's ability to recognize similar features. This module is based on the low-rank bilinear model to capture the inter-layer feature relations. It integrates the location details from low-level features and semantic information from high-level features. Various semantics can be represented more accurately, which effectively improves feature representation. The proposed network achieves state-of-the-art performance on cataract image segmentation dataset CataSeg and robotic image segmentation dataset EndoVis 2018.
Zhen-Liang Ni, Guibin Bian, Zhen Li 0049, Xiao-Hu Zhou, Rui-Qi Li, Zeng-Guang Hou
IEEE J. Biomed. Health Informatics4
2022 TR-GAN: Multi-Session Future MRI Prediction With Temporal Recurrent Generative Adversarial Network
abstract
Magnetic Resonance Imaging (MRI) has been proven to be an efficient way to diagnose Alzheimer's disease (AD). Recent dramatic progress on deep learning greatly promotes the MRI analysis based on data-driven CNN methods using a large-scale longitudinal MRI dataset. However, most of the existing MRI datasets are fragmented due to unexpected quits of volunteers. To tackle this problem, we propose a novel Temporal Recurrent Generative Adversarial Network (TR-GAN) to complete missing sessions of MRI datasets. Unlike existing GAN-based methods, which either fail to generate future sessions or only generate fixed-length sessions, TR-GAN takes all past sessions to recurrently and smoothly generate future ones with variant length. Specifically, TR-GAN adopts recurrent connection to deal with variant input sequence length and flexibly generate future variant sessions. Besides, we also design a multiple scale & location (MSL) module and a SWAP module to encourage the model to better focus on detailed information, which helps to generate high-quality MRI data. Compared with other popular GAN architectures, TR-GAN achieved the best performance in all evaluation metrics of two datasets. After expanding the Whole MRI dataset, the balanced accuracy of AD vs. cognitively normal (CN) vs. mild cognitive impairment (MCI) and stable MCI vs. progressive MCI classification can be increased by 3.61% and 4.00%, respectively.
Chen-Chen Fan, Hongjun Yang, Xiao-Hu Zhou, Zhen-Liang Ni, Guan'an Wang, Yan-Jie Zhou, Zeng-Guang Hou
IEEE Trans. Medical Imaging5
2021 A Real-Time Multi-Task Framework for Guidewire Segmentation and Endpoint Localization in Endovascular Interventions
abstract
Real-time guidewire segmentation and endpoint localization play a pivotal role in robot-assisted minimally invasive surgery, which is helpful to reduce radiation dose and procedure time. Nevertheless, the tasks often come with the challenge of limited computational resources. For this purpose, a real-time multi-task framework with two stages is developed. In the first stage, a Fast Attention-fused Network (FAD-Net) is proposed to obtain accurate guidewire segmentation masks. In the second stage, a lightweight localization network and a post-processing algorithm are designed to robustly predict the guidewire endpoint position. Quantitative and qualitative evaluations on intraoperative X-ray sequences from 30 patients demonstrate that the developed framework outperforms the previously-published results for the tasks, achieving state-of-the-art performance. Moreover, the inference rate of the developed framework is approximately 10.6 FPS, which meets the real-time requirement of X-ray fluoroscopy. These results indicate the proposed approach has the potential to be integrated into the robotic navigation framework for endovascular interventions, enabling robotic-assisted minimally invasive surgery.
Yan-Jie Zhou, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Guan'an Wang, Zeng-Guang Hou, Rui-Qi Li, Zhen-Liang Ni, Chen-Chen Fan
ICRA4
2021 Vessel Width Estimation via Convolutional Regression
Rui-Qi Li, Guibin Bian, Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Liang Ni, Yan-Jie Zhou, Yuhan Wang 0017, Zeng-Guang Hou
MICCAI (6)3
2021 Real-Time Multi-Guidewire Endpoint Localization in Fluoroscopy Images
abstract
The real-time localization of the guidewire endpoints is a stepping stone to computer-assisted percutaneous coronary intervention (PCI). However, methods for multi-guidewire endpoint localization in fluoroscopy images are still scarce. In this paper, we introduce a framework for real-time multi-guidewire endpoint localization in fluoroscopy images. The framework consists of two stages, first detecting all guidewire instances in the fluoroscopy image, and then locating the endpoints of each single guidewire instance. In the first stage, a YOLOv3 detector is used for guidewire detection, and a post-processing algorithm is proposed to refine the guidewire detection results. In the second stage, a Segmentation Attention-hourglass (SA-hourglass) network is proposed to predict the endpoint locations of each single guidewire instance. The SA-hourglass network can be generalized to the keypoint localization of other surgical instruments. In our experiments, the SA-hourglass network is applied not only on a guidewire dataset but also on a retinal microsurgery dataset, reaching the mean pixel error (MPE) of 2.20 pixels on the guidewire dataset and the MPE of 5.30 pixels on the retinal microsurgery dataset, both achieving the state-of-the-art localization results. Besides, the inference rate of our framework is at least 20FPS, which meets the real-time requirement of fluoroscopy images (6-12FPS).
Rui-Qi Li, Xiaoliang Xie, Xiao-Hu Zhou, Shiqi Liu 0004, Zhen-Liang Ni, Yan-Jie Zhou, Guibin Bian, Zeng-Guang Hou
IEEE Trans. Medical Imaging3
2020 Pyramid Attention Aggregation Network for Semantic Segmentation of Surgical Instruments
abstract
Semantic segmentation of surgical instruments plays a critical role in computer-assisted surgery. However, specular reflection and scale variation of instruments are likely to occur in the surgical environment, undesirably altering visual features of instruments, such as color and shape. These issues make semantic segmentation of surgical instruments more challenging. In this paper, a novel network, Pyramid Attention Aggregation Network, is proposed to aggregate multi-scale attentive features for surgical instruments. It contains two critical modules: Double Attention Module and Pyramid Upsampling Module. Specifically, the Double Attention Module includes two attention blocks (i.e., position attention block and channel attention block), which model semantic dependencies between positions and channels by capturing joint semantic information and global contexts, respectively. The attentive features generated by the Double Attention Module can distinguish target regions, contributing to solving the specular reflection issue. Moreover, the Pyramid Upsampling Module extracts local details and global contexts by aggregating multi-scale attentive features. It learns the shape and size features of surgical instruments in different receptive fields and thus addresses the scale variation issue. The proposed network achieves state-of-the-art performance on various datasets. It achieves a new record of 97.10% mean IOU on Cata7. Besides, it comes first in the MICCAI EndoVis Challenge 2017 with 9.90% increase on mean IOU.
Zhen-Liang Ni, Guibin Bian, Guan'an Wang, Xiao-Hu Zhou, Zeng-Guang Hou, Hua-Bin Chen, Xiaoliang Xie
AAAI4
2020 CAU-net: A Novel Convolutional Neural Network for Coronary Artery Segmentation in Digital Substraction Angiography
Rui-Qi Li, Guibin Bian, Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Liang Ni, Zeng-Guang Hou
ICONIP (1)3
2020 A Novel Vascular Robotic System: Performance Evaluation
Si-Yi Wei, Xiao-Bo Sun, Xiao-Hu Zhou, Zeng-Guang Hou
ICONIP (2)3
2020 Attention-Guided Lightweight Network for Real-Time Segmentation of Robotic Surgical Instruments
abstract
The real-time segmentation of surgical instruments plays a crucial role in robot-assisted surgery. However, it is still a challenging task to implement deep learning models to do real-time segmentation for surgical instruments due to their high computational costs and slow inference speed. In this paper, we propose an attention-guided lightweight network (LWANet), which can segment surgical instruments in real-time. LWANet adopts encoder-decoder architecture, where the encoder is the lightweight network MobileNetV2, and the decoder consists of depthwise separable convolution, attention fusion block, and transposed convolution. Depthwise separable convolution is used as the basic unit to construct the decoder, which can reduce the model size and computational costs. Attention fusion block captures global contexts and encodes semantic dependencies between channels to emphasize target regions, contributing to locating the surgical instrument. Transposed convolution is performed to upsample feature maps for acquiring refined edges. LWANet can segment surgical instruments in real-time while takes little computational costs. Based on 960x544 inputs, its inference speed can reach 39 fps with only 3.39 GFLOPs. Also, it has a small model size and the number of parameters is only 2.06 M. The proposed network is evaluated on two datasets. It achieves state-of-the- art performance 94.10% mean IOU on Cata7 and obtains a new record on EndoVis 2017 with a 4.10% increase on mean IOU.
Zhen-Liang Ni, Guibin Bian, Zeng-Guang Hou, Xiao-Hu Zhou, Xiaoliang Xie, Zhen Li 0049
ICRA4
2020 A Multilayer-Multimodal Fusion Architecture for Pattern Recognition of Natural Manipulations in Percutaneous Coronary Interventions
abstract
The increasingly-used robotic systems can provide precise delivery and reduce X-ray radiation to medical staff in percutaneous coronary interventions (PCI), but natural manipulations of interventionalists are forgone in most robot-assisted procedures. Therefore, it is necessary to explore natural manipulations to design more advanced human-robot interfaces (HRI). In this study, a multilayer-multimodal fusion architecture is proposed to recognize six typical subpatterns of guidewire manipulations in conventional PCI. The synchronously acquired multimodal behaviors from ten subjects are used as the inputs of the fusion architecture. Six classification-based and two rule-based fusion algorithms are evaluated for performance comparisons. Experimental results indicate that the multimodal fusion brings significant accuracy improvement in comparison with single-modal schemes. Furthermore, the proposed architecture can achieve the overall accuracy of 96.90%, much higher than that of a singlelayer recognition architecture (92.56%). These results have indicated the potential of the proposed method for facilitating the development of HRI for robot-assisted PCI.
Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Qiu Feng, Zeng-Guang Hou, Guibin Bian, Rui-Qi Li, Zhen-Liang Ni, Shiqi Liu 0004, Yan-Jie Zhou
ICRA1
2020 BARNet: Bilinear Attention Network with Adaptive Receptive Fields for Surgical Instrument Segmentation
abstract
Surgical instrument segmentation is crucial for computer-assisted surgery. Different from common object segmentation, it is more challenging due to the large illumination variation and scale variation in the surgical scenes. In this paper, we propose a bilinear attention network with adaptive receptive fields to address these two issues. To deal with the illumination variation, the bilinear attention module models global contexts and semantic dependencies between pixels by capturing second-order statistics. With them, semantic features in challenging areas can be inferred from their neighbors, and the distinction of various semantics can be boosted. To adapt to the scale variation, our adaptive receptive field module aggregates multi-scale features and selects receptive fields adaptively. Specifically, it models the semantic relationships between channels to choose feature maps with appropriate scales, changing the receptive field of subsequent convolutions. The proposed network achieves the best performance 97.47% mean IoU on Cata7. It also takes the first place on EndoVis 2017, exceeding the second place by 10.10% mean IoU.
Zhen-Liang Ni, Guibin Bian, Guan'an Wang, Xiao-Hu Zhou, Zeng-Guang Hou, Xiaoliang Xie, Zhen Li 0049, Yuhan Wang 0017
IJCAI4
2020 Lightweight Double Attention-Fused Networks for Intraoperative Stent Segmentation
Yan-Jie Zhou, Xiaoliang Xie, Zeng-Guang Hou, Xiao-Hu Zhou, Guibin Bian, Shiqi Liu 0004
MICCAI (6)4
2020 An Interventionalist-Behavior-Based Data Fusion Framework for Guidewire Tracking in Percutaneous Coronary Intervention
abstract
Guidewire tracking is a clinical challenge in percutaneous coronary intervention (PCI). The current practice of image-based and sensor-based tracking techniques is still limited by radiation exposure, contrast injection, device sterilization, and procedure safety. In this paper, an interventionalist-behavior-based data fusion framework is developed to provide a novel strategy for tracking guidewire motions in PCI. Four types of natural behavior were acquired from ten interventionalists while performing guidewire translation and rotation based on a simulation platform. Different numbers of behaviors are fused by a hierarchical framework with six local tracking models and three ensemble algorithms. After Gaussian mixture regression-based ensemble fusion, a three-behavior scheme can achieve average tracking errors of 1.07 ± 0.17 mm for guidewire translation, and 20.05 ± 3.36° for guidewire rotation. Relevant statistical analysis further reveals that this scheme outperforms the cases using fewer behaviors, and ensemble fusion brings significant error reduction compared with only local fusion. These meaningful results indicate the great potential of the proposed framework for promoting the improvement of guidewire tracking in PCI.
Xiao-Hu Zhou, Guibin Bian, Xiaoliang Xie, Zeng-Guang Hou
IEEE Trans. Syst. Man Cybern. Syst.1
2019 RAUNet: Residual Attention U-Net for Semantic Segmentation of Cataract Surgical Instruments
Zhen-Liang Ni, Guibin Bian, Xiao-Hu Zhou, Zeng-Guang Hou, Xiaoliang Xie, Chen Wang 0122, Yan-Jie Zhou, Rui-Qi Li, Zhen Li 0049
ICONIP (2)3
2019 Real-Time Guidewire Segmentation and Tracking in Endovascular Aneurysm Repair
Yan-Jie Zhou, Xiaoliang Xie, Guibin Bian, Zeng-Guang Hou, Zhi-Chao Lai, Xinkai Qu, Shiqi Liu 0004, Xiao-Hu Zhou
ICONIP (1)9
2019 Fully Automatic Dual-Guidewire Segmentation for Coronary Bifurcation Lesion
abstract
Interventional therapy for coronary bifurcation lesion has always been an intractable problem in percutaneous coronary intervention (PCI). Dual-guidewire detection can greatly assist physicians in interventional therapy of bifurcated lesions. Nevertheless, this task often comes with the challenges of X-ray images with low signal noise ratio (SNR) as well as the thinner structure of the guidewire compared to other interventional tools. In this paper, a fully automatic detection method based on an improved U-Net and the modified focal loss is proposed for dual-guidewire segmentation in 2D X-ray fluoroscopy, which accomplishes accurate and robust segmentation. The main contributions of this paper are twofold: (1) the proposed method not only addresses the extreme foreground-background class imbalance generated by the slender guidewire structure, but also solve the problem of misclassified examples caused by the guidewire-like structures and contrast agents; (2) the running speed is about 8 frames per second, which reaches near-real-time processing speed. Furthermore, data augmentation algorithm and transfer learning are used to further improve the performance. The proposed method was verified on clinical 2D X-ray image sequences of 30 patients, in which F1-score reached 0.932. The experiment results indicated that our approach is promising for assisting bifurcation lesion surgery.
Yan-Jie Zhou, Xiaoliang Xie, Guibin Bian, Zeng-Guang Hou, Yu-Dong Wu, Shiqi Liu 0004, Xiao-Hu Zhou, Jiaxing Wang 0001
IJCNN7
2019 A Two-Stage Framework for Real-Time Guidewire Endpoint Localization
Rui-Qi Li, Guibin Bian, Xiao-Hu Zhou, Xiaoliang Xie, Zhen-Liang Ni, Zeng-Guang Hou
MICCAI (5)3
2017 Prediction of natural guidewire rotation using an sEMG-based NARX neural network
abstract
For the treatment of cardiovascular diseases, clinical success of percutaneous coronary intervention is highly dependent on natural technical skills and dexterous manipulation strategies of surgeons. However, the increasing used robotic surgical systems have been designed without considering manipulation techniques, especially surgical behaviors and motion patterns. This has driven research towards exploitation of natural manipulation skills in recent years. In this paper, natural guidewire manipulations are analyzed and predicted using an sEMG-based nonlinear autoregressive neural network with exogenous inputs. The relationship between natural endovascular manipulation and guidewire rotation is built through the network. Two experiments at different rotational speed were performed to verify the effectiveness and robustness of the applied model. The experimental results show that the average predictive root mean error of five subjects is 15.61° at the low speed and 21.85° at the high speed. These favorable results could be of interest to improve existing robotic surgical systems.
Xiao-Hu Zhou, Guibin Bian, Xiaoliang Xie, Zeng-Guang Hou, Jian-Long Hao
IJCNN1