Xingru Huang

dblp:266/7207 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
25since 2021 · last 2027
0000-0003-3971-8434ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 SVR-UNet: Frequency-aware view-routed analysis-synthesis sampling for 3D medical image segmentation
Shuanghua Ye, Wenwen Tang, Huiyu Zhou 0001, Jin Liu 0025, Xiaoshuai Zhang, Xingru Huang
Expert Syst. Appl.8
2026 Wavefront-Constrained Passive Obscured Object Detection
abstract
Accurately localizing and segmenting obscured objects from faint light patterns beyond the field of view is highly challenging due to multiple scattering and medium-induced perturbations. Most existing methods, based on real-valued modeling or local convolutional operations, are inadequate for capturing the underlying physics of coherent light propagation. Moreover, under low signal-to-noise conditions, these methods often converge to non-physical solutions, severely compromising the stability and reliability of the observation. To address these challenges, we propose a novel physics-driven Wavefront Propagating Compensation Network (WavePCNet) to simulate wavefront propagation and enhance the perception of obscured objects. This WavePCNet integrates the Tri-Phase Wavefront Complex-Propagation Reprojection (TriWCP) to incorporate complex amplitude transfer operators to precisely constrain coherent propagation behavior, along with a momentum memory mechanism to effectively suppress the accumulation of perturbations. Additionally, a High-frequency Cross-layer Compensation Enhancement is introduced to construct frequency-selective pathways with multi-scale receptive fields and dynamically models structural consistency across layers, further boosting the model’s robustness and interpretability under complex environmental conditions. Extensive experiments conducted on four physically collected datasets demonstrate that WavePCNet consistently outperforms state-of-the-art methods across both accuracy and robustness.
Yiwei Ouyang, Xiaoshuai Zhang, Huiyu Zhou 0001, Wenwen Tang, Shaowei Jiang, Jin Liu 0025, Xingru Huang
AAAI10
2026 HMareN: Hierarchical Malicious Attack Representation Embedding Network for Web Attack Detection
Yiwen Qin, Xiaoshuai Zhang, Haipeng Qu, Wenwen Tang, Jin Liu 0025, Zhiju Yang, Xingru Huang
ICC8
2026 Electroencephalographic biomarker-guided early detection of Alzheimer's disease via cortically subdivided neurodynamic PINN
Zhengliang Zhang, Yachen Wei, Xin Rao, Liyang Yu, Ruixue Li, Xiaoshuai Zhang, Xingru Huang
Expert Syst. Appl.9
2026 P3R: Polymodal palpebral progressive refinement via symmetry aware latent diffusion for precision guided prediction of postoperative blepharoptosis morphology
Shuaixuan Zhou, Xingru Huang, Zhaoyang Xu, Huiyu Zhou 0001, Guangyuan Zhang, Wenwen Tang, Wenbin Zhang 0002, Jin Liu 0025, Lixia Lou, Xiaoshuai Zhang
Expert Syst. Appl.2
2026 Lightweight multi-scale weight pruning network for salient object detection
abstract
Salient object detection (SOD) is fundamental to computer vision, yet deep learning approaches often suffer from high computational costs, limiting deployment on resource-constrained devices. We propose a Lightweight Multi-scale Weight Pruning Network (LMWP-Net) to balance high performance with low complexity. LMWP-Net employs an encoder–decoder architecture featuring two key components: a Multi-scale Weight Pruning Module (MWPM) for efficient feature extraction and redundancy reduction, and a Multi-scale Attention Fusion Module (MAFM) for effective integration via attention mechanisms. Extensive experiments on public datasets demonstrate that LMWP-Net consistently outperforms existing lightweight methods and achieves competitive accuracy against state-of-the-art models. Remarkably, compared to the prominent BANet, LMWP-Net achieves a 94.6% reduction in parameters and a 99.5% reduction in FLOPs, validating its superior efficiency and effectiveness for real-time applications. The implemented code is publicly available at https://github.com/IMOP-lab/LMWP-Net .
Xichun Sheng, Yaoqi Sun, Gaopeng Huang, Ya-Hong Chen, Jin Liu 0025, Xiaoshuai Zhang, Xingru Huang
J. Vis. Commun. Image Represent.11
2026 TriFTM-Net: Tri-Path Fourier-Temporal Modulation Network for macular edema pathology segmentation and reconstruction in high-precision intraoperative navigation
abstract
Ophthalmic diseases such significantly impair the vision of numerous individuals globally. Accurate and real-time 3D reconstruction of macular edema and retinal tears is crucial for improving surgical efficiency and success rates. However, lesion areas often exhibit considerable noise and high heterogeneity, and the imaging devices employed may introduce electronic noise and artifacts. Current 2D medical image segmentation techniques fail to achieve optimal outcomes. To overcome these challenges, we propose the Tri-Path Fourier-Temporal Modulation Network (TriFTM-Net). TriFTM-Net synergistically integrates spatial, frequency, and spatiotemporal features. This design effectively augments both feature representation and extraction. TriFTM-Net comprises three critical modules: the Tri-Path Spectral Hierarchical Encoder (TPSHE), which amplifies feature representation by integrating tri-path features; the Feature Re-Modulation (FRM), which reduces noise interference and enhances feature extraction; and the Hierarchical Feature Reconstruction Module (HFRM), which improves detail preservation in upsampled images. Comparative analysis with thirteen baseline methods demonstrates that our approach achieves the highest Dice scores, IoU, and Kappa coefficient on the OIMHS dataset.Our code is publicly available at https://github.com/IMOP-lab/TriFTM-Net.
Xingru Huang, Shuaibin Chen, Gaopeng Huang, Zhaoyang Xu, Wenbin Zhang 0002, Jian Huang 0015, Jin Liu 0025, Xiaoshuai Zhang, Shaowei Jiang, Huiyu Zhou 0001, Yaoqi Sun
Neural Networks1
2026 MICCAI 2023 STS Challenge: A retrospective study of semi-supervised approaches for teeth segmentation
abstract
Computer-aided diagnosis greatly enhances personalized treatment planning and diagnostic efficiency by providing accurate dental anatomy through teeth segmentation. However, it still constrained by the scarcity of high-quality annotated dental datasets. To address this issue, this paper presents a dataset combining both 2D panoramic X-rays with over 6,500 images and 3D CBCT with over 580 volumes (88,500+ slices) to support the Semi-supervised Teeth Segmentation (STS) Challenge, which includes partially meticulous annotations and covers all age groups. Moreover, multi-phase semi-supervised teeth segmentation algorithms and high-confidence pseudo-labels refinement strategies were proposed by competitors during this challenge. Algorithms were verified on this proposed dataset and good segmentation performance were achieved, over 93+ and 80+ Dice score were obtained for top three 2D and 3D participants, demonstrating the high quality of this proposed dataset. This paper also summarizes the diverse methods employed by the top-ranking teams in the MICCAI 2023 STS Challenge. Our dataset is publicly accessible through Zenodo ( https://zenodo.org/records/10597292 ), and the participants’ code is hosted on GitHub ( https://github.com/ricoleehduu/STS-Challenge ).
Yaqi Wang 0002, Shuai Wang 0003, Dahong Qian, Hongyuan Zhang 0002, Ruilong Dan, Qianni Zhang, Xingru Huang, Jun Liu 0027, Zhean Ma, Weiwei Cui 0003, Shan Luo 0003, Chengkai Wang, Jiaxue Ni, Dongyun Liu, Zhouhao Lin, Chunshi Wang, Qiupu Chen, Mingqian Li, Huiyu Zhou 0001, Qun Jin
Pattern Recognit.11
2026 TriVLLo: Tri-View Dynamic Architecture and Unified Cross-Modal Representation for Efficient Fine-Grained Vision-Language Understanding
abstract
This study tackles computational bottlenecks, training instability, and insufficient cross-modal semantic alignment in high-resolution multimodal image processing. We propose TriVLLo, an innovative multi-scale vision-language modeling framework. Our main contributions are: First, we introduce factorized 2D positional encoding and a dynamically configurable modular architecture. This approach decouples height and width position information. It improves spatial localization reliability for images with extreme aspect ratios. It also reduces computational parameters and alleviates training instability. Second, we design a unified multi-scale feature extraction and modality interaction mechanism. This uses adaptive image processing and multi-perspective feature pyramids. It enhances robustness to inputs of any resolution. It also achieves fine-grained alignment of vision-language features through a shared embedding space. Third, we build a high-quality dataset with 11K samples for fine-grained reasoning. This dataset supports improvements in visual ranking, semantic alignment, and narrative reasoning. Experiments show that TriVLLo achieves 87.4% of GPT-4V's performance on the MM-Vet benchmark. It demonstrates a 93.0 percentage-point improvement over Emu2 in spatial cognition tasks. It attains 89.2% accuracy on knowledge generation tasks. These results significantly outperform state-of-the-art methods.
Liang Kou, Wenlong Fan, Xingru Huang, Bai Lin, Yun Lin 0005
IEEE Trans. Reliab.3
2025 Volumetric Axial Disentanglement Enabling Advancing in Medical Image Segmentation
abstract
Information retrieved from three dimensions is treated uniformly in CNN-based volumetric segmentation methods. However, such neglect of axial disparities fails to capture true spatio-temporal variations. This paper introduces the volumetric axial disentanglement to address the disparities in spatial information along different axial dimensions. Building on this concept, we propose the Post-Axial Refiner (PaR) module to refine segmentation masks by implementing axial disentanglement on the specific axis of the volumetric medical sequences. As a plug-and-play enhancement to existing volumetric segmentation architecture, PaR further utilizes specialized attention approaches to learn disentangled post-decoding features, enhancing spatial representation and structural detail. Validation on various datasets demonstrates PaR's consistent elevation of segmentation precision and boundary clarity across 11 baselines and different imaging modalities, achieving state-of-the-art performance on multiple datasets. Experimental tests demonstrate the ability of volumetric axial disentanglement to refine the segmentation of volumetric medical images. Code is released at https://github.com/IMOP-lab/PaR-Pytorch.
Xingru Huang, Jian Huang 0015, Tianyun Zhang, Yaqi Wang 0002, Ruipu Tang, Shaowei Jiang, Jin Liu 0025, Renjie Ruan, Xiaoshuai Zhang
IJCAI1
2025 DEFN: Dual-Encoder Fourier Group Harmonics Network for three-dimensional indistinct-boundary object segmentation
Xiaohua Jiang, Jian Huang 0015, Meiyi Luo, Zhaoyang Xu, Qianni Zhang, Xingru Huang, Shaowei Jiang, Mang Xiao
Expert Syst. Appl.8
2025 LiGu-LVM: Linguistic-Guided Generative Large Vision Model for IoMT Clinical Ocular Disease Screening via Morphology Dissection
abstract
The early detection of ocular disorders, including Graves’ disease, myasthenia gravis, conjunctival hyperemia, conjunctivitis, and keratitis, which critically impair the vision of millions worldwide, necessitates large-scale screening predicated on ocular appearance measurements as a crucial diagnostic component. The emerging Internet of Medical Things (IoMT) introduces new avenues for local clinics to embrace portable and extensive diagnostics. However, the inherent heterogeneity and blurriness of ocular images, compounded by environmental noise, and the computational resource constraint hinder the high-precision diagnostics on IoMT devices. In response to these challenges, a linguistic-guided generative large vision model (LiGu-LVM) has been formulated to assist and enhance the diagnostic capability of IoMT-enabled ocular scanners, integrating a dynamically allocated high-speed quantization system (DAHSQS), a linguistic-guided generative local-isolation module (LiGu), an oculo visio transformatrix segmentum-analytica modulorum (OVT-SAM), and a multiscale recursive attention segmentation engine (MuRASE). DAHSQS enables the flexible aggregation and transmission of patient imagery to shift heavy diagnostic tasks from IoMT-enabled mobile ocular scanners to computational clusters, facilitating rapid facial measurements and preliminary screening via dynamic task allocation and scalable server clusters. The LiGu module employs natural language guidance to generate key image locations, using extensive prior knowledge embedded within linguistic models for precise semantic isolation. OVT-SAM synthesizes multilevel features from the large vision model, extracting intermediate characteristic information and addressing global features alongside deep semantic understanding in natural images collected from IoMT-enabled ocular scanners. MuRASE achieves high-fidelity segmentation of ocular images by incorporating contextual recursive attention mechanisms and skip connections with layer-wise reverse connectivity. Extensive experiments show proposed method surpassing 80% Intersection Over Union (IoU) in ocular semantic segmentation on the CelebA-HQ dataset, achieving an IoU of 82.9%, thus exceeding the performance of existing models by 4.9%.
Xingru Huang, Tianyun Zhang, Jian Huang 0015, Gaopeng Huang, Lou Zhao, Shaowei Jiang, Jin Liu 0025, Guan Gui 0001, Xiaoshuai Zhang
IEEE Internet Things J.1
2025 PricoMS: Prior-coordinated multiscale synthesis network for self-supervised-aided vessel segmentation in intravascular ultrasound image amidst label scarcity
Xingru Huang, Shuaibin Chen, Shaowei Jiang, Retesh Bajaj, Nathan Angelo Lecaros Yap, Murat Çap, Xiaoshuai Zhang, Xingwei He 0007, Anantharaman Ramasamy, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang
Knowl. Based Syst.1
2025 Multidimensional Directionality-Enhanced Segmentation via large vision model
Xingru Huang, Changpeng Yue, Jian Huang 0015, Zhengyao Jiang, Mingkuan Wang, Zhaoyang Xu, Guangyuan Zhang, Jin Liu 0025, Tianyun Zhang, Xiaoshuai Zhang, Shaowei Jiang, Yaoqi Sun
Medical Image Anal.1
2025 Deep Learning-based Pathological Image Analysis Algorithm for Early Diagnosis of Breast Cancer
abstract
This paper presents an innovative deep neural network algorithm designed to improve the automatic recognition accuracy of epithelial and mesenchymal tissues in breast cancer pathological images. The key innovation centers around the development of two novel algorithms: the Tissue Enhancement Vision Fusion technique and the Comprehensive Gray-Scale Threshold Intelligent Judgment method. These algorithms optimize the analysis and diagnosis process of breast cancer pathological images by introducing new approaches to feature enhancement and quantitative assessment that have not been explored in existing literature. The Tissue Enhancement Vision Fusion technique enhances the intuitive observation of key features in images through heatmap generation, while the Comprehensive Gray-Scale Threshold Intelligent Judgment method quantitatively assesses image characteristics, thereby enhancing the precision and efficiency of the model in recognizing epithelial and mesenchymal tissues. By strategically adjusting the EfficientNet model structure in conjunction with these two newly proposed algorithms, the proposed deep learning framework significantly improves the performance of breast cancer pathological image processing. The experimental results demonstrate that the enhanced model achieves recognition accuracies of 0.8548 and 0.8125 for epithelial and mesenchymal tissues, respectively, outperforming existing methods in both accuracy and efficiency. This breakthrough not only validates the effectiveness of the proposed algorithms but also provides robust technical support for early detection and precise diagnosis of breast cancer.
Xingru Huang, Haibing Yin
Neural Process. Lett.2
2025 Few-Shot Specific Emitter Identification: A Knowledge, Data, and Model-Driven Fusion Framework
abstract
In the Industrial Internet of Things (IIoT) context, ensuring secure communication is essential. Specific Emitter Identification (SEI), which leverages subtle differences in radio frequency signals to identify distinct emitters, is key to enhancing communication security. However, traditional SEI methods often rely on large labeled datasets and complex signal processing techniques, which limit their practical applicability due to data acquisition challenges and inefficiency. To address these limitations, we propose a novel Few-shot Specific Emitter Identification (FS-SEI) approach named KDM. This method fuses deep learning with multi-modal data processing, utilizing a hybrid neural network architecture that combines handcrafted features, self-supervised learning, and few-shot learning techniques. Our framework improves learning efficiency and accuracy, especially in data-scarce scenarios. We evaluate KDM using open-source Wi-Fi and ADS-B datasets, and the results demonstrate that our method consistently outperforms existing state-of-the-art few-shot SEI approaches. For example, on the ADS-B dataset, KDM boosts accuracy from 60.99% to 75.34% as the sample count increases from 5-shot to 10-shot, surpassing other methods by over 10%. Similarly, on the Wi-Fi dataset, KDM achieves an impressive 88.94% accuracy in low-sample (5-shot) scenarios. The codes are available athttps://github.com/tengmouren/KDM2SEI.
Minhong Sun, Jiazhong Teng, Wei Wang 0527, Xingru Huang
IEEE Trans. Inf. Forensics Secur.5
2024 Upping the Game: How 2D U-Net Skip Connections Flip 3D Segmentation
abstract
In the present study, we introduce an innovative structure for 3D medical image segmentation that effectively integrates 2D U-Net-derived skip connections into the architecture of 3D convolutional neural networks (3D CNNs). Conventional 3D segmentation techniques predominantly depend on isotropic 3D convolutions for the extraction of volumetric features, which frequently engenders inefficiencies due to the varying information density across the three orthogonal axes in medical imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI). This disparity leads to a decline in axial-slice plane feature extraction efficiency, with slice plane features being comparatively underutilized relative to features in the time-axial. To address this issue, we introduce the U-shaped Connection (uC), utilizing simplified 2D U-Net in place of standard skip connections to augment the extraction of the axial-slice plane features while concurrently preserving the volumetric context afforded by 3D convolutions. Based on uC, we further present uC 3DU-Net, an enhanced 3D U-Net backbone that integrates the uC approach to facilitate optimal axial-slice plane feature utilization. Through rigorous experimental validation on five publicly accessible datasets—FLARE2021, OIMHS, FeTA2021, AbdomenCT-1K, and BTCV, the proposed method surpasses contemporary state-of-the-art models. Notably, this performance is achieved while reducing the number of parameters and computational complexity. This investigation underscores the efficacy of incorporating 2D convolutions within the framework of 3D CNNs to overcome the intrinsic limitations of volumetric segmentation, thereby potentially expanding the frontiers of medical image analysis. Our implementation is available at https://github.com/IMOP-lab/U-Shaped-Connection.
Xingru Huang, Jian Huang 0015, Tianyun Zhang, Shaowei Jiang, Yaoqi Sun
NeurIPS1
2024 PGKD-Net: Prior-guided and Knowledge Diffusive Network for Choroid Segmentation
abstract
The thickness of the choroid is considered to be an important indicator of clinical diagnosis. Therefore, accurate choroid segmentation in retinal OCT images is crucial for monitoring various ophthalmic diseases. However, this is still challenging due to the blurry boundaries and interference from other lesions. To address these issues, we propose a novel prior-guided and knowledge diffusive network (PGKD-Net) to fully utilize retinal structural information to highlight choroidal region features and boost segmentation performance. Specifically, it is composed of two parts: a Prior-mask Guided Network (PG-Net) for coarse segmentation and a Knowledge Diffusive Network (KD-Net) for fine segmentation. In addition, we design two novel feature enhancement modules, Multi-Scale Context Aggregation (MSCA) and Multi-Level Feature Fusion (MLFF). The MSCA module captures the long-distance dependencies between features from different receptive fields and improves the model's ability to learn global context. The MLFF module integrates the cascaded context knowledge learned from PG-Net to benefit fine-level segmentation. Comprehensive experiments are conducted to evaluate the performance of the proposed PGKD-Net. Experimental results show that our proposed method achieves superior segmentation accuracy over other state-of-the-art methods. Our code is made up publicly available at: https://github.com/yzh-hdu/choroid-segmentation.
Yaqi Wang 0002, Zehua Yang, Xindi Liu, Dechao Chen, Gangyong Jia, Juan Ye, Xingru Huang
Artif. Intell. Medicine12
2024 Sketch-Supervised Histopathology Tumour Segmentation: Dual CNN-Transformer With Global Normalised CAM
abstract
Deep learning methods are frequently used in segmenting histopathology images with high-quality annotations nowadays. Compared with well-annotated data, coarse, scribbling-like labelling is more cost-effective and easier to obtain in clinical practice. The coarse annotations provide limited supervision, so employing them directly for segmentation network training remains challenging. We present a sketch-supervised method, called DCTGN-CAM, based on a dual CNN-Transformer network and a modified global normalised class activation map. By modelling global and local tumour features simultaneously, the dual CNN-Transformer network produces accurate patch-based tumour classification probabilities by training only on lightly annotated data. With the global normalised class activation map, more descriptive gradient-based representations of the histopathology images can be obtained, and inference of tumour segmentation can be performed with high accuracy. Additionally, we collect a private skin cancer dataset named BSS, which contains fine and coarse annotations for three types of cancer. To facilitate reproducible performance comparison, experts are also invited to label coarse annotations on the public liver cancer dataset PAIP2019. On the BSS dataset, our DCTGN-CAM segmentation outperforms the state-of-the-art methods and achieves 76.68 % IOU and 86.69 % Dice scores on the sketch-based tumour segmentation task. On the PAIP2019 dataset, our method achieves a Dice gain of 8.37 % compared with U-Net as the baseline network.
Yilong Li 0002, Linyan Wang, Xingru Huang, Yaqi Wang 0002, Ruiquan Ge, Huiyu Zhou 0001, Juan Ye, Qianni Zhang
IEEE J. Biomed. Health Informatics3
2024 SASAN: Spectrum-Axial Spatial Approach Networks for Medical Image Segmentation
abstract
Ophthalmic diseases such as central serous chorioretinopathy (CSC) significantly impair the vision of millions of people globally. Precise segmentation of choroid and macular edema is critical for diagnosing and treating these conditions. However, existing 3D medical image segmentation methods often fall short due to the heterogeneous nature and blurry features of these conditions, compounded by medical image clarity issues and noise interference arising from equipment and environmental limitations. To address these challenges, we propose the Spectrum Analysis Synergy Axial-Spatial Network (SASAN), an approach that innovatively integrates spectrum features using the Fast Fourier Transform (FFT). SASAN incorporates two key modules: the Frequency Integrated Neural Enhancer (FINE), which mitigates noise interference, and the Axial-Spatial Elementum Multiplier (ASEM), which enhances feature extraction. Additionally, we introduce the Self-Adaptive Multi-Aspect Loss (LSM), which balances image regions, distribution, and boundaries, adaptively updating weights during training. We compiled and meticulously annotated the Choroid and Macular Edema OCT Mega Dataset (CMED-18k), currently the world’s largest dataset of its kind. Comparative analysis against 13 baselines shows our method surpasses these benchmarks, achieving the highest Dice scores and lowest HD95 in the CMED and OIMHS datasets. Our code is publicly available at https://github.com/IMOP-lab/SASAN-Pytorch.
Xingru Huang, Jian Huang 0015, Tianyun Zhang, Changpeng Yue, Xuanbin Chen, Qianni Zhang, Ying Fu 0001, Yangyundou Wang
IEEE Trans. Medical Imaging1
2023 D-AE: A Discriminant Encode-Decode Nets for Data Generation
Gongju Wang, Yulun Song, Mingjian Ni, Quanda Wang, Xingru Huang
CollaborateCom (2)9
2023 GOMPS: Global Attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System
abstract
Accurate measurements of ophthalmic parameters and postoperative appearance prediction are essential for the diagnosis and treatment of many ophthalmic diseases. Nevertheless, it remains challenging due to (1) inconsistent ophthalmic image sampling standards, including ocular-camera distance, facial angle, and patient number, (2) complicated ocular morphology, such as subconjunctival hemorrhage, ocular movements, lighting effects, and morphological aging. It is difficult for a model to measure parameters and make predictions in variable sampling methods and morphology conditions. Therefore, the Global attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System (GOMPS) is proposed, which quantifies ophthalmic image parameters to diagnose disease and simultaneously predict postoperative appearance of blepharoptosis. By perceiving the global structure of the ophthalmic image, GOMPS makes logical inference predictions of the sclera and cornea morphology, to overcome the above difficulties. Concretely, a global attention unit (GAU) and a novel global attention structure-aware network (GASA-Net) are designed to enhance GOMPS’s global structure awareness ability to perform logical reasoning. Extensive experimental results on our collected ophthalmic dataset for diagnosis & prediction (OD2P) demonstrate that GOMPS surpasses the state-of-the-art methods in segmentation accuracy and achieves the current optimal performance in measurement and postoperative prediction under many clinical scenes.
Xingru Huang, Lixia Lou, Ruilong Dan, Lingxiao Chen, Guodong Zeng, Gangyong Jia, Qun Jin, Juan Ye, Yaqi Wang 0002
Expert Syst. Appl.1
2023 POST-IVUS: A perceptual organisation-aware selective transformer framework for intravascular ultrasound segmentation
abstract
Intravascular ultrasound (IVUS) is recommended in guiding coronary intervention. The segmentation of coronary lumen and external elastic membrane (EEM) borders in IVUS images is a key step, but the manual process is time-consuming and error-prone, and suffers from inter-observer variability. In this paper, we propose a novel perceptual oganisation-aware selective transformer framework that can achieve accurate and robust segmentation of the vessel walls in IVUS images. In this framework, temporal context-based feature encoders extract efficient motion features of vessels. Then, a perceptual oganisation-aware selective transformer module is proposed to extract accurate boundary information, supervised by a dedicated boundary loss. The obtained EEM and lumen segmentation results will be fused in a temporal constraining and fusion module, to determine the most likely correct boundaries with robustness to morphology. Our proposed methods are extensively evaluated in non-selected IVUS sequences, including normal, bifurcated, and calcified vessels with shadow artifacts. The results show that the proposed methods outperform the state-of-the-art, with a Jaccard measure of 0.92 for lumen and 0.94 for EEM on the IVUS 2011 open challenge dataset. This work has been integrated into a software QCU-CMS2 to automatically segment IVUS images in a user-friendly environment.
Xingru Huang, Retesh Bajaj, Yilong Li 0002, Xin Ye 0006, Ji Lin 0004, Francesca Pugliese, Anantharaman Ramasamy, Yaqi Wang 0002, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang
Medical Image Anal.1
2023 Stimulus-guided adaptive transformer network for retinal blood vessel segmentation in fundus images
abstract
Automated retinal blood vessel segmentation in fundus images provides important evidence to ophthalmologists in coping with prevalent ocular diseases in an efficient and non-invasive way. However, segmenting blood vessels in fundus images is a challenging task, due to the high variety in scale and appearance of blood vessels and the high similarity in visual features between the lesions and retinal vascular. Inspired by the way that the visual cortex adaptively responds to the type of stimulus, we propose a Stimulus-Guided Adaptive Transformer Network (SGAT-Net) for accurate retinal blood vessel segmentation. It entails a Stimulus-Guided Adaptive Module (SGA-Module) that can extract local-global compound features based on inductive bias and self-attention mechanism. Alongside a light-weight residual encoder (ResEncoder) structure capturing the relevant details of appearance, a Stimulus-Guided Adaptive Pooling Transformer (SGAP-Former) is introduced to reweight the maximum and average pooling to enrich the contextual embedding representation while suppressing the redundant information. Moreover, a Stimulus-Guided Adaptive Feature Fusion (SGAFF) module is designed to adaptively emphasize the local details and global context and fuse them in the latent space to adjust the receptive field (RF) based on the task. The evaluation is implemented on the largest fundus image dataset (FIVES) and three popular retinal image datasets (DRIVE, STARE, CHASEDB1). Experimental results show that the proposed method achieves a competitive performance over the other existing method, with a clear advantage in avoiding errors that commonly happen in areas with highly similar visual features. The sourcecode is publicly available at: https://github.com/Gins-07/SGAT.
Ji Lin 0004, Xingru Huang, Huiyu Zhou 0001, Yaqi Wang 0002, Qianni Zhang
Medical Image Anal.2
2021 SRPN: similarity-based region proposal networks for nuclei and cells detection in histology images
Yibao Sun, Xingru Huang, Huiyu Zhou 0001, Qianni Zhang
Medical Image Anal.2