Shuai Wang 0003

dblp:42/1503-3 · DBLP profile ↗
← Back
45ranked-venue papers
9as first author
36since 2021 · last 2026
0000-0003-3730-6401ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-annotation agreement and prediction consistency networks: Improving semi-supervised segmentation of medical images with ambiguous boundaries
Shuai Wang 0003, Tengjin Weng, Yang Shen 0011, Zhidong Zhao, Yixiu Liu, Pengfei Jiao, Zhiming Cheng, Yaoqi Sun, Yaqi Wang 0002
Artif. Intell. Medicine1
2026 A Comprehensive Survey on the Research and Development of RGB-T Salient Object Detection
abstract
Salient object detection (SOD) aims to mimic human visual perception by identifying the most eye-catching objects within a scene, and has remained a popular research topic for many years. The introduction of thermal (T) images offers additional information for challenging scenarios such as those with low light and complex backgrounds, and thus enhance performance when combined with RGB images. In this paper, we have, to the best of our ability, conducted the first comprehensive survey of dual-modality RGB-T SOD. We summarize and categorize published RGB-T SOD models, emphasizing their characteristics and features. Important components of these models are classified and elaborated, such as feature extraction, modality fusion, and loss function design. Following this, we analyze existing RGB-T SOD datasets and evaluation metrics. We evaluate a selection of representative SOD models using unified protocols and statistical analysis. We also present comparative experiments on the impact of the choice of loss function and the usage of datasets on performance. Finally, we consider several key issues and potential solutions in RGB-T SOD research, revealing promising directions for future efforts. We hope this survey will offer an effective way to understand the current state of the technology and, more importantly, stimulate discussion within the community.
Hongfa Wen, Qiang Zhao 0005, Junbo Ma, Zongpeng Li, Shuai Wang 0003, Chenggang Yan 0001
Comput. Vis. Media6
2026 Graph Neural Networks Based Analog Circuit Link Prediction
Guanyuan Pan, Tiansheng Zhou, Jianxiang Zhao, Yugui Lin, Bingtao Ma, Yaqi Wang 0002, Pietro Liò, Shuai Wang 0003
Eng. Appl. Artif. Intell.9
2026 Weakly supervised object localization via frequency guidance with consistency awareness
Bingfeng Li, Erdong Shi, Boxiang Lv, Shuai Wang 0003
Inf. Sci.7
2026 MICCAI STS 2024 challenge: Semi-supervised instance-level tooth segmentation in panoramic X-ray and CBCT images
abstract
Orthopantomogram (OPGs) and Cone-Beam Computed Tomography (CBCT) are vital for dentistry, but creating large datasets for automated tooth segmentation is hindered by the labor-intensive process of manual instance-level annotation. This research aimed to benchmark and advance semi-supervised learning (SSL) as a solution for this data scarcity problem. We organized the 2nd Semi-supervised Teeth Segmentation (STS 2024) Challenge at MICCAI 2024. We provided a large-scale dataset comprising over 90,000 2D images and 3D axial slices, which includes 2380 OPG images and 330 CBCT scans, all featuring detailed instance-level FDI annotations on part of the data. The challenge attracted 114 (OPG) and 106 (CBCT) registered teams. To ensure algorithmic excellence and full transparency, we rigorously evaluated the valid, open-source submissions from the top 10 (OPG) and top 5 (CBCT) teams, respectively. All successful submissions were deep learning-based SSL methods. The winning semi-supervised models demonstrated impressive performance gains over a fully-supervised nnU-Net baseline trained only on the labeled data. For the 2D OPG track, the top method improved the Instance Affinity (IA) score by over 44 percentage points. For the 3D CBCT track, the winning approach boosted the Instance Dice score by 61 percentage points. This challenge demonstrates the potential benefit benefit of SSL for complex, instance-level medical image segmentation tasks where labeled data is scarce. The most effective approaches consistently leveraged hybrid semi-supervised frameworks that combined knowledge from foundational models like SAM with multi-stage, coarse-to-fine refinement pipelines. Both the challenge dataset and the participants' submitted code have been made publicly available on GitHub (https://github.com/ricoleehduu/STS-Challenge-2024), ensuring transparency and reproducibility.
Yaqi Wang 0002, Jun Liu 0027, Jiaxue Ni, Hongyuan Zhang 0002, Jin Liu 0025, Can Han, Kaiwen Fu, Changkai Ji, Xinxu Cai, Junqiang Chen, Qianni Zhang, Dahong Qian, Shuai Wang 0003, Huiyu Zhou 0001
Medical Image Anal.22
2026 MICCAI 2023 STS Challenge: A retrospective study of semi-supervised approaches for teeth segmentation
abstract
Computer-aided diagnosis greatly enhances personalized treatment planning and diagnostic efficiency by providing accurate dental anatomy through teeth segmentation. However, it still constrained by the scarcity of high-quality annotated dental datasets. To address this issue, this paper presents a dataset combining both 2D panoramic X-rays with over 6,500 images and 3D CBCT with over 580 volumes (88,500+ slices) to support the Semi-supervised Teeth Segmentation (STS) Challenge, which includes partially meticulous annotations and covers all age groups. Moreover, multi-phase semi-supervised teeth segmentation algorithms and high-confidence pseudo-labels refinement strategies were proposed by competitors during this challenge. Algorithms were verified on this proposed dataset and good segmentation performance were achieved, over 93+ and 80+ Dice score were obtained for top three 2D and 3D participants, demonstrating the high quality of this proposed dataset. This paper also summarizes the diverse methods employed by the top-ranking teams in the MICCAI 2023 STS Challenge. Our dataset is publicly accessible through Zenodo ( https://zenodo.org/records/10597292 ), and the participants’ code is hosted on GitHub ( https://github.com/ricoleehduu/STS-Challenge ).
Yaqi Wang 0002, Shuai Wang 0003, Dahong Qian, Hongyuan Zhang 0002, Ruilong Dan, Qianni Zhang, Xingru Huang, Jun Liu 0027, Zhean Ma, Weiwei Cui 0003, Shan Luo 0003, Chengkai Wang, Jiaxue Ni, Dongyun Liu, Zhouhao Lin, Chunshi Wang, Qiupu Chen, Mingqian Li, Huiyu Zhou 0001, Qun Jin
Pattern Recognit.4
2026 SGM-Net: 3D Point Cloud Class-Incremental Segmentation via Semantic-Aware Global Modeling
Jinshuo Liu, Bingtao Ma, Zhidong Zhao, Chenggang Yan 0001, Shuai Wang 0003
IEEE Signal Process. Lett.5
2026 Empirical Study on Fusion Strategy in RGB-T Salient Object Detection
abstract
In the research field of RGB-Thermal saliency object detection (RGB-T SOD), the effective exploitation of the complementary characteristics of the two modalities represents a major challenge for enhancing detection performance. Current fusion methodologies can be roughly classified into early fusion and middle fusion strategies, with prevalent techniques primarily encompassing concatenation, summation, and multiplication of the two modalities. To in depth assess the efficacy of these fusion strategies, we took an empirical investigation on them. Our findings demonstrate that the concatenation of middle features constitutes a more advantageous fusion strategy, yielding superior performance and demonstrating enhanced stability. Furthermore, observing the unique properties of thermal (T) images, we introduced gamma correction as a novel data augmentation methodology to RGB-T SOD. We subsequently evaluated the responses across varying correction parameter ranges, revealing that while the response to this data augmentation technique differs across various models, data augmentation is found to be effective in general. Building upon these findings, we proposed the Gamma Correction Network (GaCNet). Specifically, we also integrated image pyramid mechanism in a lightweight manner, which facilitates a more effective recovery of fine-grained image details. Significant improvement was achieved on commonly used RGB-T testing datasets, especially in VT821 dataset, manifesting the effectiveness of our method.
Shuai Wang 0003, Qiang Zhao 0005, Junbo Ma, Xichun Sheng, Yaoqi Sun, Hongfa Wen, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Region-Based Text-Consistent Augmentation for Multimodal Medical Segmentation
Kunyan Cai, Chenggang Yan 0001, Liangqiong Qu, Shuai Wang 0003, Tao Tan 0002
MICCAI (3)5
2025 Weakly supervised object localization via foreground generation with foreground-background constraints
Bingfeng Li, Erdong Shi, Haohao Ruan, Zhanshuo Jiang, Xinwei Li 0002, Keping Wang, Shuai Wang 0003
Expert Syst. Appl.7
2025 Lightweight three-stream encoder-decoder network for multi-modal salient object detection
Junzhe Lu 0002, Tingyu Wang 0002, Bin Wan, Qiang Zhao 0005, Shuai Wang 0003, Yaoqi Sun, Yang Zhou 0052, Chenggang Yan 0001
J. Vis. Commun. Image Represent.5
2025 Dynamic domain generalization for medical image segmentation
Zhiming Cheng, Mingxia Liu 0001, Chenggang Yan 0001, Shuai Wang 0003
Neural Networks4
2025 Domain generalization for image classification with dynamic decision boundary
Zhiming Cheng, Mingxia Liu 0001, Defu Yang, Zhidong Zhao, Chenggang Yan 0001, Shuai Wang 0003
Pattern Recognit.6
2025 GLA: Global-Local Awareness for 3D Point Cloud Class-Incremental Semantic Segmentation
abstract
Semantic segmentation of 3D point clouds has garnered considerable attention in academic research and industrial applications. However, existing methods typically assume fixed semantic classes, which is unrealistic for practical scenarios where new classes emerge incrementally. This leads to catastrophic forgetting of previous knowledge and semantic shifts where annotations treat previous classes as background in new tasks. Moreover, the unordered and unstructured properties of point clouds further exacerbate catastrophic forgetting. To address these, we propose the Global-Local Awareness for 3D point cloud class-incremental semantic segmentation (i.e., GLA). Specifically, we introduce a global-aware modeling network (GAMN) based on the state space model, providing robust feature representations. Furthermore, to alleviate catastrophic forgetting, we propose a Graph Attention Knowledge Distillation (GAKD) module. GAKD adaptively integrates local geometric features through an attention mechanism, emphasizing the distillation of geometric structural knowledge. Notably, our approach integrates global context-awareness and local attention through the state space model and GAKD module, respectively. In addition, we introduce pseudo-labeling to alleviate semantic shift. Experiments on the S3DIS dataset validate the superiority of our approach.
Bingtao Ma, Shitong Zhang, Chenggang Yan 0001, Shuai Wang 0003
IEEE Signal Process. Lett.5
2025 A Novel Multi-Modal Population-Graph Based Framework for Patients of Esophageal Squamous Cell Cancer Prognostic Risk Prediction
abstract
Prognostic risk prediction is pivotal for clinicians to appraise the patient's esophageal squamous cell cancer (ESCC) progression status precisely and tailor individualized therapy treatment plans. Currently, CT-based multi-modal prognostic risk prediction methods have gradually attracted the attention of researchers for their universality, which is also able to be applied in scenarios of preoperative prognostic risk assessment in the early stages of cancer. However, much of the current work focuses only on CT images of the primary tumor, ignoring the important role that CT images of lymph nodes play in prognostic risk prediction. Additionally, it is important to consider and explore the inter-patient feature similarity in prognosis when developing models. To solve these problems, we proposed a novel multi-modal population-graph based framework leveraging CT images including primary tumor and lymph nodes combined with clinical, hematology, and radiomics data for ESCC prognostic risk prediction. A patient population graph was constructed to excavate the homogeneity and heterogeneity of inter-patient feature embedding. Moreover, a novel node-level multi-task joint loss was proposed for graph model optimization through a supervised-based task and an unsupervised-based task. Sufficient experimental results show that our model achieved state-of-the-art performance compared with other baseline models as well as the gold standard on discriminative ability, risk stratification, and clinical utility.
Shuai Wang 0003, Yaqi Wang 0002, Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang
IEEE J. Biomed. Health Informatics2
2025 A Novel Spatio-Temporal Hub Identification in Brain Networks by Learning Dynamic Graph Embedding on Grassmannian Manifolds
abstract
Mounting evidence has revealed that functional brain networks are intrinsically dynamic, undergoing changes over time, even in the resting-state environment. Notably, recent studies have highlighted the existence of a small number of critical brain regions within each functional brain network that exhibit a flexible role in adapting the geometric pattern of brain connectivity over time, referred to as "temporal hub" regions. Therefore, the identification of these temporal hubs becomes pivotal for comprehending the mechanisms that underlie the dynamic evolution of brain connectivity. However, existing spatio-temporal hub identification methods rely on static network-based approaches, wherein each temporal hub region is independently inferred from individual time-segmented networks without considering their temporal consistency and consequently fails to align the evolution of hubs with the dynamic changes in brain states. To address this limitation, we propose a novel spatio-temporal hub identification method that fully leverages dynamic graph embedding to distinguish temporal hubs from peripheral nodes, in which dynamic graph embeddings are learned from both spatial and temporal dimensions. Specifically, to preserve the temporal consistency of evolving networks, we model the dynamic graph embedding as a physical model of time, where the network-to-network transition is mathematically expressed as a total variation of dynamic graph embedding with respect to time. Furthermore, a Grassmannian manifold optimization scheme is introduced to enhance graph embedding learning and capture the time-varying topology of brain networks. Experimental results on both synthetic and real fMRI data demonstrate superior temporal consistency in hub identification, surpassing conventional approaches.
Defu Yang, Minghan Chen 0001, Shuai Wang 0003, Jiazhou Chen 0001, Hongmin Cai, Guorong Wu 0001, Wentao Zhu 0002
IEEE Trans. Medical Imaging4
2025 Unpaired semantic neural person image synthesis
Yixiu Liu, Pengju Si, Shangdong Zhu, Chenggang Yan 0001, Shuai Wang 0003, Haibing Yin
Vis. Comput.6
2024 Accurate Segmentation of Optic Disc and Cup from Multiple Pseudo-labels by Noise-aware Learning
abstract
Optic disc and cup segmentation plays a crucial role in automating the screening and diagnosis of optic glaucoma. While data-driven convolutional neural networks (CNNs) show promise in this area, the inherent ambiguity of segmenting objects and background boundaries in the task of optic disc and cup segmentation leads to noisy annotations that impact model performance. To address this, we propose an innovative label-denoising method of Multiple Pseudo-labels Noise-aware Network (MPNN) for accurate optic disc and cup segmentation. Specifically, the Multiple Pseudo-labels Generation and Guided Denoising (MPGGD) module generates pseudo-labels by multiple different initialization networks trained on true labels, and the pixel-level consensus information extracted from these pseudo-labels guides to differentiate clean pixels from noisy pixels. The training framework of the MPNN is constructed by a teacher-student architecture to learn segmentation from clean pixels and noisy pixels. Particularly, such a framework adeptly leverages (i) reliable and fundamental insight from clean pixels and (ii) the supplementary knowledge within noisy pixels via multiple perturbation-based unsupervised consistency. Compared to other label-denoising methods, comprehensive experimental results on the RIGA dataset demonstrate our method’s excellent performance. The code is available at https://github.com/wwwtttjjj/MPNN.
Tengjin Weng, Yang Shen 0011, Zhidong Zhao, Zhiming Cheng, Shuai Wang 0003
CSCWD5
2024 MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang, Yaqi Wang 0002, Shuai Wang 0003
MICCAI (5)7
2024 Stochastic Context Consistency Reasoning for Domain Adaptive Object Detection
Liang Li 0003, Chenggang Yan 0001, Hongkui Wang, Shuai Wang 0003, Heng Jin
ACM Multimedia6
2024 Superpixel-based Efficient Sampling for Learning Neural Fields from Large Input
abstract
In recent years, neural field-based methods for synthesizing novel views have gained popularity due to their exceptional rendering quality and fast training speed. However, the computational cost of volumetric rendering has significantly increased with the advancement of camera technology and the subsequent rise in average camera resolution. Despite extensive efforts to accelerate the training process, the training duration remains unacceptable for high-resolution inputs. Therefore, it's crucial to develop efficient sampling methods to optimize the learning process of neural fields from large inputs. In this paper, we present a new technique called Superpixel-based Efficient Sampling (SES) to improve the learning efficiency of neural fields. Our approach optimizes pixel-level ray sampling by segmenting the error map into multiple superpixels and dynamically updating their errors during training to increase ray sampling in superpixel areas with higher rendering errors. Compared with other methods, our approach leverages the flexibility of superpixels, effectively reducing redundant sampling while considering local information. Our method not only speeds up the learning process but also enhances the rendering quality learned from large inputs. We conduct extensive experiments to evaluate the effectiveness of our method across several baselines and datasets. The code will be released.
Zhongwei Xuan, Zunjie Zhu, Shuai Wang 0003, Haibing Yin, Hongkui Wang, Ming Lu 0002
ACM Multimedia3
2024 ADNet: Anti-noise dual-branch network for road defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001
Eng. Appl. Artif. Intell.6
2024 GINet:Graph interactive network with semantic-guided spatial refinement for salient object detection in optical remote sensing images
Chenwei Zhu, Xiaofei Zhou 0003, Liuxin Bao, Hongkui Wang, Shuai Wang 0003, Zunjie Zhu, Chenggang Yan 0001, Jiyong Zhang 0001
J. Vis. Commun. Image Represent.5
2024 Deep Joint Semantic Adaptation Network for Multi-source Unsupervised Domain Adaptation
Zhiming Cheng, Shuai Wang 0003, Defu Yang, Mang Xiao, Chenggang Yan 0001
Pattern Recognit.2
2024 MFFNet: Multi-Modal Feature Fusion Network for V-D-T Salient Object Detection
abstract
This article discusses the limitations of single- and two-modal salient object detection (SOD) methods and the emergence of multi-modal SOD techniques that integrate Visible, Depth, or Thermal information. However, current multi-modal methods often rely on simple fusion techniques such as addition, multiplication and concatenation, to combine the different modalities, which is ineffective for challenging scenes, such as low illumination and background messy. To address this issue, we propose a novel multi-modal feature fusion network (MFFNet) for V-D-T salient object detection, where the two key points are the triple-modal deep fusion encoder and the progressive feature enhancement decoder. The MFFNet's triple-modal deep fusion (TDF) module is designed to integrate the features of the three modalities and explore their complementarity by utilizing mutual optimization during the encoding phase. In addition, the progressive feature enhancement decoder consists of the weighted context-enhanced feature (WCF) module, region optimization (RO) module and boundary perception (BP) module to produce region-aware and contour-aware features. After that, a multi-scale fusion (MF) module is proposed to integrate these features and generate high-quality saliency maps. We conduct extensive experiments on the VDT-2048 dataset, and our results show that the proposed MFFNet outperforms 12 state-of-the-art multi-modal methods.
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001
IEEE Trans. Multim.6
2023 Spatiotemporal Hub Identification in Brain Network by Learning Dynamic Graph Embedding on Grassmannian Manifold
Defu Yang, Minghan Chen 0001, Yitian Xue, Shuai Wang 0003, Guorong Wu 0001, Wentao Zhu 0002
MICCAI (2)5
2023 Parsing is All You Need for Accurate Gait Recognition in the Wild
abstract
Binary silhouettes and keypoint-based skeletons have dominated human gait recognition studies for decades since they are easy to extract from video frames. Despite their success in gait recognition for in-the-lab environments, they usually fail in real-world scenarios due to their low information entropy for gait representations. To achieve accurate gait recognition in the wild, this paper presents a novel gait representation, named Gait Parsing Sequence (GPS). GPSs are sequences of fine-grained human segmentation, i.e., human parsing, extracted from video frames, so they have much higher information entropy to encode the shapes and dynamics of fine-grained human parts during walking. Moreover, to effectively explore the capability of the GPS representation, we propose a novel human parsing-based gait recognition framework, named ParsingGait. ParsingGait contains a Convolutional Neural Network (CNN)-based backbone and two light-weighted heads. The first head extracts global semantic features from GPSs, while the other one learns mutual information of part-level features through Graph Convolutional Networks to model the detailed dynamics of human walking. Furthermore, due to the lack of suitable datasets, we build the first parsing-based dataset for gait recognition in the wild, named Gait3D-Parsing, by extending the large-scale and challenging Gait3D dataset. Based on Gait3D-Parsing, we comprehensively evaluate our method and existing gait recognition methods. Specifically, ParsingGait achieves a 17.5% Rank-1 increase compared with the state-of-the-art silhouette-based method. In addition, by replacing silhouettes with GPSs, current gait recognition methods achieve about 12.5% ~ 19.2% improvements in Rank-1 accuracy. The experimental results show a significant improvement in accuracy brought by the GPS representation and the superiority of ParsingGait.
Jinkai Zheng, Xinchen Liu, Shuai Wang 0003, Chenggang Yan 0001, Wu Liu 0005
ACM Multimedia3
2023 Transformer-Based Multi-Scale Feature Integration Network for Video Saliency Prediction
abstract
Most cutting-edge video saliency prediction models rely on spatiotemporal features extracted by 3D convolutions due to its local contextual cues acquirement ability. However, the shortage of 3D convolutions is that it cannot effectively capture long-term spatiotemporal dependencies in videos. To address this limitation, we propose a novel Transformer-based Multi-scale Feature Integration Network (TMFI-Net) for video saliency prediction, where the proposed TMFI-Net consists of a semantic-guided encoder and a hierarchical decoder. Firstly, embarking on the Transformer-based multi-level spatiotemporal features, the semantic-guided encoder enhances the features by inserting the high-level feature into each level feature via a top-down pathway and a longitudinal connection, which endows the multi-level spatiotemporal features with rich contextual information. In this way, the features are steered to give more concerns to saliency regions. Secondly, the hierarchical decoder employs a multi-dimensional attention (MA) module to elevate features along channel, temporal, and spatial dimensions jointly. Successively, the hierarchical decoder deploys a progressive decoding block to conduct an initial saliency prediction, which provides a coarse localization of saliency regions. Lastly, considering the complementarity of different saliency predictions, we integrate all initial saliency prediction results into the final saliency map. Comprehensive experimental results on four video saliency datasets firmly demonstrate that our model achieves superior performance when compared with the state-of-the-art video saliency models. The code is available athttps://github.com/wusonghe/TMFI-Net.
Xiaofei Zhou 0003, Songhe Wu, Bolun Zheng, Shuai Wang 0003, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 CDNet: Contrastive Disentangled Network for Fine-Grained Image Categorization of Ocular B-Scan Ultrasound
abstract
Precise and rapid categorization of images in the B-scan ultrasound modality is vital for diagnosing ocular diseases. Nevertheless, distinguishing various diseases in ultrasound still challenges experienced ophthalmologists. Thus a novel contrastive disentangled network (CDNet) is developed in this work, aiming to tackle the fine-grained image categorization (FGIC) challenges of ocular abnormalities in ultrasound images, including intraocular tumor (IOT), retinal detachment (RD), posterior scleral staphyloma (PSS), and vitreous hemorrhage (VH). Three essential components of CDNet are the weakly-supervised lesion localization module (WSLL), contrastive multi-zoom (CMZ) strategy, and hyperspherical contrastive disentangled loss (HCD-Loss), respectively. These components facilitate feature disentanglement for fine-grained recognition in both the input and output aspects. The proposed CDNet is validated on our ZJU Ocular Ultrasound Dataset (ZJUOUSD), consisting of 5213 samples. Furthermore, the generalization ability of CDNet is validated on two public and widely-used chest X-ray FGIC benchmarks. Quantitative and qualitative results demonstrate the efficacy of our proposed CDNet, which achieves state-of-the-art performance in the FGIC task.
Ruilong Dan, Gangyong Jia, Shuai Wang 0003, Ruiquan Ge, Guiping Qian, Qun Jin, Juan Ye, Yaqi Wang 0002
IEEE J. Biomed. Health Informatics6
2023 Breast Tumor Segmentation in DCE-MRI With Tumor Sensitive Synthesis
abstract
Segmenting breast tumors from dynamic contrast-enhanced magnetic resonance (DCE-MR) images is a critical step for early detection and diagnosis of breast cancer. However, variable shapes and sizes of breast tumors, as well as inhomogeneous background, make it challenging to accurately segment tumors in DCE-MR images. Therefore, in this article, we propose a novel tumor-sensitive synthesis module and demonstrate its usage after being integrated with tumor segmentation. To suppress false-positive segmentation with similar contrast enhancement characteristics to true breast tumors, our tumor-sensitive synthesis module can feedback differential loss of the true and false breast tumors. Thus, by following the tumor-sensitive synthesis module after the segmentation predictions, the false breast tumors with similar contrast enhancement characteristics to the true ones will be effectively reduced in the learned segmentation model. Moreover, the synthesis module also helps improve the boundary accuracy while inaccurate predictions near the boundary will lead to higher loss. For the evaluation, we build a very large-scale breast DCE-MR image dataset with 422 subjects from different patients, and conduct comprehensive experiments and comparisons with other algorithms to justify the effectiveness, adaptability, and robustness of our proposed method.
Shuai Wang 0003, Li Wang 0026, Liangqiong Qu, Fuhua Yan, Qian Wang 0001, Dinggang Shen
IEEE Trans. Neural Networks Learn. Syst.1
2022 Global-Local attention network with multi-task uncertainty loss for abnormal lymph node detection in MR images
Shuai Wang 0003, Yingying Zhu 0003, Sungwon Lee 0003, Daniel C. Elton, Thomas C. Shen, Youbao Tang, Yifan Peng 0002, Zhiyong Lu, Ronald M. Summers
Medical Image Anal.1
2022 AGMB-Transformer: Anatomy-Guided Multi-Branch Transformer Network for Automated Evaluation of Root Canal Therapy
abstract
Accurate evaluation of the treatment result on X-ray images is a significant and challenging step in root canal therapy since the incorrect interpretation of the therapy results will hamper timely follow-up which is crucial to the patients’ treatment outcome. Nowadays, the evaluation is performed in a manual manner, which is time-consuming, subjective, and error-prone. In this article, we aim to automate this process by leveraging the advances in computer vision and artificial intelligence, to provide an objective and accurate method for root canal therapy result assessment. A novel anatomy-guided multi-branch Transformer (AGMB-Transformer) network is proposed, which first extracts a set of anatomy features and then uses them to guide a multi-branch Transformer network for evaluation. Specifically, we design a polynomial curve fitting segmentation strategy with the help of landmark detection to extract the anatomy features. Moreover, a branch fusion module and a multi-branch structure including our progressive Transformer and Group Multi-Head Self-Attention (GMHSA) are designed to focus on both global and local features for an accurate diagnosis. To facilitate the research, we have collected a large-scale root canal therapy evaluation dataset with 245 root canal therapy X-ray images, and the experiment results show that our AGMB-Transformer can improve the diagnosis accuracy from 57.96% to 90.20% compared with the baseline network. The proposed AGMB-Transformer can achieve a highly accurate evaluation of root canal therapy. To our best knowledge, our work is the first to perform automatic root canal therapy evaluation and has important clinical value to reduce the workload of endodontists.
Guodong Zeng, Jun Wang 0041, Qun Jin, Lingling Sun, Qianni Zhang, Qisi Lian, Guiping Qian, Neng Xia, Ruizi Peng, Shuai Wang 0003, Yaqi Wang 0002
IEEE J. Biomed. Health Informatics13
2021 Asymmetric multi-task attention network for prostate bed segmentation in computed tomography images
Xuanang Xu, Chunfeng Lian, Shuai Wang 0003, Ronald C. Chen, Andrew Z. Wang, Trevor J. Royce, Pew-Thian Yap, Dinggang Shen, Jun Lian
Medical Image Anal.3
2021 Multiscale Attention Guided Network for COVID-19 Diagnosis Using Chest X-Ray Images
abstract
Coronavirus disease 2019 (COVID-19) is one of the most destructive pandemic after millennium, forcing the world to tackle a health crisis. Automated lung infections classification using chest X-ray (CXR) images could strengthen diagnostic capability when handling COVID-19. However, classifying COVID-19 from pneumonia cases using CXR image is a difficult task because of shared spatial characteristics, high feature variation and contrast diversity between cases. Moreover, massive data collection is impractical for a newly emerged disease, which limited the performance of data thirsty deep learning models. To address these challenges, Multiscale Attention Guided deep network with Soft Distance regularization (MAG-SD) is proposed to automatically classify COVID-19 from pneumonia CXR images. In MAG-SD, MA-Net is used to produce prediction vector and attention from multiscale feature maps. To improve the robustness of trained model and relieve the shortage of training data, attention guided augmentations along with a soft distance regularization are posed, which aims at generating meaningful augmentations and reduce noise. Our multiscale attention model achieves better classification performance on our pneumonia CXR image dataset. Plentiful experiments are proposed for MAG-SD which demonstrates its unique advantage in pneumonia classification over cutting-edge models. The code is available at https://github.com/JasonLeeGHub/MAG-SD.
Jingxiong Li, Yaqi Wang 0002, Shuai Wang 0003, Jun Wang 0041, Jun Liu 0027, Qun Jin, Lingling Sun
IEEE J. Biomed. Health Informatics3
2021 Multi-Scale Context-Guided Deep Network for Automated Lesion Segmentation With Endoscopy Images of Gastrointestinal Tract
abstract
Accurate lesion segmentation based on endoscopy images is a fundamental task for the automated diagnosis of gastrointestinal tract (GI Tract) diseases. Previous studies usually use hand-crafted features for representing endoscopy images, while feature definition and lesion segmentation are treated as two standalone tasks. Due to the possible heterogeneity between features and segmentation models, these methods often result in sub-optimal performance. Several fully convolutional networks have been recently developed to jointly perform feature learning and model training for GI Tract disease diagnosis. However, they generally ignore local spatial details of endoscopy images, as down-sampling operations (e.g., pooling and convolutional striding) may result in irreversible loss of image spatial information. To this end, we propose a multi-scale context-guided deep network (MCNet) for end-to-end lesion segmentation of endoscopy images in GI Tract, where both global and local contexts are captured as guidance for model training. Specifically, one global subnetwork is designed to extract the global structure and high-level semantic context of each input image. Then we further design two cascaded local subnetworks based on output feature maps of the global subnetwork, aiming to capture both local appearance information and relatively high-level semantic information in a multi-scale manner. Those feature maps learned by three subnetworks are further fused for the subsequent task of lesion segmentation. We have evaluated the proposed MCNet on 1,310 endoscopy images from the public EndoVis-Ab and CVC-ClinicDB datasets for abnormal segmentation and polyp segmentation, respectively. Experimental results demonstrate that MCNet achieves [Formula: see text] and [Formula: see text] mean intersection over union (mIoU) on two datasets, respectively, outperforming several state-of-the-art approaches in automated lesion segmentation with endoscopy images of GI Tract.
Shuai Wang 0003, Yang Cong, Hancan Zhu, Xianyi Chen, Liangqiong Qu, Huijie Fan, Qiang Zhang 0008, Mingxia Liu 0001
IEEE J. Biomed. Health Informatics1
2021 Boundary Coding Representation for Organ Segmentation in Prostate Cancer Radiotherapy
abstract
Accurate segmentation of the prostate and organs at risk (OARs, e.g., bladder and rectum) in male pelvic CT images is a critical step for prostate cancer radiotherapy. Unfortunately, the unclear organ boundary and large shape variation make the segmentation task very challenging. Previous studies usually used representations defined directly on unclear boundaries as context information to guide segmentation. Those boundary representations may not be so discriminative, resulting in limited performance improvement. To this end, we propose a novel boundary coding network (BCnet) to learn a discriminative representation for organ boundary and use it as the context information to guide the segmentation. Specifically, we design a two-stage learning strategy in the proposed BCnet: 1) Boundary coding representation learning. Two sub-networks under the supervision of the dilation and erosion masks transformed from the manually delineated organ mask are first separately trained to learn the spatial-semantic context near the organ boundary. Then we encode the organ boundary based on the predictions of these two sub-networks and design a multi-atlas based refinement strategy by transferring the knowledge from training data to inference. 2) Organ segmentation. The boundary coding representation as context information, in addition to the image patches, are used to train the final segmentation network. Experimental results on a large and diverse male pelvic CT dataset show that our method achieves superior performance compared with several state-of-the-art methods.
Shuai Wang 0003, Mingxia Liu 0001, Jun Lian, Dinggang Shen
IEEE Trans. Medical Imaging1
2020 Asymmetrical Multi-task Attention U-Net for the Segmentation of Prostate Bed in CT Image
Xuanang Xu, Chunfeng Lian, Shuai Wang 0003, Andrew Z. Wang, Trevor J. Royce, Ronald C. Chen, Jun Lian, Dinggang Shen
MICCAI (4)3
2020 CT Male Pelvic Organ Segmentation via Hybrid Loss Network With Incomplete Annotation
abstract
Sufficient data with complete annotation is essential for training deep models to perform automatic and accurate segmentation of CT male pelvic organs, especially when such data is with great challenges such as low contrast and large shape variation. However, manual annotation is expensive in terms of both finance and human effort, which usually results in insufficient completely annotated data in real applications. To this end, we propose a novel deep framework to segment male pelvic organs in CT images with incomplete annotation delineated in a very user-friendly manner. Specifically, we design a hybrid loss network derived from both voxel classification and boundary regression, to jointly improve the organ segmentation performance in an iterative way. Moreover, we introduce a label completion strategy to complete the labels of the rich unannotated voxels and then embed them into the training data to enhance the model capability. To reduce the computation complexity and improve segmentation performance, we locate the pelvic region based on salient bone structures to focus on the candidate segmentation organs. Experimental results on a large planning CT pelvic organ dataset show that our proposed method with incomplete annotation achieves comparable segmentation performance to the state-of-the-art methods with complete annotation. Moreover, our proposed method requires much less effort of manual contouring from medical professionals such that an institutional specific model can be more easily established.
Shuai Wang 0003, Dong Nie, Liangqiong Qu, Yeqin Shao, Jun Lian, Qian Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging1
2019 Regression Convolutional Neural Network for Automated Pediatric Bone Age Assessment From Hand Radiograph
abstract
Skeletal bone age assessment is a common clinical practice to investigate endocrinology, and genetic and growth disorders of children. However, clinical interpretation and bone age analyses are time-consuming, labor intensive, and often subject to inter-observer variability. This advocates the need of a fully automated method for bone age assessment. We propose a regression convolutional neural network (CNN) to automatically assess the pediatric bone age from hand radiograph. Our network is specifically trained to place more attention to those bone age related regions in the X-ray images. Specifically, we first adopt the attention module to process all images and generate the coarse/fine attention maps as inputs for the regression network. Then, the regression CNN follows the supervision of the dynamic attention loss during training; thus, it can estimate the bone age of the hard (or "outlier") images more accurately. The experimental results show that our method achieves an average discrepancy of 5.2-5.3 months between clinical and automatic bone age evaluations on two large datasets. In conclusion, we propose a fully automated deep learning solution to process X-ray images of the hand for bone age assessment, with the accuracy comparable to human experts but with much better efficiency.
Xuhua Ren, Xiujun Yang, Shuai Wang 0003, Sahar Ahmad, Lei Xiang 0001, Shaun Richard Stone, Yiqiang Zhan, Dinggang Shen, Qian Wang 0001
IEEE J. Biomed. Health Informatics4
2017 Deep learning of directional truncated signed distance function for robust 3D object recognition
abstract
In this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression tasks. To achieve this, we first adopt the Truncated Signed Distance Function (TSDF) to encode the point cloud and extract low compact discriminative feature via unsupervised deep learning network. This approach can not only eliminate the dense scale sampling for offline model training but also reduce the distortion by mapping the 3D shape to the 2D plane and overcome the dependence on color cues. Then, we train a Hough forests to achieve multi-object detection and 6-DoF pose estimation simultaneously. In addition, we propose a robust multilevel verification strategy that effectively reduces the unpredictability of output values which occurs in the hough regression module. Experiments on public datasets demonstrate that our approach provides effective results comparable to the state-of-the-arts.
Hongsen Liu, Yang Cong, Shuai Wang 0003, Huijie Fan, Dongying Tian, Yandong Tang
IROS3
2017 Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy Diagnosis
abstract
Successful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data.
Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao
ACM Trans. Multim. Comput. Commun. Appl.1
2016 Scalable gastroscopic video summarization via similar-inhibition dictionary selection
Shuai Wang 0003, Yang Cong, Jun Cao 0002, Yunsheng Yang, Yandong Tang, Huaici Zhao
Artif. Intell. Medicine1
2016 UDSFS: Unsupervised deep sparse feature selection
Yang Cong, Shuai Wang 0003, Baojie Fan, Yunsheng Yang
Neurocomputing2
2015 Computer aided endoscope diagnosis via weakly labeled data mining
abstract
In comparison to most computer aided endoscope diagnosis methods using pixel-wise groundtruth by physicians manually, it is easy to get lots of endoscope images with corresponding diagnostic reports. In this paper, we intend to mine pixel-wise label information from these reports with weak frame-level labels automatically. To achieve this, we formulate our computer aided diagnosis problem as a Multiple Instance Learning (MIL) issue, where we represent each image as superpixels. Each image and each superpixel is cast as bag and instance, respectively. We then evaluate and select the most positive instances from positive bags automatically which helps us transform the frame-level classification problem into a standard supervised learning problem. In the experiment, we build a new gastroscopic image dataset with more than 3000 weakly labeled images, and ours outperforms the state-of-the-art methods, which verifies the effectiveness of our model.
Shuai Wang 0003, Yang Cong, Huijie Fan, Yunsheng Yang, Yandong Tang, Huaici Zhao
ICIP1
2015 Deep sparse feature selection for computer aided endoscopy diagnosis
Yang Cong, Shuai Wang 0003, Ji Liu 0002, Jun Cao 0002, Yunsheng Yang, Jiebo Luo 0001
Pattern Recognit.2