VLDB 2026 Research / reviewers in the wild / expert
Meng Wang 0038
dblp:93/6765-38
· DBLP profile ↗
21ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0001-7882-1747ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | YoloSeg: You only label once for medical image segmentationabstractAcquiring pixel-level annotations for medical images is an extremely time-consuming and labor-intensive task, typically occupying the majority of the development cycle for medical image segmentation models. While existing semi-supervised methods have achieved promising results, they generally still require annotations for 10%-30% of the samples to effectively guide learning from unlabeled data, which remains a substantial burden for real-world applications. In this study, we propose YoloSeg, a novel framework for medical image segmentation under extreme label scarcity, where only a single labeled image is available. YoloSeg integrates Segment Anything Model 2 to propagate labels from the labeled image to unlabeled images, thereby expanding the labeled data pool. To address the inherent noise in pseudo-labels, we employ multi-view label propagation, decomposing pseudo-labels into consensus and divergence regions. We introduce a dual-component loss to handle these regions separately, facilitating more robust pseudo-label learning for segmentation models. Additionally, we propose a cross-patch data augmentation strategy to generate new samples with stronger semantic consistency, further enhancing the stability of training and improving model generalization. We validate our method on ten diverse medical image segmentation datasets, encompassing a wide range of segmentation targets including organs, vessels, and lesions. Experimental results show that YoloSeg achieves performance comparable to fully-supervised baselines, with an average Dice score difference of only 3.08% across all tasks, and significantly outperforms other state-of-the-art semi-supervised and one-shot methods. YoloSeg significantly improves the feasibility and cost-effectiveness of deep learning in scenarios with severely limited annotation budgets. This approach holds promise for enabling the rapid development and deployment of custom segmentation models across diverse medical centers, thereby supporting the broader adoption of intelligent medical technologies. Code is available at https://github.com/iMED-Lab/YoloSeg. Mingen Zhang, Meng Wang 0038, Lei Mou, Jingfeng Zhang, Yitian Zhao |
Medical Image Anal. | 3 |
| 2026 | Unsupervised domain adaptation via style-aware self-intermediate domain
Lianyu Wang, Meng Wang 0038, Daoqiang Zhang, Huazhu Fu |
Pattern Recognit. | 2 |
| 2026 | MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAMabstractThe Medical Segment Anything Model (MedSAM) has demonstrated strong performance in medical image segmentation, attracting increasing attention in the medical imaging domain. However, as with many prompt-based segmentation models, its performance is highly sensitive to the type and location of input prompts. This sensitivity often leads to suboptimal segmentation outcomes and necessitates labor-intensive manual prompt tuning, which hampers both efficiency and robustness. To address this challenge, this paper proposes MedSAM-U, an uncertainty-guided framework designed to automatically refine prompt inputs and enhance segmentation reliability. Specifically, a Multi-Prompt Adapter is integrated into MedSAM, resulting in MPA-MedSAM, which enables the model to effectively accommodate diverse multi-prompt inputs. An uncertainty estimation module is then introduced to evaluate the reliability of the prompts and their initial segmentation results. Based on this, a novel uncertainty-guided prompt adaptation strategy is applied to automatically generate refined prompts and more accurate segmentation outputs. The proposed MedSAM-U framework is evaluated across multiple medical imaging modalities. Experimental results on five diverse datasets demonstrate that MedSAM-U achieves consistent performance improvements ranging from 1.7% to 20.5% over the baseline MedSAM, confirming its effectiveness and practicality for robust and efficient medical image segmentation. Ke Zou, Mengting Luo, Linchao He, Meng Wang 0038, Yi Zhang 0018, Hu Chen 0002, Huazhu Fu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | SuperMag: Vision-based Tactile Data Guided High-resolution Tactile Shape Reconstruction for Magnetic Tactile SensorsabstractMagnetic-based tactile sensors (MBTS) combine the advantages of compact design and high-frequency operation but suffer from limited spatial resolution due to their sparse taxel arrays. This paper proposes SuperMag, a tactile shape reconstruction method that addresses this limitation by leveraging high-resolution vision-based tactile sensor (VBTS) data to supervise MBTS super-resolution. Co-designed, open-source VBTS and MBTS with identical contact modules enable synchronized data collection of high-resolution shapes and magnetic signals via a symmetric calibration setup. We frame tactile shape reconstruction as a conditional generative problem, employing a conditional variational auto-encoder to infer high-resolution shapes from low-resolution MBTS inputs. The MBTS achieves a sampling frequency of 125 Hz, whereas the shape reconstruction sustains an inference time within 2.5 ms. This cross-modality synergy advances tactile perception of the MBTS, potentially unlocking its new capabilities in high-precision robotic tasks. Peiyao Hou, Danning Sun, Meng Wang 0038, Zeyu Zhang 0001, Hangxin Liu, Wanlin Li, Ziyuan Jiao |
IROS | 3 |
| 2025 | Fairness-Aware vCDR-Controlled Generation for Glaucoma Diagnosis
Shuran Yang, Feixiang Zhou, Meng Wang 0038, Yitian Zhao, Yalin Zheng, Yanda Meng |
MICCAI (9) | 8 |
| 2025 | Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng |
MICCAI (15) | 5 |
| 2025 | Uncertainty-Aware Medical Diagnostic Phrase Identification and GroundingabstractMedical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG)-which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, < $\mathtt {BOX}$BOX >, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels. Ke Zou, Yang Bai 0011, Bo Liu 0113, Zhihao Chen 0004, Yang Zhou 0017, Xuedong Yuan, Meng Wang 0038, Xiaojing Shen, Xiaochun Cao, Huazhu Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Reliable Federated Disentangling Network for Non-IID Domain FeatureabstractFederated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability. Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Big Data | 1 |
| 2025 | Toward Reliable Medical Image Segmentation by Modeling Evidential Calibrated UncertaintyabstractMedical image segmentation is critical for disease diagnosis and treatment assessment. However, concerns regarding the reliability of segmentation regions persist among clinicians, mainly attributed to the absence of confidence assessment, robustness, and calibration to accuracy. To address this, we introduce deep evidential segmentation model (DEviS), an easily implementable foundational model that seamlessly integrates into various medical image segmentation networks. DEviS not only enhances the calibration and robustness of baseline segmentation accuracy but also provides high-efficiency uncertainty estimation for reliable predictions. By leveraging subjective logic theory, we explicitly model probability and uncertainty for medical image segmentation. Here, the Dirichlet distribution parameterizes the distribution of probabilities for different classes of the segmentation results. To generate calibrated predictions and uncertainty, we develop a trainable calibrated uncertainty penalty. Furthermore, DEviS incorporates an uncertainty-aware filtering (UAF) module, which designs the metric of uncertainty-calibrated error to filter out-of-distribution (OOD) data. We conducted validation studies on publicly available datasets, including ISIC2018, KiTS2021, LiTS2017, and BraTS2019, to assess the accuracy and robustness of different backbone segmentation models enhanced by DEviS, as well as the efficiency and reliability of uncertainty estimation. Additionally, two potential clinical trials were conducted using the UAF module. The clinical application conducted on the Johns Hopkins OCT and Duke OCT-DME datasets demonstrated the effectiveness of the model in filtering OOD data. The second trial evaluated its efficacy in filtering high-quality data on the FIVES datasets. At last, the proposed DEviS method was extended to semi-supervised medical image segmentation, where it exhibited strong robustness under noisy conditions. Our code has been released in https://github.com/Cocofeat/DEviS. Ke Zou, Ling Huang 0003, Xuedong Yuan, Xiaojing Shen, Meng Wang 0038, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Cybern. | 7 |
| 2025 | SpectralDINO: Dual Mixture-of-Subspaces Low-Rank Adaptation for Cross-Domain Hyperspectral Image Few-Shot ClassificationabstractRecently, few-shot learning-based methods have achieved impressive results in cross-domain hyperspectral classification. However, existing approaches often ignore differences in spectral information caused by varing spectral range across different datasets, and encounter limitations due to the constrained parameter size of the models. Furthermore, the substantial differences between RGB and hyperspectral images present significant challenges in applying foundation models (e.g., SAM, DINOv2) to the hyperspectral domain. This paper proposes a novel framework for cross-domain few-shot hyperspectral classification that leverages parameter-efficient fine-tuning, which we apply to DINOv2 to construct SpectralDINO. Different spectral ranges reflect different physical properties of the target. To enhance the consistency of the spectral features extracted by the model from different domains, we introduce a source domain spectral alignment (SDSA) strategy to align the spectral ranges and bands of the source domain data to the target domain. SpectralDINO employs the visual foundation model to enhance its ability to generalize cross-domain knowledge. Additionally, we propose a dual mixture-of-subspaces low-rank adaptation (Dual-MoS LoRA) method to address the structural limitation of the low-rank adaptation methods in distinguishing domain-specific features from multi-domain inputs. Only 1.14% of the 21.37M parameters need to be trained to perform fine-tuning. Extensive experimental results on three public datasets demonstrate the superiority of SpectralDINO. Baocheng Chen, Tieqiao Chen, Meng Wang 0038, Jia Liu 0014, Yihao Wang 0003, Renhao Zhang, Bingliang Hu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound DevicesabstractHandheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA. Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Masked Vascular Structure Segmentation and Completion in Retinal ImagesabstractEarly retinal vascular changes in diseases such as diabetic retinopathy often occur at a microscopic level. Accurate evaluation of retinal vascular networks at a micro-level could significantly improve our understanding of angiopathology and potentially aid ophthalmologists in disease assessment and management. Multiple angiogram-related retinal imaging modalities, including fundus, optical coherence tomography angiography, and fluorescence angiography, project continuous, inter-connected retinal microvascular networks into imaging domains. However, extracting the microvascular network, which includes arterioles, venules, and capillaries, is challenging due to the limited contrast and resolution. As a result, the vascular network often appears as fragmented segments. In this paper, we propose a backbone-agnostic Masked Vascular Structure Segmentation and Completion (MaskVSC) method to reconstruct the retinal vascular network. MaskVSC simulates missing sections of blood vessels and uses this simulation to train the model to predict the missing parts and their connections. This approach simulates highly heterogeneous forms of vessel breaks and mitigates the need for massive data labeling. Accordingly, we introduce a connectivity loss function that penalizes interruptions in the vascular network. Our findings show that masking 40% of the segments yields optimal performance in reconstructing the interconnected vascular network. We test our method on three different types of retinal images across five separate datasets. The results demonstrate that MaskVSC outperforms state-of-the-art methods in maintaining vascular network completeness and segmentation accuracy. Furthermore, MaskVSC has been introduced to different segmentation backbones and has successfully improved performance. The code and 2PFM data are available at: https://github.com/Zhouyi-Zura/MaskVSC. Yi Zhou 0024, Thiara Sana Ahmed, Meng Wang 0038, Eric A. Newman, Leopold Schmetterer, Huazhu Fu, Jun Cheng 0003, Bingyao Tan |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Say No to Freeloader: Protecting Intellectual Property of Your Deep ModelabstractModel intellectual property (IP) protection has gained attention due to the significance of safeguarding intellectual labor and computational resources. Ensuring IP safety for trainers and owners is critical, especially when ownership verification and applicability authorization are required. A notable approach involves preventing the transfer of well-trained models from authorized to unauthorized domains. We introduce a novel Compact Un-transferable Pyramid Isolation Domain (CUPI-Domain) which serves as a barrier against illegal transfers from authorized to unauthorized domains. Inspired by human transitive inference, the CUPI-Domain emphasizes distinctive style features of the authorized domain, leading to failure in recognizing irrelevant private style features on unauthorized domains. To this end, we propose CUPI-Domain generators, which select features from both authorized and CUPI-Domain as anchors. These generators fuse the style features and semantic features to create labeled, style-rich CUPI-Domain. Additionally, we design external Domain-Information Memory Banks (DIMB) for storing and updating labeled pyramid features to obtain stable domain class features and domain class-wise style features. Based on the proposed whole method, the novel style and discriminative loss functions are designed to effectively enhance the distinction in style and discriminative features between authorized and unauthorized domains. We offer two solutions for utilizing CUPI-Domain based on whether the unauthorized domain is known: target-specified CUPI-Domain and target-free CUPI-Domain. Comprehensive experiments on various public datasets demonstrate the effectiveness of our CUPI-Domain approach with different backbone models, providing an efficient solution for model intellectual property protection. Lianyu Wang, Meng Wang 0038, Huazhu Fu, Daoqiang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | A Multi-Scale Fusion and Transformer Based Registration Guided Speckle Noise Reduction for OCT ImagesabstractOptical coherence tomography (OCT) images are inevitably affected by speckle noise because OCT is based on low-coherence interference. Multi-frame averaging is one of the effective methods to reduce speckle noise. Before averaging, the misalignment between images must be calibrated. In this paper, in order to reduce misalignment between images caused during the acquisition, a novel multi-scale fusion and Transformer based (MsFTMorph) method is proposed for deformable retinal OCT image registration. The proposed method captures global connectivity and locality with convolutional vision transformer and also incorporates a multi-resolution fusion strategy for learning the global affine transformation. Comparative experiments with other state-of-the-art registration methods demonstrate that the proposed method achieves higher registration accuracy. Guided by the registration, subsequent multi-frame averaging shows better results in speckle noise reduction. The noise is suppressed while the edges can be preserved. In addition, our proposed method has strong cross-domain generalization, which can be directly applied to images acquired by different scanners with different modes. Zhiwei Tan, Yi Zhou 0024, Meng Wang 0038, Ming Liu 0030, Xinjian Chen 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Graph Attention U-Net for Retinal Layer Surface Detection and Choroid Neovascularization Segmentation in OCT ImagesabstractChoroidal neovascularization (CNV) is a typical symptom of age-related macular degeneration (AMD) and is one of the leading causes for blindness. Accurate segmentation of CNV and detection of retinal layers are critical for eye disease diagnosis and monitoring. In this paper, we propose a novel graph attention U-Net (GA-UNet) for retinal layer surface detection and CNV segmentation in optical coherence tomography (OCT) images. Due to retinal layer deformation caused by CNV, it is challenging for existing models to segment CNV and detect retinal layer surfaces with the correct topological order. We propose two novel modules to address the challenge. The first module is a graph attention encoder (GAE) in a U-Net model that automatically integrates topological and pathological knowledge of retinal layers into the U-Net structure to achieve effective feature embedding. The second module is a graph decorrelation module (GDM) that takes reconstructed features by the decoder of the U-Net as inputs, it then decorrelates and removes information unrelated to retinal layer for improved retinal layer surface detection. In addition, we propose a new loss function to maintain the correct topological order of retinal layers and the continuity of their boundaries. The proposed model learns graph attention maps automatically during training and performs retinal layer surface detection and CNV segmentation simultaneously with the attention maps during inference. We evaluated the proposed model on our private AMD dataset and another public dataset. Experiment results show that the proposed model outperformed the competing methods for retinal layer surface detection and CNV segmentation and achieved new state of the arts on the datasets. Yuhe Shen, Jiang Li 0001, Weifang Zhu, Kai Yu 0009, Meng Wang 0038, Yi Zhou 0024, Liling Guan, Xinjian Chen 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Speckle Noise Reduction for OCT Images Based on Image Style Transfer and Conditional GANabstractRaw optical coherence tomography (OCT) images typically are of low quality because speckle noise blurs retinal structures, severely compromising visual quality and degrading performances of subsequent image analysis tasks. In our previous study (Ma et al., 2018), we have developed a Conditional Generative Adversarial Network (cGAN) for speckle noise removal in OCT images collected by several commercial OCT scanners, which we collectively refer to as scanner T. In this paper, we improve the cGAN model and apply it to our in-house OCT scanner (scanner B) for speckle noise suppression. The proposed model consists of two steps: 1) We train a Cycle-Consistent GAN (CycleGAN) to learn style transfer between two OCT image datasets collected by different scanners. The purpose of the CycleGAN is to leverage the ground truth dataset created in our previous study. 2) We train a mini-cGAN model based on the PatchGAN mechanism with the ground truth dataset to suppress speckle noise in OCT images. After training, we first apply the CycleGAN model to convert raw images collected by scanner B to match the style of the images from scanner T, and subsequently use the mini-cGAN model to suppress speckle noise in the style transferred images. We evaluate the proposed method on a dataset collected by scanner B. Experimental results show that the improved model outperforms our previous method and other state-of-the-art models in speckle noise removal, retinal structure preservation and contrast enhancement. Yi Zhou 0024, Kai Yu 0009, Meng Wang 0038, Yuhui Ma, Zhongyue Chen, Weifang Zhu, Xinjian Chen 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | MsTGANet: Automatic Drusen Segmentation From Retinal OCT ImagesabstractDrusen is considered as the landmark for diagnosis of AMD and important risk factor for the development of AMD. Therefore, accurate segmentation of drusen in retinal OCT images is crucial for early diagnosis of AMD. However, drusen segmentation in retinal OCT images is still very challenging due to the large variations in size and shape of drusen, blurred boundaries, and speckle noise interference. Moreover, the lack of OCT dataset with pixel-level annotation is also a vital factor hindering the improvement of drusen segmentation accuracy. To solve these problems, a novel multi-scale transformer global attention network (MsTGANet) is proposed for drusen segmentation in retinal OCT images. In MsTGANet, which is based on U-Shape architecture, a novel multi-scale transformer non-local (MsTNL) module is designed and inserted into the top of encoder path, aiming at capturing multi-scale non-local features with long-range dependencies from different layers of encoder. Meanwhile, a novel multi-semantic global channel and spatial joint attention module (MsGCS) between encoder and decoder is proposed to guide the model to fuse different semantic features, thereby improving the model's ability to learn multi-semantic global contextual information. Furthermore, to alleviate the shortage of labeled data, we propose a novel semi-supervised version of MsTGANet (Semi-MsTGANet) based on pseudo-labeled data augmentation strategy, which can leverage a large amount of unlabeled data to further improve the segmentation performance. Finally, comprehensive experiments are conducted to evaluate the performance of the proposed MsTGANet and Semi-MsTGANet. The experimental results show that our proposed methods achieve better segmentation accuracy than other state-of-the-art CNN-based methods. Meng Wang 0038, Weifang Zhu, Jinzhu Su, Haoyu Chen 0002, Kai Yu 0009, Yi Zhou 0024, Zhongyue Chen, Xinjian Chen 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | High-Resolution Hierarchical Adversarial Learning for OCT Speckle Noise Reduction
Yi Zhou 0024, Jiang Li 0001, Meng Wang 0038, Weifang Zhu, Zhongyue Chen, Lianyu Wang, Chenpu Yao, Xinjian Chen 0001 |
MICCAI (6) | 3 |
| 2021 | Automatic Staging for Retinopathy of Prematurity With Deep Feature Fusion and Ordinal Classification StrategyabstractRetinopathy of prematurity (ROP) is a retinal disease which frequently occurs in premature babies with low birth weight and is considered as one of the major preventable causes of childhood blindness. Although automatic and semi-automatic diagnoses of ROP based on fundus image have been researched, most of the previous studies focused on plus disease detection and ROP screening. There are few studies focusing on ROP staging, which is important for the severity evaluation of the disease. To be consistent with clinical 5-level ROP staging, a novel and effective deep neural network based 5-level ROP staging network is proposed, which consists of multi-stream based parallel feature extractor, concatenation based deep feature fuser and clinical practice based ordinal classifier. First, the three-stream parallel framework including ResNet18, DenseNet121 and EfficientNetB2 is proposed as the feature extractor, which can extract rich and diverse high-level features. Second, the features from three streams are deeply fused by concatenation and convolution to generate a more effective and comprehensive feature. Finally, in the classification stage, an ordinal classification strategy is adopted, which can effectively improve the ROP staging performance. The proposed ROP staging network was evaluated with per-image and per-examination strategies. For per-image ROP staging, the proposed method was evaluated on 635 retinal fundus images from 196 examinations, including 303 Normal, 26 Stage 1, 127 Stage 2, 106 Stage 3, 61 Stage 4 and 12 Stage 5, which achieves 0.9055 for weighted recall, 0.9092 for weighted precision, 0.9043 for weighted F1 score, 0.9827 for accuracy with 1 (ACC1) and 0.9786 for Kappa, respectively. While for per-examination ROP staging, 1173 examinations with a 4-fold cross validation strategy were used to evaluate the effectiveness of the proposed method, which prove the validity and advantage of the proposed method. Weifang Zhu, Zhongyue Chen, Meng Wang 0038, Le Geng, Kai Yu 0009, Yi Zhou 0024, Daoman Xiang, Xinjian Chen 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Semi-Supervised Capsule cGAN for Speckle Noise Reduction in Retinal OCT ImagesabstractSpeckle noise is the main cause of poor optical coherence tomography (OCT) image quality. Convolutional neural networks (CNNs) have shown remarkable performances for speckle noise reduction. However, speckle noise denoising still meets great challenges because the deep learning-based methods need a large amount of labeled data whose acquisition is time-consuming or expensive. Besides, many CNNs-based methods design complex structure based networks with lots of parameters to improve the denoising performance, which consume hardware resources severely and are prone to overfitting. To solve these problems, we propose a novel semi-supervised learning based method for speckle noise denoising in retinal OCT images. First, to improve the model's ability to capture complex and sparse features in OCT images, and avoid the problem of a great increase of parameters, a novel capsule conditional generative adversarial network (Caps-cGAN) with small number of parameters is proposed to construct the semi-supervised learning system. Then, to tackle the problem of retinal structure information loss in OCT images caused by lack of detailed guidance during unsupervised learning, a novel joint semi-supervised loss function composed of unsupervised loss and supervised loss is proposed to train the model. Compared with other state-of-the-art methods, the proposed semi-supervised method is suitable for retinal OCT images collected from different OCT devices and can achieve better performance even only using half of the training data. Meng Wang 0038, Weifang Zhu, Kai Yu 0009, Zhongyue Chen, Yi Zhou 0024, Yuhui Ma, Dengsen Bao, Shuanglang Feng, Dehui Xiang, Xinjian Chen 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2020 | CPFNet: Context Pyramid Fusion Network for Medical Image SegmentationabstractAccurate and automatic segmentation of medical images is a crucial step for clinical diagnosis and analysis. The convolutional neural network (CNN) approaches based on the U-shape structure have achieved remarkable performances in many different medical image segmentation tasks. However, the context information extraction capability of single stage is insufficient in this structure, due to the problems such as imbalanced class and blurred boundary. In this paper, we propose a novel Context Pyramid Fusion Network (named CPFNet) by combining two pyramidal modules to fuse global/multi-scale context information. Based on the U-shape structure, we first design multiple global pyramid guidance (GPG) modules between the encoder and the decoder, aiming at providing different levels of global context information for the decoder by reconstructing skip-connection. We further design a scale-aware pyramid fusion (SAPF) module to dynamically fuse multi-scale context information in high-level features. These two pyramidal modules can exploit and fuse rich context information progressively. Experimental results show that our proposed method is very competitive with other state-of-the-art methods on four different challenging tasks, including skin lesion segmentation, retinal linear lesion segmentation, multi-class segmentation of thoracic organs at risk and multi-class segmentation of retinal edema lesions. Shuanglang Feng, Heming Zhao, Xuena Cheng, Meng Wang 0038, Yuhui Ma, Dehui Xiang, Weifang Zhu, Xinjian Chen 0001 |
IEEE Trans. Medical Imaging | 5 |