VLDB 2026 Research / reviewers in the wild / expert
Dong Wei 0004
dblp:34/4292-4
· DBLP profile ↗
54ranked-venue papers
5as first author
46since 2021 · last 2026
0000-0001-5969-6987ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 41 · 4 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Text-Modulated Semantic Guidance for Low-Light Endoscopic Image EnhancementabstractLow light conditions in endoscopic imaging would lead to poor visibility, reduced contrast, and increased noise, which may hinder accurate diagnosis and surgical guidance. Against this low-light endoscopic image enhancement (LLEIE) task, inspired by the remarkable performance of pretrained CLIP in downstream vision tasks, in this paper, we carefully investigate the pretrained priors of CLIP and embed them into a text-modulated semantic-aware discriminator (TMSD). Through the adversarial learning mechanism, the discriminator can be easily integrated into different low-light enhancement baselines for helping them accomplish better visual restoration effects without incurring any extra inference cost. Specifically, to make the foundation model CLIP suitable for the LLEIE task, we initially propose a prompt learning procedure to obtain the text embedding and image semantics corresponding to the normal-light endoscopic imaging scenario. Building upon the acquired text prior and image semantic priors, we devise a text modulator to synergize these two priors, yielding a richer semantic representation. Leveraging the convolutional modulation and cross-attention mechanisms, we blend this semantic guidance information into the discriminator, thereby fostering the fine-grained distribution learning of normal-light endoscopic images in visual semantics and guiding different enhancement baselines achieving higher visual quality. Based on five public benchmark datasets, including three synthetic datasets, one real clinical dataset, and one clinical downstream segmentation dataset, we comprehensively evaluate the effectiveness of our proposed TMSD. Extensive experiments substantiate that the integration of the proposed TMSD enables seven representative baselines to obtain better perceptual quality, especially in the cross-domain clinical generalization scenario. Besides, the downstream segmentation accuracy can be evidently improved, showing the favorable application potential of the proposed TMSD. Moreover, to comprehensively evaluate the generality of our TMSD framework, we successfully apply it to a new and classic metal artifact reduction task. It is worth mentioning that our TMSD does not incur any extra computational cost during inference. Hong Wang 0021, Zhijian Wu, Haodu Fang, Dong Wei 0004, Jinghan Sun, Yefeng Zheng 0001, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | GM-ABS: Promptable Generalist Model Drives Active Barely Supervised Training in Specialist Model for 3D Medical Image SegmentationabstractSemi-supervised learning (SSL) has greatly advanced 3D medical image segmentation by alleviating the need for intensive labeling by radiologists. While previous efforts focused on model-centric advancements, the emergence of foundational generalist models like the Segment Anything Model (SAM) is expected to reshape the SSL landscape. Although these generalists usually show performance gaps relative to previous specialists in medical imaging, they possess impressive zero-shot segmentation abilities with manual prompts. Thus, this capability could serve as "free lunch" for training specialists, offering future SSL a promising data-centric perspective, especially revolutionizing both pseudo and expert labeling strategies to enhance the data pool. In this regard, we propose the Generalist Model-driven Active Barely Supervised (GM-ABS) learning paradigm, for developing specialized 3D segmentation models under extremely limited (barely) annotation budgets, e.g., merely cross-labeling three slices per selected scan. In specific, building upon a basic mean-teacher SSL framework, GM-ABS modernizes the SSL paradigm with two key data-centric designs: (i) Specialist-generalist collaboration, where the in-training specialist leverages class-specific positional prompts derived from class prototypes to interact with the frozen class-agnostic generalist across multiple views to achieve noisy-yet-effective label augmentation. Then, the specialist robustly assimilates the augmented knowledge via noise-tolerant collaborative learning. (ii) Expert-model collaboration that promotes active cross-labeling with notably low labeling efforts. This design progressively furnishes the specialist with informative and efficient supervision via a human-in-the-loop manner, which in turn benefits the quality of class-specific prompts. Extensive experiments on three benchmark datasets highlight the promising performance of GM-ABS over recent SSL approaches under extremely constrained labeling resources. Zhe Xu 0012, Cheng Chen 0013, Donghuan Lu, Jinghan Sun, Dong Wei 0004, Yefeng Zheng 0001, Quanzheng Li, Raymond Kai-Yu Tong |
IEEE Trans. Medical Imaging | 5 |
| 2025 | UNiFS: Unified Multi-Contrast MRI Reconstruction via Frequency-Spatial FusionabstractRecently, Multi-Contrast MR Reconstruction (MCMR) has emerged as a hot research topic that leverages high-quality auxiliary modalities to reconstruct undersampled target modalities of interest. However, existing methods often struggle to generalize across different k-space undersampling patterns, requiring the training of a separate model for each specific pattern, which limits their practical applicability. To address this challenge, we propose UniFS, a Unified Frequency-Spatial Fusion model designed to handle multiple k-space undersampling patterns for MCMR tasks without any need for retraining. UniFS integrates three key modules: a Cross-Modal Frequency Fusion module, an Adaptive Mask-Based Prompt Learning module, and a Dual-Branch Complementary Refinement module. These modules work together to extract domain-invariant features from diverse k-space undersampling patterns while dynamically adapt to their own variations. Another limitation of existing MCMR methods is their tendency to focus solely on spatial information while neglect frequency characteristics, or extract only shallow frequency features, thus failing to fully leverage complementary cross-modal frequency information. To relieve this issue, UniFS introduces an adaptive prompt-guided frequency fusion module for k-space learning, significantly enhancing the model's generalization performance. We evaluate our model on the BraTS and HCP datasets with various k-space undersampling patterns and acceleration factors, including previously unseen patterns, to comprehensively assess UniFS's generalizability. Experimental results across multiple scenarios demonstrate that UniFS achieves state-of-the-art performance. Our code is available at https://github.com/LIKP0/UniFS. Yiwei Ren, Kai Pan, Dong Wei 0004, Pujin Cheng, Xian Wu 0001, Xiaoying Tang 0001 |
BIBM | 4 |
| 2025 | Improving Instance-Based Whole Slide Image Classification with Logit-Based Log-Sum-Exp Aggregator and Contextual AwarenessabstractCancer has become a leading cause of death worldwide, making the development of intelligent and automatic whole slide image (WSI) analysis tools crucial for diagnosis and treatment decision-making. However, the gigapixel size of WSIs poses significant challenges for annotation and analysis, motivating researchers to develop both label- and computational-efficient algorithms. While existing instance-based methods have shown promise in computational efficiency and patch-wise prediction, they often struggle with classification performance and lesion localization capabilities on pathological images. In this paper, we delve into the limitations of current instance-based approaches and attribute such inferior performance to (i) the overcontribution of normal patches and (ii) the absence of contextual information. To this end, we propose a simple yet effective logit-based log-sum-exp aggregator to modulate the contribution of normal patches and highlight the tumorous patches' contribution in the slide-wise prediction, and introduce a context-aware feature extraction module to capture the contextual patterns from neighborhood patches. Our method based on the above two components showcases the superior classification performance and lesion localization ability with low computational complexity on CAMELYON16, TCGA-NSCLC, and BRACS compared to existing methods. Wentao Pan 0001, Donghuan Lu, Jiangpeng Yan, Zhe Xu 0012, Conghao Xiong, Dong Wei 0004, Xian Wu 0001, Yixuan Yuan |
BIBM | 6 |
| 2025 | OpenDUN: To Discover Unknown Number of Visual CategoriesabstractOpen-Set methods have relaxed the underlying assumption made by most image recognition studies that all samples in the test and training datasets belong to the same classes by considering only a part of classes are known in training dataset. However, most of these approaches require a known or predefined number of novel classes, which is often not the case in real applications. In this study, we aim at a more difficult but practical scenario, where the number of novel classes is unknown. By merging the unlabeled samples into clusters instead of directly assigning categorical labels to them, the proposed end-to-end framework can simultaneously estimate the number of novel classes and learn the appropriate division of unlabeled samples. In addition, a cluster scattering strategy is introduced such that the erroneous merging can be alleviated. Comprehensive experiments on three benchmark datasets are conducted to demonstrate the superiority of the proposed method in both the estimation of the novel class number and the classification of unlabeled samples. Sik Chit Wu, Munan Ning, Dong Wei 0004, Yefeng Zheng 0001, Donghuan Lu, Li Yuan 0007 |
ICME | 3 |
| 2025 | RRG-DPO: Direct Preference Optimization for Clinically Accurate Radiology Report Generation
Dong Wei 0004, Zhe Xu 0012, Xian Wu 0001, Yefeng Zheng 0001, Liansheng Wang 0002 |
MICCAI (5) | 2 |
| 2025 | D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide ImagesabstractDiffusion-based virtual staining methods of histopathology images have demonstrated outstanding potential for stain normalization and cross-dye staining (e.g., hematoxylin-eosin to immunohistochemistry). However, achieving pathology-correct cross-dye virtual staining with versatile tone controls poses significant challenges due to the difficulty of decoupling the given pathology and tone conditions. This issue would cause non-pathologic regions to be mistakenly stained like pathologic ones, and vice versa, which we term “pathology leakage.” To address this issue, we propose diffusion virtual staining Transformer (D-VST), a new framework with versatile tone control for cross-dye virtual staining. Specifically, we introduce a pathology encoder in conjunction with a tone encoder, combined with a two-stage curriculum learning scheme that decouples pathology and tone conditions, to enable tone control while eliminating pathology leakage. Further, to extend our method for billion-pixel whole slide image (WSI) staining, we introduce a novel frequency-aware adaptive patch sampling strategy for high-quality yet efficient inference of ultra-high resolution images in a zero-shot manner. Integrating these two innovative components facilitates a pathology-correct, tone-controllable, cross-dye WSI virtual staining process. Extensive experiments on three virtual staining tasks that involve translating between four different dyes demonstrate the superiority of our approach in generating high-quality and pathologically accurate images compared to existing methods based on generative adversarial networks and diffusion models. Our code and trained models will be released. Shurong Yang, Dong Wei 0004, Yihuang Hu, Qiong Peng, Yawen Huang, Xian Wu 0001, Yefeng Zheng 0001, Liansheng Wang 0002 |
NeurIPS | 2 |
| 2025 | Federated modality-specific encoders and partially personalized fusion decoder for multimodal brain tumor segmentation
Dong Wei 0004, Qian Dai, Xian Wu 0001, Yefeng Zheng 0001, Liansheng Wang 0002 |
Medical Image Anal. | 2 |
| 2025 | Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report GenerationabstractAnatomical abnormality detection and report generation of chest X-ray (CXR) are two essential tasks in clinical practice. The former aims at localizing and characterizing cardiopulmonary radiological findings in CXRs, while the latter summarizes the findings in a detailed report for further diagnosis and treatment. Existing methods often focused on either task separately, ignoring their correlation. This work proposes a co-evolutionary abnormality detection and report generation (CoE-DG) framework. The framework utilizes both fully labeled (with bounding box annotations and clinical reports) and weakly labeled (with reports only) data to achieve mutual promotion between the abnormality detection and report generation tasks. Specifically, we introduce a bi-directional information interaction strategy with generator-guided information propagation (GIP) and detector-guided information propagation (DIP). For semi-supervised abnormality detection, GIP takes the informative feature extracted by the generator as an auxiliary input to the detector and uses the generator's prediction to refine the detector's pseudo labels. We further propose an intra-image-modal self-adaptive non-maximum suppression module (SA-NMS). This module dynamically rectifies pseudo detection labels generated by the teacher detection model with high-confidence predictions by the student. Inversely, for report generation, DIP takes the abnormalities' categories and locations predicted by the detector as input and guidance for the generator to improve the generated reports. Finally, a co-evolutionary training strategy is implemented to iteratively conduct GIP and DIP and consistently improve both tasks' performance. Experimental results on two public CXR datasets demonstrate CoE-DG's superior performance to several up-to-date object detection, report generation, and unified models. Our code is available at https://github.com/jinghanSunn/CoE-DG. Jinghan Sun, Dong Wei 0004, Zhe Xu 0012, Donghuan Lu, Hong Wang 0021, Sotirios A. Tsaftaris, Steven McDonagh 0001, Yefeng Zheng 0001, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Adaptive Weighting Based Metal Artifact Reduction in CT ImagesabstractAgainst the metal artifact reduction (MAR) task in computed tomography (CT) imaging, most of the existing deep-learning-based approaches generally select a single Hounsfield unit (HU) window followed by a normalization operation to preprocess CT images. However, in practical clinical scenarios, different body tissues and organs are often inspected under varying window settings for good contrast. The methods trained on a fixed single window would lead to insufficient removal of metal artifacts when being transferred to deal with other windows. To alleviate this problem, few works have proposed to reconstruct the CT images under multiple-window configurations. Albeit achieving good reconstruction performance for different windows, they adopt to directly supervise each window learning in an equal weighting way based on the training set. To improve the learning flexibility and model generalizability, in this paper, we propose an adaptive weighting algorithm, called AdaW, for the multiple-window metal artifact reduction, which can be applied to different deep MAR network backbones. Specifically, we first formulate the multiple window learning task as a bi-level optimization problem. Then we derive an adaptive weighting optimization algorithm where the learning process for MAR under each window is automatically weighted via a learning-to-learn paradigm based on the training set and validation set. This rationality is finely substantiated through theoretical analysis. Based on different network backbones, experimental comparisons executed on five datasets with different body sites comprehensively validate the effectiveness of AdaW in helping improve the generalization performance as well as its good applicability. We will release the code at https://github.com/hongwang01/AdaW. Hong Wang 0021, Dong Wei 0004, Xian Wu 0001, Jianhua Ma 0001, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Federated Modality-Specific Encoders and Multimodal Anchors for Personalized Brain Tumor SegmentationabstractMost existing federated learning (FL) methods for medical image analysis only considered intramodal heterogeneity, limiting their applicability to multimodal imaging applications. In practice, it is not uncommon that some FL participants only possess a subset of the complete imaging modalities, posing inter-modal heterogeneity as a challenge to effectively training a global model on all participants’ data. In addition, each participant would expect to obtain a personalized model tailored for its local data characteristics from the FL in such a scenario. In this work, we propose a new FL framework with federated modality-specific encoders and multimodal anchors (FedMEMA) to simultaneously address the two concurrent issues. Above all, FedMEMA employs an exclusive encoder for each modality to account for the inter-modal heterogeneity in the first place. In the meantime, while the encoders are shared by the participants, the decoders are personalized to meet individual needs. Specifically, a server with full-modal data employs a fusion decoder to aggregate and fuse representations from all modality-specific encoders, thus bridging the modalities to optimize the encoders via backpropagation reversely. Meanwhile, multiple anchors are extracted from the fused multimodal representations and distributed to the clients in addition to the encoder parameters. On the other end, the clients with incomplete modalities calibrate their missing-modal representations toward the global full-modal anchors via scaled dot-product cross-attention, making up the information loss due to absent modalities while adapting the representations of present ones. FedMEMA is validated on the BraTS 2020 benchmark for multimodal brain tumor segmentation. Results show that it outperforms various up-to-date methods for multimodal and personalized FL and that its novel designs are effective. Our code is available. Qian Dai, Dong Wei 0004, Jinghan Sun, Liansheng Wang 0002, Yefeng Zheng 0001 |
AAAI | 2 |
| 2024 | Learning to Segment Multiple Organs from Multimodal Partially Labeled Datasets
Dong Wei 0004, Donghuan Lu, Jinghan Sun, Hao Zheng 0008, Yefeng Zheng 0001, Liansheng Wang 0002 |
MICCAI (9) | 2 |
| 2024 | MoME: Mixture of Multimodal Experts for Cancer Survival Prediction
Conghao Xiong, Hao Chen 0011, Hao Zheng 0008, Dong Wei 0004, Yefeng Zheng 0001, Joseph J. Y. Sung, Irwin King |
MICCAI (4) | 4 |
| 2024 | TAKT: Target-Aware Knowledge Transfer for Whole Slide Image Classification
Conghao Xiong, Yi Lin 0009, Hao Chen 0011, Hao Zheng 0008, Dong Wei 0004, Yefeng Zheng 0001, Joseph J. Y. Sung, Irwin King |
MICCAI (4) | 5 |
| 2024 | FM-ABS: Promptable Foundation Model Drives Active Barely Supervised Learning for 3D Medical Image Segmentation
Zhe Xu 0012, Cheng Chen 0013, Donghuan Lu, Jinghan Sun, Dong Wei 0004, Yefeng Zheng 0001, Quanzheng Li, Raymond Kai-Yu Tong |
MICCAI (8) | 5 |
| 2024 | Triplet-branch network with contrastive prior-knowledge embedding for disease grading
Yuexiang Li, Yawen Huang, Jingxin Liu 0005, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001 |
Artif. Intell. Medicine | 7 |
| 2024 | Simultaneous alignment and surface regression using hybrid 2D-3D networks for 3D coherent layer segmentation of retinal OCT images with full and sparse annotations
Dong Wei 0004, Donghuan Lu, Xiaoying Tang 0001, Liansheng Wang 0002, Yefeng Zheng 0001 |
Medical Image Anal. | 2 |
| 2024 | Hybrid unsupervised representation learning and pseudo-label supervised self-distillation for rare disease imaging phenotype classification with dispersion-aware imbalance correction
Jinghan Sun, Dong Wei 0004, Liansheng Wang 0002, Yefeng Zheng 0001 |
Medical Image Anal. | 2 |
| 2024 | Self-supervised learning for medical image data with anatomy-oriented imaging planes
Dong Wei 0004, Mengmeng Zhua, Shi Gu, Yefeng Zheng 0001 |
Medical Image Anal. | 2 |
| 2024 | Relational Experience Replay: Continual Learning by Adaptively Tuning Task-Wise RelationshipabstractContinual learning is a promising machine learning paradigm to learn new tasks while retaining previously learned knowledge over streaming training data. Till now,rehearsal-basedmethods, keeping a small part of data from old tasks as a memory buffer, have shown good performance in mitigating catastrophic forgetting for previously learned knowledge. However, most of these methods typically treat each new task equally, which may not adequately consider the relationship or similarity between old and new tasks. Furthermore, these methods commonly neglect sample importance in the continual training process and result in sub-optimal performance on certain tasks. To address this challenging problem, we propose Relational Experience Replay (RER), a bi-level learning framework, to adaptively tune task-wise relationships and sample importance within each task to achieve a better ‘stability’ and ‘plasticity’ trade-off. As such, the proposed method is capable of accumulating new knowledge while consolidating previously learned old knowledge during continual learning. Extensive experiments conducted on three benchmark image datasets (CIFAR-10, CIFAR-100, and Tiny ImageNet) and two text datasets (20News and DBpedia) show that the proposed method can consistently improve the performance of all baselines and surpass current state-of-the-art methods. Quanziang Wang, Renzhen Wang, Yuexiang Li, Dong Wei 0004, Hong Wang 0021, Kai Ma 0002, Yefeng Zheng 0001, Deyu Meng |
IEEE Trans. Multim. | 4 |
| 2023 | M3AE: Multimodal Representation Learning for Brain Tumor Segmentation with Missing ModalitiesabstractMultimodal magnetic resonance imaging (MRI) provides complementary information for sub-region analysis of brain tumors. Plenty of methods have been proposed for automatic brain tumor segmentation using four common MRI modalities and achieved remarkable performance. In practice, however, it is common to have one or more modalities missing due to image corruption, artifacts, acquisition protocols, allergy to contrast agents, or simply cost. In this work, we propose a novel two-stage framework for brain tumor segmentation with missing modalities. In the first stage, a multimodal masked autoencoder (M3AE) is proposed, where both random modalities (i.e., modality dropout) and random patches of the remaining modalities are masked for a reconstruction task, for self-supervised learning of robust multimodal representations against missing modalities. To this end, we name our framework M3AE. Meanwhile, we employ model inversion to optimize a representative full-modal image at marginal extra cost, which will be used to substitute for the missing modalities and boost performance during inference. Then in the second stage, a memory-efficient self distillation is proposed to distill knowledge between heterogenous missing-modal situations while fine-tuning the model for supervised segmentation. Our M3AE belongs to the ‘catch-all’ genre where a single model can be applied to all possible subsets of modalities, thus is economic for both training and deployment. Extensive experiments on BraTS 2018 and 2020 datasets demonstrate its superior performance to existing state-of-the-art methods with missing modalities, as well as the efficacy of its components. Our code is available at: https://github.com/ccarliu/m3ae. Dong Wei 0004, Donghuan Lu, Jinghan Sun, Liansheng Wang 0002, Yefeng Zheng 0001 |
AAAI | 2 |
| 2023 | You've Got Two Teachers: Co-evolutionary Image and Report Distillation for Semi-supervised Anatomical Abnormality Detection in Chest X-Ray
Jinghan Sun, Dong Wei 0004, Zhe Xu 0012, Donghuan Lu, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (1) | 2 |
| 2023 | MEPNet: A Model-Driven Equivariant Proximal Network for Joint Sparse-View Reconstruction and Metal Artifact Reduction in CT Images
Hong Wang 0021, Dong Wei 0004, Yuexiang Li, Yefeng Zheng 0001 |
MICCAI (10) | 3 |
| 2023 | Category-Level Regularized Unlabeled-to-Labeled Learning for Semi-supervised Prostate Segmentation with Multi-site Unlabeled Data
Zhe Xu 0012, Donghuan Lu, Jiangpeng Yan, Jinghan Sun, Jie Luo 0003, Dong Wei 0004, Sarah F. Frisken, Quanzheng Li, Yefeng Zheng 0001, Raymond Kai-Yu Tong |
MICCAI (4) | 6 |
| 2023 | A Model-Agnostic Framework for Universal Anomaly Detection of Multi-organ and Multi-modal Images
Donghuan Lu, Munan Ning, Liansheng Wang 0002, Dong Wei 0004, Yefeng Zheng 0001 |
MICCAI (3) | 5 |
| 2023 | A deep weakly semi-supervised framework for endoscopic lesion segmentation
Hong Wang 0021, Haoqin Ji, Yuexiang Li, Nanjun He, Dong Wei 0004, Yawen Huang, Xinrong Chen, Yefeng Zheng 0001, Hongmeng Yu |
Medical Image Anal. | 7 |
| 2023 | MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation SegmentationabstractUnsupervised domain adaption has been widely adopted in tasks with scarce annotated data. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data, leading to inferior performance. To address this issue, we first propose to introduce active sample selection to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, both source and target domains can be better characterized as multimodal distributions, in which way more complementary and informative samples are selected from the target domain. With only a little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, achieving a large performance gain. In addition, a powerful semi-supervised domain adaptation strategy is proposed to alleviate the long-tail distribution problem and further improve the segmentation performance. Extensive experiments are conducted on public datasets, and the results demonstrate that the proposed approach outperforms state-of-the-art methods by large margins and achieves similar performance to the fully-supervised upperbound, i.e., 71.4% mIoU on GTA5 and 71.8% mIoU on SYNTHIA. The effectiveness of each component is also verified by thorough ablation studies. Munan Ning, Donghuan Lu, Yujia Xie, Dongdong Chen 0001, Dong Wei 0004, Yefeng Zheng 0001, Yonghong Tian 0001, Shuicheng Yan, Li Yuan 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Improving Medical Vision-Language Contrastive Pretraining With Semantics-Aware TriageabstractMedical contrastive vision-language pretraining has shown great promise in many downstream tasks, such as data-efficient/zero-shot recognition. Current studies pretrain the network with contrastive loss by treating the paired image-reports as positive samples and the unpaired ones as negative samples. However, unlike natural datasets, many medical images or reports from different cases could have large similarity especially for the normal cases, and treating all the unpaired ones as negative samples could undermine the learned semantic structure and impose an adverse effect on the representations. Therefore, we design a simple yet effective approach for better contrastive learning in medical vision-language field. Specifically, by simplifying the computation of similarity between medical image-report pairs into the calculation of the inter-report similarity, the image-report tuples are divided into positive, negative, and additional neutral groups. With this better categorization of samples, more suitable contrastive loss is constructed. For evaluation, we perform extensive experiments by applying the proposed model-agnostic strategy to two state-of-the-art pretraining frameworks. The consistent improvements on four common downstream tasks, including cross-modal retrieval, zero-shot/data-efficient image classification, and image segmentation, demonstrate the effectiveness of the proposed strategy in medical field. Bo Liu 0113, Donghuan Lu, Dong Wei 0004, Xian Wu 0001, Yan Wang 0015, Yu Zhang 0185, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Boost Supervised Pretraining for Visual Transfer Learning: Implications of Self-Supervised Contrastive Representation LearningabstractUnsupervised pretraining based on contrastive learning has made significant progress recently and showed comparable or even superior transfer learning performance to traditional supervised pretraining on various tasks. In this work, we first empirically investigate when and why unsupervised pretraining surpasses supervised counterparts for image classification tasks with a series of control experiments. Besides the commonly used accuracy, we further analyze the results qualitatively with the class activation maps and assess the learned representations quantitatively with the representation entropy and uniformity. Our core finding is that it is the amount of information effectively perceived by the learning model that is crucial to transfer learning, instead of absolute size of the dataset. Based on this finding, we propose Classification Activation Map guided contrastive (CAMtrast) learning which better utilizes the label supervsion to strengthen supervised pretraining, by making the networks perceive more information from the training images. CAMtrast is evaluated with three fundamental visual learning tasks: image recognition, object detection, and semantic segmentation, on various public datasets. Experimental results show that our CAMtrast effectively improves the performance of supervised pretraining, and that its performance is superior to both unsupervised counterparts and a recent related work which similarly attempted improving supervised pretraining. Jinghan Sun, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
AAAI | 2 |
| 2022 | Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation
Xinyu Shi 0003, Dong Wei 0004, Yu Zhang 0185, Donghuan Lu, Munan Ning, Jiashun Chen, Kai Ma 0002, Yefeng Zheng 0001 |
ECCV (20) | 2 |
| 2022 | Deformer: Towards Displacement Field Learning for Unsupervised Medical Image Registration
Jiashun Chen, Donghuan Lu, Yu Zhang 0185, Dong Wei 0004, Munan Ning, Xinyu Shi 0003, Zhe Xu 0012, Yefeng Zheng 0001 |
MICCAI (6) | 4 |
| 2022 | Point Beyond Class: A Benchmark for Weakly Semi-supervised Abnormality Localization in Chest X-Rays
Haoqin Ji, Yuexiang Li, Jinheng Xie, Nanjun He, Yawen Huang, Dong Wei 0004, Xinrong Chen, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (3) | 7 |
| 2022 | Lesion Guided Explainable Few Weak-Shot Medical Report Generation
Jinghan Sun, Dong Wei 0004, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (5) | 2 |
| 2022 | An Inclusive Task-Aware Framework for Radiology Report Generation
Lin Wang 0026, Munan Ning, Donghuan Lu, Dong Wei 0004, Yefeng Zheng 0001, Jie Chen 0001 |
MICCAI (8) | 4 |
| 2022 | Denoising for Relaxing: Unsupervised Domain Adaptive Fundus Image Segmentation Without Source Data
Zhe Xu 0012, Donghuan Lu, Yixin Wang 0003, Jie Luo 0003, Dong Wei 0004, Yefeng Zheng 0001, Raymond Kai-Yu Tong |
MICCAI (5) | 5 |
| 2022 | Multiscale Unsupervised Retinal Edema Area Segmentation in OCT Images
Wenguang Yuan, Donghuan Lu, Dong Wei 0004, Munan Ning, Yefeng Zheng 0001 |
MICCAI (2) | 3 |
| 2022 | mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001 |
MICCAI (5) | 5 |
| 2022 | Mix-and-Interpolate: A Training Strategy to Deal With Source-Biased Medical DataabstractTill March 31st, 2021, the coronavirus disease 2019 (COVID-19) had reportedly infected more than 127 million people and caused over 2.5 million deaths worldwide. Timely diagnosis of COVID-19 is crucial for management of individual patients as well as containment of the highly contagious disease. Having realized the clinical value of non-contrast chest computed tomography (CT) for diagnosis of COVID-19, deep learning (DL) based automated methods have been proposed to aid the radiologists in reading the huge quantities of CT exams as a result of the pandemic. In this work, we address an overlooked problem for training deep convolutional neural networks for COVID-19 classification using real-world multi-source data, namely, the data source bias problem. The data source bias problem refers to the situation in which certain sources of data comprise only a single class of data, and training with such source-biased data may make the DL models learn to distinguish data sources instead of COVID-19. To overcome this problem, we propose MIx-aNd-Interpolate (MINI), a conceptually simple, easy-to-implement, efficient yet effective training strategy. The proposed MINI approach generates volumes of the absent class by combining the samples collected from different hospitals, which enlarges the sample space of the original source-biased dataset. Experimental results on a large collection of real patient data (1,221 COVID-19 and 1,520 negative CT images, and the latter consisting of 786 community acquired pneumonia and 734 non-pneumonia) from eight hospitals and health institutions show that: 1) MINI can improve COVID-19 classification performance upon the baseline (which does not deal with the source bias), and 2) MINI is superior to competing methods in terms of the extent of improvement. Yuexiang Li, Jiawei Chen 0009, Dong Wei 0004, Yanchun Zhu, Junfeng Xiong, Yadong Gang, Tianyi Qian, Kai Ma 0002, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Domain Adaptation Meets Zero-Shot Learning: An Annotation-Efficient Approach to Multi-Modality Medical Image SegmentationabstractDue to the lack of properly annotated medical data, exploring the generalization capability of the deep model is becoming a public concern. Zero-shot learning (ZSL) has emerged in recent years to equip the deep model with the ability to recognize unseen classes. However, existing studies mainly focus on natural images, which utilize linguistic models to extract auxiliary information for ZSL. It is impractical to apply the natural image ZSL solutions directly to medical images, since the medical terminology is very domain-specific, and it is not easy to acquire linguistic models for the medical terminology. In this work, we propose a new paradigm of ZSL specifically for medical images utilizing cross-modality information. We make three main contributions with the proposed paradigm. First, we extract the prior knowledge about the segmentation targets, called relation prototypes, from the prior model and then propose a cross-modality adaptation module to inherit the prototypes to the zero-shot model. Second, we propose a relation prototype awareness module to make the zero-shot model aware of information contained in the prototypes. Last but not least, we develop an inheritance attention module to recalibrate the relation prototypes to enhance the inheritance process. The proposed framework is evaluated on two public cross-modality datasets including a cardiac dataset and an abdominal dataset. Extensive experiments show that the proposed framework significantly outperforms the state of the arts. Cheng Bian, Chenglang Yuan, Kai Ma 0002, Dong Wei 0004, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Alternative Baselines for Low-Shot 3D Medical Image Segmentation - An Atlas PerspectiveabstractLow-shot (one/few-shot) segmentation has attracted increasing attention as it works well with limited annotation. State-of-the-art low-shot segmentation methods on natural images usually focus on implicit representation learning for each novel class, such as learning prototypes, deriving guidance features via masked average pooling, and segmenting using cosine similarity in feature space. We argue that low-shot segmentation on medical images should step further to explicitly learn dense correspondences between images to utilize the anatomical similarity. The core ideas are inspired by the classical practice of multi-atlas segmentation, where the indispensable parts of atlas-based segmentation, i.e., registration, label propagation, and label fusion are unified into a single framework in our work. Specifically, we propose two alternative baselines, i.e., the Siamese-Baseline and Individual-Difference-Aware Baseline, where the former is targeted at anatomically stable structures (such as brain tissues), and the latter possesses a strong generalization ability to organs suffering large morphological variations (such as abdominal organs). In summary, this work sets up a benchmark for low-shot 3D medical image segmentation and sheds light on further understanding of atlas-based few-shot segmentation. Shilei Cao 0001, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Deyu Meng, Yefeng Zheng 0001 |
AAAI | 3 |
| 2021 | Multi-Anchor Active Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaption has proven to be an effective approach for alleviating the intensive workload of manual annotation by aligning the synthetic source-domain data and the real-world target-domain samples. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data. To this end, we firstly propose to introduce a novel multi-anchor based active learning strategy to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, the source domain can be better characterized as a multimodal distribution, thus more representative and complimentary samples are selected from the target domain. With little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, resulting in a large performance gain. The multi-anchor strategy is additionally employed to model the target-distribution. By regularizing the latent representation of the target samples compact around multiple anchors through a novel soft alignment loss, more precise segmentation can be achieved. Extensive experiments are conducted on public datasets to demonstrate that the proposed approach outperforms state-of-the-art methods significantly, along with thorough ablation study to verify the effectiveness of each component. The code will be released soon at https://github.com/munanning/MADA. Munan Ning, Donghuan Lu, Dong Wei 0004, Cheng Bian, Chenglang Yuan, Kai Ma 0002, Yefeng Zheng 0001 |
ICCV | 3 |
| 2021 | Triplet-Branch Network with Prior-Knowledge Embedding for Fatigue Fracture Grading
Yuexiang Li, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001 |
MICCAI (5) | 5 |
| 2021 | Simultaneous Alignment and Surface Regression Using Hybrid 2D-3D Networks for 3D Coherent Layer Segmentation of Retina OCT Images
Dong Wei 0004, Donghuan Lu, Yuexiang Li, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (8) | 2 |
| 2021 | Unsupervised Representation Learning Meets Pseudo-Label Supervised Self-Distillation: A New Approach to Rare Disease Classification
Jinghan Sun, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (5) | 2 |
| 2021 | Training Automatic View Planner for Cardiac MR Imaging via Self-supervision by Spatial Relationship Between Views
Dong Wei 0004, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (6) | 1 |
| 2021 | A Unified Framework for Generalized Low-Shot Medical Image Segmentation With Scarce DataabstractMedical image segmentation has achieved remarkable advancements using deep neural networks (DNNs). However, DNNs often need big amounts of data and annotations for training, both of which can be difficult and costly to obtain. In this work, we propose a unified framework for generalized low-shot (one- and few-shot) medical image segmentation based on distance metric learning (DML). Unlike most existing methods which only deal with the lack of annotations while assuming abundance of data, our framework works with extreme scarcity of both, which is ideal for rare diseases. Via DML, the framework learns a multimodal mixture representation for each category, and performs dense predictions based on cosine distances between the pixels' deep embeddings and the category representations. The multimodal representations effectively utilize the inter-subject similarities and intraclass variations to overcome overfitting due to extremely limited data. In addition, we propose adaptive mixing coefficients for the multimodal mixture distributions to adaptively emphasize the modes better suited to the current input. The representations are implicitly embedded as weights of the fc layer, such that the cosine distances can be computed efficiently via forward propagation. In our experiments on brain MRI and abdominal CT datasets, the proposed framework achieves superior performances for low-shot segmentation towards standard DNN-based (3D U-Net) and classical registration-based (ANTs) methods, e.g., achieving mean Dice coefficients of 81%/69% for brain tissue/abdominal multi-organ segmentation using a single training sample, as compared to 52%/31% and 72%/35% by the U-Net and ANTs, respectively. Hengji Cui, Dong Wei 0004, Kai Ma 0002, Shi Gu, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | LT-Net: Label Transfer by Learning Reversible Voxel-Wise Correspondence for One-Shot Medical Image SegmentationabstractWe introduce a one-shot segmentation method to alleviate the burden of manual annotation for medical images. The main idea is to treat one-shot segmentation as a classical atlas-based segmentation problem, where voxel-wise correspondence from the atlas to the unlabelled data is learned. Subsequently, segmentation label of the atlas can be transferred to the unlabelled data with the learned correspondence. However, since ground truth correspondence between images is usually unavailable, the learning system must be well-supervised to avoid mode collapse and convergence failure. To overcome this difficulty, we resort to the forward-backward consistency, which is widely used in correspondence problems, and additionally learn the backward correspondences from the warped atlases back to the original atlas. This cycle-correspondence learning design enables a variety of extra, cycle-consistency-based supervision signals to make the training process stable, while also boost the performance. We demonstrate the superiority of our method over both deep learning-based one-shot segmentation methods and a classical multi-atlas segmentation method via thorough experiments. Shilei Cao 0001, Dong Wei 0004, Renzhen Wang, Kai Ma 0002, Liansheng Wang 0002, Deyu Meng, Yefeng Zheng 0001 |
CVPR | 3 |
| 2020 | Superpixel-Guided Label Softening for Medical Image Segmentation
Dong Wei 0004, Shilei Cao 0001, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (4) | 2 |
| 2020 | Learning and Exploiting Interclass Visual Correlations for Medical Image Classification
Dong Wei 0004, Shilei Cao 0001, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (1) | 1 |
| 2020 | Efficient and Effective Training of COVID-19 Classification Networks With Self-Supervised Dual-Track Learning to RankabstractCoronavirus Disease 2019 (COVID-19) has rapidly spread worldwide since first reported. Timely diagnosis of COVID-19 is crucial both for disease control and patient care. Non-contrast thoracic computed tomography (CT) has been identified as an effective tool for the diagnosis, yet the disease outbreak has placed tremendous pressure on radiologists for reading the exams and may potentially lead to fatigue-related mis-diagnosis. Reliable automatic classification algorithms can be really helpful; however, they usually require a considerable number of COVID-19 cases for training, which is difficult to acquire in a timely manner. Meanwhile, how to effectively utilize the existing archive of non-COVID-19 data (the negative samples) in the presence of severe class imbalance is another challenge. In addition, the sudden disease outbreak necessitates fast algorithm development. In this work, we propose a novel approach for effective and efficient training of COVID-19 classification networks using a small number of COVID-19 CT exams and an archive of negative samples. Concretely, a novel self-supervised learning method is proposed to extract features from the COVID-19 and negative samples. Then, two kinds of soft-labels ('difficulty' and 'diversity') are generated for the negative samples by computing the earth mover's distances between the features of the negative and COVID-19 samples, from which data 'values' of the negative samples can be assessed. A pre-set number of negative samples are selected accordingly and fed to the neural network for training. Experimental results show that our approach can achieve superior performance using about half of the negative samples, substantially reducing model training time. Yuexiang Li, Dong Wei 0004, Jiawei Chen 0009, Shilei Cao 0001, Yanchun Zhu, Lan Lan 0002, Tianyi Qian, Kai Ma 0002, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Conquering Data Variations in Resolution: A Slice-Aware Multi-Branch Decoder NetworkabstractFully convolutional neural networks have made promising progress in joint liver and liver tumor segmentation. Instead of following the debates over 2D versus 3D networks (for example, pursuing the balance between large-scale 2D pretraining and 3D context), in this paper, we novelly identify the wide variation in the ratio between intra- and inter-slice resolutions as a crucial obstacle to the performance. To tackle the mismatch between the intra- and inter-slice information, we propose a slice-aware 2.5D network that emphasizes extracting discriminative features utilizing not only in-plane semantics but also out-of-plane coherence for each separate slice. Specifically, we present a slice-wise multi-input multi-output architecture to instantiate such a design paradigm, which contains a Multi-Branch Decoder (MD) with a Slice-centric Attention Block (SAB) for learning slice-specific features and a Densely Connected Dice (DCD) loss to regularize the inter-slice predictions to be coherent and continuous. Based on the aforementioned innovations, we achieve state-of-the-art results on the MICCAI 2017 Liver Tumor Segmentation (LiTS) dataset. Besides, we also test our model on the ISBI 2019 Segmentation of THoracic Organs at Risk (SegTHOR) dataset, and the result proves the robustness and generalizability of the proposed method in other segmentation tasks. Shilei Cao 0001, Zhizhong Chai, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2013 | Three-dimensional segmentation of the left ventricle in late gadolinium enhanced MR images of chronic infarction combining long- and short-axis information
Dong Wei 0004, Ying Sun 0001, Sim Heng Ong, Ping Chai, Lynette L. Teo, Adrian F. Low |
Medical Image Anal. | 1 |
| 2011 | Myocardial Segmentation of Late Gadolinium Enhanced MR Images by Propagation of Contours from Cine MR Images
Dong Wei 0004, Ying Sun 0001, Ping Chai, Adrian F. Low, Sim Heng Ong |
MICCAI (3) | 1 |
| 2010 | MTMR: A conceptual interior design framework integrating Mixed Reality with the Multi-Touch tabletop interfaceabstractThis paper introduces a conceptual interior design framework - Multi-Touch Mixed Reality (MTMR), which integrates mixed reality with the multi-touch tabletop interface, to provide an intuitive and efficient interface for collaborative design and an augmented 3D view to users at the same time. Under this framework, multiple designers can carry out design work simultaneously on the top view displayed on the tabletop, while live video of the ongoing design work is captured and augmented by overlaying virtual 3D furniture models to their 2D virtual counterparts, and shown on a vertical screen in front of the tabletop. Meanwhile, the remote client's camera view of the physical room is augmented with the interior design layout in real time, that is, as the designers place, move, and modify the virtual furniture models on the tabletop, the client sees the corresponding life-size 3D virtual furniture models residing, moving, and changing in the physical room through the camera view on his/her screen. By adopting MTMR, which we argue may also apply to other kinds of collaborative work, the designers can expect a good working experience in terms of naturalness and intuitiveness, while the client can be involved in the design process and view the design result without moving around heavy furniture. By presenting MTMR, we hope to provide reliable and precise freehand interactions to mixed reality systems, with multi-touch inputs on tabletop interfaces. Dong Wei 0004, Steven Zhiying Zhou, Du Xie |
ISMAR | 1 |