EDBT 2026 Demo / reviewers in the wild / expert
Cheng Chen 0013
dblp:10/217-13
· DBLP profile ↗
36ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0002-6040-6833ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 7 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular BarriersabstractMulti-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining these models on large-scale unlabeled image-tabular data, aiming to learn discriminative representations. However, existing SSL methods for image-tabular representation learning are often confined to specific data cohorts, mainly due to their rigid tabular modeling mechanisms when modeling heterogeneous tabular data. This inter-tabular barrier hinders the multi-modal SSL methods from effectively learning transferrable medical knowledge shared across diverse cohorts. In this paper, we propose a novel SSL framework, namely CITab, designed to learn powerful multi-modal feature representations in a cross-tabular manner. We design the tabular modeling mechanism from a semantic-awareness perspective by integrating column headers as semantic cues, which facilitates transferrable knowledge learning and the scalability in utilizing multiple data sources for pretraining. Additionally, we propose a prototype-guided mixture-of-linear layer (P-MoLin) module for tabular feature specialization, empowering the model to effectively handle the heterogeneity of tabular data and explore the underlying medical concepts. We conduct comprehensive evaluations on Alzheimer's disease diagnosis task across three publicly available data cohorts containing 4,461 subjects. Experimental results demonstrate that CITab outperforms state-of-the-art approaches, paving the way for effective and scalable cross-tabular multi-modal learning. Yibing Fu, Zhitao Zeng, Cheng Chen 0013, Yueming Jin |
AAAI | 4 |
| 2026 | Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task ProjectionabstractRecent advances in parameter-efficient transfer learning have demonstrated the utility of composing LoRA adapters from libraries of pretrained modules. However, most existing approaches rely on simple retrieval heuristics or uniform averaging, which overlook the latent structure of task relationships in representation space. We propose a new framework for adapter reuse that moves beyond retrieval, formulating adapter composition as a geometry-aware sparse reconstruction problem. Specifically, we represent each task by a latent prototype vector derived from the base model’s encoder and aim to approximate the target task prototype as a sparse linear combination of retrieved reference prototypes, under an L1-regularized optimization objective. The resulting combination weights are then used to blend the corresponding LoRA adapters, yielding a composite adapter tailored to the target task. This formulation not only preserves the local geometric structure of the task representation manifold, but also promotes interpretability and efficient reuse by selecting a minimal set of relevant adapters. We demonstrate the effectiveness of our approach across multiple domains—including medical image segmentation, medical report generation and image synthesis. Our results highlight the benefit of coupling retrieval with latent geometry-aware optimization for improved zero-shot generalization. Pengfei Jin, Peng Shu, Sifan Song, Sekeun Kim, Qing Xiao 0003, Cheng Chen 0013, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li |
AAAI | 6 |
| 2026 | Tackling Dual-stage Missing Modalities in Brain Tumor Segmentation via Robust Modality Reconstruction and Prompt-guided Modality AdaptationabstractAddressing missing modalities is a critical challenge in multimodal brain tumor segmentation. Most existing approaches merely handle modality-incomplete inputs during inference, assuming a full set of modalities for all training samples. However, this unrealistic assumption limits the usage of abundant modality-incomplete data commonly observed in clinical practice. In this paper, we explore a more practical task of tackling missing modalities during both training and inference. We propose a universal model featuring robust modality reconstruction and prompt-guided modality adaptation. Our mask-reconstruction pre-training enables robust modality-invariant representation learning, during which we design a novel distribution approximation method that supervises the reconstruction of absent modalities without requiring full-modal training data. Afterwards, when adapting our model to the segmentation task, we introduce the complete-then-distill (CTD) paradigm, which first estimates missing modalities in training samples from the available ones, and then distills the knowledge from the reconstructed full-modal representations to enhance learning from modality-incomplete data. Moreover, we propose prompt-guided modality adaptation to personalize a subset of model parameters during CTD, enabling the model to adapt to each distinct modality input scenario by using prompts with rich visual-textual information. Extensive experiments on two brain tumor segmentation benchmarks show our method consistently surpasses previous state-of-the-art approaches under dual-stage missing modality settings across various missing ratios. Cheng Chen 0013, Qing You Pang, Yibing Fu, Quanzheng Li, Carol Tang, Beng-Ti Ang, Yueming Jin |
AAAI | 2 |
| 2026 | SAM-driven cross prompting with adaptive sampling consistency for semi-supervised medical image segmentationabstractSemi-supervised learning (SSL) has achieved notable progress in medical image segmentation. To achieve effective SSL, a model needs to be able to efficiently learn from limited labeled data and effectively exploit knowledge from abundant unlabeled data. Recent developments in visual foundation models, such as the Segment Anything Model (SAM), have demonstrated remarkable adaptability with improved sample efficiency. To seamlessly harness foundation models in SSL, we propose a SAM-driven cross prompting framework with adaptive sampling and prompt consistency for semi-supervised medical image segmentation, named CPAC-SAM. Our method employs SAM's unique prompt design and innovates a cross prompting strategy within a dual-branch framework to automatically generate prompts and supervision across two decoder branches, enabling effective learning from both scarce labeled and valuable unlabeled data. To ensure the quality of prompts for unlabeled data and provide meaningful supervision in the cross prompting scheme, we propose an innovative prototype-guided grid sampling strategy with adaptive intervals to simultaneously improve the reliability of the prompt selection area and ensure both adequate prompt density and complete target coverage. We further design a novel prompt consistency regularization to reduce SAM's prompt sensitivity and to enhance the output invariance under different prompts. We validate our method on five medical image segmentation tasks, encompassing both 2D and 3D scenarios. The extensive experiments with different labeled-data ratios and modalities demonstrate the superiority of our proposed method over the state-of-the-art SSL methods, with more than 4.1% and 3.8% Dice improvement on the breast cancer segmentation task and left atrium segmentation task, respectively. Our code is available at: https://github.com/JuzhengMiao/CPAC-SAM. Juzheng Miao, Cheng Chen 0013, Yuchen Yuan, Quanzheng Li, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2026 | TD-SAM: Temporal and Distance-Guided Adaptations of SAM for Accurate Surgical Instrument SegmentationabstractAccurate automatic surgical instrument segmentation plays a crucial role in robot-assisted surgery, but analyzing surgical videos remains challenging due to factors such as rapid instrument movements, high inter-category similarity, and frequent object occlusions. Current surgical instrument segmentation models struggle to capture both inter-frame variations and intra-frame details in complex surgical scenarios. The Segment Anything Model (SAM) has shown significant potential in various segmentation tasks. However, it has not fully addressed the unique challenges posed by surgical videos. To tackle these issues, we propose a Temporal and Distance-Guided SAM model (TD-SAM) for accurate surgical instrument segmentation. Specifically, we introduce a dynamic cross-frame attention module that effectively captures temporal information across frames, allowing the model to track the dynamic changes of surgical instruments and their environment, thus improving segmentation accuracy. In addition, we present a distance-guided instance refinement module, which enhances the model's ability to distinguish between similar categories, mitigating the class ambiguity caused by inter-category similarity. Extensive experiments on the EndoVis18 and EndoVis17 datasets show that the proposed TD-SAM model outperforms existing models, achieving state-of-the-art performance without using any prompts. Cheng Xue 0003, Danqiong Wang, Cheng Chen 0013, Guanyu Yang 0001, Yang Chen 0008 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | GM-ABS: Promptable Generalist Model Drives Active Barely Supervised Training in Specialist Model for 3D Medical Image SegmentationabstractSemi-supervised learning (SSL) has greatly advanced 3D medical image segmentation by alleviating the need for intensive labeling by radiologists. While previous efforts focused on model-centric advancements, the emergence of foundational generalist models like the Segment Anything Model (SAM) is expected to reshape the SSL landscape. Although these generalists usually show performance gaps relative to previous specialists in medical imaging, they possess impressive zero-shot segmentation abilities with manual prompts. Thus, this capability could serve as "free lunch" for training specialists, offering future SSL a promising data-centric perspective, especially revolutionizing both pseudo and expert labeling strategies to enhance the data pool. In this regard, we propose the Generalist Model-driven Active Barely Supervised (GM-ABS) learning paradigm, for developing specialized 3D segmentation models under extremely limited (barely) annotation budgets, e.g., merely cross-labeling three slices per selected scan. In specific, building upon a basic mean-teacher SSL framework, GM-ABS modernizes the SSL paradigm with two key data-centric designs: (i) Specialist-generalist collaboration, where the in-training specialist leverages class-specific positional prompts derived from class prototypes to interact with the frozen class-agnostic generalist across multiple views to achieve noisy-yet-effective label augmentation. Then, the specialist robustly assimilates the augmented knowledge via noise-tolerant collaborative learning. (ii) Expert-model collaboration that promotes active cross-labeling with notably low labeling efforts. This design progressively furnishes the specialist with informative and efficient supervision via a human-in-the-loop manner, which in turn benefits the quality of class-specific prompts. Extensive experiments on three benchmark datasets highlight the promising performance of GM-ABS over recent SSL approaches under extremely constrained labeling resources. Zhe Xu 0012, Cheng Chen 0013, Donghuan Lu, Jinghan Sun, Dong Wei 0004, Yefeng Zheng 0001, Quanzheng Li, Raymond Kai-Yu Tong |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport
Qinkai Yu, Jianyang Xie, Yitian Zhao, Cheng Chen 0013, Jun Cheng 0003, Lu Liu 0001, Yalin Zheng, Yanda Meng |
MICCAI (15) | 4 |
| 2025 | Multi-Organ Segmentation From Partially Labeled and Unaligned Multi-Modal MRI in Thyroid-Associated OrbitopathyabstractThyroid-associated orbitopathy (TAO) is a prevalent inflammatory autoimmune disorder, leading to orbital disfigurement and visual disability. Automatic comprehensive segmentation tailored for quantitative multi-modal MRI assessment of TAO holds enormous promise but is still lacking. In this paper, we propose a novel method, named cross-modal attentive self-training (CMAST), for the multi-organ segmentation in TAO using partially labeled and unaligned multi-modal MRI data. Our method first introduces a dedicatedly designed cross-modal pseudo label self-training scheme, which leverages self-training to refine the initial pseudo labels generated by cross-modal registration, so as to complete the label sets for comprehensive segmentation. With the obtained pseudo labels, we further devise a learnable attentive fusion module to aggregate multi-modal knowledge based on learned cross-modal feature attention, which relaxes the requirement of pixel-wise alignment across modalities. A prototypical contrastive learning loss is further incorporated to facilitate cross-modal feature alignment. We evaluate our method on a large clinical TAO cohort with 100 cases of multi-modal orbital MRI. The experimental results demonstrate the promising performance of our method in achieving comprehensive segmentation of TAO-affected organs on both T1 and T1c modalities, outperforming previous methods by a large margin. Our code is available at: https://github.com/cchen-cc/CMAST. Cheng Chen 0013, Yuan Zhong 0003, Jinyue Cai, Karen Kar Wun Chan, Qi Dou 0001, Kelvin Kam Lung Chong, Pheng-Ann Heng, Winnie Chiu-Wing Chu |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | MediViSTA: Medical Video Segmentation Via Temporal Fusion SAM Adaptation for EchocardiographyabstractDespite achieving impressive results in general-purpose semantic segmentation with strong generalization on natural images, the Segment Anything Model (SAM) has shown less precision and stability in medical image segmentation. In particular, the original SAM architecture is designed for 2D natural images and is therefore not support to handle three-dimensional information, which is particularly important for medical imaging modalities that are often volumetric or video data. In this paper, we introduce MediViSTA, a parameter-efficient fine-tuning method designed to adapt the vision foundation model for medical video, with a specific focus on echocardiography segmentation. To achieve spatial adaptation, we propose a frequency feature fusion technique that injects spatial frequency information from a CNN branch. For temporal adaptation, we integrate temporal adapters within the transformer blocks of the image encoder. Using a fine-tuning strategy, only a small subset of pre-trained parameters is updated, allowing efficient adaptation to echocardiography data. The effectiveness of our method has been comprehensively evaluated on three datasets, comprising two public datasets and one multi-center in-house dataset. Our method consistently outperforms various state-of-the-art approaches without using any prompts. Furthermore, our model exhibits strong generalization capabilities on unseen datasets, surpassing the second-best approach by 2.15% in Dice and 0.09 in temporal consistency. The results demonstrate the potential of MediViSTA to significantly advance echocardiography video segmentation, offering improved accuracy and robustness in cardiac assessment applications. Sekeun Kim, Pengfei Jin, Cheng Chen 0013, Kyung Sang Kim, Zhiliang Lyu, Hui Ren 0001, Zhengliang Liu, Aoxiao Zhong, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | EchoFM: Foundation Model for Generalizable Echocardiogram AnalysisabstractEchocardiography is the first-line non-invasive cardiac imaging modality, providing rich spatio-temporal information on cardiac anatomy and physiology. Recently, foundation model trained on extensive and diverse datasets has shown strong performance in various downstream tasks. However, translating foundation models into the medical imaging domain remains challenging due to domain differences between medical and natural images, the lack of diverse patient and disease datasets. In this paper, we introduce EchoFM, a general-purpose vision foundation model for echocardiography trained on a large-scale dataset of over 20 million echocardiographic images from 6,500 patients. To enable effective learning of rich spatio-temporal representations from periodic videos, we propose a novel self-supervised learning framework based on a masked autoencoder with a spatio-temporal consistent masking strategy and periodic-driven contrastive learning. The learned cardiac representations can be readily adapted and fine-tuned for a wide range of downstream tasks, serving as a strong and flexible backbone model. We validate EchoFM through experiments across key downstream tasks in the clinical echocardiography workflow, leveraging public and multi-center internal datasets. EchoFM consistently outperforms SOTA methods, demonstrating superior generalization capabilities and flexibility. The code and checkpoints are available at: https://github.com/SekeunKim/EchoFM.git. Sekeun Kim, Pengfei Jin, Sifan Song, Cheng Chen 0013, Yiwei Li 0002, Hui Ren 0001, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medical Image Segmentation
Juzheng Miao, Cheng Chen 0013, Keli Zhang, Jie Chuai, Quanzheng Li, Pheng-Ann Heng |
MICCAI (11) | 2 |
| 2024 | FM-OSD: Foundation Model-Enabled One-Shot Detection of Anatomical Landmarks
Juzheng Miao, Cheng Chen 0013, Keli Zhang, Jie Chuai, Quanzheng Li, Pheng-Ann Heng |
MICCAI (11) | 2 |
| 2024 | FM-ABS: Promptable Foundation Model Drives Active Barely Supervised Learning for 3D Medical Image Segmentation
Zhe Xu 0012, Cheng Chen 0013, Donghuan Lu, Jinghan Sun, Dong Wei 0004, Yefeng Zheng 0001, Quanzheng Li, Raymond Kai-Yu Tong |
MICCAI (8) | 2 |
| 2024 | MA-SAM: Modality-agnostic SAM adaptation for 3D medical image segmentation
Cheng Chen 0013, Juzheng Miao, Dufan Wu, Aoxiao Zhong, Zhiling Yan, Sekeun Kim, Zhengliang Liu, Lichao Sun 0001, Xiang Li 0001, Tianming Liu 0001, Pheng-Ann Heng, Quanzheng Li |
Medical Image Anal. | 1 |
| 2024 | Causal Effect Estimation on Imaging and Clinical Data for Treatment Decision Support of Aneurysmal Subarachnoid HemorrhageabstractAneurysmal subarachnoid hemorrhage is a medical emergency of brain that has high mortality and poor prognosis. Causal effect estimation of treatment strategies on patient outcomes is crucial for aneurysmal subarachnoid hemorrhage treatment decision-making. However, most existing studies on treatment decision-making support of this disease are unable to simultaneously compare the potential outcomes of different treatments for a patient. Furthermore, these studies fail to harmoniously integrate the imaging data with non-imaging clinical data, both of which are useful in clinical scenarios. In this paper, we estimate the causal effect of various treatments on patients with aneurysmal subarachnoid hemorrhage by integrating plain CT with non-imaging clinical data, which is represented using structured tabular data. Specifically, we first propose a novel scheme that uses multi-modality confounders distillation architecture to predict the treatment outcome and treatment assignment simultaneously. With these distilled confounder features, we design an imaging and non-imaging interaction representation learning strategy to use the complementary information extracted from different modalities to balance the feature distribution of different treatment groups. We have conducted extensive experiments using a clinical dataset of 656 subarachnoid hemorrhage cases, which was collected from the Hospital Authority Data Collaboration Laboratory in Hong Kong. Our method shows consistent improvements on the evaluation metrics of treatment effect estimation, achieving state-of-the-art results over strong competitors. Code is released at https://github.com/med-air/TOP-aSAH. Wenao Ma, Cheng Chen 0013, Yuqi Gong, Nga Yan Chan, Meirui Jiang, Calvin Hoi-Kwan Mak, Jill M. Abrigo, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Federated Semi-Supervised Medical Image Segmentation via Prototype-Based Pseudo-Labeling and Contrastive LearningabstractExisting federated learning works mainly focus on the fully supervised training setting. In realistic scenarios, however, most clinical sites can only provide data without annotations due to the lack of resources or expertise. In this work, we are concerned with the practical yet challenging federated semi-supervised segmentation (FSSS), where labeled data are only with several clients and other clients can just provide unlabeled data. We take an early attempt to tackle this problem and propose a novel FSSS method with prototype-based pseudo-labeling and contrastive learning. First, we transmit a labeled-aggregated model, which is obtained based on prototype similarity, to each unlabeled client, to work together with the global model for debiased pseudo labels generation via a consistency- and entropy-aware selection strategy. Second, we transfer image-level prototypes from labeled datasets to unlabeled clients and conduct prototypical contrastive learning on unlabeled models to enhance their discriminative power. Finally, we perform the dynamic model aggregation with a designed consistency-aware aggregation strategy to dynamically adjust the aggregation weights of each local model. We evaluate our method on COVID-19 X-ray infected region segmentation, COVID-19 CT infected region segmentation and colorectal polyp segmentation, and experimental results consistently demonstrate the effectiveness of our proposed method. Codes areavailable at https://github.com/zhangbaiming/FedSemiSeg. Huisi Wu, Baiming Zhang, Cheng Chen 0013, Harry Qin |
IEEE Trans. Medical Imaging | 3 |
| 2023 | CauSSL: Causality-inspired Semi-supervised Learning for Medical Image SegmentationabstractSemi-supervised learning (SSL) has recently demonstrated great success in medical image segmentation, significantly enhancing data efficiency with limited annotations. However, despite its empirical benefits, there are still concerns in the literature about the theoretical foundation and explanation of semi-supervised segmentation. To explore this problem, this study first proposes a novel causal diagram to provide a theoretical foundation for the mainstream semi-supervised segmentation methods. Our causal diagram takes two additional intermediate variables into account, which are neglected in previous work. Drawing from this proposed causal diagram, we then introduce a causality-inspired SSL approach on top of co-training frameworks called CauSSL, to improve SSL for medical image segmentation. Specifically, we first point out the importance of algorithmic independence between two networks or branches in SSL, which is often overlooked in the literature. We then propose a novel statistical quantification of the uncomputable algorithmic independence and further enhance the independence via a min-max optimization process. Our method can be flexibly incorporated into different existing SSL methods to improve their performance. Our method has been evaluated on three challenging medical image segmentation tasks using both 2D and 3D network architectures and has shown consistent improvements over state-of-the-art methods. Our code is publicly available at: https://github.com/JuzhengMiao/CauSSL. Juzheng Miao, Cheng Chen 0013, Furui Liu, Pheng-Ann Heng |
ICCV | 2 |
| 2023 | Contrastive Masked Image-Text Modeling for Medical Visual Representation Learning
Cheng Chen 0013, Aoxiao Zhong, Dufan Wu, Jie Luo 0003, Quanzheng Li |
MICCAI (5) | 1 |
| 2023 | Treatment Outcome Prediction for Intracerebral Hemorrhage via Generative Prognostic Model with Imaging and Tabular Data
Wenao Ma, Cheng Chen 0013, Jill M. Abrigo, Calvin Hoi-Kwan Mak, Yuqi Gong, Nga Yan Chan, Chu Han, Zaiyi Liu, Qi Dou 0001 |
MICCAI (5) | 2 |
| 2023 | SATTA: Semantic-Aware Test-Time Adaptation for Cross-Domain Medical Image Segmentation
Yuhan Zhang 0001, Cheng Chen 0013, Qiang Chen 0004, Pheng-Ann Heng |
MICCAI (2) | 3 |
| 2023 | Uncertainty Estimation for Safety-critical Scene Segmentation via Fine-grained Reward MaximizationabstractUncertainty estimation plays an important role for future reliable deployment of deep segmentation models in safety-critical scenarios such as medical applications. However, existing methods for uncertainty estimation have been limited by the lack of explicit guidance for calibrating the prediction risk and model confidence. In this work, we propose a novel fine-grained reward maximization (FGRM) framework, to address uncertainty estimation by directly utilizing an uncertainty metric related reward function with a reinforcement learning based model tuning algorithm. This would benefit the model uncertainty estimation with direct optimization guidance for model calibration. Specifically, our method designs a new uncertainty estimation reward function using the calibration metric, which is maximized to fine-tune an evidential learning pre-trained segmentation model for calibrating prediction risk. Importantly, we innovate an effective fine-grained parameter update scheme, which imposes fine-grained reward-weighting of each network parameter according to the parameter importance quantified by the fisher information matrix. To the best of our knowledge, this is the first work exploring reward optimization for model uncertainty estimation in safety-critical vision tasks. The effectiveness of our method is demonstrated on two large safety-critical surgical scene segmentation datasets under two different uncertainty estimation settings. With real-time one forward pass at inference, our method outperforms state-of-the-art methods by a clear margin on all the calibration metrics of uncertainty estimation, while maintaining a high task accuracy for the segmentation results. Code is available at https://github.com/med-air/FGRM. Hongzheng Yang, Cheng Chen 0013, Yueyao Chen, Markus Scheppach, Hon-Chi Yip, Qi Dou 0001 |
NeurIPS | 2 |
| 2023 | IOP-FL: Inside-Outside Personalization for Federated Medical Image SegmentationabstractFederated learning (FL) allows multiple medical institutions to collaboratively learn a global model without centralizing client data. It is difficult, if possible at all, for such a global model to commonly achieve optimal performance for each individual client, due to the heterogeneity of medical images from various scanners and patient demographics. This problem becomes even more significant when deploying the global model to unseen clients outside the FL with unseen distributions not presented during federated training. To optimize the prediction accuracy of each individual client for medical imaging tasks, we propose a novel unified framework for both Inside and Outside model Personalization in FL (IOP-FL). Our inside personalization uses a lightweight gradient-based approach that exploits the local adapted model for each client, by accumulating both the global gradients for common knowledge and the local gradients for client-specific optimization. Moreover, and importantly, the obtained local personalized models and the global model can form a diverse and informative routing space to personalize an adapted model for outside FL clients. Hence, we design a new test-time routing scheme using the consistency loss with a shape constraint to dynamically incorporate the models, given the distribution information conveyed by the test data. Our extensive experimental results on two medical image segmentation tasks present significant improvements over SOTA methods on both inside and outside personalization, demonstrating the potential of our IOP-FL scheme for clinical practice. Code is available at https://github.com/med-air/IOP-FL. Meirui Jiang, Hongzheng Yang, Cheng Chen 0013, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Continual Nuclei Segmentation via Prototype-Wise Relation Distillation and Contrastive LearningabstractDeep learning models have achieved remarkable success in multi-type nuclei segmentation. These models are mostly trained at once with the full annotation of all types of nuclei available, while lack the ability of continually learning new classes due to the problem of catastrophic forgetting. In this paper, we study the practical and important class-incremental continual learning problem, where the model is incrementally updated to new classes without accessing to previous data. We propose a novel continual nuclei segmentation method, to avoid forgetting knowledge of old classes and facilitate the learning of new classes, by achieving feature-level knowledge distillation with prototype-wise relation distillation and contrastive learning. Concretely, prototype-wise relation distillation imposes constraints on the inter-class relation similarity, encouraging the encoder to extract similar class distribution for old classes in the feature space. Prototype-wise contrastive learning with a hard sampling strategy enhances the intra-class compactness and inter-class separability of features, improving the performance on both old and new classes. Experiments on two multi-type nuclei segmentation benchmarks, i.e., MoNuSAC and CoNSeP, demonstrate the effectiveness of our method with superior performance over many competitive methods. Codes are available at https://github.com/zzw-szu/CoNuSeg. Huisi Wu, Zhaoze Wang, Zebin Zhao 0004, Cheng Chen 0013, Harry Qin |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Single-Domain Generalization in Medical Image Segmentation via Test-Time Adaptation from Shape DictionaryabstractDomain generalization typically requires data from multiple source domains for model learning. However, such strong assumption may not always hold in practice, especially in medical field where the data sharing is highly concerned and sometimes prohibitive due to privacy issue. This paper studies the important yet challenging single domain generalization problem, in which a model is learned under the worst-case scenario with only one source domain to directly generalize to different unseen target domains. We present a novel approach to address this problem in medical image segmentation, which extracts and integrates the semantic shape prior information of segmentation that are invariant across domains and can be well-captured even from single domain data to facilitate segmentation under distribution shifts. Besides, a test-time adaptation strategy with dual-consistency regularization is further devised to promote dynamic incorporation of these shape priors under each unseen domain to improve model generalizability. Extensive experiments on two medical image segmentation tasks demonstrate the consistent improvements of our method across various unseen domains, as well as its superiority over state-of-the-art approaches in addressing domain generalization under the worst-case scenario. Quande Liu, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
AAAI | 2 |
| 2022 | Online Reflective Learning for Robust Medical Image Segmentation
Yuhao Huang 0001, Xin Yang 0009, Xiaoqiong Huang, Jiamin Liang, Cheng Chen 0013, Haoran Dou, Xindi Hu, Yan Cao 0002, Dong Ni 0001 |
MICCAI (8) | 6 |
| 2022 | Test-Time Adaptation with Calibration of Medical Image Classification Nets for Label Distribution Shift
Wenao Ma, Cheng Chen 0013, Harry Qin, Huimao Zhang, Qi Dou 0001 |
MICCAI (3) | 2 |
| 2022 | Learning With Privileged Multimodal Knowledge for Unimodal SegmentationabstractMultimodal learning usually requires a complete set of modalities during inference to maintain performance. Although training data can be well-prepared with high-quality multiple modalities, in many cases of clinical practice, only one modality can be acquired and important clinical evaluations have to be made based on the limited single modality information. In this work, we propose a privileged knowledge learning framework with the 'Teacher-Student' architecture, in which the complete multimodal knowledge that is only available in the training data (called privileged information) is transferred from a multimodal teacher network to a unimodal student network, via both a pixel-level and an image-level distillation scheme. Specifically, for the pixel-level distillation, we introduce a regularized knowledge distillation loss which encourages the student to mimic the teacher's softened outputs in a pixel-wise manner and incorporates a regularization factor to reduce the effect of incorrect predictions from the teacher. For the image-level distillation, we propose a contrastive knowledge distillation loss which encodes image-level structured information to enrich the knowledge encoding in combination with the pixel-level distillation. We extensively evaluate our method on two different multi-class segmentation tasks, i.e., cardiac substructure segmentation and brain tumor segmentation. Experimental results on both tasks demonstrate that our privileged knowledge learning is effective in improving unimodal segmentation and outperforms previous methods. Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Quande Liu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Exploring Intra- and Inter-Video Relation for Surgical Semantic Scene SegmentationabstractAutomatic surgical scene segmentation is fundamental for facilitating cognitive intelligence in the modern operating theatre. Previous works rely on conventional aggregation modules (e.g., dilated convolution, convolutional LSTM), which only make use of the local context. In this paper, we propose a novel framework STswinCL that explores the complementary intra- and inter-video relations to boost segmentation performance, by progressively capturing the global context. We firstly develop a hierarchy Transformer to capture intra-video relation that includes richer spatial and temporal cues from neighbor pixels and previous frames. A joint space-time window shift scheme is proposed to efficiently aggregate these two cues into each pixel embedding. Then, we explore inter-video relation via pixel-to-pixel contrastive learning, which well structures the global embedding space. A multi-source contrast training objective is developed to group the pixel embeddings across videos with the ground-truth guidance, which is crucial for learning the global property of the whole data. We extensively validate our approach on two public surgical video benchmarks, including EndoVis18 Challenge and CaDIS dataset. Experimental results demonstrate the promising performance of our method, which consistently exceeds previous state-of-the-art approaches. Code is available at https://github.com/YuemingJin/STswinCL. Yueming Jin, Yang Yu 0070, Cheng Chen 0013, Pheng-Ann Heng, Danail Stoyanov |
IEEE Trans. Medical Imaging | 3 |
| 2022 | DLTTA: Dynamic Learning Rate for Test-Time Adaptation on Cross-Domain Medical ImagesabstractTest-time adaptation (TTA) has increasingly been an important topic to efficiently tackle the cross-domain distribution shift at test time for medical images from different institutions. Previous TTA methods have a common limitation of using a fixed learning rate for all the test samples. Such a practice would be sub-optimal for TTA, because test data may arrive sequentially therefore the scale of distribution shift would change frequently. To address this problem, we propose a novel dynamic learning rate adjustment method for test-time adaptation, called DLTTA, which dynamically modulates the amount of weights update for each test image to account for the differences in their distribution shift. Specifically, our DLTTA is equipped with a memory bank based estimation scheme to effectively measure the discrepancy of a given test sample. Based on this estimated discrepancy, a dynamic learning rate adjustment strategy is then developed to achieve a suitable degree of adaptation for each test sample. The effectiveness and general applicability of our DLTTA is extensively demonstrated on three tasks including retinal optical coherence tomography (OCT) segmentation, histopathological image classification, and prostate 3D MRI segmentation. Our method achieves effective and fast test-time adaptation with consistent performance improvement over current state-of-the-art test-time adaptation methods. Code is available at https://github.com/med-air/DLTTA. Hongzheng Yang, Cheng Chen 0013, Meirui Jiang, Quande Liu, Jianfeng Cao, Pheng-Ann Heng, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency SpaceabstractFederated learning allows distributed medical institutions to collaboratively learn a shared prediction model with privacy protection. While at clinical deployment, the models trained in federated learning can still suffer from performance drop when applied to completely unseen hospitals outside the federation. In this paper, we point out and solve a novel problem setting of federated domain generalization (FedDG), which aims to learn a federated model from multiple distributed source domains such that it can directly generalize to unseen target domains. We present a novel approach, named as Episodic Learning in Continuous Frequency Space (ELCFS), for this problem by enabling each client to exploit multi-source data distributions under the challenging constraint of data decentralization. Our approach transmits the distribution information across clients in a privacy-protecting way through an effective continuous frequency space interpolation mechanism. With the transferred multi-source distributions, we further carefully design a boundary-oriented episodic learning paradigm to expose the local learning to domain distribution shifts and particularly meet the challenges of model generalization in medical image segmentation scenario. The effectiveness of our method is demonstrated with superior performance over state-of-the-arts and in-depth ablation experiments on two medical image segmentation tasks. The code is available at https://github.com/liuquande/FedDG-ELCFS. Quande Liu, Cheng Chen 0013, Harry Qin, Qi Dou 0001, Pheng-Ann Heng |
CVPR | 2 |
| 2021 | Source-Free Domain Adaptive Fundus Image Segmentation with Denoised Pseudo-Labeling
Cheng Chen 0013, Quande Liu, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 1 |
| 2021 | Temporal Memory Relation Network for Workflow Recognition From Surgical VideoabstractAutomatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal information, or separately learned visual and long temporal cues. In this paper, we propose a novel end-to-end temporal memory relation network (TMRNet) for relating long-range and multi-scale temporal patterns to augment the present features. We establish a long-range memory bank to serve as a memory cell storing the rich supportive information. Through our designed temporal variation layer, the supportive cues are further enhanced by multi-scale temporal-only convolutions. To effectively incorporate the two types of cues without disturbing the joint learning of spatio-temporal features, we introduce a non-local bank operator to attentively relate the past to the present. In this regard, our TMRNet enables the current feature to view the long-range temporal dependency, as well as tolerate complex temporal extents. We have extensively validated our approach on two benchmark surgical video datasets, M2CAI challenge dataset and Cholec80 dataset. Experimental results demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin (e.g., 67.0% v.s. 78.9% Jaccard on Cholec80 dataset). Yueming Jin, Yonghao Long 0001, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Unsupervised Bidirectional Cross-Modality Adaptation via Deeply Synergistic Image and Feature Alignment for Medical Image SegmentationabstractUnsupervised domain adaptation has increasingly gained interest in medical image computing, aiming to tackle the performance degradation of deep neural networks when being deployed to unseen data with heterogeneous characteristics. In this work, we present a novel unsupervised domain adaptation framework, named as Synergistic Image and Feature Alignment (SIFA), to effectively adapt a segmentation network to an unlabeled target domain. Our proposed SIFA conducts synergistic alignment of domains from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features by leveraging adversarial learning in multiple aspects and with a deeply supervised mechanism. The feature encoder is shared between both adaptive perspectives to leverage their mutual benefits via end-to-end learning. We have extensively evaluated our method with cardiac substructure segmentation and abdominal multi-organ segmentation for bidirectional cross-modality adaptation between MRI and CT images. Experimental results on two different tasks demonstrate that our SIFA method is effective in improving segmentation performance on unlabeled target images, and outperforms the state-of-the-art domain adaptation approaches by a large margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image SegmentationabstractThis paper presents a novel unsupervised domain adaptation framework, called Synergistic Image and Feature Adaptation (SIFA), to effectively tackle the problem of domain shift. Domain adaptation has become an important and hot topic in recent studies on deep learning, aiming to recover performance degradation when applying the neural networks to new testing domains. Our proposed SIFA is an elegant learning diagram which presents synergistic fusion of adaptations from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features towards the segmentation task. The feature encoder layers are shared by both perspectives to grasp their mutual benefits during the end-to-end learning procedure. Without using any annotation from the target domain, the learning of our unified model is guided by adversarial losses, with multiple discriminators employed from various aspects. We have extensively validated our method with a challenging application of crossmodality medical image segmentation of cardiac structures. Experimental results demonstrate that our SIFA model recovers the degraded performance from 17.2% to 73.0%, and outperforms the state-of-the-art methods by a significant margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 1 |
| 2019 | Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion
Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 1 |
| 2018 | Unsupervised Cross-Modality Domain Adaptation of ConvNets for Biomedical Image Segmentations with Adversarial LossabstractConvolutional networks (ConvNets) have achieved great successes in various challenging vision tasks. However, the performance of ConvNets would degrade when encountering the domain shift. The domain adaptation is more significant while challenging in the field of biomedical image analysis, where cross-modality data have largely different distributions. Given that annotating the medical data is especially expensive, the supervised transfer learning approaches are not quite optimal. In this paper, we propose an unsupervised domain adaptation framework with adversarial learning for cross-modality biomedical image segmentations. Specifically, our model is based on a dilated fully convolutional network for pixel-wise prediction. Moreover, we build a plug-and-play domain adaptation module (DAM) to map the target input to features which are aligned with source domain feature space. A domain critic module (DCM) is set up for discriminating the feature space of both domains. We optimize the DAM and DCM via an adversarial loss without using any target domain label. Our proposed method is validated by adapting a ConvNet trained with MRI images to unpaired CT data for cardiac structures segmentations, and achieved very promising results. Qi Dou 0001, Cheng Ouyang, Cheng Chen 0013, Hao Chen 0011, Pheng-Ann Heng |
IJCAI | 3 |