EDBT 2026 Demo / reviewers in the wild / expert
Zhen Chen 0013
dblp:11/1266-13
· DBLP profile ↗
44ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0003-0255-6435ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 7 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SP-Det: Self-prompted dual-text fusion for generalized multi-label lesion detectionabstractAutomated lesion detection in chest X-rays has demonstrated significant potential for improving clinical diagnosis by precisely localizing pathological abnormalities. While recent promptable detection frameworks have achieved remarkable accuracy in target localization, existing methods typically rely on manual annotations as prompts, which are labor-intensive and impractical for clinical applications. To address this limitation, we propose SP-Det, a novel self-prompted detection framework that automatically generates rich textual context to guide multi-label lesion detection without requiring expert annotations. Specifically, we introduce an expert-free dual-text prompt generator (DTPG) that leverages two complementary textual modalities: semantic context prompts that capture global pathological patterns and disease beacon prompts that focus on disease-specific manifestations. Moreover, we devise a bidirectional feature enhancer (BFE) that synergistically integrates comprehensive diagnostic context with disease-specific embeddings to significantly improve feature representation and detection accuracy. Extensive experiments on two chest X-ray datasets with diverse thoracic disease categories demonstrate that our SP-Det framework outperforms state-of-the-art detection methods while completely eliminating the dependency on expert-annotated prompts compared to existing promptable architectures. Qing Xu 0014, Yanqian Wang, Xiangjian He, Yixuan Zhang 0006, Rong Qu, Wenting Duan, Zhen Chen 0013 |
Knowl. Based Syst. | 8 |
| 2026 | Harnessing Lightweight Transformer With Contextual Synergic Enhancement for Efficient 3D Medical Image SegmentationabstractTransformers have shown remarkable performance in 3D medical image segmentation, but their high computational requirements and need for large amounts of labeled data limit their applicability. To address these challenges, we consider two crucial aspects: model efficiency and data efficiency. Specifically, we propose Light-UNETR, a lightweight transformer designed to achieve model efficiency. Light-UNETR features a Lightweight Dimension Reductive Attention (LIDR) module, which reduces spatial and channel dimensions while capturing both global and local features via multi-branch attention. Additionally, we introduce a Compact Gated Linear Unit (CGLU) to selectively control channel interaction with minimal parameters. Furthermore, we introduce a Contextual Synergic Enhancement (CSE) learning strategy, which aims to boost the data efficiency of Transformers. It first leverages the extrinsic contextual information to support the learning of unlabeled data with Attention-Guided Replacement, then applies Spatial Masking Consistency that utilizes intrinsic contextual information to enhance the spatial context reasoning for unlabeled data. Extensive experiments on various benchmarks demonstrate the superiority of our approach in both performance and efficiency. For example, with only 10% labeled data on the Left Atrial Segmentation dataset, our method surpasses BCP by 1.43% Jaccard while drastically reducing the FLOPs by 90.8% and parameters by 85.8%. Xinyu Liu 0001, Zhen Chen 0013, Wuyang Li, Chenxin Li, Yixuan Yuan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | De-LightSAM: Modality-Decoupled Lightweight SAM for Generalizable Medical SegmentationabstractThe universality of deep neural networks across different modalities and their generalization capabilities to unseen domains play an essential role in medical image segmentation. The recent segment anything model (SAM) has demonstrated strong adaptability across diverse natural scenarios. However, the huge computational costs, demand for manual annotations as prompts and conflict-prone decoding process of SAM degrade its generalization capabilities in medical scenarios. To address these limitations, we propose a modality-decoupled lightweight SAM for domain-generalized medical image segmentation, named De-LightSAM. Specifically, we first devise a lightweight domain-controllable image encoder (DC-Encoder) that produces discriminative visual features for diverse modalities. Further, we introduce the self-patch prompt generator (SP-Generator) to automatically generate high-quality dense prompt embeddings for guiding segmentation decoding. Finally, we design the query-decoupled modality decoder (QM-Decoder) that leverages a one-to-one strategy to provide an independent decoding channel for every modality, preventing mutual knowledge interference of different modalities. Moreover, we design a multi-modal decoupled knowledge distillation (MDKD) strategy to leverage robust common knowledge to complement domain-specific medical feature representations. Extensive experiments indicate that De-LightSAM outperforms state-of-the-arts in diverse medical imaging segmentation tasks, displaying superior modality universality and generalization capabilities. Especially, De-LightSAM uses only 2.0% parameters compared to SAM-H. The source code is available at https://github.com/xq141839/De-LightSAM. Qing Xu 0014, Xiangjian He, Chenxin Li, Fiseha B. Tesema, Wenting Duan, Zhen Chen 0013, Rong Qu, Jonathan M. Garibaldi, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Co-Seg++: Mutual Prompt-Guided Collaborative Learning for Versatile Medical SegmentationabstractMedical image analysis is critical yet challenged by the need of jointly segmenting organs or tissues, and numerous instances for anatomical structures and tumor microenvironment analysis. Existing studies typically formulated different segmentation tasks in isolation, which overlooks the fundamental interdependencies between these tasks, leading to suboptimal segmentation performance and insufficient medical image understanding. To address this issue, we propose a Co-Seg++ framework for versatile medical segmentation. Specifically, we introduce a novel co-segmentation paradigm, allowing semantic and instance segmentation tasks to mutually enhance each other. We first devise a spatio-sequential prompt encoder (SSP-Encoder) to capture long-range spatial and sequential relationships between segmentation regions and image embeddings as prior spatial constraints. Moreover, we devise a multi-task collaborative decoder (MTC-Decoder) that leverages cross-guidance to strengthen the contextual consistency of both tasks, jointly computing semantic and instance segmentation masks. Extensive experiments on diverse CT and histopathology datasets demonstrate that the proposed Co-Seg++ outperforms state-of-the-arts in the semantic, instance, and panoptic segmentation of dental anatomical structures, histopathology tissues, and nuclei instances. The source code is available at https://github.com/xq141839/Co-Seg-Plus. Qing Xu 0014, Yuxiang Luo, Wenting Duan, Zhen Chen 0013 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | E$^{3}$3-Net: Efficient E(3)-Equivariant Normal Estimation NetworkabstractPoint cloud normal estimation is a fundamental task in 3D geometry processing, playing a crucial role in applications such as 3D reconstruction, object recognition, and surface analysis. While recent learning-based methods achieve notable advancements in normal prediction, they often overlook the critical aspect of equivariance. This oversight leads to inefficient learning of symmetric patterns inherent in geometric data. To address this issue, we propose E$^{3}$3-Net, an innovative neural network architecture designed to inherently achieve equivariance for normal estimation. We introduce an efficient random frame method, which significantly reduces the training resources required for this task to just 1/8 of previous work, while simultaneously enhancing prediction accuracy. Furthermore, we design a Gaussian-weighted loss function and a receptive-aware inference strategy that effectively leverage the local properties of point clouds, ensuring more precise and reliable normal estimation. Our method demonstrates superior performance across both synthetic and real-world datasets, consistently outperforming current state-of-the-art techniques by a substantial margin. Specifically, we achieve a 4% improvement in RMSE on the PCPNet dataset, 2.67% on the SceneNN dataset, and 2.44% on the FamousShape dataset, highlighting the robustness and scalability of E$^{3}$3-Net in diverse environments. Mingyang Zhao 0001, Weize Quan, Zhen Chen 0013, Dong-Ming Yan 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | U-KAN Makes Strong Backbone for Medical Image Segmentation and GenerationabstractU-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as well as the deficient interpretability. To address these challenges, our intuition is inspired by the impressive results of the Kolmogorov-Arnold Networks (KANs) in terms of accuracy and interpretability, which reshape the neural network learning via the stack of non-linear learnable activation functions derived from the Kolmogorov-Anold representation theorem. Specifically, in this paper, we explore the untapped potential of KANs in improving backbones for vision tasks. We investigate, modify and re-design the established U-Net pipeline by integrating the dedicated KAN layers on the tokenized intermediate representation, termed U-KAN. Rigorous medical image segmentation benchmarks verify the superiority of UKAN by higher accuracy even with less computation cost. We further delved into the potential of U-KAN as an alternative U-Net noise predictor in diffusion models, demonstrating its applicability in generating task-oriented model architectures. Chenxin Li, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Hengyu Liu 0007, Yifan Liu 0010, Zhen Chen 0013, Yixuan Yuan |
AAAI | 7 |
| 2025 | Polyp-Gen: Realistic and Diverse Polyp Image Generation for Endoscopic Dataset ExpansionabstractAutomated diagnostic systems (ADS) have shown significant potential in the early detection of polyps during endoscopic examinations, thereby reducing the incidence of colorectal cancer. However, due to high annotation costs and strict privacy concerns, acquiring high-quality endoscopic images poses a considerable challenge in the development of ADS. Despite recent advancements in generating synthetic images for dataset expansion, existing endoscopic image generation algorithms failed to accurately generate the details of polyp boundary regions and typically required medical priors to specify plausible locations and shapes of polyps, which limited the realism and diversity of the generated images. To address these limitations, we present Polyp-Gen, the first full-automatic diffusion-based endoscopic image generation framework. Specifically, we devise a spatial-aware diffusion training scheme with a lesion-guided loss to enhance the structural context of polyp boundary regions. Moreover, to capture medical priors for the localization of potential polyp areas, we introduce a hierarchical retrieval-based sampling strategy to match similar fine-grained spatial features. In this way, our Polyp-Gen can generate realistic and diverse endoscopic images for building reliable ADS. Extensive experiments demonstrate the state-of-the-art generation quality, and the synthetic images can improve the downstream polyp detection task. Additionally, our Polyp-Gen has shown remarkable zeroshot generalizability on other datasets. The source code is available at https://github.com/CUHK-AIM-Group/Polyp-Gen. Zhen Chen 0013, Qiushi Yang, Weihao Yu 0005, Di Dong, Jiancong Hu, Yixuan Yuan |
ICRA | 2 |
| 2025 | Relation-Guided Versatile Regularization for Federated Semi-Supervised LearningabstractAbstract Federated semi-supervised learning (FSSL) target to address the increasing privacy concerns for the practical scenarios, where data holders are limited in labeling capability. Latest FSSL approaches leverage the prediction consistency between the local model and global model to exploit knowledge from partially labeled or completely unlabeled clients. However, they merely utilize data-level augmentation for prediction consistency and simply aggregate model parameters through the weighted average at the server, which leads to biased classifiers and suffers from skewed unlabeled clients. To remedy these issues, we present a novel FSSL framework, Relation-guided Versatile Regularization (FedRVR), consisting of versatile regularization at clients and relation-guided directional aggregation strategy at the server. In versatile regularization, we propose the model-guided regularization together with the data-guided one, and encourage the prediction of the local model invariant to two extreme global models with different abilities, which provides richer consistency supervision for local training. Moreover, we devise a relation-guided directional aggregation at the server, in which a parametric relation predictor is introduced to yield pairwise model relation and obtain a model ranking. In this manner, the server can provide a superior global model by aggregating relative dependable client models, and further produce an inferior global model via reverse aggregation to promote the versatile regularization at clients. Extensive experiments on three FSSL benchmarks verify the superiority of FedRVR over state-of-the-art counterparts across various federated learning settings. Qiushi Yang, Zhen Chen 0013, Zhe Peng, Yixuan Yuan |
Int. J. Comput. Vis. | 2 |
| 2025 | NuSegDG: Integration of heterogeneous space and Gaussian kernel for domain-generalized nuclei segmentation
Zhenye Lou, Qing Xu 0014, Zekun Jiang, Xiangjian He, Chenxin Li, Zhen Chen 0013, Yi Wang 0037, Maggie M. He, Wenting Duan |
Knowl. Based Syst. | 6 |
| 2025 | UN-SAM: Domain-adaptive self-prompt segmentation for universal nuclei images
Zhen Chen 0013, Qing Xu 0014, Xinyu Liu 0001, Yixuan Yuan |
Medical Image Anal. | 1 |
| 2025 | Unified Multi-Modal Diagnostic Framework With Reconstruction Pre-Training and Heterogeneity-Combat TuningabstractMedical multi-modal pre-training has revealed promise in computer-aided diagnosis by leveraging large-scale unlabeled datasets. However, existing methods based on masked autoencoders mainly rely on data-level reconstruction tasks, but lack high-level semantic information. Furthermore, two significant heterogeneity challenges hinder the transfer of pre-trained knowledge to downstream tasks, i.e., the distribution heterogeneity between pre-training data and downstream data, and the modality heterogeneity within downstream data. To address these challenges, we propose a Unified Medical Multi-modal Diagnostic (UMD) framework with tailored pre-training and downstream tuning strategies. Specifically, to enhance the representation abilities of vision and language encoders, we propose the Multi-level Reconstruction Pre-training (MR-Pretrain) strategy, including a feature-level and data-level reconstruction, which guides models to capture the semantic information from masked inputs of different modalities. Moreover, to tackle two kinds of heterogeneities during the downstream tuning, we present the heterogeneity-combat downstream tuning strategy, which consists of a Task-oriented Distribution Calibration (TD-Calib) and a Gradient-guided Modality Coordination (GM-Coord). In particular, TD-Calib fine-tunes the pre-trained model regarding the distribution of downstream datasets, and GM-Coord adjusts the gradient weights according to the dynamic optimization status of different modalities. Extensive experiments on five public medical datasets demonstrate the effectiveness of our UMD framework, which remarkably outperforms existing approaches on three kinds of downstream tasks. Li Pan 0004, Qiushi Yang, Tan Li 0002, Zhen Chen 0013 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Exploring Contrastive Pre-Training for Domain Connections in Medical Image SegmentationabstractUnsupervised domain adaptation (UDA) in medical image segmentation aims to improve the generalization of deep models by alleviating domain gaps caused by inconsistency across equipment, imaging protocols, and patient conditions. However, existing UDA works remain insufficiently explored and present great limitations: 1) Exhibit cumbersome designs that prioritize aligning statistical metrics and distributions, which limits the model's flexibility and generalization while also overlooking the potential knowledge embedded in unlabeled data; 2) More applicable in a certain domain, lack the generalization capability to handle diverse shifts encountered in clinical scenarios. To overcome these limitations, we introduce MedCon, a unified framework that leverages general unsupervised contrastive pre-training to establish domain connections, effectively handling diverse domain shifts without tailored adjustments. Specifically, it initially explores a general contrastive pre-training to establish domain connections by leveraging the rich prior knowledge from unlabeled images. Thereafter, the pre-trained backbone is fine-tuned using source-based images to ultimately identify per-pixel semantic categories. To capture both intra- and inter-domain connections of anatomical structures, we construct positive-negative pairs from a hybrid aspect of both local and global scales. In this regard, a shared-weight encoder-decoder is employed to generate pixel-level representations, which are then mapped into hyper-spherical space using a non-learnable projection head to facilitate positive pair matching. Comprehensive experiments on diverse medical image datasets confirm that MedCon outperforms previous methods by effectively managing a wide range of domain shifts and showcasing superior generalization capabilities. Zequn Zhang, Yunnan Wang, Baao Xie, Yuhang Li 0005, Zhen Chen 0013, Xin Jin 0014, Wenjun Zeng 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma GradingabstractRecently, multimodal deep learning, which integrates histopathology slides and molecular biomarkers, has achieved a promising performance in glioma grading. Despite great progress, due to the intra-modality complexity and intermodality heterogeneity, existing studies suffer from inadequate histopathology representation learning and inefficient molecular-pathology knowledge alignment. These two issues hinder existing methods to precisely interpret diagnostic molecular-pathology features, thereby limiting their grading performance. Moreover, the real-world applicability of existing multimodal approaches is significantly restricted as molecular biomarkers are not always available during clinical deployment. To address these problems, we introduce a novel Focus on Focus (FoF) framework with paired pathology-genomic training and applicable pathology-only inference, enhancing molecular-pathology representation effectively. Specifically, we propose a Focus-oriented Representation Learning (FRL) module to encourage the model to identify regions positively or negatively related to glioma grading and guide it to focus on the diagnostic areas with a consistency constraint. To effectively link the molecular biomarkers to morphological features, we propose a Multi-view Cross-modal Alignment (MCA) module that projects histopathology representations into molecular subspaces, aligning morphological features with corresponding molecular biomarker status by supervised contrastive learning. Experiments on the TCGA GBMLGG dataset demonstrate that our FoF framework significantly improves the glioma grading. Remarkably, our FoF achieves superior performance using only histopathology slides compared to existing multimodal methods. The source code is available at https://github.com/peterlipan/FoF. Li Pan 0004, Qiushi Yang, Tan Li 0002, Xiaohan Xing, Maximus C. F. Yeung, Zhen Chen 0013 |
BIBM | 7 |
| 2024 | Communication-Efficient Multi-Modal Federated Learning via Dynamic Client-Modality MatchingabstractMulti-modal federated learning (MFL) offers the advantage of aggregating models from diverse data modalities to obtain a more powerful fused model while preserving data privacy. However, MFL faces three key challenges: 1) Communication overhead - only a limited number of clients can participate in training due to communication budget constraints; 2) Modality heterogeneity - different modalities contribute unequally to the fused model; 3) Client heterogeneity - clients exhibit variations in data quantity and quality across modalities. To address these challenges, we formulate a joint client-modality selection problem under communication budget constraints. The goal is to determine the participating clients and their uploaded modalities in each communication round, maximizing the performance of the fused model given a limited communication budget. We propose a dynamic many-to-many matching algorithm with two quota budgeting strategies: 1) Round-aware Modality Budgeting (RMB) determines the total number of uploaded modality models per round based on the current training process (i.e., how close the model is to convergence). 2) Modality-aware Client Allocation (MCB) adaptively allocates client quota for each modality by balancing the modality’s contribution to the fusion model against its model size. After quota budgeting, we construct preference lists for clients and modalities to find a stable many-to-many matching of (client, modality) pairs. Experiments demonstrate that our algorithm achieves better model performance than baselines under the same communication budget, validating the benefits of dynamic budget allocation and client scheduling. Tan Li 0002, Yanming Gong, Hai Liu 0001, Zhen Chen 0013, Linqi Song |
IEEE Big Data | 4 |
| 2024 | Accelerated Multi-contrast MRI Reconstruction via Frequency and Spatial Mutual Learning
Qi Chen 0014, Xiaohan Xing, Zhen Chen 0013, Zhiwei Xiong |
MICCAI (7) | 3 |
| 2024 | 👦 Endora: Video Generation Models as Endoscopy Simulators
Chenxin Li, Hengyu Liu 0007, Yifan Liu 0010, Brandon Yushan Feng, Wuyang Li, Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan |
MICCAI (6) | 7 |
| 2024 | MOST: Multi-formation Soft Masking for Semi-supervised Medical Image Segmentation
Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan |
MICCAI (11) | 2 |
| 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAMabstractAs the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradicting the consensus requirement for the robustness of a model. While some established works have been dedicated to stabilizing and fortifying the prediction of SAM, this paper takes a unique path to explore how this flaw can be inverted into an advantage when modeling inherently ambiguous data distributions. We introduce an optimization framework based on a conditional variational autoencoder, which jointly models the prompt and the granularity of the object with a latent probability distribution. This approach enables the model to adaptively perceive and represent the real ambiguous label distribution, taming SAM to produce a series of diverse, convincing, and reasonable segmentation outputs controllably. Extensive experiments on several practical deployment scenarios involving ambiguity demonstrates the exceptional performance of our framework. Project page: \url{https://a-sa-m.github.io/}. Chenxin Li, Yuzhi Huang, Wuyang Li, Hengyu Liu 0007, Xinyu Liu 0001, Qing Xu 0014, Zhen Chen 0013, Yue Huang 0001, Yixuan Yuan |
NeurIPS | 7 |
| 2024 | Mask-aware transformer with structure invariant loss for CT translation
Wenting Chen, Wei Zhao 0040, Zhen Chen 0013, Tianming Liu 0001, Li Liu 0017, Jun Liu 0007, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2024 | Comprehensive learning and adaptive teaching: Distilling multi-modal knowledge for pathological glioma grading
Xiaohan Xing, Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2023 | Combat Long-Tails in Medical Classification with Relation-Aware Consistency and Virtual Features Compensation
Li Pan 0004, Qiushi Yang, Tan Li 0002, Zhen Chen 0013 |
MICCAI (6) | 5 |
| 2023 | Gradient and Feature Conformity-Steered Medical Image Classification with Noisy Labels
Xiaohan Xing, Zhen Chen 0013, Zhifan Gao, Yixuan Yuan |
MICCAI (6) | 2 |
| 2023 | Generalized Gradient Flow Based Saliency for Pruning Deep Convolutional Neural Networks
Xinyu Liu 0001, Baopu Li, Zhen Chen 0013, Yixuan Yuan |
Int. J. Comput. Vis. | 3 |
| 2023 | Medical federated learning with joint graph purification for noisy label learning
Zhen Chen 0013, Wuyang Li, Xiaohan Xing, Yixuan Yuan |
Medical Image Anal. | 1 |
| 2023 | Gradient modulated contrastive distillation of low-rank multi-modal knowledge for disease diagnosis
Xiaohan Xing, Zhen Chen 0013, Yuenan Hou, Yixuan Yuan |
Medical Image Anal. | 2 |
| 2023 | Hierarchical Bias Mitigation for Semi-Supervised Medical Image ClassificationabstractSemi-supervised learning (SSL) has demonstrated remarkable advances on medical image classification, by harvesting beneficial knowledge from abundant unlabeled samples. The pseudo labeling dominates current SSL approaches, however, it suffers from intrinsic biases within the process. In this paper, we retrospect the pseudo labeling and identify three hierarchical biases: perception bias, selection bias and confirmation bias, at feature extraction, pseudo label selection and momentum optimization stages, respectively. In this regard, we propose a HierArchical BIas miTigation (HABIT) framework to amend these biases, which consists of three customized modules including Mutual Reconciliation Network (MRNet), Recalibrated Feature Compensation (RFC) and Consistency-aware Momentum Heredity (CMH). Firstly, in the feature extraction, MRNet is devised to jointly utilize convolution and permutator-based paths with a mutual information transfer module to exchanges features and reconcile spatial perception bias for better representations. To address pseudo label selection bias, RFC adaptively recalibrates the strong and weak augmented distributions to be a rational discrepancy and augments features for minority categories to achieve the balanced training. Finally, in the momentum optimization stage, in order to reduce the confirmation bias, CMH models the consistency among different sample augmentations into network updating process to improve the dependability of the model. Extensive experiments on three semi-supervised medical image classification datasets demonstrate that HABIT mitigates three biases and achieves state-of-the-art performance. Our codes are available at https://github.com/CityU-AIM-Group/HABIT. Qiushi Yang, Zhen Chen 0013, Yixuan Yuan |
IEEE Trans. Medical Imaging | 2 |
| 2023 | FedDM: Federated Weakly Supervised Segmentation via Annotation Calibration and Gradient De-ConflictingabstractWeakly supervised segmentation (WSS) aims to exploit weak forms of annotations to achieve the segmentation training, thereby reducing the burden on annotation. However, existing methods rely on large-scale centralized datasets, which are difficult to construct due to privacy concerns on medical data. Federated learning (FL) provides a cross-site training paradigm and shows great potential to address this problem. In this work, we represent the first effort to formulate federated weakly supervised segmentation (FedWSS) and propose a novel Federated Drift Mitigation (FedDM) framework to learn segmentation models across multiple sites without sharing their raw data. FedDM is devoted to solving two main challenges (i.e., local drift on client-side optimization and global drift on server-side aggregation) caused by weak supervision signals in FL setting via Collaborative Annotation Calibration (CAC) and Hierarchical Gradient De-conflicting (HGD). To mitigate the local drift, CAC customizes a distal peer and a proximal peer for each client via a Monte Carlo sampling strategy, and then employs inter-client knowledge agreement and disagreement to recognize clean labels and correct noisy labels, respectively. Moreover, in order to alleviate the global drift, HGD online builds a client hierarchy under the guidance of history gradient of the global model in each communication round. Through de-conflicting clients under the same parent nodes from bottom layers to top layers, HGD achieves robust gradient aggregation at the server side. Furthermore, we theoretically analyze FedDM and conduct extensive experiments on public datasets. The experimental results demonstrate the superior performance of our method compared with state-of-the-art approaches. The source code is available at https://github.com/CityU-AIM-Group/FedDM. Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Discrepancy and Gradient-Guided Multi-modal Knowledge Distillation for Pathological Glioma Grading
Xiaohan Xing, Zhen Chen 0013, Meilu Zhu, Yuenan Hou, Zhifan Gao, Yixuan Yuan |
MICCAI (5) | 2 |
| 2022 | Semi-supervised Medical Image Classification with Temporal Knowledge-Aware Regularization
Qiushi Yang, Xinyu Liu 0001, Zhen Chen 0013, Bulat Ibragimov, Yixuan Yuan |
MICCAI (8) | 3 |
| 2022 | Instance importance-Aware graph convolutional network for 3D medical diagnosis
Zhen Chen 0013, Jie Liu 0044, Meilu Zhu, Yat Ming Peter Woo, Yixuan Yuan |
Medical Image Anal. | 1 |
| 2022 | Non-equivalent images and pixels: Confidence-aware resampling with meta-learning mixup for polyp segmentation
Xiaoqing Guo, Zhen Chen 0013, Jun Liu 0007, Yixuan Yuan |
Medical Image Anal. | 2 |
| 2022 | Source free domain adaptation for medical image segmentation with fourier style mining
Chen Yang 0026, Xiaoqing Guo, Zhen Chen 0013, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2022 | Personalized Retrogress-Resilient Federated Learning Toward Imbalanced Medical DataabstractClinically oriented deep learning algorithms, combined with large-scale medical datasets, have significantly promoted computer-aided diagnosis. To address increasing ethical and privacy issues, Federated Learning (FL) adopts a distributed paradigm to collaboratively train models, rather than collecting samples from multiple institutions for centralized training. Despite intensive research on FL, two major challenges are still existing when applying FL in the real-world medical scenarios, including the performance degradation (i.e., retrogress) after each communication and the intractable class imbalance. Thus, in this paper, we propose a novel personalized FL framework to tackle these two problems. For the retrogress problem, we first devise a Progressive Fourier Aggregation (PFA) at the server side to gradually integrate parameters of client models in the frequency domain. Then, at the client side, we design a Deputy-Enhanced Transfer (DET) to smoothly transfer global knowledge to the personalized local model. For the class imbalance problem, we propose the Conjoint Prototype-Aligned (CPA) loss to facilitate the balanced optimization of the FL framework. Considering the inaccessibility of private local data to other participants in FL, the CPA loss calculates the global conjoint objective based on global imbalance, and then adjusts the client-side local training through the prototype-aligned refinement to eliminate the imbalance gap with such a balanced goal. Extensive experiments are performed on real-world dermoscopic and prostate MRI FL datasets. The experimental results demonstrate the advantages of our FL framework in real-world medical scenarios, by outperforming state-of-the-art FL methods with a large margin. The source code is available at https://github.com/CityU-AIM-Group/PRR-Imbalancehttps://github.com/CityU-AIM-Group/PRR-Imbalance. Zhen Chen 0013, Chen Yang 0026, Meilu Zhu, Zhe Peng, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2022 | D2-Net: Dual Disentanglement Network for Brain Tumor Segmentation With Missing ModalitiesabstractMulti-modal Magnetic Resonance Imaging (MRI) can provide complementary information for automatic brain tumor segmentation, which is crucial for diagnosis and prognosis. While missing modality data is common in clinical practice and it can result in the collapse of most previous methods relying on complete modality data. Current state-of-the-art approaches cope with the situations of missing modalities by fusing multi-modal images and features to learn shared representations of tumor regions, which often ignore explicitly capturing the correlations among modalities and tumor regions. Inspired by the fact that modality information plays distinct roles to segment different tumor regions, we aim to explicitly exploit the correlations among various modality-specific information and tumor-specific knowledge for segmentation. To this end, we propose a Dual Disentanglement Network (D2-Net) for brain tumor segmentation with missing modalities, which consists of amodality disentanglement stage(MD-Stage) and atumor-region disentanglement stage(TD-Stage). In the MD-Stage, a spatial-frequency joint modality contrastive learning scheme is designed to directly decouple the modality-specific information from MRI data. To decompose tumor-specific representations and extract discriminative holistic features, we propose an affinity-guided dense tumor-region knowledge distillation mechanism in the TD-Stage through aligning the features of a disentangled binary teacher network with a holistic student network. By explicitly discovering relations among modalities and tumor regions, our model can learn sufficient information for segmentation even if some modalities are missing. Extensive experiments on the public BraTS-2018 database demonstrate the superiority of our framework over state-of-the-art methods in missing modalities situations. Codes are available athttps://github.com/CityU-AIM-Group/D2Net. Qiushi Yang, Xiaoqing Guo, Zhen Chen 0013, Yat Ming Peter Woo, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Diagnose Like A Pathologist: Weakly-Supervised Pathologist-Tree Network for Slide-Level Immunohistochemical ScoringabstractThe immunohistochemistry (IHC) test of biopsy tissue is crucial to develop targeted treatment and evaluate prognosis for cancer patients. The IHC staining slide is usually digitized into the whole-slide image (WSI) with gigapixels for quantitative image analysis. To perform a whole image prediction (e.g., IHC scoring, survival prediction, and cancer grading) from this kind of high-dimensional image, algorithms are often developed based on multi-instance learning (MIL) framework. However, the multi-scale information of WSI and the associations among instances are not well explored in existing MIL based studies. Inspired by the fact that pathologists jointly analyze visual fields at multiple powers of objective for diagnostic predictions, we propose a Pathologist-Tree Network (PTree-Net) to sparsely model the WSI efficiently in multi-scale manner. Specifically, we propose a Focal-Aware Module (FAM) that can approximately estimate diagnosis-related regions with an extractor trained using the thumbnail of WSI. With the initial diagnosis-related regions, we hierarchically model the multi-scale patches in a tree structure, where both the global and local information can be captured. To explore this tree structure in an end-to-end network, we propose a patch Relevance-enhanced Graph Convolutional Network (RGCN) to explicitly model the correlations of adjacent parent-child nodes, accompanied by patch relevance to exploit the implicit contextual information among distant nodes. In addition, tree-based self-supervision is devised to improve representation learning and suppress irrelevant instances adaptively. Extensive experiments are performed on a large-scale IHC HER2 dataset. The ablation study confirms the effectiveness of our design, and our approach outperforms state-of-the-art by a large margin. Zhen Chen 0013, Jun Zhang 0018, Shuanlong Che, Junzhou Huang, Xiao Han 0011, Yixuan Yuan |
AAAI | 1 |
| 2021 | Personalized Retrogress-Resilient Framework for Real-World Medical Federated Learning
Zhen Chen 0013, Meilu Zhu, Chen Yang 0026, Yixuan Yuan |
MICCAI (3) | 1 |
| 2021 | Exploring Gradient Flow Based Saliency for DNN Model CompressionabstractModel pruning aims to reduce the deep neural network (DNN) model size or computational overhead. Traditional model pruning methods such as l-1 pruning that evaluates the channel significance for DNN pay too much attention to the local analysis of each channel and make use of the magnitude of the entire feature while ignoring its relevance to the batch normalization (BN) and ReLU layer after each convolutional operation. To overcome these problems, we propose a new model pruning method from a new perspective of gradient flow in this paper. Specifically, we first theoretically analyze the channel's influence based on Taylor expansion by integrating the effects of BN layer and ReLU activation function. Then, the incorporation of the first-order Talyor polynomial of the scaling parameter and the shifting parameter in the BN layer is suggested to effectively indicate the significance of a channel in a DNN. Comprehensive experiments on both image classification and image denoising tasks demonstrate the superiority of the proposed novel theory and scheme. Code is available at https://github.com/CityU-AIM-Group/GFBS. Xinyu Liu 0001, Baopu Li, Zhen Chen 0013, Yixuan Yuan |
ACM Multimedia | 3 |
| 2021 | HTD: Heterogeneous Task Decoupling for Two-Stage Object DetectionabstractDecoupling the sibling head has recently shown great potential in relieving the inherent task-misalignment problem in two-stage object detectors. However, existing works design similar structures for the classification and regression, ignoring task-specific characteristics and feature demands. Besides, the shared knowledge that may benefit the two branches is neglected, leading to potential excessive decoupling and semantic inconsistency. To address these two issues, we propose Heterogeneous task decoupling (HTD) framework for object detection, which utilizes a Progressive Graph (PGraph) module and a Border-aware Adaptation (BA) module for task-decoupling. Specifically, we first devise a Semantic Feature Aggregation (SFA) module to aggregate global semantics with image-level supervision, serving as the shared knowledge for the task-decoupled framework. Then, the PGraph module performs progressive graph reasoning, including local spatial aggregation and global semantic interaction, to enhance semantic representations of region proposals for classification. The proposed BA module integrates multi-level features adaptively, focusing on the low-level border activation to obtain representations with spatial and border perception for regression. Finally, we utilize the aggregated knowledge from SFA to keep the instance-level semantic consistency (ISC) of decoupled frameworks. Extensive experiments demonstrate that HTD outperforms existing detection works by a large margin, and achieves single-model 50.4%AP and 33.2% APs on COCO test-dev set using ResNet-101-DCN backbone, which is the best entry among state-of-the-arts under the same configuration. Our code is available at https://github.com/CityU-AIM-Group/HTD. Wuyang Li, Zhen Chen 0013, Baopu Li, Dingwen Zhang, Yixuan Yuan |
IEEE Trans. Image Process. | 2 |
| 2021 | Super-Resolution Enhanced Medical Image Diagnosis With Sample Affinity InteractionabstractThe degradation in image resolution harms the performance of medical image diagnosis. By inferring high-frequency details from low-resolution (LR) images, super-resolution (SR) techniques can introduce additional knowledge and assist high-level tasks. In this paper, we propose a SR enhanced diagnosis framework, consisting of an efficient SR network and a diagnosis network. Specifically, a Multi-scale Refined Context Network (MRC-Net) with Refined Context Fusion (RCF) is devised to leverage global and local features for SR tasks. Instead of learning from scratch, we first develop a recursive MRC-Net with temporal context, and then propose a recursion distillation scheme to enhance the performance of MRC-Net from the knowledge of the recursive one and reduce the computational cost. The diagnosis network jointly utilizes the reliable original images and more informative SR images by two branches, with the proposed Sample Affinity Interaction (SAI) blocks at different stages to effectively extract and integrate discriminative features towards diagnosis. Moreover, two novel constraints, sample affinity consistency and sample affinity regularization, are devised to refine the features and achieve the mutual promotion of these two branches. Extensive experiments of synthetic and real LR cases are conducted on wireless capsule endoscopy and histopathology images, verifying that our proposed method is significantly effective for medical image diagnosis. Zhen Chen 0013, Xiaoqing Guo, Yat Ming Peter Woo, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2021 | DSI-Net: Deep Synergistic Interaction Network for Joint Classification and Segmentation With Endoscope ImagesabstractAutomatic classification and segmentation of wireless capsule endoscope (WCE) images are two clinically significant and relevant tasks in a computer-aided diagnosis system for gastrointestinal diseases. Most of existing approaches, however, considered these two tasks individually and ignored their complementary information, leading to limited performance. To overcome this bottleneck, we propose a deep synergistic interaction network (DSI-Net) for joint classification and segmentation with WCE images, which mainly consists of the classification branch (C-Branch), the coarse segmentation (CS-Branch) and the fine segmentation branches (FS-Branch). In order to facilitate the classification task with the segmentation knowledge, a lesion location mining (LLM) module is devised in C-Branch to accurately highlight lesion regions through mining neglected lesion areas and erasing misclassified background areas. To assist the segmentation task with the classification prior, we propose a category-guided feature generation (CFG) module in FS-Branch to improve pixel representation by leveraging the category prototypes of C-Branch to obtain the category-aware features. In such way, these modules enable the deep synergistic interaction between these two tasks. In addition, we introduce a task interaction loss to enhance the mutual supervision between the classification and segmentation tasks and guarantee the consistency of their predictions. Relying on the proposed deep synergistic interaction mechanism, DSI-Net achieves superior classification and segmentation performance on public dataset in comparison with state-of-the-art methods. The source code is available at https://github.com/CityU-AIM-Group/DSI-Net. Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Joint Spatial-Wavelet Dual-Stream Network for Super-Resolution
Zhen Chen 0013, Xiaoqing Guo, Chen Yang 0026, Bulat Ibragimov, Yixuan Yuan |
MICCAI (5) | 1 |
| 2019 | Exploiting Weight-Level Sparsity in Channel Pruning with Low-Rank ApproximationabstractAcceleration and compression on Deep Neural Networks (DNNs) have become a critical problem to develop intelligence on resource-constrained hardware, especially on Internet of Things (IoT) devices. Previous works based on channel pruning can be easily deployed and accelerated without specialized hardware and software. However, weight-level sparsity is not well explored in channel pruning, which results in relatively low compression rate. In this work, we propose a framework that combines channel pruning with low-rank decomposition to tackle this problem. First, the low-rank decomposition is utilized to eliminate redundancy within filter, and achieves acceleration in shallow layers. Then, we apply channel pruning on the decomposed network in a global way, and obtains further acceleration in deep layers. In addition, a spectral norm-based indicator is proposed to balance low-rank approximation and channel pruning. We conduct a series of ablation experiments and prove that low-rank decomposition can effectively improve channel pruning by generating small and compact filters. To further demonstrate the hardware compatibility, we deploy the pruned networks on the FPGA, and the networks produced by our method have obviously low latency. Zhen Chen 0013, Sen Liu 0001, Zhibo Chen 0001, Weiping Li 0003 |
ISCAS | 1 |
| 2019 | Importance-Aware Filter Selection for Convolutional Neural Network AccelerationabstractConvolutional Neural Networks(CNNs) are widely used in many fields, including artificial intelligence, computer vision and video coding. However, CNNs are typically over-parameterized and contain significant redundancy. Traditional model acceleration methods mainly rely on specific manual rules. This usually leads to sub-optimal results with relatively limited compression ratio. Recent works have deployed the self-learning agent on the layer-level acceleration but still combined with human-designed criterias. In this paper, we proposed a filter-based model acceleration method to directly and automatically decide which filters should be pruned with the reinforcement learning method DDPG. We designed a novel reward function with the reward shaping technique for the training process. Our method is utilized on the models trained on MNIST and CIFAR-10 datasets and achieves both higher acceleration ratio and less accuracy loss than the conventional methods simultaneously. Zikun Liu 0002, Zhen Chen 0013, Weiping Li 0003 |
VCIP | 2 |
| 2018 | Decouple and Stretch: A Boost to Channel PruningabstractDeep Neural Networks (DNNs) have shown superior performance on a variety of artificial intelligence problems. Reducing the resource usage of DNN is critical to adding intelligence on Internet of Things (IoT) devices. Channel pruning based network compression shows effective reduction simultaneously on storage, memory and computation without specialized software on general platforms. But limited by pruning flexibility, channel pruning methods have relatively low compression rate for a given target performance. In this paper, we demonstrate that channel pruning becomes more robust to decision errors by reducing the granularity of filters. Then we propose a Decouple and Stretch (DS) scheme to enhance channel pruning. Under this scheme, each filter in a specific layer is decoupled into two small spatial-wise filters, and the spatial-wise filters are stretched into two successive convolutional layers. Our scheme obtains up to 49% improvement on compression and 35% improvement on acceleration. To further demonstrate hardware compatibility, we deploy pruned networks on the FPGA, and the network produced by Decouple and Stretch scheme is more hardware-friendly with latency reduced by 42%. Zhen Chen 0013, Sen Liu 0001, Jun Xia 0003, Weiping Li 0003 |
IPCCC | 1 |