VLDB 2026 Research / reviewers in the wild / expert
Zengqiang Yan
dblp:172/4640
· DBLP profile ↗
43ranked-venue papers
7as first author
35since 2021 · last 2026
0000-0002-2039-3863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedRNC: Addressing Spatio-Temporal Label Misalignment in Federated Noisy Class-Incremental LearningabstractFederated class-incremental learning (FCIL) aims to incrementally learn new classes across decentralized clients under non-IID data distributions. However, the pervasive challenge of label noise in FCIL has been completely overlooked. In this work, we introduce federated noisy class-incremental learning (FNCIL) and, for the first time, identify a novel form of label noise—spatio-temporal label misalignment—where samples from unseen classes are entirely mislabeled as known classes, with their correctly labeled counterparts appearing in latter tasks or other clients. This phenomenon undermines the effectiveness of existing centralized denoising strategies and creates a clear requirement for noise-robust methods in real-world FNCIL scenarios. To tackle this issue, we propose FedRNC, a dual-phase framework that leverages feature-space associations to establish spatio-temporal correspondences between clean global prototypes and noisy cached samples for progressive label correction. Experiments on standard benchmarks demonstrate FedRNC's superiority against existing baselines, along with its plug-and-play capability to upgrade FCIL systems for FNCIL. Xingwei Huang, Zhaobin Sun, Xin Yang 0008, Zengqiang Yan |
AAAI | 5 |
| 2026 | Selective intra- and inter-slice interaction for efficient anisotropic medical image segmentation
Xian Lin, Xiayu Guo, Zengqiang Yan, Li Yu 0003 |
Pattern Recognit. | 3 |
| 2026 | FedHAC: Towards Robust Federated Multi-Lesion Segmentation With Heterogeneous Annotation CompletenessabstractFederated learning (FL) has emerged as a promising paradigm for collaborative medical image segmentation across institutions while preserving data privacy. Despite great efforts in addressing cross-client annotation heterogeneity FL, the prevalent annotation completeness heterogeneity in clinical practice due to varying diagnostic priorities has been completely overlooked, hindering the deployment of FL. In this paper, we formulate such a challenge and propose FedHAC for incompleteness-robust medical image segmentation. FedHAC consists of three modules, i.e., Global Class Prototype Alignment (GCPA), Annotation Completeness-Aware Aggregation (ACAA), and GMM-driven Progressive Correction (GPC). Specifically, GCPA constructs a noise-resilient warm-up model through proximal-term regularization and prototype alignment. ACAA estimates client-wise annotation completeness and dynamically prioritizes high-quality clients. GPC groups clients into "noisy" and "clean" via GMM for progressive annotation correction to minimize error propagation. Extensive comparison experiments and ablation studies on public datasets demonstrate the superiority of FedHAC over state-of-the-art methods under various levels of annotation incompleteness. Yangyang Xiang, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Addressing Imbalanced Modal Incompleteness in Realistic Multi-Modal Medical Image Segmentation via Hierarchical Gradient AlignmentabstractDespite the promising potential of multi-modal learning in medical image segmentation, real-world applications often encounter modal incompleteness sourced from diverse domains and institutions, sparking significant discussions on incomplete multi-modal learning. Existing approaches either train a unified model for all or develop individual models for specific multi-modal combinations to ensure model fairness and robustness during inference. However, the assumption of complete multi-modal data for training is unrealistic and infeasible in clinical practice. In this paper, we thoroughly formulate such a challenging setting and propose hierarchical gradient alignment (HGA) to address uni- and multi-modal imbalance. Specifically, gradient direction is aligned through sequential meta learning for multi-modal combinations and multi-level self-distillation for uni-modals within each combination. Gradient magnitude is aligned based on relative preference estimation to balance the dominance of each modal during training. Extensive experiments on five public benchmarks (BraTS2018, BraTS2020, BraTS2023, MyoPS2020, and MSSEG2016) demonstrate that HGA consistently outperforms state-of-the-art incomplete and imbalanced multi-modal learning methods, as well as representative multi-task learning optimization techniques. More importantly, HGA is validated to work as plug-and-play modules for consistent performance improvement across different backbones. Code is available at https://github.com/Jun-Jie-Shi/HGA. Zhaobin Sun, Li Yu 0003, Xin Yang 0008, Zengqiang Yan |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Towards Robust Medical Image Referring Segmentation with Incomplete Textual Prompts
Qijie Wang, Xian Lin, Zengqiang Yan |
MICCAI (7) | 3 |
| 2025 | ROXSI: Robust Cross-Sequence Semantic Interaction for Brain Tumor Segmentation on Multi-Sequence MR ImagesabstractDeep learning-based brain tumor segmentation on multi-sequence magnetic resonance imaging (MRI) has gained widespread attention due to its great potential in supporting brain disease diagnosis. Although, compared to single-sequence images, more information is available from multi-sequence MR images, noise and artifacts on any given MR sequence can result in significant performance degradations. As in clinical routine, it is not always possible to maintain high imaging quality across all MR sequences (e.g., foreign bodies, ventricular drainage, shunts, involuntary patient motion, etc.), ensuring robustness of brain tumor segmentation from multi-sequence MR images is of great importance in clinical practice, but rarely explored. Accordingly, in this paper, we propose a robust brain tumor segmentation framework to mitigate the performance degradation caused by noise and artifacts on multi-sequence MR images. Specifically, based on semantic affinity, we propose a unique cross-sequence semantic interaction module (CSSI) to exploit inter-sequence correlations and extract noise-resilient features. In addition, we incorporate a batch-level covariance mechanism to suppress the redundant background information and improve the semantic enhancement effect of the CSSI module. In order to further improve segmentation performance, we also incorporate a sequence-level variance regularization mechanism to exploit sequence-specific features. To validate the robustness of ROXSI, brain tumor segmentation performance was evaluated under the existence of four common artifacts, at five different perturbation levels. We further performed a blinded qualitative clinical evaluation with two experienced neuro-radiologists, evaluating results from ROXSI and other popular CNN and Transformer-based segmentation models. Experimental results on two benchmark datasets demonstrate the superior robustness of ROXSI over other state-of-the-art segmentation methods. Zhuo Kuang, Zengqiang Yan, Aly Abayazeed, Franca Wagner, Li Yu 0003, Mauricio Reyes 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | SAMCT: Segment Any CT Allowing Labor-Free Task-Indicator PromptsabstractSegment anything model (SAM), a foundation model with superior versatility and generalization across diverse segmentation tasks, has attracted widespread attention in medical imaging. However, it has been proved that SAM would encounter severe performance degradation due to the lack of medical knowledge in training and local feature encoding. Though several SAM-based models have been proposed for tuning SAM in medical imaging, they still suffer from insufficient feature extraction and highly rely on high-quality prompts. In this paper, we propose a powerful foundation model SAMCT allowing labor-free prompts and train it on a collected large CT dataset consisting of 1.1M CT images and 5M masks from public datasets. Specifically, based on SAM, SAMCT is further equipped with a U-shaped CNN image encoder, a cross-branch interaction module, and a task-indicator prompt encoder. The U-shaped CNN image encoder works in parallel with the ViT image encoder in SAM to supplement local features. Cross-branch interaction enhances the feature expression capability of the CNN image encoder and the ViT image encoder by exchanging global perception and local features from one to the other. The task-indicator prompt encoder is a plug-and-play component to effortlessly encode task-related indicators into prompt embeddings. In this way, SAMCT can work in an automatic manner in addition to the semi-automatic interactive strategy in SAM. Extensive experiments demonstrate the superiority of SAMCT against the state-of-the-art task-specific and SAM-based medical foundation models on various tasks. The code, data, and model checkpoints are available at https://github.com/xianlin7/SAMCT. Xian Lin, Yangyang Xiang, Zhehao Wang, Kwang-Ting Cheng, Zengqiang Yan, Li Yu 0003 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | DTMFormer: Dynamic Token Merging for Boosting Transformer-Based Medical Image SegmentationabstractDespite the great potential in capturing long-range dependency, one rarely-explored underlying issue of transformer in medical image segmentation is attention collapse, making it often degenerate into a bypass module in CNN-Transformer hybrid architectures. This is due to the high computational complexity of vision transformers requiring extensive training data while well-annotated medical image data is relatively limited, resulting in poor convergence. In this paper, we propose a plug-n-play transformer block with dynamic token merging, named DTMFormer, to avoid building long-range dependency on redundant and duplicated tokens and thus pursue better convergence. Specifically, DTMFormer consists of an attention-guided token merging (ATM) module to adaptively cluster tokens into fewer semantic tokens based on feature and dependency similarity and a light token reconstruction module to fuse ordinary and semantic tokens. In this way, as self-attention in ATM is calculated based on fewer tokens, DTMFormer is of lower complexity and more friendly to converge. Extensive experiments on publicly-available datasets demonstrate the effectiveness of DTMFormer working as a plug-n-play module for simultaneous complexity reduction and performance improvement. We believe it will inspire future work on rethinking transformers in medical image segmentation. Code: https://github.com/iam-nacl/DTMFormer. Zhehao Wang, Xian Lin, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
AAAI | 6 |
| 2024 | FedA3I: Annotation Quality-Aware Aggregation for Federated Medical Image Segmentation against Heterogeneous Annotation NoiseabstractFederated learning (FL) has emerged as a promising paradigm for training segmentation models on decentralized medical data, owing to its privacy-preserving property. However, existing research overlooks the prevalent annotation noise encountered in real-world medical datasets, which limits the performance ceilings of FL. In this paper, we, for the first time, identify and tackle this problem. For problem formulation, we propose a contour evolution for modeling non-independent and identically distributed (Non-IID) noise across pixels within each client and then extend it to the case of multi-source data to form a heterogeneous noise model (i.e., Non-IID annotation noise across clients). For robust learning from annotations with such two-level Non-IID noise, we emphasize the importance of data quality in model aggregation, allowing high-quality clients to have a greater impact on FL. To achieve this, we propose Federated learning with Annotation quAlity-aware AggregatIon, named FedA3I, by introducing a quality factor based on client-wise noise estimation. Specifically, noise estimation at each client is accomplished through the Gaussian mixture model and then incorporated into model aggregation in a layer-wise manner to up-weight high-quality clients. Extensive experiments on two real-world medical image segmentation datasets demonstrate the superior performance of FedA3I against the state-of-the-art approaches in dealing with cross-client annotation noise. The code is available at https://github.com/wnn2000/FedAAAI. Zhaobin Sun, Zengqiang Yan, Li Yu 0003 |
AAAI | 3 |
| 2024 | From Optimization to Generalization: Fair Federated Learning against Quality Shift via Inter-Client Sharpness Matching
Zhuo Kuang, Zengqiang Yan, Li Yu 0003 |
IJCAI | 3 |
| 2024 | Revisiting Self-attention in Medical Transformers via Dependency Sparsification
Xian Lin, Zhehao Wang, Zengqiang Yan, Li Yu 0003 |
MICCAI (11) | 3 |
| 2024 | Beyond Adapting SAM: Towards End-to-End Ultrasound Image Segmentation via Auto Prompting
Xian Lin, Yangyang Xiang, Li Yu 0003, Zengqiang Yan |
MICCAI (8) | 4 |
| 2024 | FedMLP: Federated Multi-label Medical Image Classification Under Task Heterogeneity
Zhaobin Sun, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (10) | 6 |
| 2024 | FedIA: Federated Medical Image Segmentation with Heterogeneous Annotation Completeness
Yangyang Xiang, Li Yu 0003, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (10) | 6 |
| 2024 | PASSION: Towards Effective Incomplete Multi-Modal Medical Image Segmentation with Imbalanced Missing Rates
Caozhi Shang, Zhaobin Sun, Li Yu 0003, Xin Yang 0008, Zengqiang Yan |
ACM Multimedia | 6 |
| 2024 | Boosting integral-based human pose estimation through implicit heatmap learning
Congju Du, Zengqiang Yan, Zixiang Xiong, Li Yu 0003 |
Neural Networks | 2 |
| 2024 | UCTNet: Uncertainty-guided CNN-Transformer hybrid networks for medical image segmentation
Xiayu Guo, Xian Lin, Xin Yang 0008, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
Pattern Recognit. | 6 |
| 2024 | MFTrans: Modality-Masked Fusion Transformer for Incomplete Multi-Modality Brain Tumor SegmentationabstractBrain tumor segmentation is a fundamental task and existing approaches usually rely on multi-modality magnetic resonance imaging (MRI) images for accurate segmentation. However, the common problem of missing/incomplete modalities in clinical practice would severely degrade their segmentation performance, and existing fusion strategies for incomplete multi-modality brain tumor segmentation are far from ideal. In this work, we propose a novel framework named M$^{2}$FTrans to explore and fuse cross-modality features through modality-masked fusion transformers under various incomplete multi-modality settings. Considering vanilla self-attention is sensitive to missing tokens/inputs, both learnable fusion tokens and masked self-attention are introduced to stably build long-range dependency across modalities while being more flexible to learn from incomplete modalities. In addition, to avoid being biased toward certain dominant modalities, modality-specific features are further re-weighted through spatial weight attention and channel-wise fusion transformers for feature redundancy reduction and modality re-balancing. In this way, the fusion strategy in M$^{2}$FTrans is more robust to missing modalities. Experimental results on the widely-used BraTS2018, BraTS2020, and BraTS2021 datasets demonstrate the effectiveness of M$^{2}$FTrans, outperforming the state-of-the-art approaches with large margins under various incomplete modalities for brain tumor segmentation. Li Yu 0003, Qimin Cheng, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | FedIOD: Federated Multi-Organ Segmentation From Partial Labels by Exploring Inter-Organ DependencyabstractMulti-organ segmentation is a fundamental task and existing approaches usually rely on large-scale fully-labeled images for training. However, data privacy and incomplete/partial labels make those approaches struggle in practice. Federated learning is an emerging tool to address data privacy but federated learning with partial labels is under-explored. In this work, we explore generating full supervision by building and aggregating inter-organ dependency based on partial labels and propose a single-encoder-multi-decoder framework named FedIOD. To simulate the annotation process where each organ is labeled by referring to other closely-related organs, a transformer module is introduced and the learned self-attention matrices modeling pairwise inter-organ dependency are used to build pseudo full labels. By using those pseudo-full labels for regularization in each client, the shared encoder is trained to extract rich and complete organ-related features rather than being biased toward certain organs. Then, each decoder in FedIOD projects the shared organ-related features into a specific space trained by the corresponding partial labels. Experimental results based on five widely-used datasets, including LiTS, KiTS, MSD, BCTV, and ACDC, demonstrate the effectiveness of FedIOD, outperforming the state-of-the-art approaches under in-federation evaluation and achieving the second-best performance under out-of-federation evaluation for multi-organ segmentation from partial labels. Qin Wan 0002, Zengqiang Yan, Li Yu 0003 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | APCAFlow: All-Pairs Cost Volume Aggregation for Optical Flow EstimationabstractOptical flow estimation is a fundamental task in computer vision. The all-pairs correlation volume has enabled state-of-the-art performance in many optical flow estimation methods. However, all-pairs correlations provide only local matching clues, and lack global context, which could lead to mismatches in textureless and occluded regions. In this paper, we propose a novel all-pairs correlation volume aggregation (APCA) method which includes two key innovations. The first is a cost volume splitting and reassembling approach which partitions the full cost volume into smaller blocks and re-arranges those blocks to allow the use of 2D and 3D convolutions for cost volume aggregation. The second is hierarchical aggregation which performs 2D convolutions within blocks for local matching aggregation and 3D convolutions across blocks for global matching aggregation. We further design a novel optical flow estimation network APCAFlow based on APCA. APCAFlow achieves comparable performance to the most advanced approach, FlowFormer, but with significantly lower complexity. Specifically, APCAFlow reduces the model parameters, inference time, and memory consumption by 24.1%, 35.5%, and 21.6%, respectively, compared to FlowFormer. Furthermore, APCA can be easily integrated into several existing all-pairs cost volume-based methods for performance improvement. Code is available athttps://github.com/MiaoJieF/APCAFlow. Miaojie Feng, Zengqiang Yan, Xin Yang 0008 |
IEEE Trans. Multim. | 3 |
| 2023 | FedNoRo: Towards Noise-Robust Federated Learning by Addressing Class Imbalance and Label Noise HeterogeneityabstractFederated noisy label learning (FNLL) is emerging as a promising tool for privacy-preserving multi-source decentralized learning. Existing research, relying on the assumption of class-balanced global data, might be incapable to model complicated label noise, especially in medical scenarios. In this paper, we first formulate a new and more realistic federated label noise problem where global data is class-imbalanced and label noise is heterogeneous, and then propose a two-stage framework named FedNoRo for noise-robust federated learning. Specifically, in the first stage of FedNoRo, per-class loss indicators followed by Gaussian Mixture Model are deployed for noisy client identification. In the second stage, knowledge distillation and a distance-aware aggregation function are jointly adopted for noise-robust federated model updating. Experimental results on the widely-used ICH and ISIC2019 datasets demonstrate the superiority of FedNoRo against the state-of-the-art FNLL methods for addressing class imbalance and label noise heterogeneity in real-world FL scenarios. Li Yu 0003, Xuefeng Jiang 0001, Kwang-Ting Cheng, Zengqiang Yan |
IJCAI | 5 |
| 2023 | ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation
Xian Lin, Zengqiang Yan, Xianbo Deng, Chuansheng Zheng, Li Yu 0003 |
MICCAI (4) | 2 |
| 2023 | FedIIC: Towards Robust Federated Learning for Class-Imbalanced Medical Image Classification
Li Yu 0003, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (2) | 5 |
| 2023 | MMA-Net: Multi-view mixed attention mechanism for facial action unit detection
Ziqiao Shang, Congju Du, Bingyin Li, Zengqiang Yan, Li Yu 0003 |
Pattern Recognit. Lett. | 4 |
| 2023 | Hierarchical Associative Encoding and Decoding for Bottom-Up Human Pose EstimationabstractBottom-up human pose estimation decouples computational complexity from the number of people but requires additional operations to match the detected keypoints to each human instance. Existing approaches treat all keypoints equally while ignoring the relationships among keypoints, which in turn limit the performance ceilings. In this work, we propose a hierarchical associative encoding and decoding framework for bottom-up human pose estimation by introducing additional prior knowledge. Specifically, in addition to keypoint-level and instance-level associations, we further divide keypoints into groups and explore group-level associations. This way, prior knowledge is incorporated to determine the keypoint groups for better associative encoding. To deal with complex poses, we introduce a focal pulling loss to focus more on the hard-to-associate keypoints. Moreover, instead of using a pre-defined order for keypoint grouping, we propose a progressive associative decoding method to dynamically determine the order of keypoints for grouping, which helps reduce isolated keypoints. Experimental results on the MS-COCO, CrowdPose and MPII datasets show superior performance of our proposed associative encoding and decoding algorithms. More importantly, we prove, through validation, that hierarchical associative encoding and decoding can be used as a plug-n-play module for performance improvement regardless of backbone architecture. Our source code and pretrained models are available athttps://github.com/ducongju/HAE. Congju Du, Zengqiang Yan, Li Yu 0003, Zixiang Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Cluster-Re-Supervision: Bridging the Gap Between Image-Level and Pixel-Wise Labels for Weakly Supervised Medical Image SegmentationabstractWeakly supervised learning, releasing deep learning from highly labor-intensive pixel-wise annotations, has gained great attention, especially for medical image segmentation. With only image-level labels, pixel-wise segmentation/localization usually is achieved based on class activation maps (CAMs) containing the most discriminative regions. One common consequence of CAM-based approaches is incomplete foreground segmentation, i.e. under-segmentation/false negatives. Meanwhile, suffering from relatively limited medical imaging data, class-irrelevant tissues can hardly be suppressed during classification, resulting in incorrect background identification, i.e. over-segmentation/false positives. The above two issues are determined by the loose-constraint nature of image-level labels penalizing on the entire image space, and thus how to develop pixel-wise constraints based on image-level labels is the key for performance improvement which is under-explored. In this paper, based on unsupervised clustering, we propose a new paradigm called cluster-re-supervision to evaluate the contribution of each pixel in CAMs to final classification and thus generate pixel-wise supervision (i.e., clustering maps) for CAMs refinement on both over- and under-segmentation reduction. Furthermore, based on self-supervised learning, an inter-modality image reconstruction module, together with random masking, is designed to complement local information in feature learning which helps stabilize clustering. Experimental results on two popular public datasets demonstrate the superior performance of the proposed weakly-supervised framework for medical image segmentation. More importantly, cluster-re-supervision is independent of specific tasks and highly extendable to other applications. Zhuo Kuang, Zengqiang Yan, Huiyu Zhou 0001, Li Yu 0003 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | BATFormer: Towards Boundary-Aware Lightweight Transformer for Efficient Medical Image SegmentationabstractOBJECTIVE: Transformers, born to remedy the inadequate receptive fields of CNNs, have drawn explosive attention recently. However, the daunting computational complexity of global representation learning, together with rigid window partitioning, hinders their deployment in medical image segmentation. This work aims to address the above two issues in transformers for better medical image segmentation. METHODS: We propose a boundary-aware lightweight transformer (BATFormer) that can build cross-scale global interaction with lower computational complexity and generate windows flexibly under the guidance of entropy. Specifically, to fully explore the benefits of transformers in long-range dependency establishment, a cross-scale global transformer (CGT) module is introduced to jointly utilize multiple small-scale feature maps for richer global features with lower computational complexity. Given the importance of shape modeling in medical image segmentation, a boundary-aware local transformer (BLT) module is constructed. Different from rigid window partitioning in vanilla transformers which would produce boundary distortion, BLT adopts an adaptive window partitioning scheme under the guidance of entropy for both computational complexity reduction and shape preservation. RESULTS: BATFormer achieves the best performance in Dice of 92.84 %, 91.97 %, 90.26 %, and 96.30 % for the average, right ventricle, myocardium, and left ventricle respectively on the ACDC dataset and the best performance in Dice, IoU, and ACC of 90.76 %, 84.64 %, and 96.76 % respectively on the ISIC 2018 dataset. More importantly, BATFormer requires the least amount of model parameters and the lowest computational complexity compared to the state-of-the-art approaches. CONCLUSION AND SIGNIFICANCE: Our results demonstrate the necessity of developing customized transformers for efficient and better medical image segmentation. We believe the design of BATFormer is inspiring and extendable to other applications/frameworks. Xian Lin, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Affinity Feature Strengthening for Accurate, Complete and Robust Vessel SegmentationabstractVessel segmentation is crucial in many medical image applications, such as detecting coronary stenoses, retinal vessel diseases and brain aneurysms. However, achieving high pixel-wise accuracy, complete topology structure and robustness to various contrast variations are critical and challenging, and most existing methods focus only on achieving one or two of these aspects. In this paper, we present a novel approach, the affinity feature strengthening network (AFN), which jointly models geometry and refines pixel-wise segmentation features using a contrast-insensitive, multiscale affinity approach. Specifically, we compute a multiscale affinity field for each pixel, capturing its semantic relationships with neighboring pixels in the predicted mask image. This field represents the local geometry of vessel segments of different sizes, allowing us to learn spatial- and scale-aware adaptive weights to strengthen vessel features. We evaluate our AFN on four different types of vascular datasets: X-ray angiography coronary vessel dataset (XCAD), portal vein dataset (PV), digital subtraction angiography cerebrovascular vessel dataset (DSA) and retinal vessel dataset (DRIVE). Extensive experimental results demonstrate that our AFN outperforms the state-of-the-art methods in terms of both higher accuracy and topological metrics, while also being more robust to various contrast changes. Xiaohuan Ding, Wei Zhou 0068, Zengqiang Yan, Xiang Bai, Xin Yang 0008 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | The Lighter the Better: Rethinking Transformers in Medical Image Segmentation Through Adaptive PruningabstractVision transformers have recently set off a new wave in the field of medical image analysis due to their remarkable performance on various computer vision tasks. However, recent hybrid-/transformer-based approaches mainly focus on the benefits of transformers in capturing long-range dependency while ignoring the issues of their daunting computational complexity, high training costs, and redundant dependency. In this paper, we propose to employ adaptive pruning to transformers for medical image segmentation and propose a lightweight and effective hybrid network APFormer. To our best knowledge, this is the first work on transformer pruning for medical image analysis tasks. The key features of APFormer are self-regularized self-attention (SSA) to improve the convergence of dependency establishment, Gaussian-prior relative position embedding (GRPE) to foster the learning of position information, and adaptive pruning to eliminate redundant computations and perception information. Specifically, SSA and GRPE consider the well-converged dependency distribution and the Gaussian heatmap distribution separately as the prior knowledge of self-attention and position embedding to ease the training of transformers and lay a solid foundation for the following pruning operation. Then, adaptive transformer pruning, both query-wise and dependency-wise, is performed by adjusting the gate control parameters for both complexity reduction and performance improvement. Extensive experiments on two widely-used datasets demonstrate the prominent segmentation performance of APFormer against the state-of-the-art methods with much fewer parameters and lower GFLOPs. More importantly, we prove, through ablation studies, that adaptive pruning can work as a plug-n-play module for performance improvement on other hybrid-/transformer-based methods. Code is available at https://github.com/xianlin7/APFormer. Xian Lin, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
IEEE Trans. Medical Imaging | 4 |
| 2023 | FedMix: Mixed Supervised Federated Learning for Medical Image SegmentationabstractThe purpose of federated learning is to enable multiple clients to jointly train a machine learning model without sharing data. However, the existing methods for training an image segmentation model have been based on an unrealistic assumption that the training set for each local client is annotated in a similar fashion and thus follows the same image supervision level. To relax this assumption, in this work, we propose a label-agnostic unified federated learning framework, named FedMix, for medical image segmentation based on mixed image labels. In FedMix, each client updates the federated model by integrating and effectively making use of all available labeled data ranging from strong pixel-level labels, weak bounding box labels, to weakest image-level class labels. Based on these local models, we further propose an adaptive weight assignment procedure across local clients, where each client learns an aggregation weight during the global model update. Compared to the existing methods, FedMix not only breaks through the constraint of a single level of image supervision but also can dynamically adjust the aggregation weight of each local client, achieving rich yet discriminative feature representations. Experimental results on multiple publicly-available datasets validate that the proposed FedMix outperforms the state-of-the-art methods by a large margin. In addition, we demonstrate through experiments that FedMix is extendable to multi-class medical image segmentation and much more feasible in clinical scenarios. The code is available at: https://github.com/Jwicaksana/FedMix. Jeffry Wicaksana, Zengqiang Yan, Xijie Huang, Huimin Wu 0001, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Symmetry-Aware Deep Learning for Cerebral Ventricle Segmentation With Intra-Ventricular HemorrhageabstractCerebral ventricles are one of the prominent structures in the brain, segmenting which can provide rich information for brain-related disease diagnosis. Unfortunately, cerebral ventricle segmentation in complex clinical cases, such as in the coexistence with other lesions/hemorrhages, remains unexplored. In this paper, we, for the first time, focus on cerebral ventricle segmentation with the presence of intra-ventricular hemorrhages (IVH). To overcome the occlusions formed by IVH, we propose a symmetry-aware deep learning approach inspired by contrastive self-supervised learning. Specifically, for each slice, we jointly employ the raw slice and the horizontally flipped slice as inputs and penalize the consistency loss between the corresponding segmentation maps in addition to their segmentation losses. In this way, the symmetry of cerebral ventricles is enforced to eliminate the occlusions brought by IVH. Extensive experimental results show that the proposed symmetry-aware deep learning approach achieves consistent performance improvements for ventricle segmentation in both normal (i.e. without IVH) and challenging cases (i.e. with IVH). Through evaluation of multiple backbone networks, we demonstrate the architecture-independence of the proposed approach for performance improvements. Moreover, we re-design an end-to-end version of symmetry-aware deep learning, making it more extendable to other approaches for brain-related analysis. Yineng Hua, Zengqiang Yan, Zhuo Kuang, Xianbo Deng, Li Yu 0003 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Uncertainty-Aware Deep Learning With Cross-Task Supervision for PHE Segmentation on CT ImagesabstractPerihematomal edema (PHE) volume, surrounding spontaneous intracerebral hemorrhage (SICH), is an important biomarker for the presence of SICH-associated diseases. However, due to irregular shapes and extremely low contrast of PHE on CT images, manually annotating PHE in pixel-wise is time-consuming and labour intensive even for experienced experts, which makes it almost infeasible to deploy current supervised deep learning approaches for automated PHE segmentation. How to develop annotation-efficient deep learning to achieve accurate PHE segmentation is an open problem. In this paper, we, for the first time, propose a cross-task supervised framework by introducing slice-level PHE labels and pixel-wise SICH annotations, which are more accessible in clinical scenarios compared to pixel-wise PHE annotations. Specifically, we first train a multi-level classifier based on slice-level PHE labels to produce high-quality class activation maps (CAMs) as pseudo PHE annotations. Then, we train a deep learning model to produce accurate PHE segmentation by iteratively refining the pseudo annotations via an uncertainty-aware corrective training strategy for noise removal and a distance-aware loss for background compression. Experimental results demonstrate that, the proposed framework achieves a comparative performance with the fully supervised methods on PHE segmentation, and largely improves the baseline performance where only pseudo PHE labels are used for training. We believe the findings from this study of using cross-task supervision for annotation-efficient deep learning can be applied to other medical imaging applications. Zhuo Kuang, Zengqiang Yan, Li Yu 0003, Xianbo Deng, Yineng Hua, Shuyun Li |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Customized Federated Learning for Multi-Source Decentralized Medical Image ClassificationabstractThe performance of deep networks for medical image analysis is often constrained by limited medical data, which is privacy-sensitive. Federated learning (FL) alleviates the constraint by allowing different institutions to collaboratively train a federated model without sharing data. However, the federated model is often suboptimal with respect to the characteristics of each client's local data. Instead of training a single global model, we propose Customized FL (CusFL), for which each client iteratively trains a client-specific/private model based on a federated global model aggregated from all private models trained in the immediate previous iteration. Two overarching strategies employed by CusFL lead to its superior performance: 1) the federated model is mainly for feature alignment and thus only consists of feature extraction layers; 2) the federated feature extractor is used to guide the training of each private model. In that way, CusFL allows each client to selectively learn useful knowledge from the federated model to improve its personalized model. We evaluated CusFL on multi-source medical image datasets for the identification of clinically significant prostate cancer and the classification of skin lesions. Jeffry Wicaksana, Zengqiang Yan, Xin Yang 0008, Yang Liu 0165, Lixin Fan, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Exploring intermediate representation for monocular vehicle pose estimationabstractWe present a new learning-based framework to recover vehicle pose in SO(3) from a single RGB image. In contrast to previous works that map local appearance to observation angles, we explore a progressive approach by extracting meaningful Intermediate Geometrical Representations (IGRs) to estimate egocentric vehicle orientation. This approach features a deep model that transforms perceived intensities to IGRs, which are mapped to a 3D representation encoding object orientation in the camera coordinate system. Core problems are what IGRs to use and how to learn them more effectively. We answer the former question by designing IGRs based on an interpolated cuboid that derives from primitive 3D annotation readily. The latter question motivates us to incorporate geometry knowledge with a new loss function based on a projective invariant. This loss function allows unlabeled data to be used in the training stage to improve representation learning. Without additional labels, our system outperforms previous monocular RGB-based methods for joint vehicle detection and pose estimation on the KITTI benchmark, achieving performance even comparable to stereo methods. Code and pre-trained models are available at this HTTPS URL1. Shichao Li 0002, Zengqiang Yan, Kwang-Ting Cheng |
CVPR | 2 |
| 2021 | Variation-Aware Federated Learning With Multi-Source Decentralized Medical Image DataabstractPrivacy concerns make it infeasible to construct a large medical image dataset by fusing small ones from different sources/institutions. Therefore, federated learning (FL) becomes a promising technique to learn from multi-source decentralized data with privacy preservation. However, the cross-client variation problem in medical image data would be the bottleneck in practice. In this paper, we propose a variation-aware federated learning (VAFL) framework, where the variations among clients are minimized by transforming the images of all clients onto a common image space. We first select one client with the lowest data complexity to define the target image space and synthesize a collection of images through a privacy-preserving generative adversarial network, called PPWGAN-GP. Then, a subset of those synthesized images, which effectively capture the characteristics of the raw images and are sufficiently distinct from any raw image, is automatically selected for sharing with other clients. For each client, a modified CycleGAN is applied to translate its raw images to the target image space defined by the shared synthesized images. In this way, the cross-client variation problem is addressed with privacy preservation. We apply the framework for automated classification of clinically significant prostate cancer and evaluate it using multi-source decentralized apparent diffusion coefficient (ADC) image data. Experimental results demonstrate that the proposed VAFL framework stably outperforms the current horizontal FL framework. As VAFL is independent of deep learning architectures for classification, we believe that the proposed framework is widely applicable to other medical image classification tasks. Zengqiang Yan, Jeffry Wicaksana, Zhiwei Wang 0002, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Enabling a Single Deep Learning Model for Accurate Gland Instance Segmentation: A Shape-Aware Adversarial Learning FrameworkabstractSegmenting gland instances in histology images is highly challenging as it requires not only detecting glands from a complex background but also separating each individual gland instance with accurate boundary detection. However, due to the boundary uncertainty problem in manual annotations, pixel-to-pixel matching based loss functions are too restrictive for simultaneous gland detection and boundary detection. State-of-the-art approaches adopted multi-model schemes, resulting in unnecessarily high model complexity and difficulties in the training process. In this paper, we propose to use one single deep learning model for accurate gland instance segmentation. To address the boundary uncertainty problem, instead of pixel-to-pixel matching, we propose a segment-level shape similarity measure to calculate the curve similarity between each annotated boundary segment and the corresponding detected boundary segment within a fixed searching range. As the segment-level measure allows location variations within a fixed range for shape similarity calculation, it has better tolerance to boundary uncertainty and is more effective for boundary detection. Furthermore, by adjusting the radius of the searching range, the segment-level shape similarity measure is able to deal with different levels of boundary uncertainty. Therefore, in our framework, images of different scales are down-sampled and integrated to provide both global and local contextual information for training, which is helpful in segmenting gland instances of different sizes. To reduce the variations of multi-scale training images, by referring to adversarial domain adaptation, we propose a pseudo domain adaptation framework for feature alignment. By constructing loss functions based on the segment-level shape similarity measure, combining with the adversarial loss function, the proposed shape-aware adversarial learning framework enables one single deep learning model for gland instance segmentation. Experimental results on the 2015 MICCAI Gland Challenge dataset demonstrate that the proposed framework achieves state-of-the-art performance with one single deep learning model. As the boundary uncertainty problem widely exists in medical image segmentation, it is broadly applicable to other applications. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 1 |
| 2019 | A Three-Stage Deep Learning Model for Accurate Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is a fundamental step in the diagnosis of eye-related diseases, in which both thick vessels and thin vessels are important features for symptom detection. All existing deep learning models attempt to segment both types of vessels simultaneously by using a unified pixel-wise loss that treats all vessel pixels with equal importance. Due to the highly imbalanced ratio between thick vessels and thin vessels (namely the majority of vessel pixels belong to thick vessels), the pixel-wise loss would be dominantly guided by thick vessels and relatively little influence comes from thin vessels, often leading to low segmentation accuracy for thin vessels. To address the imbalance problem, in this paper, we explore to segment thick vessels and thin vessels separately by proposing a three-stage deep learning model. The vessel segmentation task is divided into three stages, namely thick vessel segmentation, thin vessel segmentation, and vessel fusion. As better discriminative features could be learned for separate segmentation of thick vessels and thin vessels, this process minimizes the negative influence caused by their highly imbalanced ratio. The final vessel fusion stage refines the results by further identifying nonvessel pixels and improving the overall vessel thickness consistency. The experiments on public datasets DRIVE, STARE, and CHASE_DB1 clearly demonstrate that the proposed three-stage deep learning model outperforms the current state-of-the-art vessel segmentation methods. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | A Deep Model with Shape-Preserving Loss for Gland Instance Segmentation
Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
MICCAI (2) | 1 |
| 2018 | Describing Upper-Body Motions Based on Labanotation for Learning-from-Observation RobotsabstractWe have been developing a paradigm that we call learning-from-observation for a robot to automatically acquire a robot program to conduct a series of operations, or for a robot to understand what to do, through observing humans performing the same operations. Since a simple mimicking method to repeat exact joint angles or exact end-effector trajectories does not work well because of the kinematic and dynamic differences between a human and a robot, the proposed method employs intermediate symbolic representations, tasks, for conceptually representing what-to-do through observation. These tasks are subsequently mapped to appropriate robot operations depending on the robot hardware. In the present work, task models for upper-body operations of humanoid robots are presented, which are designed on the basis of Labanotation. Given a series of human operations, we first analyze the upper-body motions and extract certain fixed poses from key frames. These key poses are translated into tasks represented by Labanotation symbols. Then, a robot performs the operations corresponding to those task models. Because tasks based on Labanotation are independent of robot hardware, different robots can share the same observation module, and only different task-mapping modules specific to robot hardware are required. The system was implemented and demonstrated that three different robots can automatically mimic human upper-body operations with a satisfactory level of resemblance. Katsushi Ikeuchi, Zhaoyuan Ma, Zengqiang Yan, Shunsuke Kudoh, Minako Nakamura |
Int. J. Comput. Vis. | 3 |
| 2018 | A Skeletal Similarity Metric for Quality Evaluation of Retinal Vessel SegmentationabstractThe most commonly used evaluation metrics for quality assessment of retinal vessel segmentation are sensitivity, specificity, and accuracy, which are based on pixel-to-pixel matching. However, due to the inter-observer problem that vessels annotated by different observers vary in both thickness and location, pixel-to-pixel matching is too restrictive to fairly evaluate the results of vessel segmentation. In this paper, the proposed skeletal similarity metric is constructed by comparing the skeleton maps generated from the reference and the source vessel segmentation maps. To address the inter-observer problem, instead of using a pixel-to-pixel matching strategy, each skeleton segment in the reference skeleton map is adaptively assigned with a searching range whose radius is determined based on its vessel thickness. Pixels in the source skeleton map located within the searching range are then selected for similarity calculation. The skeletal similarity consists of a curve similarity, which measures the structural similarity between the reference and the source skeleton maps and a thickness similarity, which measures the thickness consistency between the reference and the source vessel segmentation maps. In contrast to other metrics that provide a global score for the overall performance, we modify the definitions of true positive, false negative, true negative, and false positive based on the skeletal similarity, based on which sensitivity, specificity, accuracy, and other objective measurements can be constructed. More importantly, the skeletal similarity metric has better potential to be used as a pixelwise loss function for training deep learning models for retinal vessel segmentation. Through comparison of a set of examples, we demonstrate that the redefined metrics based on the skeletal similarity are more effective for quality evaluation, especially with greater tolerance to the inter-observer problem. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Texture edge-guided depth recovery for structured light-based depth sensor
Huiping Deng, Zengqiang Yan, Li Yu 0003 |
Multim. Tools Appl. | 4 |
| 2015 | Large-area depth recovery for RGB-D cameraabstractIn this paper, a large-area depth recovery method for RGB-D camera is proposed. Considering that pixels along edges between different regions usually share similar depth values, we first select reliable pixels along edges of large-area depth missing regions and project them into the world coordinate system. Then, by examining the distribution of these pixels, we apply a weighted least squares method to approximate the surface function. With the help of the surface function, missing depth values can be recovered correctly. To the best of our knowledge, this is the first recovery method focusing on large areas of missing depth information. Qualitative evaluation demonstrates the effectiveness of the proposed method. Zengqiang Yan, Li Yu 0003, Zixiang Xiong |
ICIP | 1 |
| 2015 | Texture-free large-area depth recovery for planar surfacesabstractThis paper presents a texture-free depth enhancement method for large-area depth recovery. The proposed algorithm identifies a large-area depth missing region, and iteratively segments its contour by setting different initial pixels in each iteration. Coordinate transformation is used to analyze the distribution of each contour segment. By examining distributions of all contour segments, statistical histogram analysis is applied in our approach to select contour pixels. Then, selected pixels are projected into the world coordinate system, and multiple linear regression is utilized for surface function approximation. Missing depth values of a large-area depth missing region can be recovered with guidance of the approximated surface function. Quantitative and qualitative evaluations over state-of-the-art depth enhancement methods demonstrate the effectiveness and superiority of our method. Being texture-free, the proposed method has the flexibility of being merged into traditional depth enhancement methods. Zengqiang Yan, Li Yu 0003, Zixiang Xiong |
MMSP | 1 |