EDBT 2026 Demo / reviewers in the wild / expert
Xiaoying Tang 0001
dblp:134/9714-1
· DBLP profile ↗
47ranked-venue papers
0as first author
42since 2021 · last 2026
0000-0002-7549-6560ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 15 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hessian-Aware Zeroth-Order Optimization for Quantized Large Language Models
Zhangchi Wang, Yidu Wu, Yijin Huang, Pujin Cheng, Qinghai Guo, Xiaoying Tang 0001 |
ICIC (5) | 6 |
| 2026 | Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs
Yidu Wu, Kejie Zhao, Zhangchi Wang, Qinghai Guo, Xiaoying Tang 0001 |
ICIC (16) | 6 |
| 2026 | MIRAGE: Medical image-text pre-training for robustness against noisy environments
Pujin Cheng, Yijin Huang, Li Lin 0006, Junyan Lyu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
Medical Image Anal. | 6 |
| 2026 | AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware TransformerabstractIn the era of foundation models, achieving a unified understanding of different dynamic objects through a single network has the potential to empower stronger spatial intelligence. Moreover, accurate estimation of animal pose and shape across diverse species is essential for quantitative analysis in biological research. However, this topic remains underexplored due to the limited network capacity of previous methods and the scarcity of comprehensive multi-species datasets. To address these limitations, we introduce AniMer+, an extended version of our scalable AniMer framework. In this paper, we focus on a unified approach for reconstructing mammals (mammalia) and birds (aves). A key innovation of AniMer+ is its high-capacity, family-aware Vision Transformer (ViT) incorporating a Mixture-of-Experts (MoE) design. Its architecture partitions network layers into taxa-specific components (for mammalia and aves) and taxa-shared components, enabling efficient learning of both distinct and common anatomical features within a single model. To overcome the critical shortage of 3D training data, especially for birds, we introduce a diffusion-based conditional image generation pipeline. This pipeline produces two large-scale synthetic datasets: CtrlAni3D for quadrupeds (about 10 k images with pixel-aligned SMAL labels) and CtrlAVES3D (about 7 k images with pixel-aligned AVES labels). To note, CtrlAVES3D is the first large-scale, 3D-annotated dataset for birds, which is crucial for resolving single-view depth ambiguities. Trained on an aggregated collection of 41.3 k mammalian and 12.4 k avian images (combining real and synthetic data), our method demonstrates superior performance over existing approaches across a wide range of benchmarks, including the challenging out-of-domain Animal Kingdom dataset. Ablation studies confirm the effectiveness of both our novel network architecture and the generated synthetic datasets in enhancing real-world application performance. Liang An 0001, Jin Lyu, Li Lin 0006, Pujin Cheng, Yebin Liu, Xiaoying Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | SynLiTS: Phase Prompting-Driven Diffusion Synthesis and Context-Aware Fusion for Unaligned Liver Tumor SegmentationabstractGiven the rich and complementary information contained in multi-phase CT images (CTs), they play an indispensable role in liver cancer diagnosis and prognosis, wherein an important prerequisite is liver tumor segmentation. However, spatial misalignments across phases and the limited availability of high-quality multi-phase CT datasets significantly hinder the performance of liver tumor segmentation. To tackle these challenges, we here propose SynLiTS, a novel multi-phase liver tumor segmentation framework. Its core idea is to synthesize multi-phase CTs with strictly-aligned liver tumors based on pseudo-normal multi-phase CTs. Specifically, an FFC-based Inpainter is first designed to generate pseudo-normal CTs by reconstructing dilated liver tumors. The pseudo-normal CTs and randomly generated tumor masks are then combined via a phase prompting-driven diffusion model to synthesize multi-phase liver tumor CTs with diverse tumor characteristics. In this way, multi-phase CTs with perfectly-aligned liver tumor labels are obtained. We also construct a real multi-phase liver tumor dataset, named MPLiTS. Finally, the synthesized and real multi-phase CTs are used to train a liver tumor segmentation model, which incorporates a context-aware fusion module to effectively learn and integrate multi-phase information. SynLiTS is evaluated on both internal and external datasets, and the results show that it outperforms state-of-the-art methods by large margins. Code will be released at https://github.com/Chyiun/SynLiTS. Li Lin 0006, ZhiCheng Jin, Pujin Cheng, JianJian Chen, HaiDong Zhu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2025 | UNiFS: Unified Multi-Contrast MRI Reconstruction via Frequency-Spatial FusionabstractRecently, Multi-Contrast MR Reconstruction (MCMR) has emerged as a hot research topic that leverages high-quality auxiliary modalities to reconstruct undersampled target modalities of interest. However, existing methods often struggle to generalize across different k-space undersampling patterns, requiring the training of a separate model for each specific pattern, which limits their practical applicability. To address this challenge, we propose UniFS, a Unified Frequency-Spatial Fusion model designed to handle multiple k-space undersampling patterns for MCMR tasks without any need for retraining. UniFS integrates three key modules: a Cross-Modal Frequency Fusion module, an Adaptive Mask-Based Prompt Learning module, and a Dual-Branch Complementary Refinement module. These modules work together to extract domain-invariant features from diverse k-space undersampling patterns while dynamically adapt to their own variations. Another limitation of existing MCMR methods is their tendency to focus solely on spatial information while neglect frequency characteristics, or extract only shallow frequency features, thus failing to fully leverage complementary cross-modal frequency information. To relieve this issue, UniFS introduces an adaptive prompt-guided frequency fusion module for k-space learning, significantly enhancing the model's generalization performance. We evaluate our model on the BraTS and HCP datasets with various k-space undersampling patterns and acceleration factors, including previously unseen patterns, to comprehensively assess UniFS's generalizability. Experimental results across multiple scenarios demonstrate that UniFS achieves state-of-the-art performance. Our code is available at https://github.com/LIKP0/UniFS. Yiwei Ren, Kai Pan, Dong Wei 0004, Pujin Cheng, Xian Wu 0001, Xiaoying Tang 0001 |
BIBM | 7 |
| 2025 | AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerabstractQuantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformer-based backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available open-sourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce CtrlAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-ofdistribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for in-the-wild applications. Project page: https://luoxue-star.github.io/AniMer_project_page/. Jin Lyu, Yi Gu 0005, Li Lin 0006, Pujin Cheng, Yebin Liu, Xiaoying Tang 0001, Liang An 0001 |
CVPR | 7 |
| 2025 | Masked Contrastive Language-Image Modeling For Brain Segmentation
Jianwen Liang, Junyan Lyu, Yixuan Yuan, Xiaoying Tang 0001 |
MICCAI (8) | 4 |
| 2025 | Probabilistic Prior-Guided Anatomical Alignment for MRI Super-Resolution
Yiwen Luo, Xiaoying Tang 0001, Yixuan Yuan |
MICCAI (4) | 2 |
| 2025 | UniOCTSeg: Towards Universal OCT Retinal Layer Segmentation via Hierarchical Prompting and Progressive Consistency Learning
Li Lin 0006, Chaoran Miao, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
MICCAI (16) | 5 |
| 2025 | SET: Superpixel Embedded Transformer for skin lesion segmentation
Junyan Lyu, Xiaoying Tang 0001 |
Medical Image Anal. | 3 |
| 2025 | ProCNS: Progressive Prototype Calibration and Noise Suppression for Weakly-Supervised Medical Image SegmentationabstractWeakly-supervised segmentation (WSS) has emerged as a solution to mitigate the conflict between annotation cost and model performance by adopting sparse annotation formats (e.g., point, scribble, block, etc.). Typical approaches attempt to exploit anatomy and topology priors to directly expand sparse annotations into pseudo-labels. However, due to lack of attention to the ambiguous boundaries in medical images and insufficient exploration of sparse supervision, existing approaches tend to generate erroneous and overconfident pseudo proposals in noisy regions, leading to cumulative model error and performance degradation. In this work, we propose a novel WSS approach, named ProCNS, encompassing two synergistic modules devised with the principles of progressive prototype calibration and noise suppression. Specifically, we design a Prototype-based Regional Spatial Affinity (PRSA) loss to maximize the pair-wise affinities between spatial and semantic elements, providing our model of interest with more reliable guidance. The affinities are derived from the input images and the prototype-refined predictions. Meanwhile, we propose an Adaptive Noise Perception and Masking (ANPM) module to obtain more enriched and representative prototype representations, which adaptively identifies and masks noisy regions within the pseudo proposals, reducing potential erroneous interference during prototype computation. Furthermore, we generate specialized soft pseudo-labels for the noisy regions identified by ANPM, providing supplementary supervision. Extensive experiments on six medical image segmentation tasks involving different modalities demonstrate that the proposed framework significantly outperforms representative state-of-the-art methods. Yixiang Liu, Li Lin 0006, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | LF-SynthSeg: Label-Free Brain Tissue-Assisted Tumor Synthesis and SegmentationabstractUnsupervised brain tumor segmentation is pivotal in realms of disease diagnosis, surgical planning, and treatment response monitoring, with the distinct advantage of obviating the need for labeled data. Traditional methodologies in this domain, however, often fall short in fully capitalizing on the extensive prior knowledge of brain tissue, typically approaching the task merely as an anomaly detection challenge. In our research, we present an innovative strategy that effectively integrates brain tissues' prior knowledge into both the synthesis and segmentation of brain tumor from T2-weighted Magnetic Resonance Imaging scans. Central to our method is the tumor synthesis mechanism, employing randomly generated ellipsoids in conjunction with the intensity profiles of brain tissues. This methodology not only fosters a significant degree of variation in the tumor presentations within the synthesized images but also facilitates the creation of an essentially unlimited pool of abnormal T2-weighted images. These synthetic images closely replicate the characteristics of real tumor-bearing scans. Our training protocol extends beyond mere tumor segmentation; it also encompasses the segmentation of brain tissues, thereby directing the network's attention to the boundary relationship between brain tumor and brain tissue, thus improving the robustness of our method. We evaluate our approach across five widely recognized public datasets (BRATS 2019, BRATS 2020, BRATS 2021, PED and SSA), and the results show that our method outperforms state-of-the-art unsupervised tumor segmentation methods by large margins. Moreover, the proposed method achieves more than 92 of the fully supervised performance on the same testing datasets. Pengxiao Xu, Junyan Lyu, Li Lin 0006, Pujin Cheng, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | FedLPPA: Learning Personalized Prompt and Aggregation for Federated Weakly-Supervised Medical Image SegmentationabstractFederated learning (FL) effectively mitigates the data silo challenge brought about by policies and privacy concerns, implicitly harnessing more data for deep model training. However, traditional centralized FL models grapple with diverse multi-center data, especially in the face of significant data heterogeneity, notably in medical contexts. In the realm of medical image segmentation, the growing imperative to curtail annotation costs has amplified the importance of weakly-supervised techniques which utilize sparse annotations such as points, scribbles, etc. A pragmatic FL paradigm shall accommodate diverse annotation formats across different sites, which research topic remains under-investigated. In such context, we propose a novel personalized FL framework with learnable prompt and aggregation (FedLPPA) to uniformly leverage heterogeneous weak supervision for medical image segmentation. In FedLPPA, a learnable universal knowledge prompt is maintained, complemented by multiple learnable personalized data distribution prompts and prompts representing the supervision sparsity. Integrated with sample features through a dual-attention mechanism, those prompts empower each local task decoder to adeptly adjust to both the local distribution and the supervision form. Concurrently, a dual-decoder strategy, predicated on prompt similarity, is introduced for enhancing the generation of pseudo-labels in weakly-supervised learning, alleviating overfitting and noise accumulation inherent to local data, while an adaptable aggregation method is employed to customize the task decoder on a parameter-wise basis. Extensive experiments on four distinct medical image segmentation tasks involving different modalities underscore the superiority of FedLPPA, with its efficacy closely parallels that of fully supervised centralized training. Our code and data will be available at https://github.com/llmir/FedLPPA. Li Lin 0006, Yixiang Liu, Jiewei Wu, Pujin Cheng, Zhiyuan Cai, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Masked Deformation Modeling for Volumetric Brain MRI Self-Supervised Pre-TrainingabstractSelf-supervised learning (SSL) has been proposed to alleviate neural networks' reliance on annotated data and to improve downstream tasks' performance, which has obtained substantial success in several volumetric medical image segmentation tasks. However, most existing approaches are designed and pre-trained on CT or MRI datasets of non-brain organs. The lack of brain prior limits those methods' performance on brain segmentation, especially on fine-grained brain parcellation. To overcome this limitation, we here propose a novel SSL strategy for MRI of the human brain, named Masked Deformation Modeling (MDM). MDM first conducts atlas-guided patch sampling on individual brain MRI scans (moving volumes) and an MNI152 template (a fixed volume). The sampled moving volumes are randomly masked in a feature-aligned manner, and then sent into a U-Net-based network to extract latent features. An intensity head and a deformation field head are used to decode the latent features, respectively restoring the masked volume and predicting the deformation field from the moving volume to the fixed volume. The proposed MDM is fine-tuned and evaluated on three brain parcellation datasets with different granularities (JHU, Mindboggle-101, CANDI), a brain lesion segmentation dataset (ATLAS2), and a brain tumor segmentation dataset (BraTS21). Results demonstrate that MDM outperforms various state-of-the-art medical SSL methods by considerable margins, and can effectively reduce the annotation effort by at least 40%. Codes and pre-trained weights will be released at https://github.com/CRazorback/MDM. Junyan Lyu, Perry F. Bartlett, Fatima A. Nasrallah, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Boosting Memory Efficiency in Transfer Learning for High-Resolution Medical Image ClassificationabstractThe success of large-scale pretrained models has established fine-tuning as a standard method for achieving significant improvements in downstream tasks. However, fine-tuning the entire parameter set of a pretrained model is costly. Parameter-efficient transfer learning (PETL) has recently emerged as a cost-effective alternative for adapting pretrained models to downstream tasks. Despite its advantages, the increasing model size and input resolution present challenges for PETL, as the training memory consumption is not reduced as effectively as the parameter usage. In this article, we introduce fine-grained prompt tuning plus (FPT+), a PETL method designed for high-resolution medical image classification, which significantly reduces the training memory consumption compared to other PETL methods. FPT+ performs transfer learning by training a lightweight side network and accessing pretrained knowledge from a large pretrained model (LPM) through fine-grained prompts and fusion modules. Specifically, we freeze the LPM of interest and construct a learnable lightweight side network. The frozen LPM processes high-resolution images to extract fine-grained features, while the side network employs corresponding downsampled low-resolution images to minimize memory usage. To enable the side network to leverage pretrained knowledge, we propose fine-grained prompts and fusion modules, which collaborate to summarize information through the LPM's intermediate activations. We evaluate FPT+ on eight medical image datasets of varying sizes, modalities, and complexities. Experimental results demonstrate that FPT+ outperforms other PETL methods, using only 1.03% of the learnable parameters and 3.18% of the memory required for fine-tuning an entire ViT-B model. Our code is available https://github.com/YijinHuang/FPT. Yijin Huang, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Masked Modality Complementary Modeling for Brain Tumor SegmentationabstractSelf-supervised pre-training techniques based on image reconstruction have achieved substantial success in medical image analysis, allowing for the transferability of pre-trained model weights to various downstream tasks for further fine-tuning. However, current pre-training methods primarily target single-modal medical images, like CT scans, scarcely considering the multi-modal images like multi-modal brain MRI. Yet, the latter is especially pivotal for accurate tumor segmentation, given that each modality provides unique insights into the tumor’s characteristics. In this study, we introduce a self-supervised pre-training approach tailored for multi-modal brain MRI, equipped with Masked Modality Complementary Modeling (MMCM). Specifically, the proposed method involves masking a designated portion of each modality to ensure the visible parts are distinct and complementary. We assume that through the process of reconstructing such images, the model not only learns general anatomical and modality-specific characteristics but also gains insights into bridging and mapping across various modalities. Results of downstream experiments validate that our method outperforms state-of-the-art self-supervised learning methods in the tumor segmentation accuracy on the BraTS 2021 dataset. Additionally, in downstream tasks in two common scenarios: the small sample size and the single available modality, our method substantially improves the performance over the baseline model trained from scratch. The model and code are available at https://github.com/liangjianwen01/MMCM. Jianwen Liang, Li Lin 0006, Junyan Lyu, Xiaoying Tang 0001 |
BIBM | 4 |
| 2024 | Joint Super-Resolution and Modality Translation Network for Multi-Contrast Arbitrary-Scale Isotropic MRI ReconstructionabstractDue to time and cost limitations, Magnetic Resonance (MR) imaging often employs anisotropic scanning with large slice spacing and thickness. This causes blurring in views perpendicular to the slices, which adversely affects clinical diagnosis and research. Taking into account the complementary information from the reference modality, deep learning (DL) based multi-contrast methods have become a focal point of research. These methods aim to reconstruct the isotropic target MR image with the auxiliary high-resolution (HR) reference modality. However, most of the methods primarily concentrate on the structural restoration of the target low-resolution (LR) image, neglecting the crucial aspect that the coexisting structural and modality differences between target and reference modalities can impede effective restoration. Additionally, these methods are designed for a fixed upsampling scale, not accounting for the practical scenario of varying slice thickness. In this work, we propose a joint Super-resolution and Modality translation network (SMNet) for multi-contrast arbitrary-scale isotropic MRI reconstruction. The modality translation branch includes the Modality-Specific-Augmented Alignment (MSAA) block, which eliminates modality distribution disparities and enhances modality-specific regions on the reference feature before fusion. And the super-resolution branch employs the Reliability-based Spatial Fusion (RSF) block for the structural restoration of the target LR feature using a reliability prior. The outputs from these two branches are then ensembled to obtain the final reconstructed result. Extensive experiments on both a private dataset and the Brasts2021 dataset demonstrate the effectiveness and generalizability of the proposed method. Our code is available at https://github.com/11710615/smnet. Kai Pan, Li Lin 0006, Pujin Cheng, Junyan Lyu, Xiaoying Tang 0001 |
BIBM | 5 |
| 2024 | Fusion Side Tuning: A Parameter and Memory Efficient Fine-tuning Method for High-resolution Medical Image ClassificationabstractParameter-efficient fine-tuning (PEFT) has been proposed as a cost-effective approach for transferring large-scale pre-trained models (LPMs) to downstream tasks, mitigating the high costs associated with updating all parameters of LPMs. However, current PEFT methods encounter the challenge that GPU memory usage during training is not reduced as effectively as parameter usage. In this paper, we propose Fusion Side Tuning (FST), a novel memory-efficient and parameter-efficient fine-tuning method. FST significantly reduces GPU memory consumption during training, particularly at high resolutions. To achieve this, we freeze the backbone LPM and construct a learnable side fusion network that takes intermediate features from the backbone as input. The side fusion network consists of a sequence of fusion modules, which enable it to leverage the knowledge embedded in the intermediate features. Additionally, we employ an important token selection mechanism to further reduce training costs and memory requirements. We evaluate FST on eight medical image datasets of varying modalities and sizes. Experimental results demonstrate that FST outperforms existing PEFT methods, utilizing only 2% of the learnable parameters and 20% of the GPU memory required for full fine-tuning of a ViT-B encoder with an input resolution of 512 × 512. Zhangchi Wang, Yijin Huang, Yidu Wu, Pujin Cheng, Li Lin 0006, Qinghai Guo, Xiaoying Tang 0001 |
BIBM | 7 |
| 2024 | BPaCo: Balanced Parametric Contrastive Learning for Long-Tailed Medical Image Classification
Zhiyuan Cai, Tianyunxi Wei, Li Lin 0006, Hao Chen 0011, Xiaoying Tang 0001 |
MICCAI (1) | 5 |
| 2024 | Fine-Grained Prompt Tuning: A Parameter and Memory Efficient Transfer Learning Method for High-Resolution Medical Image Classification
Yijin Huang, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
MICCAI (12) | 4 |
| 2024 | Simultaneous alignment and surface regression using hybrid 2D-3D networks for 3D coherent layer segmentation of retinal OCT images with full and sparse annotations
Dong Wei 0004, Donghuan Lu, Xiaoying Tang 0001, Liansheng Wang 0002, Yefeng Zheng 0001 |
Medical Image Anal. | 4 |
| 2024 | LDDMM-Face: Large deformation diffeomorphic metric learning for cross-annotation face alignment
Junyan Lyu, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
Pattern Recognit. | 5 |
| 2024 | Dual Teacher Knowledge Distillation With Domain Alignment for Face Anti-SpoofingabstractFace recognition systems have raised concerns due to their vulnerability to different presentation attacks, and system security has become an increasingly critical concern. Although many face anti-spoofing (FAS) methods perform well in intra-dataset scenarios, their generalization remains a challenge. To address this issue, some methods adopt domain adversarial training (DAT) to extract domain-invariant features. Differently, in this paper, we propose a domain adversarial attack (DAA) method by adding perturbations to the input images, which makes them indistinguishable across domains and enables domain alignment. Moreover, since models trained on limited data and types of attacks cannot generalize well to unknown attacks, we propose a dual perceptual and generative knowledge distillation framework for face anti-spoofing that utilizes pre-trained face-related models containing rich face priors. Specifically, we adopt two different face-related models as teachers to transfer knowledge to the target student model. The pre-trained teacher models are not from the task of face anti-spoofing but from perceptual and generative tasks, respectively, which implicitly augment the data. By combining both DAA and dual-teacher knowledge distillation, we develop a dual teacher knowledge distillation with domain alignment framework (DTDA) for face anti-spoofing. The advantage of our proposed method has been verified through extensive ablation studies and comparison with state-of-the-art methods on public datasets across multiple protocols. Zhe Kong, Wentian Zhang, Tao Wang 0052, Kaihao Zhang, Yuexiang Li, Xiaoying Tang 0001, Wenhan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | SSiT: Saliency-Guided Self-Supervised Image Transformer for Diabetic Retinopathy GradingabstractSelf-supervised Learning (SSL) has been widely applied to learn image representations through exploiting unlabeled images. However, it has not been fully explored in the medical image analysis field. In this work, Saliency-guided Self-Supervised image Transformer (SSiT) is proposed for Diabetic Retinopathy (DR) grading from fundus images. We novelly introduce saliency maps into SSL, with a goal of guiding self-supervised pre-training with domain-specific prior knowledge. Specifically, two saliency-guided learning tasks are employed in SSiT: 1) Saliency-guided contrastive learning is conducted based on the momentum contrast, wherein fundus images' saliency maps are utilized to remove trivial patches from the input sequences of the momentum-updated key encoder. Thus, the key encoder is constrained to provide target representations focusing on salient regions, guiding the query encoder to capture salient features. 2) The query encoder is trained to predict the saliency segmentation, encouraging the preservation of fine-grained information in the learned representations. To assess our proposed method, four publicly-accessible fundus image datasets are adopted. One dataset is employed for pre-training, while the three others are used to evaluate the pre-trained models' performance on downstream DR grading. The proposed SSiT significantly outperforms other representative state-of-the-art SSL methods on all downstream datasets and under various evaluation settings. For example, SSiT achieves a Kappa score of 81.88% on the DDR dataset under fine-tuning evaluation, outperforming all other ViT-based SSL methods by at least 9.48%. Yijin Huang, Junyan Lyu, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Uni4Eye++: A General Masked Image Modeling Multi-Modal Pre-Training Framework for Ophthalmic Image Classification and SegmentationabstractA large-scale labeled dataset is a key factor for the success of supervised deep learning in most ophthalmic image analysis scenarios. However, limited annotated data is very common in ophthalmic image analysis, since manual annotation is time-consuming and labor-intensive. Self-supervised learning (SSL) methods bring huge opportunities for better utilizing unlabeled data, as they do not require massive annotations. To utilize as many unlabeled ophthalmic images as possible, it is necessary to break the dimension barrier, simultaneously making use of both 2D and 3D images as well as alleviating the issue of catastrophic forgetting. In this paper, we propose a universal self-supervised Transformer framework named Uni4Eye++ to discover the intrinsic image characteristic and capture domain-specific feature embedding in ophthalmic images. Uni4Eye++ can serve as a global feature extractor, which builds its basis on a Masked Image Modeling task with a Vision Transformer architecture. On the basis of our previous work Uni4Eye, we further employ an image entropy guided masking strategy to reconstruct more-informative patches and a dynamic head generator module to alleviate modality confusion. We evaluate the performance of our pre-trained Uni4Eye++ encoder by fine-tuning it on multiple downstream ophthalmic image classification and segmentation tasks. The superiority of Uni4Eye++ is successfully established through comparisons to other state-of-the-art SSL pre-training methods. Our code is available at https://github.com/Davidczy/Uni4Eye++. Zhiyuan Cai, Li Lin 0006, Huaqing He, Pujin Cheng, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsabstractContrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local alignment between medical images and reports. In contrast to standard global multi-modality alignment methods, we employ a local alignment module for fine-grained representation. Furthermore, a cross-modality conditional reconstruction module is designed to interchange information across modalities in the training phase by reconstructing masked images and reports. For reconstructing long reports, a sentence-wise prototype memory bank is constructed, enabling the network to focus on low-level localized visual and high-level clinical linguistic features. Additionally, a non-auto-regressive generation paradigm is proposed for reconstructing non-sequential reports. Experimental results on five downstream tasks, including supervised classification, zero-shot classification, image-to-text retrieval, semantic segmentation, and object detection, show the proposed method outperforms other state-of-the-art methods across multiple datasets and under different dataset size settings. The code is available at https://github.com/QtacierP/PRIOR. Pujin Cheng, Li Lin 0006, Junyan Lyu, Yijin Huang, Wenhan Luo, Xiaoying Tang 0001 |
ICCV | 6 |
| 2023 | Segmentation of Brachial Plexus Ultrasound Images Based on Modified SegNet Model
Songlin Yan, Xiujiao Chen, Xiaoying Tang 0001, Xuebing Chi |
IDEAL | 3 |
| 2023 | Learning Ontology-Based Hierarchical Structural Relationship for Whole Brain Segmentation
Junyan Lyu, Pengxiao Xu, Fatima A. Nasrallah, Xiaoying Tang 0001 |
MICCAI (4) | 4 |
| 2023 | Whole-Heart Reconstruction with Explicit Topology Integrated Learning
Roger C. Tam, Xiaoying Tang 0001 |
MICCAI (6) | 3 |
| 2023 | YoloCurvSeg: You only label one noisy skeleton for vessel-style curvilinear structure segmentation
Li Lin 0006, Linkai Peng, Huaqing He, Pujin Cheng, Jiewei Wu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
Medical Image Anal. | 7 |
| 2023 | GAMMA challenge: Glaucoma grAding from Multi-Modality imAges
Huihui Fang, Fei Li 0021, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang 0001, Qinji Yu, Sifan Song, Xinxing Xu, Yanyu Xu 0001, Wensai Wang, Shuai Lu 0003, Huiqi Li, Shihua Huang, Zhichao Lu, Chubin Ou, Xifei Wei, Bingyuan Liu, Riadh Kobbi, Xiaoying Tang 0001, Li Lin 0006, Hrvoje Bogunovic, José Ignacio Orlando, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 22 |
| 2023 | autoSMIM: Automatic Superpixel-Based Masked Image Modeling for Skin Lesion SegmentationabstractSkin lesion segmentation from dermoscopic images plays a vital role in early diagnoses and prognoses of various skin diseases. However, it is a challenging task due to the large variability of skin lesions and their blurry boundaries. Moreover, most existing skin lesion datasets are designed for disease classification, with relatively fewer segmentation labels having been provided. To address these issues, we propose a novel automatic superpixel-based masked image modeling method, named autoSMIM, in a self-supervised setting for skin lesion segmentation. It explores implicit image features from abundant unlabeled dermoscopic images. autoSMIM begins with restoring an input image with randomly masked superpixels. The policy of generating and masking superpixels is then updated via a novel proxy task through Bayesian Optimization. The optimal policy is subsequently used for training a new masked image modeling model. Finally, we finetune such a model on the downstream skin lesion segmentation task. Extensive experiments are conducted on three skin lesion segmentation datasets, including ISIC 2016, ISIC 2017, and ISIC 2018. Ablation studies demonstrate the effectiveness of superpixel-based masked image modeling and establish the adaptability of autoSMIM. Comparisons with state-of-the-art methods show the superiority of our proposed autoSMIM. The source code is available at https://github.com/Wzhjerry/autoSMIM. Junyan Lyu, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Uni4Eye: Unified 2D and 3D Self-supervised Pre-training via Masked Image Modeling Transformer for Ophthalmic Image Classification
Zhiyuan Cai, Li Lin 0006, Huaqing He, Xiaoying Tang 0001 |
MICCAI (8) | 4 |
| 2022 | DS3-Net: Difficulty-Perceived Common-to-T1ce Semi-supervised Multimodal MRI Synthesis Network
Li Lin 0006, Pujin Cheng, Kai Pan, Xiaoying Tang 0001 |
MICCAI (6) | 5 |
| 2022 | AADG: Automatic Augmentation for Domain Generalization on Retinal Image SegmentationabstractConvolutional neural networks have been widely applied to medical image segmentation and have achieved considerable performance. However, the performance may be significantly affected by the domain gap between training data (source domain) and testing data (target domain). To address this issue, we propose a data manipulation based domain generalization method, called Automated Augmentation for Domain Generalization (AADG). Our AADG framework can effectively sample data augmentation policies that generate novel domains and diversify the training set from an appropriate search space. Specifically, we introduce a novel proxy task maximizing the diversity among multiple augmented novel domains as measured by the Sinkhorn distance in a unit sphere space, making automated augmentation tractable. Adversarial training and deep reinforcement learning are employed to efficiently search the objectives. Quantitative and qualitative experiments on 11 publicly-accessible fundus image datasets (four for retinal vessel segmentation, four for optic disc and cup (OD/OC) segmentation and three for retinal lesion segmentation) are comprehensively performed. Two OCTA datasets for retinal vasculature segmentation are further involved to validate cross-modality generalization. Our proposed AADG exhibits state-of-the-art generalization performance and outperforms existing approaches by considerable margins on retinal vessel, OD/OC and lesion segmentation tasks. The learned policies are empirically validated to be model-agnostic and can transfer well to other models. The source code is available at https://github.com/CRazorback/AADG. Junyan Lyu, Yijin Huang, Li Lin 0006, Pujin Cheng, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2021 | I-SECRET: Importance-Guided Fundus Image Enhancement via Semi-supervised Contrastive Constraining
Pujin Cheng, Li Lin 0006, Yijin Huang, Junyan Lyu, Xiaoying Tang 0001 |
MICCAI (8) | 5 |
| 2021 | Lesion-Based Contrastive Learning for Diabetic Retinopathy Grading from Fundus Images
Yijin Huang, Li Lin 0006, Pujin Cheng, Junyan Lyu, Xiaoying Tang 0001 |
MICCAI (2) | 5 |
| 2021 | BSDA-Net: A Boundary Shape and Distance Aware Joint Learning Framework for Segmenting and Classifying OCTA Images
Li Lin 0006, Jiewei Wu, Yijin Huang, Junyan Lyu, Pujin Cheng, Xiaoying Tang 0001 |
MICCAI (8) | 8 |
| 2021 | A deep learning framework for pancreas segmentation with multi-atlas registration and 3D level-set
Yue Zhang 0033, Yifan Chen 0001, Ed X. Wu, Chunming Li, Xiaoying Tang 0001 |
Medical Image Anal. | 8 |
| 2021 | Brain segmentation based on multi-atlas and diffeomorphism guided 3D fully convolutional network ensembles
Xiaoying Tang 0001 |
Pattern Recognit. | 2 |
| 2021 | MI-UNet: Multi-Inputs UNet Incorporating Brain Parcellation for Stroke Lesion Segmentation From T1-Weighted Magnetic Resonance ImagesabstractStroke is a serious manifestation of various cerebrovascular diseases and one of the most dangerous diseases in the world today. Volume quantification and location detection of chronic stroke lesions provide vital biomarkers for stroke rehabilitation. Recently, deep learning has seen a rapid growth, with a great potential in segmenting medical images. In this work, unlike most deep learning-based segmentation methods utilizing only magnetic resonance (MR) images as the input, we propose and validate a novel stroke lesion segmentation approach named multi-inputs UNet (MI-UNet) that incorporates brain parcellation information, including gray matter (GM), white matter (WM) and lateral ventricle (LV). The brain parcellation is obtained from 3D diffeomorphic registration and is concatenated with the original MR image to form two-channel inputs to the subsequent MI-UNet. Effectiveness of the proposed pipeline is validated using a dataset consisting of 229 T1-weighted MR images. Experiments are conducted via a five-fold cross-validation. The proposed MI-UNet performed significantly better than UNet in both 2D and 3D settings. Our best results obtained by 3D MI-UNet has superior segmentation performance, as measured by the Dice score, Hausdorff distance, average symmetric surface distance, as well as precision, over other state-of-the-art methods. Yue Zhang 0033, Yifan Chen 0001, Ed X. Wu, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2019 | Prostate Segmentation using 2D Bridged U-netabstractIn this paper, we focus on three problems in deep learning based medical image segmentation. Firstly, U-net, as a popular model for medical image segmentation, is difficult to train when convolutional layers increase even though a deeper network usually has a better generalization ability because of more learnable parameters. Secondly, the exponential ReLU (ELU), as an alternative of ReLU, is not much different from ReLU when the network of interest gets deep. Thirdly, the Dice loss, as one of the pervasive loss functions for medical image segmentation, is not effective when the prediction is close to ground truth and will cause oscillation during training. To address the aforementioned three problems, we propose and validate a deeper network that can fit medical image datasets that are usually small in the sample size. Meanwhile, we propose a new loss function to accelerate the learning process and a combination of different activation functions to improve the network performance. Our experimental results suggest that our network is comparable or superior to state-of-the-art methods. Yue Zhang 0033, Junjun He, Yu Qiao 0001, Yifan Chen 0001, Hongjian Shi, Ed X. Wu, Xiaoying Tang 0001 |
IJCNN | 8 |
| 2019 | A Joint 3D+2D Fully Convolutional Framework for Subcortical Segmentation
Yue Zhang 0033, Xiaoying Tang 0001 |
MICCAI (3) | 3 |
| 2019 | Automated Classification of Amyotrophic Lateral Sclerosis Using Multi-level Whole-brain Volumes from Structural Magnetic Resonance ImagingabstractWe proposed and validated a fully-automated classification procedure for amyotrophic lateral sclerosis (ALS) using structural magnetic resonance imaging; specifically, T1-weighted images from 28 ALS subjects and 28 healthy control (HC) subjects were used. The raw features were obtained from a validated multi-granularity whole-brain analysis pipeline, providing multi-level whole-brain segmentation volumes. We employed the support vector machine as our classification algorithm with several feature selection techniques analyzed. According to our leave-one-out cross validation experiment results, the whole-brain structural volumes from Level 4, followed by a feature selection utilizing the standardized Wilcoxon two-sample rank sum statistic, yielded the best classification performance; overall accuracy: 83.93%, sensitivity: 85.71%, specificity: 82.14%, and the area under the receiver operating characteristic curve: 0.8380. The feature selection procedure revealed that the volumes of the thalamus, especially that on the left hemisphere, are the most important (of highest ranking) in the ALS-vs-HC discrimination. Yuanyuan Wei 0001, Siyuan Jiang, Yuanyuan Qin, Xiaoying Tang 0001 |
SMC | 4 |
| 2019 | Whole brain volume and cortical thickness based automatic classification of Wilson's diseaseabstractWilson's disease (WD) is a progressive autosomal-recessive genetic disorder of copper metabolism that can induce cognitive, physical and psychiatric symptoms. Despite the wide use of machine learning methods in neuroimaging analysis, WD-related research has been very rare. In this work, we proposed and validated an efficient pipeline for an automated WD classification based on whole brain segmentation volumes and cortical thicknesses obtained from T1-weighted magnetic resonance images (MRIs). Three well-known supervised machine learning algorithms, including support vector machine (SVM), linear discriminant analysis (LDA), and logistic regression (LR), were evaluated and compared in the setting of WD classification. A total of 51 images, including 27 acquired from WD patients and 24 from age-matched healthy controls, were used in the validation analysis. Univariate feature selection was conducted to eliminate non-relevant features and retain features contributing to the classification performance. Two nested leave-one-out cross validations were adopted, with the inner folder used for optimal parameter estimation and the outer folder for classification performance evaluation. Experimental results showed that when employing volume features, SVM significantly outperformed both LDA and LR, yielding an overall accuracy of 96.1%, a sensitivity of 92.6% and a specificity of 100%. LR could also reach such best classification performance when using a combination of volumes and thicknesses as input features. This study provides a new non-invasive tool (MRI-based) for an automated detection of WD. Jianping Chu, Xiaoying Tang 0001 |
SMC | 4 |
| 2018 | Geometry-Based Facial Expression Recognition via Large Deformation Diffeomorphic Metric Curve MappingabstractWe proposed a new geometry-based facial expression recognition (FER) system in the framework of large deformation diffeomorphic metric curve mapping. The geometry of a face was represented by 12 distinct curves, with curve-based facial deformations being used to identify two sets of geometric features in two settings. In each setting, four types of features were extracted and tested. Leave-one-out cross-validation experiments on 327 image sequences yielded accuracies as high as 94%. Furthermore, using a multi-kernel technique to combine features from the two settings has boosted the recognition accuracy to be as high as 95.4%. Pucheng Yang, Yuanyuan Wei 0001, Xiaoying Tang 0001 |
ICIP | 4 |