EDBT 2026 Demo / reviewers in the wild / expert
Yong Xia 0001
dblp:50/2433-1
· DBLP profile ↗
205ranked-venue papers
11as first author
128since 2021 · last 2026
0000-0001-9273-2847ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 105 · 81 since 2021Graphics, computer vision, multimedia, augmented reality and games · 88 · 7 first-author · 49 since 2021Artificial intelligence and machine learning · 61 · 5 first-author · 37 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Clustered Federated Learning for Distributed Channel Prediction
Zhenyu Xie, Huixiang Zhu, Yong Xia 0001, Yingyu Li |
ICC | 4 |
| 2026 | CE-CoLSM: Cloud-Edge Large and Small Models Collaborative Framework for Traffic Prediction
Xubo Li, Yong Xia 0001, Yingyu Li |
ICC | 4 |
| 2026 | Precise estimation of tissue microstructure with hybrid graph transformer
Geng Chen 0001, Jiquan Ma, Hui Cui 0002, Shu Zhang 0001, Yong Xia 0001, Pew-Thian Yap |
Artif. Intell. Medicine | 6 |
| 2026 | MRFMA: A hybrid paradigm integrating multi-receptive field network with mediator attention for 3D multi-organ segmentation
Hengfei Cui, Jiatong Li 0006, Dianrong Du, Yanning Zhang 0001, Yong Xia 0001 |
Expert Syst. Appl. | 5 |
| 2026 | SegRap2025: A benchmark of gross tumor volume and lymph node clinical target volume Segmentation for Radiotherapy Planning of nasopharyngeal carcinoma
Litingyu Wang, Chenyuan Bian, Zijun Gao, Chunbin Gu, Xin Weng, Jianghao Wu 0001, Yicheng Wu 0001, Jin Ye 0002, Linhao Li, Yiwen Ye, Yong Xia 0001, Elias Tappeiner, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Junqiang Chen, Chuanyi Huang, Lisheng Wang, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Shichuan Zhang, Shaoting Zhang 0001, Wenjun Liao, Guotai Wang |
Medical Image Anal. | 15 |
| 2026 | STAGE challenge: Structural-Functional Transition in Glaucoma Assessment
Shiqi Zhou, Yuancong Liang, Huihui Fang, Ziyang Chen 0003, Yong Xia 0001, Chubin Ou, Yubo Tan, Haojie Yin, Chengcheng Feng, Hao Zhou 0030, Hrvoje Bogunovic, Huazhu Fu, Fei Li 0021, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 7 |
| 2026 | Day-Night Adaptation: A domain adaptation framework for medical image segmentation without source data
Yiwen Ye, Yongsheng Pan, Jingfeng Zhang, Yong Xia 0001 |
Pattern Recognit. | 5 |
| 2026 | Source-free domain adaptation using prompt learning for medical image segmentation
Shishuai Hu, Zehui Liao, Yong Xia 0001 |
Pattern Recognit. | 3 |
| 2026 | A variational Bayesian algorithm for probabilistic affine and non-rigid point cloud registration
Xinke Ma, Qingjie Zeng, Mengkang Lu, Yong Xia 0001 |
Pattern Recognit. | 6 |
| 2026 | Unsupervised domain adaptation for cardiac MRI segmentation via adversarial learning in latent space
Hengfei Cui, Yong Xia 0001 |
Pattern Recognit. | 4 |
| 2026 | Decoupling Target Semantics via Text-Anchored Visual Contrast for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) provides an effective means of reducing reliance on large-scale annotated datasets by leveraging unlabeled data. However, existing SSL methods often struggle with semantic ambiguity, especially under limited supervision. Recent studies have incorporated textual information to provide contextual guidance, yet most focus on feature fusion rather than emphasizing target semantics critical for segmentation. In this paper, we proposed a novel Text-anchored Visual Decoupling (TeViD) framework for semi-supervised medical image segmentation. TeViD is built upon a teacher-student architecture with a dual-decoder design that explicitly disentangles target and background representations using both labeled and unlabeled data. For unlabeled data, a reversed cross-supervision mechanism is introduced to enhance decoder diversity and semantic separation. Furthermore, two contrastive learning objectives are proposed: a teacher-guided visual contrastive loss and a text-anchored contrastive loss, both designed to reinforce semantic disentanglement from visual and textual perspectives. Extensive experiments on five public datasets (covering X-ray, pathology, ultrasound, MRI, and CT) demonstrate that TeViD consistently outperforms both standard SSL and text-enhanced SSL methods, achieving average improvements of 5.72% in Dice and 8.15% in mIoU over the second-best competitor. The code is available at: https://github.com/jgfiuuuu/TeViD. Qingjie Zeng, Xinke Ma, Zilin Lu, Mengkang Lu, Yanning Zhang 0001, Yong Xia 0001 |
IEEE Trans. Image Process. | 8 |
| 2026 | Deformable Medical Image Registration With Effective Anatomical Structure Representation and Divide-and-Conquer NetworkabstractEffective representation of Regions of Interest (ROI) and independent alignment of these ROIs can significantly enhance the performance of deformable medical image registration (DMIR). However, current learning-based DMIR methods have limitations. Unsupervised techniques disregard ROI representation and proceed directly with aligning pairs of images, while weakly-supervised methods heavily depend on label constraints to facilitate registration. To address these issues, we introduce a weakly-supervised ROI-based registration approach named EASR-DCN. Our method represents medical images through effective ROIs and achieves independent alignment of these ROIs without requiring labels. Specifically, we first used a Gaussian mixture model for intensity analysis to represent images using multiple effective ROIs with distinct intensities. Furthermore, we propose a novel Divide-and-Conquer Network (DCN) that processes ROIs through separate channels to independently align their features. The resulting sub-deformation fields are seamlessly integrated to generate a comprehensive displacement vector field. Extensive experiments were performed on three MRI and one CT datasets to showcase the superior accuracy and deformation reduction efficacy of our EASR-DCN. Compared to VoxelMorph, our EASR-DCN achieved improvements of 10.31% in the Dice score for brain MRI, 13.01% for cardiac MRI, and 5.75% for hippocampus MRI, highlighting its promising potential for clinical applications. Xinke Ma, Yongsheng Pan, Qingjie Zeng, Mengkang Lu, Bolysbek Murat Yerzhanuly, Bazargul Matkerim, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | From Few to More: Scribble-Based Medical Image Segmentation via Masked Context Modeling and Continuous Pseudo LabelsabstractScribble-based weakly supervised segmentation methods have shown promising results in medical image segmentation, significantly reducing annotation costs. However, existing approaches often rely on auxiliary tasks to enforce semantic consistency and use hard pseudo labels for supervision, overlooking the unique challenges faced by models trained with sparse annotations. These models must predict pixel-wise segmentation maps from limited data, making it crucial to handle varying levels of annotation richness effectively. In this paper, we propose MaCo, a weakly supervised model designed for medical image segmentation, based on the principle of "from few to more." MaCo leverages Masked Context Modeling (MCM) and Continuous Pseudo Labels (CPL). MCM employs an attention-based masking strategy to perturb the input image, ensuring that the model's predictions align with those of the original image. CPL converts scribble annotations into continuous pixel-wise labels by applying an exponential decay function to distance maps, producing confidence maps that represent the likelihood of each pixel belonging to a specific category, rather than relying on hard pseudo labels. We evaluate MaCo on three public datasets, comparing it with other weakly supervised methods. Our results show that MaCo outperforms competing methods across all datasets, establishing a new record in weakly supervised medical image segmentation. Zhisong Wang, Yiwen Ye, Ziyang Chen 0003, Minglei Shu, Yanning Zhang 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and RetentionabstractHistopathology and transcriptomics are fundamental modalities in cancer diagnostics, encapsulating the morphological and molecular characteristics of the disease. Multi-modal self-supervised learning has demonstrated remarkable potential in learning pathological representations by integrating diverse data sources. Conventional multi-modal integration methods primarily emphasize modality alignment, while paying insufficient attention to retaining the modality-specific intrinsic structures. However, unlike conventional scenarios where multi-modal inputs often share highly overlapping features, histopathology and transcriptomics exhibit pronounced heterogeneity, offering orthogonal yet complementary insights. Histopathology data provides morphological and spatial context, elucidating tissue architecture and cellular topology, whereas transcriptomics data delineates molecular signatures through quantifying gene expression patterns. This inherent disparity introduces a major challenge in aligning these modalities while maintaining modality-specific fidelity. To address these challenges, we present MIRROR, a novel multi-modal representation learning framework designed to foster both modality alignment and retention. MIRROR employs dedicated encoders to extract comprehensive feature representations for each modality, which is further complemented by a modality alignment module to achieve seamless integration between phenotype patterns and molecular profiles. Furthermore, a modality retention module safeguards unique attributes from each modality, while a style clustering module mitigates redundancy and enhances disease-relevant information by modeling and aligning consistent pathological signatures within a clustering space. Extensive evaluations on The Cancer Genome Atlas (TCGA) cohorts for cancer subtyping and survival analysis highlight MIRROR's superior performance, demonstrating its effectiveness in constructing comprehensive oncological feature representations and benefiting the cancer diagnosis. Code is available at https://github.com/TianyiFranklinWang/MIRROR. Jianan Fan, Dingxin Zhang 0001, Dongnan Liu, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Harnessing Text Insights With Visual Alignment for Medical Image SegmentationabstractPre-trained vision-language models (VLMs) and language models (LMs) have recently garnered significant attention due to their remarkable ability to represent textual concepts, opening up new avenues in vision tasks. In medical image segmentation, efforts are being made to integrate text and image data using VLMs and LMs. However, current text-enhanced approaches face several challenges. First, using separate pre-trained vision and text models to encode image and text data can result in semantic shifts. Second, while VLMs can establish the correspondence between visual and textual features when pre-trained on paired image-text data, this alignment often deteriorates during segmentation tasks due to misalignment between the text and vision components in ongoing learning. In this paper, we propose TeViA, a novel approach that seamlessly integrates with various vision and text models, irrespective of their pre-training relationships. This integration is achieved through a segmentation-specific text-to-vision alignment design, ensuring both information gain and semantic consistency. Specifically, for each training data, a foreground visual representation is extracted from the segmentation head and used to supervise projection layers, thereby adjusting the textual features to better contribute to the segmentation task. Additionally, a historic visual prototype is created by aggregating target semantics from all training data and is updated using a momentum-based manner. This prototype aims to enhance the visual representation of each data instance by establishing feature-level connections, which in turn refines the textual features. The superiority of TeViA is validated on five public datasets, exhibiting over 6% Dice improvements compared to vision-only methods. Code is available at: https://github.com/jgfiuuuu/TeViA. Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Zhiyong Wang 0001, Yanning Zhang 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Gradient Alignment Improves Test-Time Adaptation for Medical Image SegmentationabstractAlthough recent years have witnessed significant advancements in medical image segmentation, the pervasive issue of domain shift among medical images from diverse centres hinders the effective deployment of pre-trained models. Many Test-time Adaptation (TTA) methods have been proposed to address this issue by fine-tuning pre-trained models with test data during inference. These methods, however, often suffer from less-satisfactory optimization due to suboptimal optimization direction (dictated by the gradient) and fixed step-size (predicated on the learning rate). In this paper, we propose the Gradient alignment-based Test-time adaptation (GraTa) method to improve both the gradient direction and learning rate in the optimization procedure. Unlike conventional TTA methods, which primarily optimize the pseudo gradient derived from a self-supervised objective, our method incorporates an auxiliary gradient with the pseudo one to facilitate gradient alignment. Such gradient alignment enables the model to excavate the similarities between different gradients and correct the gradient direction to approximate the empirical gradient related to the current segmentation task. Additionally, we design a dynamic learning rate based on the cosine similarity between the pseudo and auxiliary gradients, thereby empowering the adaptive fine-tuning of pre-trained models on diverse test data. Extensive experiments establish the effectiveness of the proposed gradient alignment and dynamic learning rate and substantiate the superiority of our GraTa method over other state-of-the-art TTA methods on a benchmark medical image segmentation task. Ziyang Chen 0003, Yiwen Ye, Yongsheng Pan, Yong Xia 0001 |
AAAI | 4 |
| 2025 | Transfer Attention-Guided Multi-Receptive Field Network for Multi-Modality Cardiac Image SegmentationabstractExisting whole heart segmentation algorithms usually combine 3D Convolutional Neural Networks (3D CNNs) with Transformers, for the purpose of capturing local and global features. However, traditional CNNs with a fixed size of receptive field cannot capture long-range contextual information. Transformers have been widely used to establish dependencies on global information, despite this, they greatly increase the computational complexity. To mitigate these challenges, we propose a hybrid paradigm, called Transfer Attention-Guided MultiReceptive Field Network (TAMRNet), to boost the representation quality for multi-modality cardiac image segmentation. In TAMRNet, the novel adaptive-scale depthwise convolution module adeptly preserves the inherent inductive biases of convolution while concurrently amplifying the network's ability to establish dependencies on long-range contextual information. Besides, a novel attention mechanism called Transfer Attention is developed to establish dependencies on global information. Transfer Attention avoids the direct similarity calculation of$Q$and$K$by introducing the Transfer tokens, and thus dramatically decreases the computational cost. The proposed TAMRNet is tested on the MM-WHS 2017 challenge dataset, achieving the average Dice scores of 93.7 % and 82.2 % on the CT and MRI datasets respectively. Extensive experimental results prove that our proposed method achieves superior performances in comparison with state-of-the-art methods. Jiatong Li 0006, Hengfei Cui, Dianrong Du, Geng Chen 0001, Yong Xia 0001 |
BIBM | 6 |
| 2025 | Optimal Grid-Battery Power Allocation for Time-Energy-Sensitive Wireless Systems
Tianzi Li, Yong Xia 0001, Yingyu Li |
GLOBECOM | 2 |
| 2025 | SMF-Net: Unlocking Multimodal Insights for Enhanced Stroke Lesion Segmentation
Meklit Mesfin Atlaw, Geng Chen 0001, Xuyun Wen, Hengfei Cui, Yong Xia 0001 |
MICCAI (3) | 6 |
| 2025 | RadioFormer: Integrating Radiologist Inductive Bias for Tumor Classification on Multi-Sequence MR Images
Yong Xia 0001 |
MICCAI (1) | 2 |
| 2025 | Personalized Federated Side-Tuning for Medical Image Classification
Jiayi Chen 0006, Benteng Ma, Yongsheng Pan, Bin Pu, Hengfei Cui, Yong Xia 0001 |
MICCAI (14) | 6 |
| 2025 | CAUDA-MI: Cross Attention-Guided Unsupervised Domain Adaptation with Mutual Information for Cardiac MRI Segmentation
Dianrong Du, Hengfei Cui, Jiatong Li 0006, Yong Xia 0001 |
MICCAI (6) | 5 |
| 2025 | Cycle Context Verification for In-Context Medical Image Segmentation
Shishuai Hu, Zehui Liao, Liangli Zhen, Huazhu Fu, Yong Xia 0001 |
MICCAI (1) | 5 |
| 2025 | Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering
Zehui Liao, Shishuai Hu, Ke Zou, Huazhu Fu, Liangli Zhen, Yong Xia 0001 |
MICCAI (5) | 6 |
| 2025 | Enjoying Information Dividend: Gaze Track-Based Medical Weakly Supervised Segmentation
Zhisong Wang, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001 |
MICCAI (10) | 4 |
| 2025 | Exploring Text-Enhanced Mixture-of-Experts for Semi-supervised Medical Image Segmentation with Composite Data
Qingjie Zeng, Xinke Ma, Zilin Lu, Yong Xia 0001 |
MICCAI (6) | 6 |
| 2025 | Hybrid Graph Mamba: Unlocking Non-Euclidean Potential for Accurate Polyp Segmentation
Yueyue Zhu, Haolin Lv, Geng Chen 0001, Yong Xia 0001 |
MICCAI (10) | 6 |
| 2025 | Towards Accurate Left Atrium and Scar Segmentation from LGE MRI with Boundary Loss Constrained Multi-Attention U-Net
Hengfei Cui, Jiatong Li 0006, Dianrong Du, Geng Chen 0001, Yong Xia 0001 |
PRCV (14) | 6 |
| 2025 | Mixture-attention Siamese transformer for video polyp segmentation
Geng Chen 0001, Junqing Yang, Xiaozhou Pu, Ge-Peng Ji, Huan Xiong, Yongsheng Pan, Hengfei Cui, Yong Xia 0001 |
Artif. Intell. Medicine | 8 |
| 2025 | Instance-dependent Label Distribution Estimation for Learning with Label Noise
Zehui Liao, Shishuai Hu, Yutong Xie 0001, Yong Xia 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Draw Sketch, Draw Flesh: Whole-Body Computed Tomography from Any X-Ray Views
Yongsheng Pan, Yiwen Ye, Yanning Zhang 0001, Yong Xia 0001, Dinggang Shen |
Int. J. Comput. Vis. | 4 |
| 2025 | PICK: Predict and Mask for Semi-supervised Medical Image Segmentation
Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Yong Xia 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | UAE: Universal Anatomical Embedding on multi-modality medical images
Fan Bai 0008, Xiaofei Huo, Jia Ge, Jingjing Lu, Xianghua Ye, Minglei Shu, Ke Yan 0006, Yong Xia 0001 |
Medical Image Anal. | 9 |
| 2025 | Unleashing the potential of open-set noisy samples against label noise for medical image classification
Zehui Liao, Shishuai Hu, Yanning Zhang 0001, Yong Xia 0001 |
Medical Image Anal. | 4 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 11 |
| 2025 | ATEC23 Challenge: Automated prediction of treatment effectiveness in ovarian cancer using histopathological images
Ching-Wei Wang, Nabila Puspita Firdi, Tzu-Chiao Chu, Mohammad Faiz Iqbal Faiz, Mohammad Zafar Iqbal, Mayur Mallya, Ali Bashashati, Fei Li 0021, Mengkang Lu, Yong Xia 0001, Tai-Kuang Chao |
Medical Image Anal. | 13 |
| 2025 | Contrastive Neuron Pruning for Backdoor DefenseabstractRecent studies have revealed that deep neural networks (DNNs) are susceptible to backdoor attacks, in which attackers insert a pre-defined backdoor into a DNN model by poisoning a few training samples. A small subset of neurons in DNN is responsible for activating this backdoor and pruning these backdoor-associated neurons has been shown to mitigate the impact of such attacks. Current neuron pruning techniques often face challenges in accurately identifying these critical neurons, and they typically depend on the availability of labeled clean data, which is not always feasible. To address these challenges, we propose a novel defense strategy called Contrastive Neuron Pruning (CNP). This approach is based on the observation that poisoned samples tend to cluster together and are distinguishable from benign samples in the feature space of a backdoored model. Given a backdoored model, we initially apply a reversed trigger to benign samples, generating multiple positive (benign-benign) and negative (benign-poisoned) feature pairs from the backdoored model. We then employ contrastive learning on these pairs to improve the separation between benign and poisoned features. Subsequently, we identify and prune neurons in the Batch Normalization layers that show significant response differences to the generated pairs. By removing these backdoor-associated neurons, CNP effectively defends against backdoor attacks while requiring the pruning of only about 1% of the total neurons. Comprehensive experiments conducted on various benchmarks validate the efficacy of CNP, demonstrating its robustness and effectiveness in mitigating backdoor attacks compared to existing methods. Benteng Ma, Dongnan Liu, Yanning Zhang 0001, Tom Weidong Cai, Yong Xia 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Active Learning Based on Temporal Difference of Gradient Flow in Thoracic Disease DiagnosisabstractGiven the significant advancements in thoracic disease diagnosis due to deep learning, there is a reliance on the availability of numerous annotated samples, which, however, can hardly be guaranteed due to the resource-intensive nature of medical image annotation. Active learning has been introduced to mitigate annotation costs by selecting a subset of uncertain samples for annotation and training. Existing active learning methods encounter two primary challenges: 1) overlooking the impact of samples on the dynamics of model training during data selection, and 2) suffering from high costs of data evaluation and selection. To tackle both issues, we propose a novel metric called Temporal Difference of Gradient Flow (TDGF) for data selection in active learning. Each round of active learning involves three steps: model training, data selection, and data annotation. First, we train a target model, a proxy model, and a historical proxy model on the labeled set. Second, the TDGF scores of unlabeled samples are evaluated based on the surrogate gradient flow, i.e., the TDGF w.r.t the final fully-connected layer between the proxy and historical proxy models, and top-K samples with the highest TDGF scores are selected. Third, the selected samples are annotated, and the labeled pool and unlabeled pool are updated. Comparative experiments have been conducted on two public chest radiograph datasets, i.e., ChestX-ray14 and CheXpert. Our results suggest that the proposed TDGF metric is prone to selecting hard and uncertain samples, and the use of proxy models and surrogate gradient flow substantially reduces the complexity of TDGF calculation. More importantly, the results also indicate that our TDGF-based method outperforms classical and state-of-the-art active learning methods in thoracic disease diagnosis. Jiayi Chen 0006, Benteng Ma, Hengfei Cui, Jingfeng Zhang, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | P2TC: A Lightweight Pyramid Pooling Transformer-CNN Network for Accurate 3D Whole Heart SegmentationabstractCardiovascular disease is a leading global cause of death, requiring accurate heart segmentation for diagnosis and surgical planning. Deep learning methods have been demonstrated to achieve superior performances in cardiac structures segmentation. However, there are still limitations in 3D whole heart segmentation, such as inadequate spatial context modeling, difficulty in capturing long-distance dependencies, high computational complexity, and limited representation of local high-level semantic information. To tackle the above problems, we propose a lightweight Pyramid Pooling Transformer-CNN (P2TC) network for accurate 3D whole heart segmentation. The proposed architecture comprises a dual encoder-decoder structure with a 3D pyramid pooling Transformer for multi-scale information fusion and a lightweight large-kernel Convolutional Neural Network (CNN) for local feature extraction. The decoder has two branches for precise segmentation and contextual residual handling. The first branch is used to generate segmentation masks for pixel-level classification based on the features extracted by the encoder to achieve accurate segmentation of cardiac structures. The second branch highlights contextual residuals across slices, enabling the network to better handle variations and boundaries. Extensive experimental results on the Multi-Modality Whole Heart Segmentation (MM-WHS) 2017 challenge dataset demonstrate that P2TC outperforms the most advanced methods, achieving the Dice scores of 92.6% and 88.1% in Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) modalities respectively, which surpasses the baseline model by 1.5% and 1.7%, and achieves state-of-the-art segmentation results. Hengfei Cui, Yifan Wang 0033, Yan Li 0129, Yanning Zhang 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Hyperbolic Geometry-Driven Robustness Enhancement for Rare Skin Disease DiagnosisabstractThe automated diagnosis of rare skin diseases using dermoscopy images, known as a few-shot learning (FSL) problem, remains challenging, since traditional FSL research tends to disregard the intrinsic hierarchical nature of rare diseases and data uncertainty. To address these issues, we propose to conduct rare skin disease diagnosis in hyperbolic space, which facilitates implicit class hierarchical structures and precise uncertainty measurement due to pivotal geometrical properties. We propose a Hyperbolic Geometry-driven Robustness Enhancement (HGRE) framework specifically tailored for diagnosing rare skin diseases. The HGRE framework uses implicit hierarchical relation in the hyperbolic space to better represent the features of rare diseases. Moreover, the framework incorporates an Adversarial Proxy Construction (APC) module to address the problem of data uncertainty. Specifically, the APC module uses the distance to the hyperbolic space origin as an indicator of uncertainty to filter and construct adversarial proxies for each uncertain prototype to achieve adversarial robust training. Leveraging the two unique geometrical properties, our HGRE framework effectively addresses the limitations of insufficient hierarchical relation utilization and data uncertainty in FSL-based rare skin disease diagnosis. This enhancement of the model's robustness in training has been corroborated by extensive empirical validation on two skin lesion datasets, where HGRE's performance notably surpassed existing state-of-the-art FSL methods. Yuanyuan Chen 0001, Xiaohan Xing, Jingfeng Zhang, Bolysbek Murat Yerzhanuly, Bazargul Matkerim, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | PathBot: A Foundation Model for Pathological Image AnalysisabstractComputational pathology has emerged as a transformative paradigm by leveraging artificial intelligence to automate and enhance diagnostic procedures. However, existing models often target narrow tasks or specific tumor types, missing opportunities to unify diverse datasets and tasks through joint learning. In this work, we introduce PathBot, a foundation model tailored for comprehensive pathological image analysis. Central to PathBot is a ViT-Giant encoder with one billion parameters, the largest model to date trained on publicly available pathological data. We pre-train this encoder using a novel Masked Distillation Network (MDN) and an integrated learning strategy that combines contrastive and generative objectives. The pre-training leverages over 30 million image patches derived from 11,765 whole slide images (WSIs) across 32 cancer types in the Cancer Genome Atlas (TCGA). To evaluate its versatility, we pair the encoder with task-specific decoders for segmentation, detection, classification, and regression. Extensive experiments across 20 downstream tasks demonstrate that PathBot achieves state-of-the-art performance in most cases, showcasing its robustness and generalizability. Mengkang Lu, Qingjie Zeng, Zilin Lu, Zhe Li 0006, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | FedDAG: Federated Domain Adversarial Generation Toward Generalizable Medical Image AnalysisabstractFederated domain generalization aims to train a global model from multiple source domains and ensure its generalization ability to unseen target domains. Due to the target domain being with unknown domain shifts, attempting to approximate these gaps by source domains may be the key to improving model generalization capability. Existing works mainly focus on sharing and recombining local domain-specific attributes to increase data diversity and simulate potential domain shifts. However, these methods may be insufficient since only the local attribute recombination can be hard to touch the out-of-distribution of global data. In this paper, we propose a simple-yet-efficient framework named Federated Domain Adversarial Generation (FedDAG). It aims to simulate the domain shift and improve the model generalization by adversarially generating novel domains different from local and global source domains. Specifically, it generates novel-style images by maximizing the instance-level feature discrepancy between original and generated images and trains a generalizable task model by minimizing their feature discrepancy. Further, we observed that FedDAG could cause different performance improvements for local models. It may be due to inherent data isolation and heterogeneity among clients, exacerbating the imbalance in their generalization contributions to the global model. Ignoring this imbalance can lead the global model's generalization ability to be sub-optimal, further limiting the novel domain generation procedure. Thus, to mitigate this imbalance, FedDAG hierarchically aggregates local models at the within-client and across-client levels by using the sharpness concept to evaluate client model generalization contributions. Extensive experiments across four medical benchmarks demonstrate FedDAG's ability to enhance generalization in federated medical scenarios. Haoxuan Che, Haibo Jin, Yong Xia 0001, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Bridging the Semantic Gap in Medical Visual Question Answering With Prompt LearningabstractMedical Visual Question Answering (Med-VQA) aims to answer questions regarding the content of medical images, crucial for enhancing diagnostics and education in healthcare. However, progress in this field is hindered by data scarcity due to the resource-intensive nature of medical data annotation. While existing Med-VQA approaches often rely on pre-training to mitigate this issue, bridging the semantic gap between pre-trained models and specific tasks remains a significant challenge. This paper presents the Dynamic Semantic-Adaptive Prompting (DSAP) framework, leveraging prompt learning to enhance model performance in Med-VQA. To this end, we introduce two prompting strategies: Semantic Alignment Prompting (SAP) and Dynamic Question-Aware Prompting (DQAP). SAP prompts multi-modal inputs during fine-tuning, reducing the semantic gap by aligning model outputs with domain-specific contexts. Simultaneously, DQAP enhances answer selection by leveraging grammatical relationships between questions and answers, thereby improving accuracy and relevance. The DSAP framework was pre-trained on three datasets-ROCO, MedICaT, and MIMIC-CXR-and comprehensively evaluated against 15 existing Med-VQA models on three public datasets: VQA-RAD, SLAKE, and PathVQA. Our results demonstrate a substantial performance improvement, with DSAP achieving a 1.9% enhancement in average results across benchmarks. These findings underscore DSAP's effectiveness in addressing critical challenges in Med-VQA and suggest promising avenues for future developments in medical AI. Zilin Lu, Qingjie Zeng, Mengkang Lu, Geng Chen 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | CADS: A Self-Supervised Learner via Cross-Modal Alignment and Deep Self-Distillation for CT Volume SegmentationabstractSelf-supervised learning (SSL) has long had great success in advancing the field of annotation-efficient learning. However, when applied to CT volume segmentation, most SSL methods suffer from two limitations, including rarely using the information acquired by different imaging modalities and providing supervision only to the bottleneck encoder layer. To address both limitations, we design a pretext task to align the information in each 3D CT volume and the corresponding 2D generated X-ray image and extend self-distillation to deep self-distillation. Thus, we propose a self-supervised learner based on Cross-modal Alignment and Deep Self-distillation (CADS) to improve the encoder's ability to characterize CT volumes. The cross-modal alignment is a more challenging pretext task that forces the encoder to learn better image representation ability. Deep self-distillation provides supervision to not only the bottleneck layer but also shallow layers, thus boosting the abilities of both. Comparative experiments show that, during pre-training, our CADS has lower computational complexity and GPU memory cost than competing SSL methods. Based on the pre-trained encoder, we construct PVT-UNet for 3D CT volume segmentation. Our results on seven downstream tasks indicate that PVT-UNet outperforms state-of-the-art SSL methods like MOCOv3 and DiRA, as well as prevalent medical image segmentation methods like nnUNet and CoTr. Code and pre-trained weight will be available at https://github.com/yeerwen/CADS. Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Segment Together: A Versatile Paradigm for Semi-Supervised Medical Image SegmentationabstractThe scarcity of annotations has become a significant obstacle in training powerful deep-learning models for medical image segmentation, limiting their clinical application. To overcome this, semi-supervised learning that leverages abundant unlabeled data is highly desirable to enhance model training. However, most existing works still focus on specific medical tasks and underestimate the potential of learning across diverse tasks and datasets. In this paper, we propose a Versatile Semi-supervised framework (VerSemi) to present a new perspective that integrates various SSL tasks into a unified model with an extensive label space, exploiting more unlabeled data for semi-supervised medical image segmentation. Specifically, we introduce a dynamic task-prompted design to segment various targets from different datasets. Next, this unified model is used to identify the foreground regions from all labeled data, capturing cross-dataset semantics. Particularly, we create a synthetic task with a CutMix strategy to augment foreground targets within the expanded label space. To effectively utilize unlabeled data, we introduce a consistency constraint that aligns aggregated predictions from various tasks with those from the synthetic task, further guiding the model to accurately segment foreground regions during training. We evaluated our VerSemi framework against seven established SSL methods on four public benchmarking datasets. Our results suggest that VerSemi consistently outperforms all competing methods, beating the second-best method with a 2.69% average Dice gain on four datasets and setting a new state of the art for semi-supervised medical image segmentation. Code is available at https://github.com/maxwell0027/VerSemi. Qingjie Zeng, Yutong Xie 0001, Zilin Lu, Mengkang Lu, Yicheng Wu 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Consistency-Guided Differential Decoding for Enhancing Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) has been proven beneficial for mitigating the issue of limited labeled data, especially on volumetric medical image segmentation. Unlike previous SSL methods which focus on exploring highly confident pseudo-labels or developing consistency regularization schemes, our empirical findings suggest that differential decoder features emerge naturally when two decoders strive to generate consistent predictions. Based on the observation, we first analyze the treasure of discrepancy in learning towards consistency, under both pseudo-labeling and consistency regularization settings, and subsequently propose a novel SSL method called LeFeD, which learns the feature-level discrepancies obtained from two decoders, by feeding such information as feedback signals to the encoder. The core design of LeFeD is to enlarge the discrepancies by training differential decoders, and then learn from the differential features iteratively. We evaluate LeFeD against eight state-of-the-art (SOTA) methods on three public datasets. Experiments show LeFeD surpasses competitors without any bells and whistles, such as uncertainty estimation and strong constraints, as well as setting a new state of the art for semi-supervised medical image segmentation. Code has been released at https://github.com/maxwell0027/LeFeD. Qingjie Zeng, Yutong Xie 0001, Zilin Lu, Mengkang Lu, Jingfeng Zhang, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | EM-Trans: Edge-Aware Multimodal Transformer for RGB-D Salient Object DetectionabstractRGB-D salient object detection (SOD) has gained tremendous attention in recent years. In particular, transformer has been employed and shown great potential. However, existing transformer models usually overlook the vital edge information, which is a major issue restricting the further improvement of SOD accuracy. To this end, we propose a novel edge-aware RGB-D SOD transformer, called EM-Trans, which explicitly models the edge information in a dual-band decomposition framework. Specifically, we employ two parallel decoder networks to learn the high-frequency edge and low-frequency body features from the low- and high-level features extracted from a two-steam multimodal backbone network, respectively. Next, we propose a cross-attention complementarity exploration module to enrich the edge/body features by exploiting the multimodal complementarity information. The refined features are then fed into our proposed color-hint guided fusion module for enhancing the depth feature and fusing the multimodal features. Finally, the resulting features are fused using our deeply supervised progressive fusion module, which progressively integrates edge and body features for predicting saliency maps. Our model explicitly considers the edge information for accurate RGB-D SOD, overcoming the limitations of existing methods and effectively improving the performance. Extensive experiments on benchmark datasets demonstrate that EM-Trans is an effective RGB-D SOD framework that outperforms the current state-of-the-art models, both quantitatively and qualitatively. A further extension to RGB-T SOD demonstrates the promising potential of our model in various kinds of multimodal SOD tasks. Geng Chen 0001, Qingyue Wang, Bo Dong 0001, Ruitao Ma, Nian Liu 0002, Huazhu Fu, Yong Xia 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | SurgicalSAM: Efficient Class Promptable Surgical Instrument SegmentationabstractThe Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we observe two problems with this naive pipeline: (1) the domain gap between natural objects and surgical instruments leads to inferior generalisation of SAM; and (2) SAM relies on precise point or box locations for accurate segmentation, requiring either extensive manual guidance or a well-performing specialist detector for prompt preparation, which leads to a complex multi-stage pipeline. To address these problems, we introduce SurgicalSAM, a novel end-to-end efficient-tuning approach for SAM to effectively integrate surgical-specific information with SAM’s pre-trained knowledge for improved generalisation. Specifically, we propose a lightweight prototype-based class prompt encoder for tuning, which directly generates prompt embeddings from class prototypes and eliminates the use of explicit prompts for improved robustness and a simpler pipeline. In addition, to address the low inter-class variance among surgical instrument categories, we propose contrastive prototype learning, further enhancing the discrimination of the class prototypes for more accurate class prompting. The results of extensive experiments on both EndoVis2018 and EndoVis2017 datasets demonstrate that SurgicalSAM achieves state-of-the-art performance while only requiring a small number of tunable parameters. The source code is available at https://github.com/wenxi-yue/SurgicalSAM. Wenxi Yue, Jing Zhang 0037, Kun Hu 0008, Yong Xia 0001, Jiebo Luo 0001, Zhiyong Wang 0001 |
AAAI | 4 |
| 2024 | Unsupervised Super-Resolution of Diffusion-Weighted Images via Deep Diffusion PriorabstractDeep learning-based super-resolution (SR) has shown great potential in improving the resolution of diffusion-weighted imaging (DWI), which is useful in clinical diagnosis and neuroscience studies of white matter. However, most existing deep learning methods for DWI SR are supervised, relying on paired low-high resolution images, which can be in practice difficult to acquire. To address this limitation, we propose an unsupervised DWI SR model, called deep diffusion prior (DDP), to learn low-level features for effective resolution enhancement of DW images using only information from low-resolution (LR) images. We incorporate structural and angular information to improve SR performance. The former is provided by structural magnetic resonance images encoding rich anatomical information. The latter is formulated based on angular neighboring constraints in the diffusion wavevector space. Extensive experiments on data from the human connectome project (HCP) show that DDP is qualitatively and quantitatively superior to competing methods in the absence of paired HR DW images. Geng Chen 0001, Hao Yang 0032, Runlin Zhang, Musa Bakarr, Yong Xia 0001, Pew-Thian Yap |
BIBM | 5 |
| 2024 | Think Twice Before Selection: Federated Evidential Active Learning for Medical Image Analysis with Domain ShiftsabstractFederated learning facilitates the collaborative learning of a global model across multiple distributed medical in-stitutions without centralizing data. Nevertheless, the ex-pensive cost of annotation on local clients remains an ob-stacle to effectively utilizing local data. To mitigate this issue, federated active learning methods suggest leveraging local and global model predictions to select a relatively small amount of informative local data for annotation. However, existing methods mainly focus on all local data sampled from the same domain, making them un-reliable in realistic medical scenarios with domain shifts among different clients. In this paper, we make the first at-tempt to assess the informativeness of local data derived from diverse domains and propose a novel methodology termed Federated Evidential Active Learning (FEAL) to calibrate the data evaluation under domain shift. Specif-ically, we introduce a Dirichlet prior distribution in both local and global models to treat the prediction as a distribution over the probability simplex and capture both aleatoric and epistemic uncertainties by using the Dirichlet-based evidential model. Then we employ the epistemic uncer-tainty to calibrate the aleatoric uncertainty. Afterward, we design a diversity relaxation strategy to reduce data re-dundancy and maintain data diversity. Extensive experi-ments and analysis on five real multi-center medical image datasets demonstrate the superiority of FEAL over the state-of-the-art active learning methods in federated sce-narios with domain shifts. The code will be available at https://github.com/JiayiChen815/FEAL. Jiayi Chen 0006, Benteng Ma, Hengfei Cui, Yong Xia 0001 |
CVPR | 4 |
| 2024 | Each Test Image Deserves A Specific Prompt: Continual Test-Time Adaptation for 2D Medical Image SegmentationabstractDistribution shift widely exists in medical images acquired from different medical centres and poses a significant obstacle to deploying the pretrained semantic segmentation model in real-world applications. Test-time adaptation has proven its effectiveness in tackling the cross-domain distribution shift during inference. However, most existing methods achieve adaptation by updating the pretrained models, rendering them susceptible to error accumulation and catastrophic forgetting when encountering a series of distribution shifts (i.e., under the continual test-time adaptation setup). To overcome these challenges caused by updating the models, in this paper, we freeze the pretrained model and propose the Visual Prompt-based Test-Time Adaptation (VPTTA) method to train a specific prompt for each test image to align the statistics in the batch normalization layers. Specifically, we present the low-frequency prompt, which is lightweight with only a few parameters and can be effectively trained in a single iteration. To enhance prompt initialization, we equip VPTTA with a memory bank to benefit the current prompt from previous ones. Additionally, we design a warm-up mechanism, which mixes source and target statistics to construct warm-up statistics, thereby facilitating the training process. Extensive experiments demonstrate the superiority of our VPTTA over other state-of-the-art methods on two medical image segmentation benchmark tasks. The code and weights of pretrained source models are available at https://github.com/Chen-Ziyang/VPTTA. Ziyang Chen 0003, Yongsheng Pan, Yiwen Ye, Mengkang Lu, Yong Xia 0001 |
CVPR | 5 |
| 2024 | PairAug: What Can Augmented Image-Text Pairs Do for Radiology?abstractCurrent vision-language pre-training (VLP) methodologies predominantly depend on paired image-text datasets, a resource that is challenging to acquire in radiology due to privacy considerations and labelling complexities. Data augmentation provides a practical solution to overcome the issue of data scarcity, however, most augmentation methods exhibit a limited focus, prioritising either image or text augmentation exclusively. Acknowledging this limitation, our objective is to devise a framework capable of concurrently augmenting medical image and text data. We design a Pairwise Augmentation (PairAug) approach that contains an Inter-patient Augmentation (InterAug) branch and an Intra-patient Augmentation (IntraAug) branch. Specifically, the InterAug branch of our approach generates radiology images using synthesised yet plausible reports derived from a Large Language Model (LLM). The generated pairs can be considered a collection of new patient cases since they are artificially created and may not exist in the original dataset. In contrast, the IntraAug branch uses newly generated reports to manipulate images. This process allows us to create new paired data for each individual with diverse medical conditions. Our extensive experiments on various downstream tasks covering medical image classification zero-shot and fine-tuning analysis demonstrate that our PairAug, concurrently expanding both image and text data, substantially outperforms image-/text-only expansion baselines and advanced medical VLP baselines. Our code is released at https://github.com/YtongXie/PairAug. Yutong Xie 0001, Qi Chen 0014, Sinuo Wang, Minh-Son To, Iris Lee, Ee Win Khoo, Kerolos Hendy, Daniel Koh, Yong Xia 0001, Qi Wu 0001 |
CVPR | 9 |
| 2024 | Continual Self-Supervised Learning: Towards Universal Multi-Modal Medical Data Representation LearningabstractSelf-supervised learning (SSL) is an efficient pre-training method for medical image analysis. However, current research is mostly confined to certain modalities, consuming considerable time and resources without achieving universality across different modalities. A straightforward solution is combining all modality data for joint SSL, which poses practical challenges. Firstly, our experiments reveal conflicts in representation learning as the number of modalities increases. Secondly, multi-modal data collected in advance cannot cover all real-world scenarios. In this paper, we reconsider versatile SSL from the perspective of continual learning and propose MedCoSS, a continuous SSL approach for multi-modal medical data. Different from joint representation learning, MedCoSS assigns varying data modalities to separate training stages, creating a multi-stage pre-training process. We propose a rehearsal- based continual learning approach to manage modal conflicts and prevent catastrophic forgetting. Specifically, we use the k-means sampling to retain and rehearse previous modality data during new modality learning. Moreover, we apply feature distillation and intra-modal mixup on buffer data for knowledge retention, bypassing pretext tasks. We conduct experiments on a large-scale multi-modal unlabeled dataset, including clinical reports, X-rays, CT, MRI, and pathological images. Experimental results demonstrate MedCoSS's exceptional generalization ability across 9 downstream datasets and its significant scalability in inte- grating new modality data. The code and pre-trained model are available at https://github.com/yeerwen/MedCoSS. Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Qi Wu 0001, Yong Xia 0001 |
CVPR | 6 |
| 2024 | FedEvi: Improving Federated Medical Image Segmentation via Evidential Weight Aggregation
Jiayi Chen 0006, Benteng Ma, Hengfei Cui, Yong Xia 0001 |
MICCAI (10) | 4 |
| 2024 | Spot the Difference: Difference Visual Question Answering with Residual Alignment
Zilin Lu, Yutong Xie 0001, Qingjie Zeng, Mengkang Lu, Qi Wu 0001, Yong Xia 0001 |
MICCAI (5) | 6 |
| 2024 | Enhancing Federated Learning Performance Fairness via Collaboration Graph-Based Reinforcement Learning
Yuexuan Xia, Benteng Ma, Qi Dou 0001, Yong Xia 0001 |
MICCAI (10) | 4 |
| 2024 | Reciprocal Collaboration for Semi-supervised Medical Image Classification
Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Mengkang Lu, Xinke Ma, Yong Xia 0001 |
MICCAI (11) | 6 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 23 |
| 2024 | VNAS: Variational Neural Architecture Search
Benteng Ma, Jing Zhang 0037, Yong Xia 0001, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2024 | Modeling annotator preference and stochastic annotation error for medical image segmentation
Zehui Liao, Shishuai Hu, Yutong Xie 0001, Yong Xia 0001 |
Medical Image Anal. | 4 |
| 2024 | Rethinking masked image modelling for medical image representationabstractMasked Image Modelling (MIM), a form of self-supervised learning, has garnered significant success in computer vision by improving image representations using unannotated data. Traditional MIMs typically employ a strategy of random sampling across the image. However, this random masking technique may not be ideally suited for medical imaging, which possesses distinct characteristics divergent from natural images. In medical imaging, particularly in pathology, disease-related features are often exceedingly sparse and localized, while the remaining regions appear normal and undifferentiated. Additionally, medical images frequently accompany reports, directly pinpointing pathological changes' location. Inspired by this, we propose Masked medical Image Modelling (MedIM), a novel approach, to our knowledge, the first research that employs radiological reports to guide the masking and restore the informative areas of images, encouraging the network to explore the stronger semantic representations from medical images. We introduce two mutual comprehensive masking strategies, knowledge-driven masking (KDM), and sentence-driven masking (SDM). KDM uses Medical Subject Headings (MeSH) words unique to radiology reports to identify symptom clues mapped to MeSH words (e.g., cardiac, edema, vascular, pulmonary) and guide the mask generation. Recognizing that radiological reports often comprise several sentences detailing varied findings, SDM integrates sentence-level information to identify key regions for masking. MedIM reconstructs images informed by this masking from the KDM and SDM modules, promoting a comprehensive and enriched medical image representation. Our extensive experiments on seven downstream tasks covering multi-label/class image classification, pneumothorax segmentation, and medical image-report analysis, demonstrate that MedIM with report-guided masking achieves competitive performance. Our method substantially outperforms ImageNet pre-training, MIM-based pre-training, and medical image-report pre-training counterparts. Codes are available at https://github.com/YtongXie/MedIM. Yutong Xie 0001, Lin Gu 0003, Tatsuya Harada, Yong Xia 0001, Qi Wu 0001 |
Medical Image Anal. | 5 |
| 2024 | ReFs: A hybrid pre-training paradigm for 3D medical image segmentation
Yutong Xie 0001, Lingqiao Liu, Hu Wang 0005, Yiwen Ye, Johan Verjans, Yong Xia 0001 |
Medical Image Anal. | 7 |
| 2024 | UniMiSS+: Universal Medical Self-Supervised Learning From Cross-Dimensional Unpaired DataabstractSelf-supervised learning (SSL) opens up huge opportunities for medical image analysis that is well known for its lack of annotations. However, aggregating massive (unlabeled) 3D medical images like computerized tomography (CT) remains challenging due to its high imaging cost and privacy restrictions. In our pilot study, we advocated bringing a wealth of 2D images like X-rays as compensation for the lack of 3D data, aiming to build a universal medical self-supervised representation learning framework, called UniMiSS. Especially, we designed a pyramid U-like medical Transformer (MiT) as the backbone to make UniMiSS possible to perform SSL with both 2D and 3D images. UniMiSS surpasses current 3D-specific SSL in effectiveness and versatility, excelling in various downstream tasks and overcoming the limitations of dimensionality. However, the initial version did not fully explore the anatomical correlations between 2D and 3D images due to the absence of paired multi-modal patient data. In this extension, we introduce UniMiSS+, which leverages digitally reconstructed radiographs (DRR) technology to simulate X-rays from CT volumes, providing access to paired data. Benefiting from the paired group, we introduce an extra pair-wise constraint to boost the cross modality correlation learning, which also can be adopted as a cross dimension regularization to further improve the representations. We conduct expensive experiments on multiple 3D/2D medical image analysis tasks, including segmentation and classification. The results show that our UniMiSS+ achieves promising performance on various downstream tasks, not only outperforming ImageNet pre-training and other advanced SSL counterparts but also improving the predecessor UniMiSS pre-training. Yutong Xie 0001, Yong Xia 0001, Qi Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Momentum recursive DARTS
Benteng Ma, Yanning Zhang 0001, Yong Xia 0001 |
Pattern Recognit. | 3 |
| 2024 | Exploratory Training for Universal Lesion Detection: Enhancing Lesion Mining Quality Through Temporal VerificationabstractUniversal lesion detection (ULD) has great value in clinical practice as it can detect various lesions across multiple organs. Deep learning-based detectors have great potential but require high-quality annotated training data. In practice, due to cost, expertise requirements, and the diverse nature of lesions, incomplete annotations are encountered. Directly training ULD detectors under this condition can yield suboptimal results. Leading pseudo-label methods rely on a dynamic lesion-mining mechanism operating at the mini-batch level to address this issue. However, the quality of mined lesions is inconsistent across different iterations, potentially limiting performance enhancement. Inspired by the observation that deep models learn concepts with increasing complexity, we propose an exploratory-training-based ULD (ET-ULD) method to assess the reliability of mined lesions over time. Our approach uses a teacher-student detection model where the teacher mines suspicious lesions, which are then combined with incomplete annotations to train the student. On top of that, we design a bounding-box bank to record the mining timestamps. Each image is trained in several rounds, allowing us to get a sequence of timestamps for the mined lesions. If a mined lesion consistently appears, it is likely to be a true lesion, otherwise, it may just be a noise. This serves as a crucial criterion for selecting reliable mined lesions for retraining. Experimental results show that ET-ULD surpass existing state-of-the-art methods on two distinct lesion image datasets. Notably, on the DeepLesion dataset, ET-ULD achieved a 5.4% improvement in Average Precision (AP) over the previous methods, demonstrating its superior performance. Geng Chen 0001, Benteng Ma, ChangYang Li, Jingfeng Zhang, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | TriLA: Triple-Level Alignment Based Unsupervised Domain Adaptation for Joint Segmentation of Optic Disc and Optic CupabstractCross-domain joint segmentation of optic disc and optic cup on fundus images is essential, yet challenging, for effective glaucoma screening. Although many unsupervised domain adaptation (UDA) methods have been proposed, these methods can hardly achieve complete domain alignment, leading to suboptimal performance. In this paper, we propose a triple-level alignment (TriLA) model to address this issue by aligning the source and target domains at the input level, feature level, and output level simultaneously. At the input level, a learnable Fourier domain adaptation (LFDA) module is developed to learn the cut-off frequency adaptively for frequency-domain translation. At the feature level, we disentangle the style and content features and align them in the corresponding feature spaces using consistency constraints. At the output level, we design a segmentation consistency constraint to emphasize the segmentation consistency across domains. The proposed model is trained on the RIGA+ dataset and widely evaluated on six different UDA scenarios. Our comprehensive results not only demonstrate that the proposed TriLA substantially outperforms other state-of-the-art UDA methods in joint segmentation of optic disc and optic cup, but also suggest the effectiveness of the triple-level alignment strategy. Ziyang Chen 0003, Yongsheng Pan, Yiwen Ye, Zhiyong Wang 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Disentangle Then Calibrate With Gradient Guidance: A Unified Framework for Common and Rare Disease DiagnosisabstractThe computer-aided diagnosis (CAD) for rare diseases using medical imaging poses a significant challenge due to the requirement of large volumes of labeled training data, which is particularly difficult to collect for rare diseases. Although Few-shot learning (FSL) methods have been developed for this task, these methods focus solely on rare disease diagnosis, failing to preserve the performance in common disease diagnosis. To address this issue, we propose the Disentangle then Calibrate with Gradient Guidance (DCGG) framework under the setting of generalized few-shot learning, i.e., using one model to diagnose both common and rare diseases. The DCGG framework consists of a network backbone, a gradient-guided network disentanglement (GND) module, and a gradient-induced feature calibration (GFC) module. The GND module disentangles the network into a disease-shared component and a disease-specific component based on gradient guidance, and devises independent optimization strategies for both components, respectively, when learning from rare diseases. The GFC module transfers only the disease-shared channels of common-disease features to rare diseases, and incorporates the optimal transport theory to identify the best transport scheme based on the semantic relationship among different diseases. Based on the best transport scheme, the GFC module calibrates the distribution of rare-disease features at the disease-shared channels, deriving more informative rare-disease features for better diagnosis. The proposed DCGG framework has been evaluated on three public medical image classification datasets. Our results suggest that the DCGG framework achieves state-of-the-art performance in diagnosing both common and rare diseases. Yuanyuan Chen 0001, Xiaoqing Guo, Yong Xia 0001, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Toward Accurate Cardiac MRI Segmentation With Variational Autoencoder-Based Unsupervised Domain AdaptationabstractAccurate myocardial segmentation is crucial in the diagnosis and treatment of myocardial infarction (MI), especially in Late Gadolinium Enhancement (LGE) cardiac magnetic resonance (CMR) images, where the infarcted myocardium exhibits a greater brightness. However, segmentation annotations for LGE images are usually not available. Although knowledge gained from CMR images of other modalities with ample annotations, such as balanced-Steady State Free Precession (bSSFP), can be transferred to the LGE images, the difference in image distribution between the two modalities (i.e., domain shift) usually results in a significant degradation in model performance. To alleviate this, an end-to-end Variational autoencoder based feature Alignment Module Combining Explicit and Implicit features (VAMCEI) is proposed. We first re-derive the Kullback-Leibler (KL) divergence between the posterior distributions of the two domains as a measure of the global distribution distance. Second, we calculate the prototype contrastive loss between the two domains, bringing closer the prototypes of the same category across domains and pushing away the prototypes of different categories within or across domains. Finally, a domain discriminator is added to the output space, which indirectly aligns the feature distribution and forces the extracted features to be more favorable for segmentation. In addition, by combining CycleGAN and VAMCEI, we propose a more refined multi-stage unsupervised domain adaptation (UDA) framework for myocardial structure segmentation. We conduct extensive experiments on the MSCMRSeg 2019, MyoPS 2020 and MM-WHS 2017 datasets. The experimental results demonstrate that our framework achieves superior performances than state-of-the-art methods. Hengfei Cui, Yan Li 0129, Yifan Wang 0033, Di Xu 0012, Lianming Wu, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Robust Stochastic Neural Ensemble Learning With Noisy Labels for Thoracic Disease ClassificationabstractChest radiography is the most common radiology examination for thoracic disease diagnosis, such as pneumonia. A tremendous number of chest X-rays prompt data-driven deep learning models in constructing computer-aided diagnosis systems for thoracic diseases. However, in realistic radiology practice, a deep learning-based model often suffers from performance degradation when trained on data with noisy labels possibly caused by different types of annotation biases. To this end, we present a novel stochastic neural ensemble learning (SNEL) framework for robust thoracic disease diagnosis using chest X-rays. The core idea of our method is to learn from noisy labels by constructing model ensembles and designing noise-robust loss functions. Specifically, we propose a fast neural ensemble method that collects parameters simultaneously across model instances and along optimization trajectories. Moreover, we propose a loss function that both optimizes a robust measure and characterizes a diversity measure of ensembles. We evaluated our proposed SNEL method on three publicly available hospital-scale chest X-ray datasets. The experimental results indicate that our method outperforms competing methods and demonstrate the effectiveness and robustness of our method in learning from noisy labels. Our code is available at https://github.com/hywang01/SNEL. Hongyu Wang 0011, Hengfei Cui, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Fusion-Embedding Siamese Network for Light Field Salient Object DetectionabstractLight field salient object detection (SOD) has shown remarkable success and gained considerable attention from the computer vision community. Existing methods usually employ a single-/two-stream network to detect saliency. However, these methods can only handle up to two different modalities at a time, preventing them from being able to fully explore the rich information in multi-modal light field derived data. To address this, we propose the first joint multi-modal learning framework, called FES-Net, for light field SOD, which can take rich inputs not limited to two modalities. Specifically, we propose an attention-aware adaptation module to first transform the multi-modal inputs for use in our joint learning framework. The transformed inputs are then fed to a Siamese network along with multiple embedded feature fusion modules to extract informative multi-modal features. Finally, we predict saliency maps from the high-level extracted features using a saliency decoder module. Our joint multi-modal learning framework effectively resolves the limitations of existing methods, providing efficient and effective multi-modal learning that can fully explore the valuable information in light field data for accurate saliency detection. Furthermore, we improve the performance by introducing the Transformer as our backbone network. To the best of our knowledge, the improved version of our model, called FES-Trans, is the first attempt to address the challenging light field SOD with the powerful Transformer technique. Extensive experiments on benchmark datasets demonstrate that our models are superior light field SOD approaches and outperform cutting-edge models remarkably. Geng Chen 0001, Huazhu Fu, Tao Zhou 0002, Guobao Xiao, Keren Fu, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | PEFAT: Boosting Semi-Supervised Medical Image Classification via Pseudo-Loss Estimation and Feature Adversarial TrainingabstractPseudo-labeling approaches have been proven beneficial for semi-supervised learning (SSL) schemes in computer vision and medical imaging. Most works are dedicated to finding samples with high-confidence pseudo-labels from the perspective of model predicted probability. Whereas this way may lead to the inclusion of incorrectly pseudo-labeled data if the threshold is not carefully adjusted. In addition, low-confidence probability samples are frequently disregarded and not employed to their full potential. In this paper, we propose a novel Pseudo-loss Estimation and Feature Adversarial Training semi-supervised framework, termed as PEFAT, to boost the performance of multi-class and multi-label medical image classification from the point of loss distribution modeling and adversarial training. Specifically, we develop a trustworthy data selection scheme to split a high-quality pseudo-labeled set, inspired by the dividable pseudo-loss assumption that clean data tend to show lower loss while noise data is the opposite. Instead of directly discarding these samples with low-quality pseudo-labels, we present a novel regularization approach to learn discriminate information from them via injecting adversarial noises at the feature-level to smooth the decision boundary. Experimental results on three medical and two natural image benchmarks validate that our PEFAT can achieve a promising performance and surpass other state-of-the-art methods. The code is available at https://github.com/maxwell0027/PEFAT. Qingjie Zeng, Yutong Xie 0001, Zilin Lu, Yong Xia 0001 |
CVPR | 4 |
| 2023 | Treasure in Distribution: A Domain Randomization Based Multi-source Domain Generalization for 2D Medical Image Segmentation
Ziyang Chen 0003, Yongsheng Pan, Yiwen Ye, Hengfei Cui, Yong Xia 0001 |
MICCAI (4) | 5 |
| 2023 | Unpaired Cross-Modal Interaction Learning for COVID-19 Segmentation on Limited CT Images
Qingbiao Guan, Yutong Xie 0001, Zhibin Liao, Qi Wu 0001, Yong Xia 0001 |
MICCAI (3) | 7 |
| 2023 | Devil is in Channels: Contrastive Single Domain Generalization for Medical Image Segmentation
Shishuai Hu, Zehui Liao, Yong Xia 0001 |
MICCAI (4) | 3 |
| 2023 | Transformer-Based Annotation Bias-Aware Medical Image Segmentation
Zehui Liao, Shishuai Hu, Yutong Xie 0001, Yong Xia 0001 |
MICCAI (4) | 4 |
| 2023 | Multi-modal Pathological Pre-training via Masked Autoencoders for Breast Cancer Diagnosis
Mengkang Lu, Yong Xia 0001 |
MICCAI (6) | 3 |
| 2023 | Revealing Anatomical Structures in PET to Generate CT for Attenuation Correction
Yongsheng Pan, Feihong Liu, Caiwen Jiang, Yong Xia 0001, Dinggang Shen |
MICCAI (10) | 5 |
| 2023 | MedIM: Boost Medical Image Representation via Radiology Report-Guided Masking
Yutong Xie 0001, Lin Gu 0003, Tatsuya Harada, Yong Xia 0001, Qi Wu 0001 |
MICCAI (1) | 5 |
| 2023 | Towards Accurate Microstructure Estimation via 3D Hybrid Graph Transformer
Junqing Yang, Tewodros Megabiaw Tassew, Jiquan Ma, Yong Xia 0001, Pew-Thian Yap, Geng Chen 0001 |
MICCAI (8) | 6 |
| 2023 | UniSeg: A Prompt-Driven Universal Segmentation Model as Well as A Strong Representation Learner
Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001 |
MICCAI (3) | 5 |
| 2023 | TPRO: Text-Prompting-Based Weakly Supervised Histopathology Tissue Segmentation
Shaoteng Zhang, Yutong Xie 0001, Yong Xia 0001 |
MICCAI (1) | 4 |
| 2023 | DAN-NucNet: A dual attention based framework for nuclei segmentation in cancer histology images under wild clinical conditions
Ibtihaj Ahmad, Yong Xia 0001, Hengfei Cui, Zain Ul Islam |
Expert Syst. Appl. | 2 |
| 2023 | Hyperspectral anomaly detection via weighted-sparsity-regularized tensor linear representationabstractAbstract Anomaly detection aims at locating the spectral different objects of a specific scene without any prior information, and has gained increasing attention. By decomposing the input hyperspectral image (HSI) into a background tensor and an anomaly tensor, the tensor approximation is an efficient tool for detecting the anomalies. Low rankness is usually utilized as the regularizer during the background reconstruction process. Different from most existing hyperspectral anomaly detection methods which compute the truncated nuclear norm of the third folding of the original HSI, a novel weighted‐sparsity‐regularized tensor linear representation (WsrTLR) method is proposed for hyperspectral anomaly detection in this paper. Tensor linear representation is utilized to formulate the background HSI by a three‐dimensional (3D) representation base and the corresponding 3D representation coefficient. Low rankness is applied to constrict the representation coefficient, an operation which avoids destroying the multi‐way structure and losing information during the matrixing process, and ensures a satisfactory detection accuracy. Meanwhile, by incorporating the weighted‐sparsity‐regularized tensor linear representation to reconstruct the background tensor, the anomalies can be easily detected by eliminating the background tensor from the original scene. In addition, to avoid negative influence caused by the redundant bands and noisy bands in the representation process, informative bands have been first selected via an optimal neighborhood reconstruction strategy. Experimental results and data analysis on four real hyperspectral datasets, which contain anomalies with different sizes, have demonstrated the effectiveness of the proposed method. Jinqiu Sun, Yong Xia 0001, Yanning Zhang 0001 |
IET Image Process. | 3 |
| 2023 | MyoPS: A benchmark of myocardial pathology segmentation combining three-sequence cardiac magnetic resonance images
Lei Li 0020, Fuping Wu, Xinzhe Luo, Carlos Martín-Isla, Shuwei Zhai, Zhen Zhang 0057, Markus J. Ankenbrand, Haochuan Jiang, Linhong Wang, Tewodros Weldebirhan Arega, Elif Altunok, Jun Ma 0016, Xiaoping Yang 0001, Élodie Puybareau, Ilkay Öksüz, Stéphanie Bricq, Weisheng Li 0001, Kumaradevan Punithakumar, Sotirios A. Tsaftaris, Laura Maria Schreiber, Guocai Liu, Yong Xia 0001, Guotai Wang, Sergio Escalera, Xiahai Zhuang |
Medical Image Anal. | 29 |
| 2023 | Dynamic feature splicing for few-shot rare disease diagnosis
Yuanyuan Chen 0001, Xiaoqing Guo, Yongsheng Pan, Yong Xia 0001, Yixuan Yuan |
Medical Image Anal. | 4 |
| 2023 | Learning From Partially Labeled Data for Multi-Organ and Tumor SegmentationabstractMedical image benchmarks for the segmentation of organs and tumors suffer from the partially labeling issue due to its intensive cost of labor and expertise. Current mainstream approaches follow the practice of one network solving one task. With this pipeline, not only the performance is limited by the typically small dataset of a single task, but also the computation cost linearly increases with the number of tasks. To address this, we propose a Transformer based dynamic on-demand network (TransDoDNet) that learns to segment organs and tumors on multiple partially labeled datasets. Specifically, TransDoDNet has a hybrid backbone that is composed of the convolutional neural network and Transformer. A dynamic head enables the network to accomplish multiple segmentation tasks flexibly. Unlike existing approaches that fix kernels after training, the kernels in the dynamic head are generated adaptively by the Transformer, which employs the self-attention mechanism to model long-range organ-wise dependencies and decodes the organ embedding that can represent each organ. We create a large-scale partially labeled Multi-Organ and Tumor Segmentation benchmark, termed MOTS, and demonstrate the superior performance of our TransDoDNet over other competitors on seven organ and tumor segmentation tasks. This study also provides a general 3D medical image segmentation model, which has been pre-trained on the large-scale MOTS benchmark and has demonstrated advanced performance over current predominant self-supervised learning methods. Yutong Xie 0001, Yong Xia 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Federated adaptive reweighting for medical image classification
Benteng Ma, Geng Chen 0001, ChangYang Li, Yong Xia 0001 |
Pattern Recognit. | 5 |
| 2023 | Inter-layer transition in neural architecture search
Benteng Ma, Jing Zhang 0037, Yong Xia 0001, Dacheng Tao |
Pattern Recognit. | 3 |
| 2023 | Reconstruction-Driven Dynamic Refinement Based Unsupervised Domain Adaptation for Joint Optic Disc and Cup SegmentationabstractGlaucoma is one of the leading causes of irreversible blindness. Segmentation of optic disc (OD) and optic cup (OC) on fundus images is a crucial step in glaucoma screening. Although many deep learning models have been constructed for this task, it remains challenging to train an OD/OC segmentation model that could be deployed successfully to different healthcare centers. The difficulties mainly comes from the domain shift issue, i.e., the fundus images collected at these centers usually vary greatly in the tone, contrast, and brightness. To address this issue, in this paper, we propose a novel unsupervised domain adaptation (UDA) method called Reconstruction-driven Dynamic Refinement Network (RDR-Net), where we employ a due-path segmentation backbone for simultaneous edge detection and region prediction and design three modules to alleviate the domain gap. The reconstruction alignment (RA) module uses a variational auto-encoder (VAE) to reconstruct the input image and thus boosts the image representation ability of the network in a self-supervised way. It also uses a style-consistency constraint to force the network to retain more domain-invariant information. The low-level feature refinement (LFR) module employs input-specific dynamic convolutions to suppress the domain-variant information in the obtained low-level features. The prediction-map alignment (PMA) module elaborates the entropy-driven adversarial learning to encourage the network to generate source-like boundaries and regions. We evaluated our RDR-Net against state-of-the-art solutions on four public fundus image datasets. Our results indicate that RDR-Net is superior to competing models in both segmentation performance and generalization ability. Ziyang Chen 0003, Yongsheng Pan, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | An Improved Combination of Faster R-CNN and U-Net Network for Accurate Multi-Modality Whole Heart SegmentationabstractDetailed information of substructures of the whole heart is usually vital in the diagnosis of cardiovascular diseases and in 3D modeling of the heart. Deep convolutional neural networks have been demonstrated to achieve state-of-the-art performance in 3D cardiac structures segmentation. However, when dealing with high-resolution 3D data, current methods employing tiling strategies usually degrade segmentation performances due to GPU memory constraints. This work develops a two-stage multi-modality whole heart segmentation strategy, which adopts an improved Combination of Faster R-CNN and 3D U-Net (CFUN+). More specifically, the bounding box of the heart is first detected by Faster R-CNN, and then the original Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) images of the heart aligned with the bounding box are input into 3D U-Net for segmentation. The proposed CFUN+ method redefines the bounding box loss function by replacing the previous Intersection over Union (IoU) loss with Complete Intersection over Union (CIoU) loss. Meanwhile, the integration of the edge loss makes the segmentation results more accurate, and also improves the convergence speed. The proposed method achieves an average Dice score of 91.1% on the Multi-Modality Whole Heart Segmentation (MM-WHS) 2017 challenge CT dataset, which is 5.2% higher than the baseline CFUN model, and achieves state-of-the-art segmentation results. In addition, the segmentation speed of a single heart has been dramatically improved from a few minutes to less than 6 seconds. Hengfei Cui, Yifan Wang 0033, Yan Li 0129, Di Xu 0010, Lei Jiang 0015, Yong Xia 0001, Yanning Zhang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Disentangle First, Then Distill: A Unified Framework for Missing Modality Imputation and Alzheimer's Disease DiagnosisabstractMulti-modality medical data provide complementary information, and hence have been widely explored for computer-aided AD diagnosis. However, the research is hindered by the unavoidable missing-data problem, i.e., one data modality was not acquired on some subjects due to various reasons. Although the missing data can be imputed using generative models, the imputation process may introduce unrealistic information to the classification process, leading to poor performance. In this paper, we propose the Disentangle First, Then Distill (DFTD) framework for AD diagnosis using incomplete multi-modality medical images. First, we design a region-aware disentanglement module to disentangle each image into inter-modality relevant representation and intra-modality specific representation with emphasis on disease-related regions. To progressively integrate multi-modality knowledge, we then construct an imputation-induced distillation module, in which a lateral inter-modality transition unit is created to impute representation of the missing modality. The proposed DFTD framework has been evaluated against six existing methods on an ADNI dataset with 1248 subjects. The results show that our method has superior performance in both AD-CN classification and MCI-to-AD prediction tasks, substantially over-performing all competing methods. Yuanyuan Chen 0001, Yongsheng Pan, Yong Xia 0001, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Domain and Content Adaptive Convolution Based Multi-Source Domain Generalization for Medical Image SegmentationabstractThe domain gap caused mainly by variable medical image quality renders a major obstacle on the path between training a segmentation model in the lab and applying the trained model to unseen clinical data. To address this issue, domain generalization methods have been proposed, which however usually use static convolutions and are less flexible. In this paper, we propose a multi-source domain generalization model based on the domain and content adaptive convolution (DCAC) for the segmentation of medical images across different modalities. Specifically, we design the domain adaptive convolution (DAC) module and content adaptive convolution (CAC) module and incorporate both into an encoder-decoder backbone. In the DAC module, a dynamic convolutional head is conditioned on the predicted domain code of the input to make our model adapt to the unseen target domain. In the CAC module, a dynamic convolutional head is conditioned on the global image features to make our model adapt to the test image. We evaluated the DCAC model against the baseline and four state-of-the-art domain generalization methods on the prostate segmentation, COVID-19 lesion segmentation, and optic cup/optic disc segmentation tasks. Our results not only indicate that the proposed DCAC model outperforms all competing methods on each segmentation task but also demonstrate the effectiveness of the DAC and CAC modules. Code is available at https://git.io/DCAC. Shishuai Hu, Zehui Liao, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Survival Prediction via Hierarchical Multimodal Co-Attention Transformer: A Computational Histology-Radiology SolutionabstractThe rapid advances in deep learning-based computational pathology and radiology have demonstrated the promise of using whole slide images (WSIs) and radiology images for survival prediction in cancer patients. However, most image-based survival prediction methods are limited to using either histology or radiology alone, leaving integrated approaches across histology and radiology relatively underdeveloped. There are two main challenges in integrating WSIs and radiology images: (1) the gigapixel nature of WSIs and (2) the vast difference in spatial scales between WSIs and radiology images. To address these challenges, in this work, we propose an interpretable, weakly-supervised, multimodal learning framework, called Hierarchical Multimodal Co-Attention Transformer (HMCAT), to integrate WSIs and radiology images for survival prediction. Our approach first uses hierarchical feature extractors to capture various information including cellular features, cellular organization, and tissue phenotypes in WSIs. Then the hierarchical radiology-guided co- attention (HRCA) in HMCAT characterizes the multimodal interactions between hierarchical histology-based visual concepts and radiology features and learns hierarchical co- attention mappings for two modalities. Finally, HMCAT combines their complementary information into a multimodal risk score and discovers prognostic features from two modalities by multimodal interpretability. We apply our approach to two cancer datasets (365 WSIs with matched magnetic resonance [MR] images and 213 WSIs with matched computed tomography [CT] images). Our results demonstrate that the proposed HMCAT consistently achieves superior performance over the unimodal approaches trained on either histology or radiology data alone, as well as other state-of-the-art methods. Zhe Li 0006, Yuming Jiang 0005, Mengkang Lu, Ruijiang Li, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Cascade Multi-Level Transformer Network for Surgical Workflow AnalysisabstractSurgical workflow analysis aims to recognise surgical phases from untrimmed surgical videos. It is an integral component for enabling context-aware computer-aided surgical operating systems. Many deep learning-based methods have been developed for this task. However, most existing works aggregate homogeneous temporal context for all frames at a single level and neglect the fact that each frame has its specific need for information at multiple levels for accurate phase prediction. To fill this gap, in this paper we propose Cascade Multi-Level Transformer Network (CMTNet) composed of cascaded Adaptive Multi-Level Context Aggregation (AMCA) modules. Each AMCA module first extracts temporal context at the frame level and the phase level and then fuses frame-specific spatial feature, frame-level temporal context, and phase-level temporal context for each frame adaptively. By cascading multiple AMCA modules, CMTNet is able to gradually enrich the representation of each frame with the multi-level semantics that it specifically requires, achieving better phase prediction in a frame-adaptive manner. In addition, we propose a novel refinement loss for CMTNet, which explicitly guides each AMCA module to focus on extracting the key context for refining the prediction of the previous stage in terms of both prediction confidence and smoothness. This further enhances the quality of the extracted context effectively. Extensive experiments on the Cholec80 and the M2CAI datasets demonstrate that CMTNet achieves state-of-the-art performance. Wenxi Yue, Hongen Liao, Yong Xia 0001, Vincent Lam, Jiebo Luo 0001, Zhiyong Wang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | FIBA: Frequency-Injection based Backdoor Attack in Medical Image AnalysisabstractIn recent years, the security of AI systems has drawn increasing research attention, especially in the medical imaging realm. To develop a secure medical image analysis (MIA) system, it is a must to study possible backdoor attacks (BAs), which can embed hidden malicious behaviors into the system. However, designing a unified BA method that can be applied to various MIA systems is challenging due to the diversity of imaging modalities (e.g., X-Ray, CT, and MRI) and analysis tasks (e.g., classification, detection, and segmentation). Most existing BA methods are designed to attack natural image classification models, which apply spatial triggers to training images and inevitably corrupt the semantics of poisoned pixels, leading to the failures of attacking dense prediction models. To address this issue, we propose a novel Frequency-Injection based Backdoor Attack method (FIBA) that is capable of delivering attacks in various MIA tasks. Specifically, FIBA leverages a trigger function in the frequency domain that can inject the low-frequency information of a trigger image into the poisoned image by linearly combining the spectral amplitude of both images. Since it preserves the semantics of the poisoned image pixels, FIBA can perform attacks on both classification and dense prediction models. Experiments on three benchmarks in MIA (i.e., ISIC-2019 [4] for skin lesion classification, KiTS-19 [17] for kidney tumor segmentation, and EAD-2019 [1] for endoscopic artifact detection), validate the effectiveness of FIBA and its superiority over stateof-the-art methods in attacking MIA models and bypassing backdoor defense. Source code will be available at code. Benteng Ma, Jing Zhang 0037, Shanshan Zhao 0001, Yong Xia 0001, Dacheng Tao |
CVPR | 5 |
| 2022 | UniMiSS: Universal Medical Self-supervised Learning via Breaking Dimensionality Barrier
Yutong Xie 0001, Yong Xia 0001, Qi Wu 0001 |
ECCV (21) | 3 |
| 2022 | Disentangle Then Calibrate: Selective Treasure Sharing for Generalized Rare Disease Diagnosis
Yuanyuan Chen 0001, Xiaoqing Guo, Yong Xia 0001, Yixuan Yuan |
MICCAI (3) | 3 |
| 2022 | Hybrid Graph Transformer for Tissue Microstructure Estimation with Undersampled Diffusion MRI Data
Geng Chen 0001, Jiannan Liu, Jiquan Ma, Hui Cui 0002, Yong Xia 0001, Pew-Thian Yap |
MICCAI (1) | 6 |
| 2022 | Domain Specific Convolution and High Frequency Reconstruction Based Unsupervised Domain Adaptation for Medical Image Segmentation
Shishuai Hu, Zehui Liao, Yong Xia 0001 |
MICCAI (8) | 3 |
| 2022 | DeSD: Self-Supervised Learning with Deep Self-Distillation for 3D Medical Image Segmentation
Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001 |
MICCAI (4) | 4 |
| 2022 | Semantically-Consistent Dynamic Blurry Image Generation for Image DeblurringabstractThe training of deep learning-based image deblurring models heavily relies on the paired sharp/blurry image dataset. Although many works verified that synthesized blurry-sharp pairs contribute to improving the deblurring performance, it is still an open problem about how to synthesize realistic and diverse dynamic blurry images. Instead of directly synthesizing blurry images, in this paper, we propose a novel method to generate semantic-aware dense dynamic motion, and employ the generated motion to synthesize blurry images. Specifically, for each sharp image, both the global motion (camera shake) and local motion (object moving) are considered given the depth information as the condition. Then, a blur creation module takes the spatial-variant motion information and the sharp image as input to synthesize a motion-blurred image. A relativistic GAN loss is employed to assure the synthesized blurry image is as realistic as possible. Experiments show that our method can generate diverse dynamic motion and visually realistic blurry images. Also, the generated image pairs can further improve the quantitative performance and generalization ability of the existing deblurring method on several test sets. Zhaohui Jing, Youjian Zhang, Daqing Liu, Yong Xia 0001 |
ACM Multimedia | 5 |
| 2022 | Deep U-Net architecture with curriculum learning for myocardial pathology segmentation in multi-sequence cardiac magnetic resonance images
Hengfei Cui, Lei Jiang 0015, Chang Yuwen, Yong Xia 0001, Yanning Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Rapid artificial intelligence solutions in a pandemic - The COVID-19-20 Lung CT Lesion Segmentation Challenge
Holger Roth, Ziyue Xu 0001, Carlos Tor-Díez, Ramon Sánchez-Jacob, Jonathan Zember, Jose Molto, Wenqi Li 0001, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Dong Yang 0005, Ahmed Harouni, Nicola Rieke, Shishuai Hu, Fabian Isensee, Claire Tang, Qinji Yu, Jan Sölter, Vitali Liauchuk, Jan Hendrik Moltz, Bruno Oliveira 0002, Yong Xia 0001, Klaus H. Maier-Hein, Qikai Li, Andreas Husch, Vassili Kovalev, Alessa Hering, João L. Vilaça, Mona Flores, Daguang Xu, Bradford J. Wood, Marius George Linguraru |
Medical Image Anal. | 24 |
| 2022 | Mutual consistency learning for semi-supervised medical image segmentation
Yicheng Wu 0001, ZongYuan Ge, Donghao Zhang 0004, Minfeng Xu, Lei Zhang 0006, Yong Xia 0001, Jianfei Cai 0001 |
Medical Image Anal. | 6 |
| 2022 | Disease-Image-Specific Learning for Diagnosis-Oriented Neuroimage Synthesis With Incomplete Multi-Modality DataabstractIncomplete data problem is commonly existing in classification tasks with multi-source data, particularly the disease diagnosis with multi-modality neuroimages, to track which, some methods have been proposed to utilize all available subjects by imputing missing neuroimages. However, these methods usually treat image synthesis and disease diagnosis as two standalone tasks, thus ignoring the specificity conveyed in different modalities, i.e., different modalities may highlight different disease-relevant regions in the brain. To this end, we propose a disease-image-specific deep learning (DSDL) framework for joint neuroimage synthesis and disease diagnosis using incomplete multi-modality neuroimages. Specifically, with each whole-brain scan as input, we first design a Disease-image-Specific Network (DSNet) with a spatial cosine module to implicitly model the disease-image specificity. We then develop a Feature-consistency Generative Adversarial Network (FGAN) to impute missing neuroimages, where feature maps (generated by DSNet) of a synthetic image and its respective real image are encouraged to be consistent while preserving the disease-image-specific information. Since our FGAN is correlated with DSNet, missing neuroimages can be synthesized in a diagnosis-oriented manner. Experimental results on three datasets suggest that our method can not only generate reasonable neuroimages, but also achieve state-of-the-art performance in both tasks of Alzheimer's disease identification and mild cognitive impairment conversion prediction. Yongsheng Pan, Mingxia Liu 0001, Yong Xia 0001, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning multi-scale synergic discriminative features for prostate image segmentation
Haozhe Jia, Tom Weidong Cai, Heng Huang 0001, Yong Xia 0001 |
Pattern Recognit. | 4 |
| 2022 | A cascaded nested network for 3T brain MR image segmentation guided by 7T labeling
Zhengwang Wu, Li Wang 0026, Toan Duc Bui, Liangqiong Qu, Pew-Thian Yap, Yong Xia 0001, Gang Li 0001, Dinggang Shen |
Pattern Recognit. | 7 |
| 2022 | Intra- and Inter-Pair Consistency for Semi-Supervised Gland SegmentationabstractAccurate gland segmentation in histology tissue images is a critical but challenging task. Although deep models have demonstrated superior performance in medical image segmentation, they commonly require a large amount of annotated data, which are hard to obtain due to the extensive labor costs and expertise required. In this paper, we propose an intra- and inter-pair consistency-based semi-supervised (I2CS) model that can be trained on both labeled and unlabeled histology images for gland segmentation. Considering that each image contains glands and hence different images could potentially share consistent semantics in the feature space, we introduce a novel intra- and inter-pair consistency module to explore such consistency for learning with unlabeled data. It first characterizes the pixel-level relation between a pair of images in the feature space to create an attention map that highlights the regions with the same semantics but on different images. Then, it imposes a consistency constraint on the attention maps obtained from multiple image pairs, and thus filters low-confidence attention regions to generate refined attention maps that are then merged with original features to improve their representation ability. In addition, we also design an object-level loss to address the issues caused by touching glands. We evaluated our model against several recent gland segmentation methods and three typical semi-supervised methods on the GlaS and CRAG datasets. Our results not only demonstrate the effectiveness of the proposed due consistency module and Obj-Dice loss, but also indicate that the proposed I2CS model achieves state-of-the-art gland segmentation performance on both benchmarks. Yutong Xie 0001, Zhibin Liao, Johan Verjans, Chunhua Shen, Yong Xia 0001 |
IEEE Trans. Image Process. | 6 |
| 2022 | MFI-Net: Multiscale Feature Interaction Network for Retinal Vessel SegmentationabstractSegmentation of retinal vessels on fundus images plays a critical role in the diagnosis of micro-vascular and ophthalmological diseases. Although being extensively studied, this task remains challenging due to many factors including the highly variable vessel width and poor vessel-background contrast. In this paper, we propose a multiscale feature interaction network (MFI-Net) for retinal vessel segmentation, which is a U-shaped convolutional neural network equipped with the pyramid squeeze-and-excitation (PSE) module, coarse-to-fine (C2F) module, deep supervision, and feature fusion. We extend the SE operator to multiscale features, resulting in the PSE module, which uses the channel attention learned at multiple scales to enhance multiscale features and enables the network to handle the vessels with variable width. We further design the C2F module to generate and re-process the residual feature maps, aiming to preserve more vessel details during the decoding process. The proposed MFI-Net has been evaluated against several public models on the DRIVE, STARE, CHASE_DB1, and HRF datasets. Our results suggest that both PSE and C2F modules are effective in improving the accuracy of MFI-Net, and also indicate that our model has superior segmentation performance and generalization ability over existing models on four public datasets. Yiwen Ye, Chengwei Pan, Yicheng Wu 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | SC2Net: A Novel Segmentation-Based Classification Network for Detection of COVID-19 in Chest X-Ray ImagesabstractThe pandemic of COVID-19 has become a global crisis in public health, which has led to a massive number of deaths and severe economic degradation. To suppress the spread of COVID-19, accurate diagnosis at an early stage is crucial. As the popularly used real-time reverse transcriptase polymerase chain reaction (RT-PCR) swab test can be lengthy and inaccurate, chest screening with radiography imaging is still preferred. However, due to limited image data and the difficulty of the early-stage diagnosis, existing models suffer from ineffective feature extraction and poor network convergence and optimisation. To tackle these issues, a segmentation-based COVID-19 classification network, namely SC2Net, is proposed for effective detection of the COVID-19 from chest x-ray (CXR) images. The SC2Net consists of two subnets: a COVID-19 lung segmentation network (CLSeg), and a spatial attention network (SANet). In order to supress the interference from the background, the CLSeg is first applied to segment the lung region from the CXR. The segmented lung region is then fed to the SANet for classification and diagnosis of the COVID-19. As a shallow yet effective classifier, SANet takes the ResNet-18 as the feature extractor and enhances high-level feature via the proposed spatial attention module. For performance evaluation, the COVIDGR 1.0 dataset is used, which is a high-quality dataset with various severity levels of the COVID-19. Experimental results have shown that, our SC2Net has an average accuracy of 84.23% and an average F1 score of 81.31% in detection of COVID-19, outperforming several state-of-the-art approaches. Huimin Zhao 0001, Zhenyu Fang, Jinchang Ren, Calum MacLellan, Yong Xia 0001, Shuo Li 0001, Meijun Sun, Kevin Ren |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Learning From Ambiguous Labels for Lung Nodule Malignancy PredictionabstractLung nodule malignancy prediction is an essential step in the early diagnosis of lung cancer. Besides the difficulties commonly discussed, the challenges of this task also come from the ambiguous labels provided by annotators, since deep learning models have in some cases been found to reproduce or amplify human biases. In this paper, we propose a multi-view 'divide-and-rule' (MV-DAR) model to learn from both reliable and ambiguous annotations for lung nodule malignancy prediction on chest CT scans. According to the consistency and reliability of their annotations, we divide nodules into three sets: a consistent and reliable set (CR-Set), an inconsistent set (IC-Set), and a low reliable set (LR-Set). The nodule in IC-Set is annotated by multiple radiologists inconsistently, and the nodule in LR-Set is annotated by only one radiologist. Although ambiguous, inconsistent labels tell which label(s) is consistently excluded by all annotators, and the unreliable labels of a cohort of nodules are largely correct from the statistical point of view. Hence, both IC-Set and LR-Set can be used to facilitate the training of MV-DAR. Our MV-DAR contains three DAR models to characterize a lung nodule from three orthographic views and is trained following a two-stage procedure. Each DAR consists of three networks with the same architecture, including a prediction network (Prd-Net), a counterfactual network (CF-Net), and a low reliable network (LR-Net), which are trained on CR-Set, IC-Set, and LR-Set respectively in the pretraining phase. In the fine-tuning phase, the image representation ability learned by CF-Net and LR-Net is transferred to Prd-Net by negative-attention module (NA-Module) and consistent-attention module (CA-Module), aiming to boost the prediction ability of Prd-Net. The MV-DAR model has been evaluated on the LIDC-IDRI dataset and LUNGx dataset. Our results indicate not only the effectiveness of the MV-DAR in learning from ambiguous labels but also its superiority over present noisy label-learning models in lung nodule malignancy prediction. Zehui Liao, Yutong Xie 0001, Shishuai Hu, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Learning Synergistic Attention for Light Field Salient Object Detection
Yi Zhang 0076, Geng Chen 0001, Yong Xia 0001, Olivier Déforges, Wassim Hamidouche, Lu Zhang 0037 |
BMVC | 5 |
| 2021 | BiCMTS: Bidirectional Coupled Multivariate Learning of Irregular Time Series with Missing ValuesabstractMultivariate time series (MTS) such as multiple medical measures in intensive care units (ICU) are irregularly acquired and hold missing values. Conducting learning tasks on such irregular MTS with missing values, e.g., predicting the mortality of ICU patients, poses significant challenge to existing MTS forecasting models and recurrent neural networks (RNNs), which capture the temporal dependencies within a time series. This work proposes a bidirectional coupled MTS learning (BiCMTS) method to represent both forward and backward value couplings within a time series by RNNs and between MTS by self-attention networks; the learned bidirectional intra- and inter-time series coupling representations are fused to estimate missing values. We test BiCMTS on both data imputation and mortality prediction for ICU patients, showing a great potential of leveraging the deep and hidden relations captured in RNNs by the BiCMTS-learned intra- and inter-time series value couplings in MTS. Qinfen Wang, Yong Xia 0001, Longbing Cao |
CIKM | 3 |
| 2021 | DoDNet: Learning To Segment Multi-Organ and Tumors From Multiple Partially Labeled DatasetsabstractDue to the intensive cost of labor and expertise in annotating 3D medical images at a voxel level, most benchmark datasets are equipped with the annotations of only one type of organs and/or tumors, resulting in the so-called partially labeling issue. To address this issue, we propose a dynamic on-demand network (DoDNet) that learns to segment multiple organs and tumors on partially labeled datasets. DoD-Net consists of a shared encoder-decoder architecture, a task encoding module, a controller for dynamic filter generation, and a single but dynamic segmentation head. The information of current segmentation task is encoded as a task-aware prior to tell the model what the task is expected to achieve. Different from existing approaches which fix kernels after training, the kernels in dynamic head are generated adaptively by the controller, conditioned on both input image and assigned task. Thus, DoDNet is able to segment multiple organs and tumors, as done by multiple networks or a multi-head network, in a much efficient and flexible manner. We created a large-scale partially labeled dataset called MOTS and demonstrated the superior performance of our DoDNet over other competitors on seven organ and tumor segmentation tasks. We also transferred the weights pre-trained on MOTS to a downstream multi-organ segmentation task and achieved state-of-the-art performance. This study provides a general 3D medical image segmentation model that has been pre-trained on a large-scale partially labeled dataset and can be extended (after fine-tuning) to downstream volumetric medical data segmentation tasks. Code and models are available at: https://git.io/DoDNet Yutong Xie 0001, Yong Xia 0001, Chunhua Shen |
CVPR | 3 |
| 2021 | Deformable Convolution and Semi-supervised Learning in Point Clouds for Aneurysm Classification and Segmentation
Erik Meijering, Yong Xia 0001, Yang Song 0001 |
ICONIP (6) | 3 |
| 2021 | Collaborative Image Synthesis and Disease Diagnosis for Classification of Neurodegenerative Disorders with Incomplete Multi-modal Neuroimages
Yongsheng Pan, Yuanyuan Chen 0001, Dinggang Shen, Yong Xia 0001 |
MICCAI (5) | 4 |
| 2021 | Predicting Symptoms from Multiphasic MRI via Multi-instance Attention Learning for Hepatocellular Carcinoma Grading
Zelin Qiu, Yongsheng Pan, Dijia Wu, Yong Xia 0001, Dinggang Shen |
MICCAI (5) | 5 |
| 2021 | Consistent Segmentation of Longitudinal Brain MR Images with Spatio-Temporal Constrained Networks
Feng Shi 0001, Zhiming Cui 0001, Yongsheng Pan, Yong Xia 0001, Dinggang Shen |
MICCAI (1) | 5 |
| 2021 | CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation
Yutong Xie 0001, Chunhua Shen, Yong Xia 0001 |
MICCAI (3) | 4 |
| 2021 | Anomaly Detection of Hyperspectral Image via Tensor CompletionabstractIn this letter, a novel method of anomaly detection for the hyperspectral image (HSI) is proposed. This method originates from two ideas. First, compared with the anomalies, the spectral curves of some (not all) backgrounds are usually easy to be accurately found. Second, the spectral curves of the missing pixels in the background can be recovered via the tensor completion technology. In this way, anomalies can be picked out via discriminating the background tensor and the original HSI. Specifically, some background pixels with low response in the detection map are first detected. Then, the selected background pixels with their spectral curves are utilized to reconstruct a three-order tensor whose elements are missing at some extent. Tensor completion technology is applied to this tensor and achieves a complete tensor, which depicts the background of the scene. Finally, the reconstructed tensor originated from the background pixels is discriminated from the original HSI to pick out the anomalies. The experimental data and performance analysis have demonstrated the effectiveness of the proposed method. Yong Xia 0001, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Triple attention learning for classification of 14 thoracic diseases using chest radiography
Hongyu Wang 0011, Shanshan Wang 0002, Zibo Qin, Yanning Zhang 0001, Ruijiang Li, Yong Xia 0001 |
Medical Image Anal. | 6 |
| 2021 | Iterative sparse and deep learning for accurate diagnosis of Alzheimer's disease
Yuanyuan Chen 0001, Yong Xia 0001 |
Pattern Recognit. | 2 |
| 2021 | Multi-View Mammographic Density Classification by Dilated and Attention-Guided Residual LearningabstractBreast density is widely adopted to reflect the likelihood of early breast cancer development. Existing methods of mammographic density classification either require steps of manual operations or achieve only moderate classification accuracy due to the limited model capacity. In this study, we present a radiomics approach based on dilated and attention-guided residual learning for the task of mammographic density classification. The proposed method was instantiated with two datasets, one clinical dataset and one publicly available dataset, and classification accuracies of 88.7 and 70.0 percent were obtained, respectively. Although the classification accuracy of the public dataset was lower than the clinical dataset, which was very likely related to the dataset size, our proposed model still achieved a better performance than the naive residual networks and several recently published deep learning-based approaches. Furthermore, we designed a multi-stream network architecture specifically targeting at analyzing the multi-view mammograms. Utilizing the clinical dataset, we validated that multi-view inputs were beneficial to the breast density classification task with an increase of at least 2.0 percent in accuracy and the different views lead to different model classification capacities. Our method has a great potential to be further developed and applied in computer-aided diagnosis systems. Our code is available at https://github.com/lich0031/Mammographic_Density_Classification. Cheng Li 0008, Jingxu Xu, Qiegen Liu, Yongjin Zhou 0002, Lisha Mou, Zuhui Pu, Yong Xia 0001, Hairong Zheng, Shanshan Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | D-UNet: A Dimension-Fusion U Shape Network for Chronic Stroke Lesion SegmentationabstractAssessing the location and extent of lesions caused by chronic stroke is critical for medical diagnosis, surgical planning, and prognosis. In recent years, with the rapid development of 2D and 3D convolutional neural networks (CNN), the encoder-decoder structure has shown great potential in the field of medical image segmentation. However, the 2D CNN ignores the 3D information of medical images, while the 3D CNN suffers from high computational resource demands. This paper proposes a new architecture called dimension-fusion-UNet (D-UNet), which combines 2D and 3D convolution innovatively in the encoding stage. The proposed architecture achieves a better segmentation performance than 2D networks, while requiring significantly less computation time in comparison to 3D networks. Furthermore, to alleviate the data imbalance issue between positive and negative samples for the network training, we propose a new loss function called Enhance Mixing Loss (EML). This function adds a weighted focal coefficient and combines two traditional loss functions. The proposed method has been tested on the ATLAS dataset and compared to three state-of-the-art methods. The results demonstrate that the proposed method achieves the best quality performance in terms of DSC = 0.5349 ± 0.2763 and precision = 0.6331 ± 0.295). Yongjin Zhou 0002, Weijian Huang, Pei Dong, Yong Xia 0001, Shanshan Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Deep Reinforcement Learning for Weakly-Supervised Lymph Node Segmentation in CT ImagesabstractAccurate and automated lymph node segmentation is pivotal for quantitatively accessing disease progression and potential therapeutics. The complex variation of lymph node morphology and the difficulty of acquiring voxel-wise manual annotations make lymph node segmentation a challenging task. Since the Response Evaluation Criteria in Solid Tumors (RECIST) annotation, which indicates the location, length, and width of a lymph node, is commonly available in hospital data archives, we advocate to use RECIST annotations as the supervision, and thus formulate this segmentation task into a weakly-supervised learning problem. In this paper, we propose a deep reinforcement learning-based lymph node segmentation (DRL-LNS) model. Based on RECIST annotations, we segment RECIST-slices in an unsupervised way to produce pseudo ground truths, which are then used to train U-Net as a segmentation network. Next, we train a DRL model, in which the segmentation network interacts with the policy network to optimize the lymph node bounding boxes and segmentation results simultaneously. The proposed DRL-LNS model was evaluated against three widely used image segmentation networks on a public thoracoabdominal Computed Tomography (CT) dataset that contains 984 3D lymph nodes, and achieves the mean Dice similarity coefficient (DSC) of 77.17% and the mean Intersection over Union (IoU) of 64.78% in the four-fold cross-validation. Our results suggest that the DRL-based bounding box prediction strategy outperforms the label propagation strategy and the proposed DRL-LNS model is able to achieve the state-of-the-art performance on this weakly-supervised lymph node segmentation task. Zhe Li 0006, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | SESV: Accurate Medical Image Segmentation by Predicting and Correcting ErrorsabstractMedical image segmentation is an essential task in computer-aided diagnosis. Despite their prevalence and success, deep convolutional neural networks (DCNNs) still need to be improved to produce accurate and robust enough segmentation results for clinical use. In this paper, we propose a novel and generic framework called Segmentation-Emendation-reSegmentation-Verification (SESV) to improve the accuracy of existing DCNNs in medical image segmentation, instead of designing a more accurate segmentation model. Our idea is to predict the segmentation errors produced by an existing model and then correct them. Since predicting segmentation errors is challenging, we design two ways to tolerate the mistakes in the error prediction. First, rather than using a predicted segmentation error map to correct the segmentation mask directly, we only treat the error map as the prior that indicates the locations where segmentation errors are prone to occur, and then concatenate the error map with the image and segmentation mask as the input of a re-segmentation network. Second, we introduce a verification network to determine whether to accept or reject the refined mask produced by the re-segmentation network on a region-by-region basis. The experimental results on the CRAG, ISIC, and IDRiD datasets suggest that using our SESV framework can improve the accuracy of DeepLabv3+ substantially and achieve advanced performance in the segmentation of gland cells, skin lesions, and retinal microaneurysms. Consistent conclusions can also be drawn when using PSPNet, U-Net, and FPN as the segmentation network, respectively. Therefore, our SESV framework is capable of improving the accuracy of different DCNNs on different medical image segmentation tasks. Yutong Xie 0001, Hao Lu 0003, Chunhua Shen, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Viral Pneumonia Screening on Chest X-Rays Using Confidence-Aware Anomaly DetectionabstractClusters of viral pneumonia occurrences over a short period may be a harbinger of an outbreak or pandemic. Rapid and accurate detection of viral pneumonia using chest X-rays can be of significant value for large-scale screening and epidemic prevention, particularly when other more sophisticated imaging modalities are not readily accessible. However, the emergence of novel mutated viruses causes a substantial dataset shift, which can greatly limit the performance of classification-based approaches. In this paper, we formulate the task of differentiating viral pneumonia from non-viral pneumonia and healthy controls into a one-class classification-based anomaly detection problem. We therefore propose the confidence-aware anomaly detection (CAAD) model, which consists of a shared feature extractor, an anomaly detection module, and a confidence prediction module. If the anomaly score produced by the anomaly detection module is large enough, or the confidence score estimated by the confidence prediction module is small enough, the input will be accepted as an anomaly case (i.e., viral pneumonia). The major advantage of our approach over binary classification is that we avoid modeling individual viral pneumonia classes explicitly and treat all known viral pneumonia cases as anomalies to improve the one-class model. The proposed model outperforms binary classification models on the clinical X-VIRAL dataset that contains 5,977 viral pneumonia (no COVID-19) cases, 37,393 non-viral pneumonia or healthy cases. Moreover, when directly testing on the X-COVID dataset that contains 106 COVID-19 cases and 107 normal controls without any fine-tuning, our model achieves an AUC of 83.61% and sensitivity of 71.70%, which is comparable to the performance of radiologists reported in the literature. Yutong Xie 0001, Guansong Pang, Zhibin Liao, Johan Verjans, Wenxing Li, Zongji Sun, Chunhua Shen, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 11 |
| 2021 | Inter-Slice Context Residual Learning for 3D Medical Image SegmentationabstractAutomated and accurate 3D medical image segmentation plays an essential role in assisting medical professionals to evaluate disease progresses and make fast therapeutic schedules. Although deep convolutional neural networks (DCNNs) have widely applied to this task, the accuracy of these models still need to be further improved mainly due to their limited ability to 3D context perception. In this paper, we propose the 3D context residual network (ConResNet) for the accurate segmentation of 3D medical images. This model consists of an encoder, a segmentation decoder, and a context residual decoder. We design the context residual module and use it to bridge both decoders at each scale. Each context residual module contains both context residual mapping and context attention mapping, the formal aims to explicitly learn the inter-slice context information and the latter uses such context as a kind of attention to boost the segmentation accuracy. We evaluated this model on the MICCAI 2018 Brain Tumor Segmentation (BraTS) dataset and NIH Pancreas Segmentation (Pancreas-CT) dataset. Our results not only demonstrate the effectiveness of the proposed 3D context residual learning scheme but also indicate that the proposed ConResNet is more accurate than six top-ranking methods in brain tumor segmentation and seven top-ranking methods in pancreas segmentation. Yutong Xie 0001, Yan Wang 0033, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Learning High-Resolution and Efficient Non-local Features for Brain Glioma Segmentation in MR Images
Haozhe Jia, Yong Xia 0001, Tom Weidong Cai, Heng Huang 0001 |
MICCAI (4) | 2 |
| 2020 | Pairwise Relation Learning for Semi-supervised Gland Segmentation
Yutong Xie 0001, Zhibin Liao, Johan Verjans, Chunhua Shen, Yong Xia 0001 |
MICCAI (5) | 6 |
| 2020 | Auto Learning AttentionabstractAttention modules have been demonstrated effective in strengthening the representation ability of a neural network via reweighting spatial or channel features or stacking both operations sequentially. However, designing the structures of different attention operations requires a bulk of computation and extensive expertise. In this paper, we devise an Auto Learning Attention (AutoLA) method, which is the first attempt on automatic attention design. Specifically, we define a novel attention module named high order group attention (HOGA) as a directed acyclic graph (DAG) where each group represents a node, and each edge represents an operation of heterogeneous attentions. A typical HOGA architecture can be searched automatically via the differential AutoLA method within 1 GPU day using the ResNet-20 backbone on CIFAR10. Further, the searched attention module can generalize to various backbones as a plug-and-play component and outperforms popular manually designed channel and spatial attentions for many vision tasks, including image classification on CIFAR100 and ImageNet, object detection and human keypoint detection on COCO dataset. The code will be released. Benteng Ma, Jing Zhang 0037, Yong Xia 0001, Dacheng Tao |
NeurIPS | 3 |
| 2020 | Autonomous deep learning: A genetic DCNN designer for image classification
Benteng Ma, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2020 | NFN+: A novel network followed network for retinal vessel segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
Neural Networks | 2 |
| 2020 | Thorax-Net: An Attention Regularized Deep Neural Network for Classification of Thoracic Diseases on Chest RadiographyabstractDeep learning techniques have been increasingly used to provide more accurate and more accessible diagnosis of thorax diseases on chest radiographs. However, due to the lack of dense annotation of large-scale chest radiograph data, this computer-aided diagnosis task is intrinsically a weakly supervised learning problem and remains challenging. In this paper, we propose a novel deep convolutional neural network called Thorax-Net to diagnose 14 thorax diseases using chest radiography. Thorax-Net consists of a classification branch and an attention branch. The classification branch serves as a uniform feature extraction-classification network to free users from the troublesome hand-crafted feature extraction and classifier construction. The attention branch exploits the correlation between class labels and the locations of pathological abnormalities via analyzing the feature maps learned by the classification branch. Feeding a chest radiograph to the trained Thorax-Net, a diagnosis is obtained by averaging and binarizing the outputs of two branches. The proposed Thorax-Net model has been evaluated against three state-of-the-art deep learning models using the patientwise official split of the ChestX-ray14 dataset and against other five deep learning models using the imagewise random data split. Our results show that Thorax-Net achieves an average per-class area under the receiver operating characteristic curve (AUC) of 0.7876 and 0.896 in both experiments, respectively, which are higher than the AUC values obtained by other deep models when they were all trained with no external data. Hongyu Wang 0011, Haozhe Jia, Le Lu 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | 3D APA-Net: 3D Adversarial Pyramid Anisotropic Convolutional Network for Prostate Segmentation in MR ImagesabstractAccurate and reliable segmentation of the prostate gland using magnetic resonance (MR) imaging has critical importance for the diagnosis and treatment of prostate diseases, especially prostate cancer. Although many automated segmentation approaches, including those based on deep learning have been proposed, the segmentation performance still has room for improvement due to the large variability in image appearance, imaging interference, and anisotropic spatial resolution. In this paper, we propose the 3D adversarial pyramid anisotropic convolutional deep neural network (3D APA-Net) for prostate segmentation in MR images. This model is composed of a generator (i.e., 3D PA-Net) that performs image segmentation and a discriminator (i.e., a six-layer convolutional neural network) that differentiates between a segmentation result and its corresponding ground truth. The 3D PA-Net has an encoder-decoder architecture, which consists of a 3D ResNet encoder, an anisotropic convolutional decoder, and multi-level pyramid convolutional skip connections. The anisotropic convolutional blocks can exploit the 3D context information of the MR images with anisotropic resolution, the pyramid convolutional blocks address both voxel classification and gland localization issues, and the adversarial training regularizes 3D PA-Net and thus enables it to generate spatially consistent and continuous segmentation results. We evaluated the proposed 3D APA-Net against several state-of-the-art deep learning-based segmentation approaches on two public databases and the hybrid of the two. Our results suggest that the proposed model outperforms the compared approaches on three databases and could be used in a routine clinical workflow. Haozhe Jia, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Heng Huang 0001, Yanning Zhang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Spatially-Constrained Fisher Representation for Brain Disease Identification With Incomplete Multi-Modal NeuroimagesabstractMulti-modal neuroimages, such as magnetic resonance imaging (MRI) and positron emission tomography (PET), can provide complementary structural and functional information of the brain, thus facilitating automated brain disease identification. Incomplete data problem is unavoidable in multi-modal neuroimage studies due to patient dropouts and/or poor data quality. Conventional methods usually discard data-missing subjects, thus significantly reducing the number of training samples. Even though several deep learning methods have been proposed, they usually rely on pre-defined regions-of-interest in neuroimages, requiring disease-specific expert knowledge. To this end, we propose a spatially-constrained Fisher representation framework for brain disease diagnosis with incomplete multi-modal neuroimages. We first impute missing PET images based on their corresponding MRI scans using a hybrid generative adversarial network. With the complete (after imputation) MRI and PET data, we then develop a spatially-constrained Fisher representation network to extract statistical descriptors of neuroimages for disease diagnosis, assuming that these descriptors follow a Gaussian mixture model with a strong spatial constraint (i.e., images from different subjects have similar anatomical structures). Experimental results on three databases suggest that our method can synthesize reasonable neuroimages and achieve promising results in brain disease identification, compared with several state-of-the-art methods. Yongsheng Pan, Mingxia Liu 0001, Chunfeng Lian, Yong Xia 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2020 | A Mutual Bootstrapping Model for Automated Skin Lesion Segmentation and ClassificationabstractAutomated skin lesion segmentation and classification are two most essential and related tasks in the computer-aided diagnosis of skin cancer. Despite their prevalence, deep learning models are usually designed for only one task, ignoring the potential benefits in jointly performing both tasks. In this paper, we propose the mutual bootstrapping deep convolutional neural networks (MB-DCNN) model for simultaneous skin lesion segmentation and classification. This model consists of a coarse segmentation network (coarse-SN), a mask-guided classification network (mask-CN), and an enhanced segmentation network (enhanced-SN). On one hand, the coarse-SN generates coarse lesion masks that provide a prior bootstrapping for mask-CN to help it locate and classify skin lesions accurately. On the other hand, the lesion localization maps produced by mask-CN are then fed into enhanced-SN, aiming to transfer the localization information learned by mask-CN to enhanced-SN for accurate lesion segmentation. In this way, both segmentation and classification networks mutually transfer knowledge between each other and facilitate each other in a bootstrapping way. Meanwhile, we also design a novel rank loss and jointly use it with the Dice loss in segmentation networks to address the issues caused by class imbalance and hard-easy pixel imbalance. We evaluate the proposed MB-DCNN model on the ISIC-2017 and PH2 datasets, and achieve a Jaccard index of 80.4% and 89.4% in skin lesion segmentation and an average AUC of 93.8% and 97.7% in skin lesion classification, which are superior to the performance of representative state-of-the-art skin lesion segmentation and classification methods. Our results suggest that it is possible to boost the performance of skin lesion segmentation and classification simultaneously via training a unified model to perform both tasks in a mutual bootstrapping way. Yutong Xie 0001, Yong Xia 0001, Chunhua Shen |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Light-Weight Hybrid Convolutional Network for Liver Tumor SegmentationabstractAutomated segmentation of liver tumors in contrast-enhanced abdominal computed tomography (CT) scans is essential in assisting medical professionals to evaluate tumor development and make fast therapeutic schedule. Although deep convolutional neural networks (DCNNs) have contributed many breakthroughs in image segmentation, this task remains challenging, since 2D DCNNs are incapable of exploring the inter-slice information and 3D DCNNs are too complex to be trained with the available small dataset. In this paper, we propose the light-weight hybrid convolutional network (LW-HCN) to segment the liver and its tumors in CT volumes. Instead of combining a 2D and a 3D networks for coarse-to-fine segmentation, LW-HCN has a encoder-decoder structure, in which 2D convolutions used at the bottom of the encoder decreases the complexity and 3D convolutions used in other layers explore both spatial and temporal information. To further reduce the complexity, we design the depthwise and spatiotemporal separate (DSTS) factorization for 3D convolutions, which not only reduces parameters dramatically but also improves the performance. We evaluated the proposed LW-HCN model against several recent methods on the LiTS and 3D-IRCADb datasets and achieved, respectively, the Dice per case of 73.0% and 94.1% for tumor segmentation, setting a new state of the art. Yutong Xie 0001, Hao Chen 0041, Yong Xia 0001, Chunhua Shen |
IJCAI | 5 |
| 2019 | HD-Net: Hybrid Discriminative Network for Prostate Segmentation in MR Images
Haozhe Jia, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai, Yong Xia 0001 |
MICCAI (2) | 5 |
| 2019 | Disease-Image Specific Generative Adversarial Network for Brain Disease Diagnosis with Incomplete Multi-modal Neuroimages
Yongsheng Pan, Mingxia Liu 0001, Chunfeng Lian, Yong Xia 0001, Dinggang Shen |
MICCAI (3) | 4 |
| 2019 | Vessel-Net: Retinal Vessel Segmentation Under Multi-path Supervision
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Dongnan Liu, Chaoyi Zhang, Tom Weidong Cai |
MICCAI (1) | 2 |
| 2019 | Deep Segmentation-Emendation Model for Gland Instance Segmentation
Yutong Xie 0001, Hao Lu 0003, Chunhua Shen, Yong Xia 0001 |
MICCAI (1) | 5 |
| 2019 | DAEimp: Denoising Autoencoder-Based Imputation of Sleep Heart Health Study for Identification of Cardiovascular Diseases
Xiaoyun Dong, Yong Xia 0001 |
PRCV (1) | 4 |
| 2019 | EMS-Net: Ensemble of Multiscale Convolutional Neural Networks for Classification of Breast Cancer Histology Images
Zhanbo Yang, Lingyan Ran, Shizhou Zhang, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 4 |
| 2019 | Affective image classification by jointly using interpretable art features and semantic annotations
Yong Xia 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Semi-supervised adversarial model for benign-malignant lung nodule classification on chest CT
Yutong Xie 0001, Yong Xia 0001 |
Medical Image Anal. | 3 |
| 2019 | Medical image classification using synergic deep learning
Yutong Xie 0001, Qi Wu 0001, Yong Xia 0001 |
Medical Image Anal. | 4 |
| 2019 | Fast-Convergent Fully Connected Deep Learning Model Using Constrained Nodes Input
Chen Ding 0002, Ying Li 0017, Lei Zhang 0054, Lu Yang 0016, Wei Wei 0008, Yong Xia 0001, Yanning Zhang 0001 |
Neural Process. Lett. | 7 |
| 2019 | M3Net: A multi-model, multi-size, and multi-view deep neural network for brain magnetic resonance image segmentation
Yong Xia 0001, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2019 | Normalized Non-Negative Sparse Encoder for Fast Image RepresentationabstractImage representation based on sparse coding generalizes the bag of words model. Although it reduces the reconstruction error for local features to achieve the state-of-the-art image classification performance, the large computational cost hinders the application of sparse coding-based image features. In this paper, we propose approximating a sparse code using the output of a simple neural network. The resulting parameter learning model for the neural network automatically incorporates non-negative and shift-invariant constraints, leading to an efficient normalized non-negative sparse coding (N3SC) sparse encoder. Without the use of the traditional iterative process to solve the sparse coding objective, the sparse encoder directly “converts” each local feature into a sparse code. We also introduce a method for training the encoder based on the auto-encoder method. In addition, we formally propose the corresponding sparse coding scheme called N3SC, which enforces both the non-negative constraint and the shift-invariant constraint in addition to the traditional sparse coding criteria. As demonstrated by several experiments, the obtained N3SC encoder requires only 3%-10% of the processing time for image feature extraction compared with the standard sparse coding scheme. At the same time, the features extracted using the exact solutions of the N3SC coding scheme and the N3SC encoder offer superior image classification accuracy compared to the accuracy of many existing sparse coding-based representations. Shizhou Zhang, Jinjun Wang, Weiwei Shi 0003, Yihong Gong, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Foreground Fisher Vector: Encoding Class-Relevant Foreground to Improve Image ClassificationabstractImage classification is an essential and challenging task in computer vision. Despite its prevalence, the combination of the deep convolutional neural network (DCNN) and the Fisher vector (FV) encoding method has limited performance since the class-irrelevant background used in the traditional FV encoding may result in less discriminative image features. In this paper, we propose the foreground FV (fgFV) encoding algorithm and its fast approximation for image classification. We try to separate implicitly the class-relevant foreground from the class-irrelevant background during the encoding process via tuning the weights of the partial gradients corresponding to each Gaussian component under the supervision of image labels and, then, use only those local descriptors extracted from the class-relevant foreground to estimate FVs. We have evaluated our fgFV against the widely used FV and improved FV (iFV) under the combined DCNN-FV framework and also compared them to several state-of-the-art image classification approaches on ten benchmark image datasets for the recognition of fine-grained natural species and artificial manufactures, categorization of course objects, and classification of scenes. Our results indicate that the proposed fgFV encoding algorithm can construct more discriminative image presentations from local descriptors than FV and iFV, and the combined DCNN-fgFV algorithm can improve the performance of image classification. Yongsheng Pan, Yong Xia 0001, Dinggang Shen |
IEEE Trans. Image Process. | 2 |
| 2019 | Knowledge-based Collaborative Deep Learning for Benign-Malignant Lung Nodule Classification on Chest CTabstractThe accurate identification of malignant lung nodules on chest CT is critical for the early detection of lung cancer, which also offers patients the best chance of cure. Deep learning methods have recently been successfully introduced to computer vision problems, although substantial challenges remain in the detection of malignant nodules due to the lack of large training data sets. In this paper, we propose a multi-view knowledge-based collaborative (MV-KBC) deep model to separate malignant from benign nodules using limited chest CT data. Our model learns 3-D lung nodule characteristics by decomposing a 3-D nodule into nine fixed views. For each view, we construct a knowledge-based collaborative (KBC) submodel, where three types of image patches are designed to fine-tune three pre-trained ResNet-50 networks that characterize the nodules' overall appearance, voxel, and shape heterogeneity, respectively. We jointly use the nine KBC submodels to classify lung nodules with an adaptive weighting scheme learned during the error back propagation, which enables the MV-KBC model to be trained in an end-to-end manner. The penalty loss function is used for better reduction of the false negative rate with a minimal effect on the overall performance of the MV-KBC model. We tested our method on the benchmark LIDC-IDRI data set and compared it to the five state-of-the-art classification approaches. Our results show that the MV-KBC model achieved an accuracy of 91.60% for lung nodule classification with an AUC of 95.70%. These results are markedly superior to the state-of-the-art approaches. Yutong Xie 0001, Yong Xia 0001, Yang Song 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Attention Residual Learning for Skin Lesion ClassificationabstractAutomated skin lesion classification in dermoscopy images is an essential way to improve the diagnostic performance and reduce melanoma deaths. Although deep convolutional neural networks (DCNNs) have made dramatic breakthroughs in many image classification tasks, accurate classification of skin lesions remains challenging due to the insufficiency of training data, inter-class similarity, intra-class variation, and the lack of the ability to focus on semantically meaningful lesion parts. To address these issues, we propose an attention residual learning convolutional neural network (ARL-CNN) model for skin lesion classification in dermoscopy images, which is composed of multiple ARL blocks, a global average pooling layer, and a classification layer. Each ARL block jointly uses the residual learning and a novel attention learning mechanisms to improve its ability for discriminative representation. Instead of using extra learnable layers, the proposed attention learning mechanism aims to exploit the intrinsic self-attention ability of DCNNs, i.e., using the feature maps learned by a high layer to generate the attention map for a low layer. We evaluated our ARL-CNN model on the ISIC-skin 2017 dataset. Our results indicate that the proposed ARL-CNN model can adaptively focus on the discriminative parts of skin lesions, and thus achieve the state-of-the-art performance in skin lesion classification. Yutong Xie 0001, Yong Xia 0001, Chunhua Shen |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Densely Connected Large Kernel Convolutional Network for Semantic Membrane Segmentation in Microscopy ImagesabstractStructural analysis of neurons can provide valuable insights of brain function. Semantic segmentation of neurons thus becomes an important technique in bioinformatics. Deep learning approaches have shown promising performance in various semantic segmentation problems. However, segmentation of neurons in Electron Microscopy (EM) images has some differences compared with typical segmentation tasks due to the image noise and the disturbance of the intracellular structures. In our work, we propose a network with a ResNet encoder and densely connected decoder with large kernels, and then refinement with simple morphological post-possessing. Two main advantages of our method are: 1) the network can prevent the loss of high-resolution information and enlarge the reception field; 2) the post-processing method is simple and can be directly applied to the probability map from the network to enhance the unconfident area. Evaluated on the ISBI2012 EM membrane segmentation challenge, the proposed method achieves competitive performance. Dongnan Liu, Donghao Zhang 0004, Siqi Liu 0001, Yang Song 0001, Haozhe Jia, David Dagan Feng, Yong Xia 0001, Tom Weidong Cai |
ICIP | 7 |
| 2018 | Synthesizing Missing PET from MRI with Cycle-consistent Generative Adversarial Networks for Alzheimer's Disease Diagnosis
Yongsheng Pan, Mingxia Liu 0001, Chunfeng Lian, Tao Zhou 0002, Yong Xia 0001, Dinggang Shen |
MICCAI (3) | 5 |
| 2018 | Multiscale Network Followed Network Model for Retinal Vessel Segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
MICCAI (2) | 2 |
| 2018 | Panoptic Segmentation with an End-to-End Cell R-CNN for Pathology Image Analysis
Donghao Zhang 0004, Yang Song 0001, Dongnan Liu, Haozhe Jia, Siqi Liu 0001, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (2) | 6 |
| 2018 | Skin Lesion Classification in Dermoscopy Images Using Synergic Deep Learning
Yutong Xie 0001, Qi Wu 0001, Yong Xia 0001 |
MICCAI (2) | 4 |
| 2018 | Deep Classification and Segmentation Model for Vessel Extraction in Retinal Images
Yicheng Wu 0001, Yong Xia 0001, Yanning Zhang 0001 |
PRCV (2) | 2 |
| 2018 | Atlas registration and ensemble deep convolutional neural network-based prostate segmentation using magnetic resonance imaging
Haozhe Jia, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 2 |
| 2018 | Pedestrian search in surveillance videos by learning discriminative deep features
Shizhou Zhang, De Cheng, Yihong Gong, Dahu Shi, Xi Qiu, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2018 | NODULe: Combining constrained multi-scale LoG filters with densely dilated 3D deep convolutional neural network for pulmonary nodule detection
Yong Xia 0001, Haoyue Zeng, Yanning Zhang 0001 |
Neurocomputing | 2 |
| 2018 | Affective image classification via semi-supervised learning from web images
Yong Xia 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Locality constrained encoding of frequency and spatial information for image classification
Yongsheng Pan, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai |
Multim. Tools Appl. | 2 |
| 2018 | VBI-MRF model for image segmentation
Yong Xia 0001, Zhe Li 0006 |
Multim. Tools Appl. | 1 |
| 2018 | Validation of right coronary artery lumen area from cardiac computed tomography against intravascular ultrasound
Hengfei Cui, Yong Xia 0001, Yanning Zhang 0001, Liang Zhong 0001 |
Mach. Vis. Appl. | 2 |
| 2018 | Classification of Medical Images in the Biomedical Literature by Jointly Using Deep and Handcrafted Visual FeaturesabstractThe classification of medical images and illustrations from the biomedical literature is important for automated literature review, retrieval, and mining. Although deep learning is effective for large-scale image classification, it may not be the optimal choice for this task as there is only a small training dataset. We propose a combined deep and handcrafted visual feature (CDHVF) based algorithm that uses features learned by three fine-tuned and pretrained deep convolutional neural networks (DCNNs) and two handcrafted descriptors in a joint approach. We evaluated the CDHVF algorithm on the ImageCLEF 2016 Subfigure Classification dataset and it achieved an accuracy of 85.47%, which is higher than the best performance of other purely visual approaches listed in the challenge leaderboard. Our results indicate that handcrafted features complement the image representation learned by DCNNs on small training datasets and improve accuracy in certain medical image classification problems. Yong Xia 0001, Yutong Xie 0001, Michael J. Fulham, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Space target recognition based on deep learningabstractAutomated recognition of spacecraft and space debris using imaging plays an important role in securing space safety and space exploration. Although deep learning is now the most successful solution for image-based object classification, it requires a myriad number of training data, which are not available for most real applications. In this paper, we investigate different single and hybrid data augmentation methods for both training and testing images, and thus propose a data augmentation-based deep learning approach to space target recognition. Experimental results on 400 synthetic space target images rendered by the Systems Tool Kit (STK) demonstrate that our proposed algorithm achieves higher accuracy than several traditional methods. Haoyue Zeng, Yong Xia 0001 |
FUSION | 2 |
| 2017 | Dictionary learning-based image compressionabstractDictionary learning based image compression has attracted a lot of research efforts due to the inherent sparsity of image contents. Most algorithms in the literature, however, suffer from two drawbacks. First, the atoms selected for image patch reconstruction scatter over the entire dictionary, which leads to a high coding cost. Second, the sparse representation of image patches is performed independently from the quantization of sparse coefficients, which may result in a sub-optimal solution. In this paper, we propose the entropy based orthogonal matching pursuit (EOMP) algorithm and quantization KSVD (QKSVD) algorithm for dictionary learning-based image compression. An entropy regularization term is utilized in EOMP to restrict atom selection, and hence reduces the coding cost, and an adaptive quantization method is incorporated into the dictionary learning procedure in QKSVD to minimize the reconstruction error and quantization error simultaneously. Experimental results on 10 standard benchmark images demonstrate that our proposed approach achieves better performance than several state-of-the-art ones at low bit rate, such as KSVD based compression approach, JPEG, and JPEG-2000. Yong Xia 0001, Zhiyong Wang 0001 |
ICIP | 2 |
| 2017 | Transferable Multi-model Ensemble for Benign-Malignant Lung Nodule Classification on Chest CT
Yutong Xie 0001, Yong Xia 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai |
MICCAI (3) | 2 |
| 2017 | A robust modified Gaussian mixture model with rough set for image segmentation
Zexuan Ji, Yong Xia 0001, Yuhui Zheng |
Neurocomputing | 3 |
| 2017 | Brain voxel classification in magnetic resonance images using niche differential evolution based Bayesian inference of variational mixture of Gaussians
Zhe Li 0006, Yong Xia 0001, Zexuan Ji, Yanning Zhang 0001 |
Neurocomputing | 2 |
| 2016 | Brain MRI image segmentation based on learning local variational Gaussian mixture models
Yong Xia 0001, Zexuan Ji, Yanning Zhang 0001 |
Neurocomputing | 1 |
| 2016 | Foreground Detection With Simultaneous Dictionary Learning and Historical Pixel MaintenanceabstractForeground detection is fundamental in surveillance video analysis and meaningful toward object tracking and higher level tasks, such as anomaly detection and activity analysis. Nevertheless, existing methods are still limited in accurately detecting the foreground due to the complex scene settings. To robustly handle the diverse background variations and foreground challenges, this paper proposes a Background REpresentation approach With Dictionary Learning and Historical Pixel Maintenance (BREW-DLHPM). Specifically, a dictionary learning problem is formulated at the frame level to adaptively represent the background signals with the varied structure information captured, while a pixel-level maintenance is exploited to grasp the dynamic nature of historical information under the help of the learned background. The simultaneous utilization of dictionary learning and historical pixel maintenance facilitates the accurate description of the background and thus guides a wise foreground detection decision. The proposed BREW-DLHPM has been evaluated on the prestigious change detection challenge data set against 11 state-of-the-art foreground detection approaches and encouraging performances have been achieved by our method. Pei Dong, Shanshan Wang 0002, Yong Xia 0001, Dong Liang 0001, David Dagan Feng |
IEEE Trans. Image Process. | 3 |
| 2015 | Semi-supervised emotional classification of color images by learning from cloudabstractClassification of images based on the feelings generated by each image in its reviewers is becoming more and more popular. Due to the difficulty of gathering training data, this task is intrinsically a small-sample learning problem. Hence, the results produced by most existing solutions are less accurate. In this paper, we propose the semi-supervised hierarchical classification (SSHC) algorithm for emotional classification of color images. We extract three groups of features for each classification task and use those features in a two-level classification model that is based on the support vector machine (SVM) and Adaboost technique. To enlarge the training dataset, we employ each training image to retrieve similar images from the Internet cloud and jointly use the manually labeled small dataset and retrieved large but unlabeled dataset to train a classifier via semi-supervised learning. We have evaluated the proposed algorithm against the fuzzy similarity-based emotional classification (FSBEC) algorithm and another supervised hierarchical classification algorithm that does not learn from online images in three bi-class classification tasks, including “warm vs. cool”, “light vs. heavy” and “static vs. dynamic”. Our pilot results suggest that, by learning from the similar images archived in the Internet cloud, the proposed SSHC algorithm can produce more accurate emotional classification of color images. Yong Xia 0001, Yuwei Xia |
ACII | 2 |
| 2015 | Robust saliency detection via regularized random walks rankingabstractIn the field of saliency detection, many graph-based algorithms heavily depend on the accuracy of the pre-processed superpixel segmentation, which leads to significant sacrifice of detail information from the input image. In this paper, we propose a novel bottom-up saliency detection approach that takes advantage of both region-based features and image details. To provide more accurate saliency estimations, we first optimize the image boundary selection by the proposed erroneous boundary removal. By taking the image details and region-based estimations into account, we then propose the regularized random walks ranking to formulate pixel-wised saliency maps from the superpixel-based background and foreground saliency estimations. Experiment results on two public datasets indicate the significantly improved accuracy and robustness of the proposed algorithm in comparison with 12 state-of-the-art saliency detection approaches. ChangYang Li, Yuchen Yuan, Tom Weidong Cai, Yong Xia 0001, David Dagan Feng |
CVPR | 4 |
| 2015 | Active contours driven by local likelihood image fitting energy for image segmentation
Zexuan Ji, Yong Xia 0001, Quan-Sen Sun, Guo Cao, Qiang Chen 0004 |
Inf. Sci. | 2 |
| 2015 | An iteratively reweighting algorithm for dynamic video summarization
Pei Dong, Yong Xia 0001, Shanshan Wang 0002, Li Zhuo 0001, David Dagan Feng |
Multim. Tools Appl. | 2 |
| 2014 | Non-sparse infinite-kernel learning for automated identification of Alzheimer's disease using PET imagingabstractMulti-kernel learning machine (MKLM) has recently been introduced to the research of computer-aided dementia identification and pathology progress tracking. Despite its good performance especially in case of using heterogeneous data, such learning schema and its variants usually utilize a L-l norm constraint that promotes sparse solutions, which may cause loss of potentially important information. In this paper, we propose the non-sparse infinite-kernel learning machine (NS-IKLM) for automated identification of Alzheimer cases from normal controls. In our approach, a modified constraint is utilized to promotes non-sparse solutions and kernel parameters are automatically tuned during the learning process. The proposed algorithm has been evaluated on a set of FDG-PET images selected from the Alzheimer's disease neuroimaing initiative (ADNI) cohort. Our results demonstrate that the proposed non-sparse NS-IKLM is able to achieve satisfying dementia identification at a relatively low computational cost. Yong Xia 0001, Shen Lu, Wei Wei 0008, David Dagan Feng, Yanning Zhang 0001 |
ICARCV | 1 |
| 2014 | Demographic information prediction based on smartphone application usageabstractDemographic information is usually treated as private data (e.g., gender and age), but has been shown great values in personalized services, advertisement, behavior study and other aspects. In this paper, we propose a novel approach to make efficient demographic prediction based on smartphone application usage. Specifically, we firstly consider to characterize the data set by building a matrix to correlate users with types of categories from the log file of smartphone applications. By considering the category-unbalance problem, we predict users' demographic information and propose an optimization method to further smooth the obtained results with category neighbors and user neighbors. The evaluation is supplemented by the dataset from real world workload. The results show advantages of the proposed prediction approach compared with baseline prediction. In particular, the proposed approach can achieve 81.21% of Accuracy in gender prediction. While in dealing with a more challenging multi-class problem, the proposed approach can still achieve good performance (e.g., 73.84% of Accuracy in the prediction of age group and 66.42% of Accuracy in the prediction of phone level). Zhen Qin 0002, Yong Xia 0001, Hongrong Cheng, Yingjie Zhou 0001, Zhengguo Sheng, Victor C. M. Leung |
SMARTCOMP | 3 |
| 2014 | Interval-valued possibilistic fuzzy C-means clustering algorithm
Zexuan Ji, Yong Xia 0001, Quan-Sen Sun, Guo Cao |
Fuzzy Sets Syst. | 2 |
| 2014 | A clonal selection based approach to statistical brain voxel classification in magnetic resonance images
Tong Zhang 0017, Yong Xia 0001, David Dagan Feng |
Neurocomputing | 2 |
| 2014 | Adaptive scale fuzzy local Gaussian mixture model for brain MR image segmentation
Zexuan Ji, Yong Xia 0001, Quan-Sen Sun, Qiang Chen 0004, David Dagan Feng |
Neurocomputing | 2 |
| 2013 | Multi-pose 3D face recognition based on 2D sparse representation
Yanning Zhang 0001, Yong Xia 0001, Zenggang Lin, Yangyu Fan, David Dagan Feng |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Discriminative two-level feature selection for realistic human action recognition
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Yong Xia 0001, Wenxiong Kang, David Dagan Feng |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Dictionary learning based impulse noise removal via L1-L1 minimization
Shanshan Wang 0002, Qiegen Liu, Yong Xia 0001, Pei Dong, Jianhua Luo, Qiu Huang, David Dagan Feng |
Signal Process. | 3 |
| 2013 | Fenchel Duality Based Dictionary Learning for Restoration of Noisy ImagesabstractDictionary learning based sparse modeling has been increasingly recognized as providing high performance in the restoration of noisy images. Although a number of dictionary learning algorithms have been developed, most of them attack this learning problem in its primal form, with little effort being devoted to exploring the advantage of solving this problem in a dual space. In this paper, a novel Fenchel duality based dictionary learning (FD-DL) algorithm has been proposed for the restoration of noise-corrupted images. With the restricted attention to the additive white Gaussian noise, the sparse image representation is formulated as an 2-1 minimization problem, whose dual formulation is constructed using a generalization of Fenchel’s duality theorem and solved under the augmented Lagrangian framework. The proposed algorithm has been compared with four state-of-the-art algorithms, including the local pixel grouping-principal component analysis, method of optimal directions, K-singular value decomposition, and beta process factor analysis, on grayscale natural images. Our results demonstrate that the FD-DL algorithm can effectively improve the image quality and its noisy image restoration ability is comparable or even superior to the abilities of the other four widely-used algorithms. Shanshan Wang 0002, Yong Xia 0001, Qiegen Liu, Pei Dong, David Dagan Feng, Jianhua Luo |
IEEE Trans. Image Process. | 2 |
| 2012 | GARDEN: Generic Addressing and Routing for Data Center NetworksabstractData centers often hold tens to hundreds of thousands of servers in order to offer cloud computing services at scale. Ethernet switching and IP routing have their own advantages and limitations in building data center networks. Recent research, such as PortLand and BCube, has proposed scalable data center network designs. A common feature of these designs is that their addressing and routing are customized to specific topologies. In this paper, we propose a generic addressing, routing and forwarding protocol for data center networks, which works on arbitrarily "layered'' network topologies. We first form the network as a multi-rooted tree. Each network node (i.e., hosts and switches) is then assigned one or more locators, and each locator encodes a downward path from the roots to this node. Data center networks often have rich path diversity, so tracking all locators of a destination node will cause switches to have very large forwarding tables. We further use a new forwarding model to reduce the forwarding states. In addition, the multiple-locator mechanism brings built-in support for multi-path routing, load balancing and fault tolerance. Evaluations based on simulations and prototype experiments demonstrate that our proposal achieves our design goals. Yong Xia 0001, Kai Chen 0005, Yanlin Luo |
IEEE CLOUD | 3 |
| 2012 | Real-Time Storyboard Generation for H.264/AVC Compressed VideosabstractVideo summarization enables convenient and efficient management of large volume of visual data. However, most existing summarization approaches are based on either the pixel domain information or conventional video compression standards. As the most recent and popular international video coding standard, H.264/AVC adopts a number of advanced techniques and brings not only opportunities but also challenges to video summarization. In this paper, we propose a real-time image storyboard generation algorithm for H.264/AVC compressed videos by using both compressed domain and pixel domain information jointly and adaptively. This algorithm extracts compressed domain information for visual content representation, video structuring and candidate representative frame selection. By fusing both compressed domain and pixel domain information, the redundancy in the candidate representative frames is further reduced. Our experimental results show that the proposed algorithm can efficiently produce image storyboards conforming to human interpretation of the essential content in generic videos. Pei Dong, Yong Xia 0001, David Dagan Feng |
ICME | 2 |
| 2012 | Gabor feature based nonlocal means filter for textured image denoising
Shanshan Wang 0002, Yong Xia 0001, Qiegen Liu, Jianhua Luo, Yue Min Zhu, David Dagan Feng |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | 2D representation of facial surfaces for multi-pose 3D face recognition
Yanning Zhang 0001, Yong Xia 0001, Zenggang Lin, David Dagan Feng |
Pattern Recognit. Lett. | 3 |
| 2012 | Fuzzy Local Gaussian Mixture Model for Brain MR Image SegmentationabstractAccurate brain tissue segmentation from magnetic resonance (MR) images is an essential step in quantitative brain image analysis. However, due to the existence of noise and intensity inhomogeneity in brain MR images, many segmentation algorithms suffer from limited accuracy. In this paper, we assume that the local image data within each voxel's neighborhood satisfy the Gaussian mixture model (GMM), and thus propose the fuzzy local GMM (FLGMM) algorithm for automated brain MR image segmentation. This algorithm estimates the segmentation result that maximizes the posterior probability by minimizing an objective energy function, in which a truncated Gaussian kernel function is used to impose the spatial constraint and fuzzy memberships are employed to balance the contribution of each GMM. We compared our algorithm to state-of-the-art segmentation approaches in both synthetic and clinical data. Our results show that the proposed algorithm can largely overcome the difficulties raised by noise, low contrast, and bias field, and substantially improve the accuracy of brain MR image segmentation. Zexuan Ji, Yong Xia 0001, Quan-Sen Sun, Qiang Chen 0004, De-Shen Xia, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2011 | Real-time moving object segmentation and tracking for H.264/AVC surveillance videosabstractWith increased use of H.264/AVC in various applications including video surveillance systems, feature extraction and knowledge representation in compressed domain are becoming attractive. A real-time H.264/AVC compressed domain moving object segmentation and tracking algorithm for surveillance videos is proposed in this paper. This algorithm consists of moving object detection, bounding box matching, spatiotemporal merge and split reasoning and trajectory smoothing, with major innovation in incorporating the information provided by the prediction modes into the framework of motion detection and trajectory construction. The experimental results on both indoor and outdoor surveillance videos demonstrate that the adaptive use of the information from motion vectors, DCT coefficients and prediction modes can substantially improve the performance of moving object segmentation and tracking. Pei Dong, Yong Xia 0001, Li Zhuo 0001, David Dagan Feng |
ICIP | 2 |
| 2011 | Scalable data center multicast using multi-class Bloom FilterabstractMulticast benefits data center group communications in saving network bandwidth and increasing application throughput. However, it is challenging to scale Multicast to support tens of thousands of concurrent group communications due to limited forwarding table memory space in the switches, particularly the low-end ones commonly used in modern data centers. Bloom Filter is an efficient tool to compress the Multicast forwarding table, but significant traffic leakage may occur when group membership testing is false positive. To reduce the Multicast traffic leakage, in this paper we bring forward a novel multi-class Bloom Filter (MBF), which extends the standard Bloom Filter by embracing element uncertainty. Specifically, MBF sets the number of hash functions in a per-element level, based on the probability for each Multicast group to be inserted into the Bloom Filter. We design a simple yet effective algorithm to calculate the number of hash functions for each Multicast group. We have prototyped a software based MBF forwarding engine on the Linux platform. Simulation and prototype evaluation results demonstrate that MBF can significantly reduce Multicast traffic leakage compared to the standard Bloom Filter, while causing little system overhead. Henggang Cui, Yong Xia 0001, Xin Wang 0001 |
ICNP | 4 |
| 2011 | QoS-Enabled Dynamic Resource Management in Multi-Cell OFDMA-Based SystemsabstractIn this paper, we propose a novel cross-layer framework, which includes a two-level resource management to efficiently allocate subchannels among cell BSs (at the super-frame level) and do resource scheduling (at frame level) among intra-cell users based on channel condition and QoS requirement, and a system-level call admission control. The coordination among a plurality of BSs in the OFDMA-based network is realized at a controller. This framework improves the system throughput while maintaining the QoS requirement of mobile users. Our evaluations show that the proposed framework can improve the global resource utilization, reduce the packet loss rate, and improve the cell-edge user performance for different deployment scenarios. In addition, the dynamic resource management algorithm decouples the functionality of the controller and the BS so it is practical to implement in a real system. Yong Xia 0001 |
VTC Spring | 3 |
| 2011 | Hybrid Genetic and Variational Expectation-Maximization Algorithm for Gaussian-Mixture-Model-Based Brain MR Image SegmentationabstractThe expectation-maximization (EM) algorithm has been widely applied to the estimation of gaussian mixture model (GMM) in brain MR image segmentation. However, the EM algorithm is deterministic and intrinsically prone to overfitting the training data and being trapped in local optima. In this paper, we propose a hybrid genetic and variational EM (GA-VEM) algorithm for brain MR image segmentation. In this approach, the VEM algorithm is performed to estimate the GMM, and the GA is employed to initialize the hyperparameters of the conjugate prior distributions of GMM parameters involved in the VEM algorithm. Since GA has the potential to achieve global optimization and VEM can steadily avoid overfitting, the hybrid GA-VEM algorithm is capable of overcoming the drawbacks of traditional EM-based methods. We compared our approach to the EM-based, VEM-based, and GA-EM based segmentation algorithms, and the segmentation routines used in the statistical parametric mapping package and FMRIB Software Library in 20 low-resolution and 17 high-resolution brain MR studies. Our results show that the proposed approach can improve substantially the performance of brain MR image segmentation. Guangjian Tian, Yong Xia 0001, Yanning Zhang 0001, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2010 | 3D face representation and recognition by Intrinsic Shape Description MapsabstractWe present a novel method for 3D face recognition, in which the 3D facial surface is first mapped into a 2D domain with specified resolution through a global optimization by constrained conformal geometric maps. The Intrinsic Shape Description Map (ISDM) is then constructed through a modeling technique capable to express geometric and appearance information of the 3D face. Hence the 3D surface matching problem can be simplified to a 2D image matching problem, which greatly reduces the computational complexity. Finally, the Intrinsic Shape Description Feature (ISDF) of ISDM and the discrimination analysis can be calculated. Experimental results implemented on GavabDB demonstrate that our proposed method significantly outperforms the existing methods with respect to pose variation. Yanning Zhang 0001, Yong Xia 0001, Zenggang Lin, David Dagan Feng |
ICASSP | 3 |
| 2010 | Dual-modality 3D brain PET-CT image segmentation based on probabilistic brain atlas and classification fusionabstractThe increasing prevalence of dual medical imaging modalities, such as PET-CT scanners, poses both challenges and opportunities to image segmentation, as they provide distinct but complementary information. In this paper, we propose a novel segmentation algorithm for 3D brain PET-CT images, which classifies each voxel by fusing the voxel's memberships estimated from four points of view using the PET information, CT information, smoothness prior, and probabilistic brain atlas. All memberships having the same dynamic range greatly facilitates weighting the contribution of the four different information sources. The probabilistic brain atlas estimated for each PET-CT image from a set of training samples provides the anatomical information to the segmentation process. We compared the proposed algorithm to three single-classifier based methods, PET-based SPM algorithm, CT-based Otsu thresholding, and PET-CT based MAP-MRF algorithm. The experimental results in 11 clinical brain PET-CT studies demonstrate that the novel algorithm is capable of providing more accurate and reliable segmentation. Yong Xia 0001, Stefan Eberl, David Dagan Feng |
ICIP | 1 |
| 2010 | Multifractal signature estimation for textured image segmentation
Yong Xia 0001, David Dagan Feng, Rongchun Zhao, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 1 |
| 2009 | A General Image Segmentation Model and its ApplicationabstractThis paper proposed a general image segmentation model, namely the energy-minimization based image segmentation (EMBIS) model. This model converts image segmentation into a controlled optimization process minimizing the weighted sum of the feature energy and spatial energy, which interpret the homogeneity restriction and spatial constraints, respectively. The EMBIS model provides a unified understanding of various existing segmentation algorithms, and can also serve as a framework for systematic generation of new segmentation algorithms. We provided four examples to illustrate that many existing segmentation algorithms are indeed specialized cases of this model with different instances of both energy functions. We also presented a case study to demonstrate how to use this model to create new algorithms and resulted in the spatial-constrained OTSU (SC-OTSU) algorithm, where segmentation can be achieved by minimizing the feature energy of the OTSU algorithm and spatial energy of the algorithm based on a simple MRF (SMRF) model. Evaluation on both synthetic and real images proved that novel segmentation algorithms derived form the proposed EMBIS model can provide accurate and efficient image segmentation. Yong Xia 0001, David Dagan Feng |
ICIG | 1 |
| 2008 | Segmentation of dual modality brain PET/CT images using the MAP-MRF modelabstractDual modality PET/CT has now essentially replaced PET in clinical practice and provided an opportunity to improve image segmentation through the high resolution, lower noise CT data. Thus far most research efforts have concentrated on segmentation of PET-only data. In this work we propose a systematic solution for the automated segmentation of brain PET/CT images into gray, white matter and CSF regions with the MAP-MRF model. Our approach takes advantage of the full information available from the combined scan. A PET/CT image pair and its segmentation result are modelled as a random field triplet, and segmentation is eventually achieved by solving a maximum a posteriori (MAP) problem using the expectation-maximization (EM) algorithm with simulated annealing. We compared the novel algorithm to two widely used PET-only based segmentation methods in the SPM5 toolbox and the VBM toolbox for simulation and patient data. Our results suggest that using the proposed approach substantially improves the accuracy of the delineation of brain structures. Yong Xia 0001, Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
MMSP | 1 |
| 2007 | Image segmentation by clustering of spatial patterns
Yong Xia 0001, David Dagan Feng, Rongchun Zhao, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 1 |
| 2006 | Morphology-based multifractal estimation for texture segmentationabstractMultifractal analysis is becoming more and more popular in image segmentation community, in which the box-counting based multifractal dimension estimations are most commonly used. However, in spite of its computational efficiency, the regular partition scheme used by various box-counting methods intrinsically produces less accurate results. In this paper, a novel multifractal estimation algorithm based on mathematical morphology is proposed and a set of new multifractal descriptors, namely the local morphological multifractal exponents is defined to characterize the local scaling properties of textures. A series of cubic structure elements and an iterative dilation scheme are utilized so that the computational complexity of the morphological operations can be tremendously reduced. Both the proposed algorithm and the box-counting based methods have been applied to the segmentation of texture mosaics and real images. The comparison results demonstrate that the morphological multifractal estimation can differentiate texture images more effectively and provide more robust segmentations. Yong Xia 0001, David Dagan Feng, Rongchun Zhao |
IEEE Trans. Image Process. | 1 |
| 2006 | Adaptive Segmentation of Textured Images by Using the Coupled Markov Random Field ModelabstractAlthough simple and efficient, traditional feature-based texture segmentation methods usually suffer from the intrinsical less inaccuracy, which is mainly caused by the oversimplified assumption that each textured subimage used to estimate a feature is homogeneous. To solve this problem, an adaptive segmentation algorithm based on the coupled Markov random field (CMRF) model is proposed in this paper. The CMRF model has two mutually dependent components: one models the observed image to estimate features, and the other models the labeling to achieve segmentation. When calculating the feature of each pixel, the homogeneity of the subimage is ensured by using only the pixels currently labeled as the same pattern. With the acquired features, the labeling is obtained through solving a maximum a posteriori problem. In our adaptive approach, the feature set and the labeling are mutually dependent on each other, and therefore are alternately optimized by using a simulated annealing scheme. With the gradual improvement of features' accuracy, the labeling is able to locate the exact boundary of each texture pattern adaptively. The proposed algorithm is compared with a simple MRF model based method in segmentation of Brodatz texture mosaics and real scene images. The satisfying experimental results demonstrate that the proposed approach can differentiate textured images more accurately. Yong Xia 0001, David Dagan Feng, Rongchun Zhao |
IEEE Trans. Image Process. | 1 |
| 2005 | Learning-based algorithm selection for image segmentation
Yong Xia 0001, David Dagan Feng, Rongchun Zhao, Maria Petrou |
Pattern Recognit. Lett. | 1 |