EDBT 2026 Demo / reviewers in the wild / expert
Xiaohan Xing
dblp:229/1085
· DBLP profile ↗
29ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0002-9992-3387ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MF2MR2: Multi-frequency fusion for accelerated multi-contrast MRI reconstruction
Lanqing Liu, Xiaohan Xing, Angelica I. Avilés-Rivero, Harry Qin |
Expert Syst. Appl. | 3 |
| 2026 | Multi-contrast low-field MRI acceleration with k-space progressive learning and image-space hybrid attention fusion
Xiaohan Xing, Qi Chen 0014, Lequan Yu, Lingting Zhu, Lei Xing 0001, Lianli Liu |
Medical Image Anal. | 1 |
| 2025 | Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsabstractThe significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gas-troenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments. © 2025 Association for Computational Linguistics. Jie Liu 0044, Wenxuan Wang 0001, Yihang Su, Yudi Zhang 0005, Cheng-Yi Li, Wenting Chen, Xiaohan Xing, Kao-Jung Chang, LinLin Shen, Michael R. Lyu |
ACL (1) | 8 |
| 2025 | OccMamba: Semantic Occupancy Prediction with State Space ModelsabstractTraining deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability in learning input-conditioned weights and long-range relationships. However, transformer-based networks are notorious for their quadratic computation complexity, seriously undermining their efficacy and deployment in semantic occupancy prediction. Inspired by the global modeling and linear computation complexity of the Mamba architecture, we present the first Mamba-based network for semantic occupancy prediction, termed OccMamba. Specifically, we first design the hierarchical Mamba module and local context processor to better aggregate global and local contextual information, respectively. Besides, to relieve the inherent domain gap between the linguistic and 3D domains, we present a simple yet effective 3D-to-1D reordering scheme, i.e., height-prioritized 2D Hilbert expansion. It can maximally retain the spatial structure of 3D voxels as well as facilitate the processing of Mamba blocks. Endowed with the aforementioned designs, our OccMamba is capable of directly and efficiently processing large volumes of dense scene grids, achieving state-of-the-art performance across three prevalent occupancy prediction benchmarks, including OpenOccupancy, SemanticKITTI, and SemanticPOSS. Notably, on OpenOc-cupancy, our OccMamba outperforms the previous state-of-the-art Co-Occ by 5.1% IoU and 4.3% mIoU, respectively. Our implementation is open-sourced and available at: https://github.com/USTCLH/OccMamba. Yuenan Hou, Xiaohan Xing, Yuexin Ma, Xiao Sun 0001, Yanyong Zhang |
CVPR | 3 |
| 2025 | WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageabstractRecent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy. Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Xiangjian He, Xiaohan Xing, Sen Yang 0006, LinLin Shen |
ICCV | 8 |
| 2025 | One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt EvolutionabstractPolyp segmentation is vital for early colorectal cancer detection, yet traditional fully supervised methods struggle with morphological variability and domain shifts, requiring frequent retraining. Additionally, reliance on large-scale annotations is a major bottleneck due to the time-consuming and error-prone nature of polyp boundary labeling. Recently, vision foundation models like Segment Anything Model (SAM) have demonstrated strong generalizability and fine-grained boundary detection with sparse prompts, effectively addressing key polyp segmentation challenges. However, SAM's prompt-dependent nature limits automation in medical applications, since manually inputting prompts for each image is labor-intensive and time-consuming. We propose OP-SAM, a One-shot Polyp segmentation framework based on SAM that automatically generates prompts from a single annotated image, ensuring accurate and generalizable segmentation without additional annotation burdens. Our method introduces Correlation-based Prior Generation (CPG) for semantic label transfer and Scale-cascaded Prior Fusion (SPF) to adapt to polyp size variations as well as filter out noisy transfers. Instead of dumping all prompts at once, we devise Euclidean Prompt Evolution (EPE) for iterative prompt refinement, progressively enhancing segmentation quality. Extensive evaluations across five datasets validate OP-SAM's effectiveness. Notably, on Kvasir, it achieves 76.93% IoU, surpassing the state-of-the-art by 11.44%. Xiaohan Xing, Jianbang Liu 0002, Fan Bai 0008, Qiang Nie, Max Q.-H. Meng |
ICCV | 2 |
| 2025 | An improved graph attention network combined with reinforcement learning for capacitated vehicle routing problem
Xiaohan Xing, Jian Wang 0005 |
Appl. Intell. | 2 |
| 2025 | MMR-Mamba: Multi-modal MRI reconstruction with Mamba and spatial-frequency information fusion
Lanqing Liu, Qi Chen 0014, Zhanli Hu, Xiaohan Xing, Harry Qin |
Medical Image Anal. | 6 |
| 2025 | Hyperbolic Geometry-Driven Robustness Enhancement for Rare Skin Disease DiagnosisabstractThe automated diagnosis of rare skin diseases using dermoscopy images, known as a few-shot learning (FSL) problem, remains challenging, since traditional FSL research tends to disregard the intrinsic hierarchical nature of rare diseases and data uncertainty. To address these issues, we propose to conduct rare skin disease diagnosis in hyperbolic space, which facilitates implicit class hierarchical structures and precise uncertainty measurement due to pivotal geometrical properties. We propose a Hyperbolic Geometry-driven Robustness Enhancement (HGRE) framework specifically tailored for diagnosing rare skin diseases. The HGRE framework uses implicit hierarchical relation in the hyperbolic space to better represent the features of rare diseases. Moreover, the framework incorporates an Adversarial Proxy Construction (APC) module to address the problem of data uncertainty. Specifically, the APC module uses the distance to the hyperbolic space origin as an indicator of uncertainty to filter and construct adversarial proxies for each uncertain prototype to achieve adversarial robust training. Leveraging the two unique geometrical properties, our HGRE framework effectively addresses the limitations of insufficient hierarchical relation utilization and data uncertainty in FSL-based rare skin disease diagnosis. This enhancement of the model's robustness in training has been corroborated by extensive empirical validation on two skin lesion datasets, where HGRE's performance notably surpassed existing state-of-the-art FSL methods. Yuanyuan Chen 0001, Xiaohan Xing, Jingfeng Zhang, Bolysbek Murat Yerzhanuly, Bazargul Matkerim, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma GradingabstractRecently, multimodal deep learning, which integrates histopathology slides and molecular biomarkers, has achieved a promising performance in glioma grading. Despite great progress, due to the intra-modality complexity and intermodality heterogeneity, existing studies suffer from inadequate histopathology representation learning and inefficient molecular-pathology knowledge alignment. These two issues hinder existing methods to precisely interpret diagnostic molecular-pathology features, thereby limiting their grading performance. Moreover, the real-world applicability of existing multimodal approaches is significantly restricted as molecular biomarkers are not always available during clinical deployment. To address these problems, we introduce a novel Focus on Focus (FoF) framework with paired pathology-genomic training and applicable pathology-only inference, enhancing molecular-pathology representation effectively. Specifically, we propose a Focus-oriented Representation Learning (FRL) module to encourage the model to identify regions positively or negatively related to glioma grading and guide it to focus on the diagnostic areas with a consistency constraint. To effectively link the molecular biomarkers to morphological features, we propose a Multi-view Cross-modal Alignment (MCA) module that projects histopathology representations into molecular subspaces, aligning morphological features with corresponding molecular biomarker status by supervised contrastive learning. Experiments on the TCGA GBMLGG dataset demonstrate that our FoF framework significantly improves the glioma grading. Remarkably, our FoF achieves superior performance using only histopathology slides compared to existing multimodal methods. The source code is available at https://github.com/peterlipan/FoF. Li Pan 0004, Qiushi Yang, Tan Li 0002, Xiaohan Xing, Maximus C. F. Yeung, Zhen Chen 0013 |
BIBM | 5 |
| 2024 | Accelerated Multi-contrast MRI Reconstruction via Frequency and Spatial Mutual Learning
Qi Chen 0014, Xiaohan Xing, Zhen Chen 0013, Zhiwei Xiong |
MICCAI (7) | 2 |
| 2024 | Comprehensive learning and adaptive teaching: Distilling multi-modal knowledge for pathological glioma grading
Xiaohan Xing, Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
Medical Image Anal. | 1 |
| 2023 | Locate before Segment: Topology-guided Retinal Layer Segmentation in Optical Coherence Tomography ImagesabstractOptical Coherence Tomography (OCT) is a non-invasive imaging technique that is instrumental in retinal disease diagnosis and treatment. Segmentation of retinal layers in OCT is an essential step, but remains challenging for common pixel-wise segmentation methods usually fail to obtain the correct layer topology. To tackle this challenge, we propose a novel Locate-to-Segment (L2S) framework to provide a layer region location guidance for pixel-wise labeling learning so as to obtain better segmentation with the correct topology and smooth boundaries. Specifically, a Structured Boundary Regression Network (SBRNet) is devised to first predict the surface positions. For effective learning on normal-size images, we design two regression branches to regress the top surface and eight layer widths separately in SBRNet to locate each layer region with absolutely correct orderings. Then, we take the prediction of SBRNet as an additional input for a common pixel-wise segmentation network to provide the guidance of correct topology. In this L2S manner, our framework takes merits of regression-based methods and pixel-wise labeling-based methods to obtain accurate segmentation with the correct topology and smooth continuous boundaries. Experimental results on a public retinal OCT dataset demonstrate the effectiveness of our method, outperforming state-of-the-art segmentation methods with the highest average Dice score of 90.29% and the lowest average MAD score of 0.782. Yutian Shen, Xiaohan Xing, Max Q.-H. Meng |
ICRA | 3 |
| 2023 | Gradient and Feature Conformity-Steered Medical Image Classification with Noisy Labels
Xiaohan Xing, Zhen Chen 0013, Zhifan Gao, Yixuan Yuan |
MICCAI (6) | 1 |
| 2023 | Medical federated learning with joint graph purification for noisy label learning
Zhen Chen 0013, Wuyang Li, Xiaohan Xing, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2023 | Gradient modulated contrastive distillation of low-rank multi-modal knowledge for disease diagnosis
Xiaohan Xing, Zhen Chen 0013, Yuenan Hou, Yixuan Yuan |
Medical Image Anal. | 1 |
| 2022 | Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited AnnotationsabstractOptical Coherence Tomography (OCT) is a rapidly growing and promising imaging technique, enabling non-invasive high-resolution visualization of biological tissues. Segmentation of tissue structures from OCT scans is essen-tial for disease diagnosis but remains challenging for the blurry boundaries and large volumes. Deep learning-based OCT segmentation algorithms always require large numbers of annotations for satisfying performance, which is hard to meet since manually labeling is time-consuming and labor-intensive. Therefore, we propose a novel semi-supervised OCT segmentation framework utilizing very few labeled scans, i.e., 5 samples, and abundant unlabeled data. Specifically, our framework con-sists of one shared encoder and two different decoder branches. For the two branches, we design a strong augmentation-consistent supervision module and a scaling transformation-consistent supervision module respectively to improve their generalization ability. Besides, cross consistency supervision with feature perturbations between two branches is proposed to incorporate their advantages for further regularization. With such multiple consistency supervision, we aim to enrich the diversity of unsupervised information so as to make full use of labeled and unlabeled data. Experimental results on a public retinal OCT dataset demonstrate the effectiveness of our method, achieving an average dice score of 87.25% in the case of only 5 labeled samples used. It outperforms the supervised baseline by 3.46% and the best semi-supervised model by 1.42% in our experiments. Yantao Shen 0002, Xiaohan Xing, Max Q.-H. Meng |
ICRA | 3 |
| 2022 | Discrepancy-Based Active Learning for Weakly Supervised Bleeding Segmentation in Wireless Capsule Endoscopy Images
Fan Bai 0008, Xiaohan Xing, Yantao Shen 0002, Max Q.-H. Meng |
MICCAI (8) | 2 |
| 2022 | Discrepancy and Gradient-Guided Multi-modal Knowledge Distillation for Pathological Glioma Grading
Xiaohan Xing, Zhen Chen 0013, Meilu Zhu, Yuenan Hou, Zhifan Gao, Yixuan Yuan |
MICCAI (5) | 1 |
| 2022 | Multi-level attention graph neural network based on co-expression gene modules for disease diagnosis and prognosisabstractMOTIVATION: Advanced deep learning techniques have been widely applied in disease diagnosis and prognosis with clinical omics, especially gene expression data. In the regulation of biological processes and disease progression, genes often work interactively rather than individually. Therefore, investigating gene association information and co-functional gene modules can facilitate disease state prediction. RESULTS: To explore the gene modules and inter-gene relational information contained in the omics data, we propose a novel multi-level attention graph neural network (MLA-GNN) for disease diagnosis and prognosis. Specifically, we format omics data into co-expression graphs via weighted correlation network analysis, and then construct multi-level graph features, finally fuse them through a well-designed multi-level graph feature fully fusion module to conduct predictions. For model interpretation, a novel full-gradient graph saliency mechanism is developed to identify the disease-relevant genes. MLA-GNN achieves state-of-the-art performance on transcriptomic data from TCGA-LGG/TCGA-GBM and proteomic data from coronavirus disease 2019 (COVID-19)/non-COVID-19 patient sera. More importantly, the relevant genes selected by our model are interpretable and are consistent with the clinical understanding. AVAILABILITYAND IMPLEMENTATION: The codes are available at https://github.com/TencentAILabHealthcare/MLA-GNN. Xiaohan Xing, Fan Yang 0081, Jun Zhang 0018, Yu Zhao 0009, Mingxuan Gao, Junzhou Huang, Jianhua Yao 0001 |
Bioinform. | 1 |
| 2021 | An Interpretable Multi-Level Enhanced Graph Attention Network for Disease Diagnosis with Gene Expression DataabstractClinical omics, especially gene expression data, have been widely studied and successfully applied for disease diagnosis using machine learning techniques. As genes often work interactively rather than individually, investigating co-functional gene modules can improve our understanding of disease mechanisms and facilitate disease state prediction. To this end, we in this paper propose a novel Multi-Level Enhanced Graph ATtention (MLE-GAT) network to explore the gene modules and intergene relational information contained in the omics data. In specific, we first format the omics data of each patient into co-expression graphs using weighted correlation network analysis (WGCNA) and then feed them to a well-designed multi-level graph feature fully fusion (MGFFF) module for disease diagnosis. For model interpretation, we develop a novel full-gradient graph saliency (FGS) mechanism to identify the disease-relevant genes. Comprehensive experiments show that our proposed MLE-GAT achieves state-of-the-art performance on transcriptomics data from TCGA-LGG/TCGA-GBM and proteomics data from COVID-19/non-COVID-19 patient sera. Xiaohan Xing, Fan Yang 0081, Jun Zhang 0018, Yu Zhao 0009, Mingxuan Gao, Junzhou Huang, Jianhua Yao 0001 |
BIBM | 1 |
| 2021 | Multibranch Learning for Angiodysplasia Segmentation with Attention-Guided Networks and Domain AdaptationabstractAs a common cause of anemia and gastrointestinal bleeding, angiodysplasia (AD) diagnosis in wireless capsule endoscopy (WCE) images is important in clinical. Current manual review requires undivided concentration of the gastroenterologists, which is laborious and time-consuming. The development of computational methods that can assist automated diagnosis of angiodysplasia is highly desirable. In this paper, we present a new approach, ADNet, for angiodysplasia segmentation using convolutional neural networks (CNNs). Compared with previous learning strategies, ADNet gains accuracy from attentionguided and domain-adversarial training via a multibranch CNN architecture. Specifically, the core branch is constructed for AD segmentation in a fully convolutional manner. Then we propose an attention module embedded in the attention branch to enhance network feature learning, which allows ADNet to focus on the most informative and AD relevant regions while processing. Furthermore, an adaptation branch is built to learn domain-invariant features by adversarial training, aiming to improve the performance when datasets are expanded while preventing the degradation induced by the variations in WCE image acquisition. ADNet is evaluated using two WCE datasets with angiodysplasia and the results show the accuracy gains we obtain, where the state-of-the-art segmentation performance on the public dataset of GIANA’17 is achieved. Xiao Jia 0005, Xiaochun Mai, Xiaohan Xing, Yantao Shen 0002, Jiankun Wang 0001, Max Q.-H. Meng |
ICRA | 3 |
| 2021 | Multi-modal Multi-instance Learning Using Weakly Correlated Histopathological Images and Tabular Clinical Information
Fan Yang 0081, Xiaohan Xing, Yu Zhao 0009, Jun Zhang 0018, Yueping Liu, Mengxue Han, Junzhou Huang, Liansheng Wang 0002, Jianhua Yao 0001 |
MICCAI (8) | 3 |
| 2021 | DT-MIL: Deformable Transformer for Multi-instance Learning on Histopathological Image
Fan Yang 0081, Yu Zhao 0009, Xiaohan Xing, Jun Zhang 0018, Mingxuan Gao, Junzhou Huang, Liansheng Wang 0002, Jianhua Yao 0001 |
MICCAI (8) | 4 |
| 2021 | Categorical Relation-Preserving Contrastive Knowledge Distillation for Medical Image Classification
Xiaohan Xing, Yuenan Hou, Yixuan Yuan, Hongsheng Li 0001, Max Q.-H. Meng |
MICCAI (5) | 1 |
| 2020 | Diagnose like a Clinician: Third-order Attention Guided Lesion Amplification Network for WCE Image ClassificationabstractWireless capsule endoscopy (WCE) is a novel imaging tool that allows the noninvasive visualization of the entire gastrointestinal (GI) tract without causing discomfort to the patients. Although convolutional neural networks (CNNs) have obtained promising performance for the automatic lesion recognition, the results of the current approaches are still limited due to the small lesions and the background interference in the WCE images. To overcome these limits, we propose a Third-order Attention guided Lesion Amplification Network (TALA-Net) for WCE image classification. The TALA-Net consists of two branches, including a global branch and an attention-aware branch. Specifically, taking the high-level features in the global branch as the input, we propose a Third-order Attention (ToA) module to generate attention maps that can indicate potential lesion regions. Then, an Attention Guided Lesion Amplification (AGLA) module is proposed to deform multiple level features in the global branch, so as to zoom in the potential lesion features. The deformed features are fused into the attention-aware branch to achieve finer-scale lesion recognition. Finally, predictions from the global and attention-aware branches are averaged to obtain the classification results. Extensive experiments show that the proposed TALA-Net outperforms state-of-the-art methods with an overall classification accuracy of 94.72% on the WCE dataset. Xiaohan Xing, Yixuan Yuan, Max Q.-H. Meng |
IROS | 1 |
| 2020 | Wireless Capsule Endoscopy: A New Tool for Cancer Screening in the Colon With Deep-Learning-Based Polyp RecognitionabstractAccurate recognition of polyps is crucial for early colorectal cancer diagnosis and treatment. Wireless capsule endoscopy (WCE) is a noninvasive, wireless imaging tool that allows direct visualization of the entire colon without discomfort to patients and has the potential to revolutionize the screening workup for colorectal diseases. However, current manual review is laborious and time consuming, requiring the undivided concentration of the gastroenterologist. Computational methods that can assist automated polyp recognition will enhance the outcome both in terms of diagnostic accuracy and efficiency of WCE. This review introduces the computer-assisted algorithms as applied to colorectal polyp screening, focusing on the successes of deep-learning-based strategies in the WCE sequences. We survey key applications of WCE polyp recognition, covering deep-learning-based image-level classification, lesion region detection, and pixel-accurate segmentation. We conclude by discussing emerging research challenges, possible trends, and future directions. Xiao Jia 0005, Xiaohan Xing, Yixuan Yuan, Lei Xing 0001, Max Q.-H. Meng |
Proc. IEEE | 2 |
| 2020 | Automatic Polyp Recognition in Colonoscopy Images Using Deep Learning and Two-Stage Pyramidal Feature PredictionabstractPolyp recognition in colonoscopy images is crucial for early colorectal cancer detection and treatment. However, the current manual review requires undivided concentration of the gastroenterologist and is prone to diagnostic errors. In this article, we present an effective, two-stage approach called PLPNet, where the abbreviation “PLP” stands for the word “polyp,” for automated pixel-accurate polyp recognition in colonoscopy images using very deep convolutional neural networks (CNNs). Compared to hand-engineered approaches and previous neural network architectures, our PLPNet model improves recognition accuracy by adding a polyp proposal stage that predicts the location box with polyp presence. Several schemes are proposed to ensure the model's performance. First of all, we construct a polyp proposal stage as an extension of the faster R-CNN, which performs as a region-level polyp detector to recognize the lesion area as a whole and constitutes stage I of PLPNet. Second, stageII of PLPNet is built in a fully convolutional fashion for pixelwise segmentation. We define a feature sharing strategy to transfer the learned semantics of polyp proposals to the segmentation task of stage II, which is proven to be highly capable of guiding the learning process and improve recognition accuracy. Additionally, we design skip schemes to enrich the feature scales and thus allow the model to generate detailed segmentation predictions. For accurate recognition, the advanced residual nets and feature pyramids are adopted to seek deeper and richer semantics at all network levels. Finally, we construct a two-stage framework for training and run our model convolutionally via a single-stream network at inference time to efficiently output the polyp mask. Experimental results on public data sets of GIANA Challenge demonstrate the accuracy gains of our approach, which surpasses previous state-of-the-art methods on the polyp segmentation task (74.7 Jaccard Index) and establishes new top results in the polyp localization challenge (81.7 recall). Xiao Jia 0005, Xiaochun Mai, Yi Cui 0002, Yixuan Yuan, Xiaohan Xing, Hyunseok Seo, Lei Xing 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2020 | Zoom in Lesions for Better Diagnosis: Attention Guided Deformation Network for WCE Image ClassificationabstractWireless capsule endoscopy (WCE) is a novel imaging tool that allows noninvasive visualization of the entire gastrointestinal (GI) tract without causing discomfort to patients. Convolutional neural networks (CNNs), though perform favorably against traditional machine learning methods, show limited capacity in WCE image classification due to the small lesions and background interference. To overcome these limits, we propose a two-branch Attention Guided Deformation Network (AGDN) for WCE image classification. Specifically, the attention maps of branch1 are utilized to guide the amplification of lesion regions on the input images of branch2, thus leading to better representation and inspection of the small lesions. What's more, we devise and insert Third-order Long-range Feature Aggregation (TLFA) modules into the network. By capturing long-range dependencies and aggregating contextual features, TLFAs endow the network with a global contextual view and stronger feature representation and discrimination capability. Furthermore, we propose a novel Deformation based Attention Consistency (DAC) loss to refine the attention maps and achieve the mutual promotion of the two branches. Finally, the global feature embeddings from the two branches are fused to make image label predictions. Extensive experiments show that the proposed AGDN outperforms state-of-the-art methods with an overall classification accuracy of 91.29% on two public WCE datasets. The source code is available at https://github.com/hathawayxxh/WCE-AGDN. Xiaohan Xing, Yixuan Yuan, Max Q.-H. Meng |
IEEE Trans. Medical Imaging | 1 |