EDBT 2026 Demo / reviewers in the wild / expert
Xiao Ma 0011
dblp:35/573-11
· DBLP profile ↗
19ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-1842-5029ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Test-time generative augmentation for medical image segmentation
Xiao Ma 0011, Yuhui Tao, Zetian Zhang, Yuhan Zhang 0001, Xi Wang 0013, Sheng Zhang 0024, Zexuan Ji, Yizhe Zhang 0001, Qiang Chen 0004, Guang Yang 0006 |
Medical Image Anal. | 1 |
| 2025 | ST-MIGD: Spatial-Temporal Domain Medical Image Generation via Deformation-Based Diffusion Models
Xiao Ma 0011, Yizhe Zhang 0001, Qiang Chen 0004 |
PRCV (14) | 3 |
| 2025 | Shared Hybrid Attention Transformer network for colon polyp segmentation
Zexuan Ji, Xiao Ma 0011 |
Neurocomputing | 3 |
| 2024 | Model-Based Label-to-Image Diffusion for Semi-Supervised Choroidal Vessel SegmentationabstractCurrent successful choroidal vessel segmentation methods rely on large amounts of voxel-level annotations on the 3D optical coherence tomography images, which are hard and time-consuming. Semi-supervised learning solves this issue by enabling model learning from both unlabeled data and a limited amount of labeled data. A challenge is the defective pseudo labels generated for the unlabeled data. In this work, we propose a model-based label-to-image diffusion (MLD) framework for semi-supervised choroidal vessel segmentation. We first generate pseudo labels from unlabeled images with a coarse correspondence using a model-based strategy. Then, we generate precisely corresponding images of pseudo labels by a hierarchical diffusion probabilistic model. We evaluated our method on myopia data with a new topological connectivity metric. The quantitative and qualitative experimental results indicate the effectiveness of the label-to-image diffusion framework and its benefit for enhancing the existing supervised choroidal segmentation methods. The code is available at: https://github.com/nicetomeetu21/MLD. Xiao Ma 0011, Songtao Yuan, Qiang Chen 0004 |
ICASSP | 2 |
| 2024 | Memory-Efficient High-Resolution OCT Volume Synthesis with Cascaded Amortized Latent Diffusion Models
Xiao Ma 0011, Yuhan Zhang 0001, Songtao Yuan, Yong Liu 0026, Qiang Chen 0004, Huazhu Fu |
MICCAI (7) | 2 |
| 2024 | Coarse-to-Fine Latent Diffusion Model for Glaucoma Forecast on Sequential Fundus Images
Yuhan Zhang 0001, Xikai Yang, Xiao Ma 0011, Ningli Wang, Xi Wang 0013, Pheng-Ann Heng |
MICCAI (5) | 4 |
| 2024 | SAVE: Encoding spatial interactions for vision transformers
Xiao Ma 0011, Zetian Zhang, Zexuan Ji, Mingchao Li 0002, Yuhan Zhang 0001, Qiang Chen 0004 |
Image Vis. Comput. | 1 |
| 2024 | Mirrored X-Net: Joint classification and contrastive learning for weakly supervised GA segmentation in SD-OCT
Zexuan Ji, Xiao Ma 0011, Theodore Leng, Daniel L. Rubin, Qiang Chen 0004 |
Pattern Recognit. | 2 |
| 2024 | Sparse Coding Inspired LSTM and Self-Attention Integration for Medical Image SegmentationabstractAccurate and automatic segmentation of medical images plays an essential role in clinical diagnosis and analysis. It has been established that integrating contextual relationships substantially enhances the representational ability of neural networks. Conventionally, Long Short-Term Memory (LSTM) and Self-Attention (SA) mechanisms have been recognized for their proficiency in capturing global dependencies within data. However, these mechanisms have typically been viewed as distinct modules without a direct linkage. This paper presents the integration of LSTM design with SA sparse coding as a key innovation. It uses linear combinations of LSTM states for SA's query, key, and value (QKV) matrices to leverage LSTM's capability for state compression and historical data retention. This approach aims to rectify the shortcomings of conventional sparse coding methods that overlook temporal information, thereby enhancing SA's ability to do sparse coding and capture global dependencies. Building upon this premise, we introduce two innovative modules that weave the SA matrix into the LSTM state design in distinct manners, enabling LSTM to more adeptly model global dependencies and meld seamlessly with SA without accruing extra computational demands. Both modules are separately embedded into the U-shaped convolutional neural network architecture for handling both 2D and 3D medical images. Experimental evaluations on downstream medical image segmentation tasks reveal that our proposed modules not only excel on four extensively utilized datasets across various baselines but also enhance prediction accuracy, even on baselines that have already incorporated contextual modules. Code is available at https://github.com/yeshunlong/SALSTM. Zexuan Ji, Shunlong Ye, Xiao Ma 0011 |
IEEE Trans. Image Process. | 3 |
| 2024 | Diverse Data Generation for Retinal Layer Segmentation With Potential Structure ModelingabstractAccurate retinal layer segmentation on optical coherence tomography (OCT) images is hampered by the challenges of collecting OCT images with diverse pathological characterization and balanced distribution. Current generative models can produce high-realistic images and corresponding labels without quantitative limitations by fitting distributions of real collected data. Nevertheless, the diversity of their generated data is still limited due to the inherent imbalance of training data. To address these issues, we propose an image-label pair generation framework that generates diverse and balanced potential data from imbalanced real samples. Specifically, the framework first generates diverse layer masks, and then generates plausible OCT images corresponding to these layer masks using two customized diffusion probabilistic models respectively. To learn from imbalanced data and facilitate balanced generation, we introduce pathological-related conditions to guide the generation processes. To enhance the diversity of the generated image-label pairs, we propose a potential structure modeling technique that transfers the knowledge of diverse sub-structures from lowly- or non-pathological samples to highly pathological samples. We conducted extensive experiments on two public datasets for retinal layer segmentation. Firstly, our method generates OCT images with higher image quality and diversity compared to other generative methods. Furthermore, based on the extensive training with the generated OCT images, downstream retinal layer segmentation tasks demonstrate improved results. The code is publicly available at: https://github.com/nicetomeetu21/GenPSM. Xiao Ma 0011, Zetian Zhang, Yuhan Zhang 0001, Songtao Yuan, Huazhu Fu, Qiang Chen 0004 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Semantic-Oriented Visual Prompt Learning for Diabetic Retinopathy Grading on Fundus ImagesabstractDiabetic retinopathy (DR) is a serious ocular condition that requires effective monitoring and treatment by ophthalmologists. However, constructing a reliable DR grading model remains a challenging and costly task, heavily reliant on high-quality training sets and adequate hardware resources. In this paper, we investigate the knowledge transferability of large-scale pre-trained models (LPMs) to fundus images based on prompt learning to construct a DR grading model efficiently. Unlike full-tuning which fine-tunes all parameters of LPMs, prompt learning only involves a minimal number of additional learnable parameters while achieving a competitive effect as full-tuning. Inspired by visual prompt tuning, we propose Semantic-oriented Visual Prompt Learning (SVPL) to enhance the semantic perception ability for better extracting task-specific knowledge from LPMs, without any additional annotations. Specifically, SVPL assigns a group of learnable prompts for each DR level to fit the complex pathological manifestations and then aligns each prompt group to task-specific semantic space via a contrastive group alignment (CGA) module. We also propose a plug-and-play adapter module, Hierarchical Semantic Delivery (HSD), which allows the semantic transition of prompt groups from shallow to deep layers to facilitate efficient knowledge mining and model convergence. Our extensive experiments on three public DR grading datasets demonstrate that SVPL achieves superior results compared to other transfer tuning and DR grading methods. Further analysis suggests that the generalized knowledge from LPMs is advantageous for constructing the DR grading model on fundus images. Yuhan Zhang 0001, Xiao Ma 0011, Mingchao Li 0002, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Adjustable Robust Transformer for High Myopia Screening in Optical Coherence Tomography
Xiao Ma 0011, Zetian Zhang, Zexuan Ji, Songtao Yuan, Qiang Chen 0004 |
MICCAI (5) | 1 |
| 2023 | CBAV-Loss: Crossover and Branch Losses for Artery-Vein Segmentation in OCTA Images
Zetian Zhang, Xiao Ma 0011, Zexuan Ji, Songtao Yuan, Qiang Chen 0004 |
PRCV (13) | 2 |
| 2023 | LAGAN: Lesion-Aware Generative Adversarial Networks for Edema Area Segmentation in SD-OCT ImagesabstractLarge volume of labeled data is a cornerstone for deep learning (DL) based segmentation methods. Medical images require domain experts to annotate, and full segmentation annotations of large volumes of medical data are difficult, if not impossible, to acquire in practice. Compared with full annotations, image-level labels are multiple orders of magnitude faster and easier to obtain. Image-level labels contain rich information that correlates with the underlying segmentation tasks and should be utilized in modeling segmentation problems. In this article, we aim to build a robust DL-based lesion segmentation model using only image-level labels (normal v.s. abnormal). Our method consists of three main steps: (1) training an image classifier with image-level labels; (2) utilizing a model visualization tool to generate an object heat map for each training sample according to the trained classifier; (3) based on the generated heat maps (as pseudo-annotations) and an adversarial learning framework, we construct and train an image generator for Edema Area Segmentation (EAS). We name the proposed method Lesion-Aware Generative Adversarial Networks (LAGAN) as it combines the merits of supervised learning (being lesion-aware) and adversarial training (for image generation). Additional technical treatments, such as the design of a multi-scale patch-based discriminator, further enhance the effectiveness of our proposed method. We validate the superior performance of LAGAN via comprehensive experiments on two publicly available datasets (i.e., AI Challenger and RETOUCH). Yuhui Tao, Xiao Ma 0011, Yizhe Zhang 0001, Zexuan Ji, Wen Fan 0003, Songtao Yuan, Qiang Chen 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | PRGAN: A Progressive Refined GAN for Lesion Localization and Segmentation on High-Resolution Retinal Fundus Photography
Xiao Ma 0011, Qiang Chen 0004, Zexuan Ji |
PRCV (2) | 2 |
| 2022 | Unsupervised Medical Image Registration Based on Multi-scale Cascade Network
Yuying Ge, Xiao Ma 0011, Qiang Chen 0004, Zexuan Ji |
PRCV (2) | 2 |
| 2022 | Self-Supervised Sequence Recovery for Semi-Supervised Retinal Layer SegmentationabstractAutomated layer segmentation plays an important role for retinal disease diagnosis in optical coherence tomography (OCT) images. However, the severe retinal diseases result in the performance degeneration of automated layer segmentation approaches. In this paper, we present a robust semi-supervised layer segmentation network to relieve the model failures on abnormal retinas. We obtain the lesion features from the labeled images with disease-balanced distribution, and utilize the unlabeled images to supplement the layer structure information. Specifically, in our method, the cross-consistency training is utilized over the predictions of different decoders, and we enforce a consistency between different decoder predictions to improve the encoder's representation. Then, we propose a sequence prediction branch based on self-supervised manner, which is designed to predict the position of each jigsaw puzzle to obtain sensory perception of the retinal layer structure. To this task, a layer spatial pyramid pooling (LSPP) module is designed to extract multi-scale layer spatial features. Furthermore, we use the optical coherence tomography angiography (OCTA) to supplement the information damaged by diseases. The experimental results illustrate that our method achieves more robust results compared with current supervised segmentation methods. Meanwhile, advanced segmentation performance can be obtained compared with state-of-the-art semi-supervised segmentation methods. Jiadong Yang, Yuhui Tao, Qiuzhuo Xu, Yuhan Zhang 0001, Xiao Ma 0011, Songtao Yuan, Qiang Chen 0004 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | LamNet: A Lesion Attention Maps-Guided Network for the Prediction of Choroidal Neovascularization Volume in SD-OCT ImagesabstractChoroidal neovascularization (CNV) volume prediction has an important clinical significance to predict the therapeutic effect and schedule the follow-up. In this paper, we propose a Lesion Attention Maps-Guided Network (LamNet) to automatically predict the CNV volume of next follow-up visit after therapy based on 3-dimentional spectral-domain optical coherence tomography (SD-OCT) images. In particular, the backbone of LamNet is a 3D convolutional neural network (3D-CNN). In order to guide the network to focus on the local CNV lesion regions, we use CNV attention maps generated by an attention map generator to produce the multi-scale local context features. Then, the multi-scale of both local and global feature maps are fused to achieve the high-precision CNV volume prediction. In addition, we also design a synergistic multi-task predictor, in which a trend-consistent loss ensures that the change trend of the predicted CNV volume is consistent with the real change trend of the CNV volume. The experiments include a total of 541 SD-OCT cubes from 68 patients with two types of CNV captured by two different SD-OCT devices. The results demonstrate that LamNet can provide the reliable and accurate CNV volume prediction, which would further assist the clinical diagnosis and design the treatment options. Yuhan Zhang 0001, Xiao Ma 0011, Mingchao Li 0002, Zexuan Ji, Songtao Yuan, Qiang Chen 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | MS-CAM: Multi-Scale Class Activation Maps for Weakly-Supervised Segmentation of Geographic Atrophy Lesions in SD-OCT ImagesabstractAs one of the most critical characteristics in advanced stage of non-exudative Age-related Macular Degeneration (AMD), Geographic Atrophy (GA) is one of the significant causes of sustained visual acuity loss. Automatic localization of retinal regions affected by GA is a fundamental step for clinical diagnosis. In this paper, we present a novel weakly supervised model for GA segmentation in Spectral-Domain Optical Coherence Tomography (SD-OCT) images. A novel Multi-Scale Class Activation Map (MS-CAM) is proposed to highlight the discriminatory significance regions in localization and detail descriptions. To extract available multi-scale features, we design a Scaling and UpSampling (SUS) module to balance the information content between features of different scales. To capture more discriminative features, an Attentional Fully Connected (AFC) module is proposed by introducing the attention mechanism into the fully connected operations to enhance the significant informative features and suppress less useful ones. Based on the location cues, the final GA region prediction is obtained by the projection segmentation of MS-CAM. The experimental results on two independent datasets demonstrate that the proposed weakly supervised model outperforms the conventional GA segmentation methods and can produce similar or superior accuracy comparing with fully supervised approaches. The source code has been released and is available on GitHub: https://github.com/ jizexuan/Multi-Scale-Class-Activation-Map-Tensorflow. Xiao Ma 0011, Zexuan Ji, Sijie Niu, Theodore Leng, Daniel L. Rubin, Qiang Chen 0004 |
IEEE J. Biomed. Health Informatics | 1 |