Xiaofeng Yang 0005

dblp:10/2453-5 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-9023-5855ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Vision foundation model for 3D magnetic resonance imaging segmentation, classification, and registration
Shansong Wang, Mojtaba Safari, Chih-Wei Chang, Richard L. J. Qiu, Justin Roper, David S. Yu, Xiaofeng Yang 0005
Medical Image Anal.8
2026 Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
abstract
Vision-language models (VLMs) have achieved impressive progress in natural image reasoning, yet their potential in medical imaging remains underexplored. Medical vision-language tasks demand precise understanding and clinically coherent answers, which are difficult to achieve due to complexity of medical data and the scarcity of high-quality expert annotations. These challenges limit the effectiveness of conventional supervised fine-tuning (SFT) and Chain-of-Thought (CoT) strategies that work well in general domains. To address these challenges, we propose Med-R1, a reinforcement learning (RL)-enhanced VLM designed to improve generalization and reliability in medical reasoning. Med-R1 adopts Group Relative Policy Optimization (GRPO) to encourage reward-guided learning beyond static annotations. We comprehensively evaluate Med-R1 across eight distinct medical imaging modalities. Med-R1 achieves a 29.94% improvement in average accuracy over its base model Qwen2-VL-2B, and even outperforms Qwen2-VL-72B-a model with $36\times $ more parameters. To assess cross-task generalization, we further evaluate Med-R1 on five question types. Med-R1 outperforms Qwen2-VL-2B by 32.06% in question-type generalization, also surpassing Qwen2-VL-72B. We further explore the thinking process in Med-R1, a crucial component of Deepseek-R1. Our results show that omitting intermediate rationales (No-Thinking Med-R1) not only improves cross-domain generalization with less training, but also challenges the common assumption that more reasoning always helps. Nevertheless, we also find that the Think-After Med-R1 variant further improves performance while maintaining interpretability. These findings suggest that, in medical VQA, the mere presence of explicit reasoning does not guarantee better performance. Instead, performance depends on the quality of the reasoning and the position where the reasoning is generated.
Yuxiang Lai, Jike Zhong, Shitian Zhao, Konstantinos Psounis, Xiaofeng Yang 0005
IEEE Trans. Medical Imaging7
2026 OpenVocabCT: Toward Universal Text-Driven CT Image Segmentation
abstract
Computed tomography (CT) is extensively used for accurate visualization and segmentation of organs and lesions. While deep learning models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have significantly improved CT image analysis, their performance often declines when applied to diverse, real-world clinical data. Although foundation models offer a broader and more adaptable solution, their potential is limited due to the challenge of obtaining large-scale, voxel-level annotations for medical images. In response to these challenges, prompting-based models using visual or text prompts have emerged. Visual-prompting methods, such as the Segment Anything Model (SAM), still require significant manual input and can introduce ambiguity when applied to clinical scenarios. Instead, foundation models that use text prompts offer a more versatile and clinically relevant approach. Notably, current text-prompt models, such as the CLIP-Driven Universal Model, are limited to text prompts already encountered during training and struggle to process the complex and diverse scenarios of real-world clinical applications. Instead of fine-tuning models trained from natural imaging, we propose OpenVocabCT, a vision-language model pretrained on large-scale 3D CT images for universal text-driven segmentation. Using the large-scale CT-RATE dataset, we decompose the diagnostic reports into fine-grained, organ-level descriptions using large language models for multi-granular contrastive learning. We evaluate our OpenVocabCT on downstream segmentation tasks across 14 public datasets and 1 institutional dataset for organ and tumor segmentation, demonstrating the superior performance of our model compared to existing methods. All code, datasets, and models will be publicly released at https://github.com/ricklisz/OpenVocabCT.
Yuxiang Lai, Maria Thor, Deborah Marshall, Zachary Buchwald, David S. Yu, Xiaofeng Yang 0005
IEEE Trans. Medical Imaging7
2025 HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
abstract
Fine-tuning large language models (LLMs) faces significant memory challenges due to the high cost of back-propagation.MeZO addresses this issue using zeroth-order (ZO) optimization, matching memory usage to inference but suffering from slow convergence due to varying curvatures across model parameters.To overcome this limitation, we propose HELENE, a scalable and memoryefficient optimizer that integrates annealed A-GNB gradients with diagonal Hessian estimation and layer-wise clipping as a second-order pre-conditioner.HELENE provably accelerates and stabilizes convergence by reducing dependence on total parameter space and scaling with the larger layer dimension.Experiments on RoBERTa-large and OPT-1.3Bdemonstrate superior performances, achieving up to 20× speedup over MeZO with an average accuracy improvement of 1.5%.HELENE also supports full and parameter-efficient fine-tuning methods, outperforming several state-of-the-art optimizers.
Huaqin Zhao, Jiaxi Li 0002, Yi Pan 0001, Shizhe Liang, Xiaofeng Yang 0005, Fei Dou, Tianming Liu 0001, Jin Lu 0001
EMNLP5
2025 Real-time volumetric CBCT reconstruction using surface and X-ray imaging for image-guided radiotherapy
Shaoyan Pan, Vanessa Su, Junbo Peng, Junyuan Li, Yuan Gao 0027, Chih-Wei Chang, Tonghe Wang, Zhen Tian 0003, Xiaofeng Yang 0005
Medical Image Anal.9
2025 CBCT Reconstruction Using Single X-Ray Projection With Cycle-Domain Geometry-Integrated Denoising Diffusion Probabilistic Models
abstract
In the sphere of Cone Beam Computed Tomography (CBCT), acquiring X-ray projections from sufficient angles is indispensable for traditional image reconstruction methods to accurately reconstruct 3D anatomical intricacies. However, this acquisition procedure for the linear accelerator-mounted CBCT systems in radiotherapy takes approximately one minute, impeding its use for ultra-fast intra-fractional motion monitoring during treatment delivery. To address this challenge, we introduce the Patient-specific Cycle-domain Geometric-integrated Denoising Diffusion Probabilistic Model (CG-DDPM). This model aims to leverage patient-specific priors from patient's CT/4DCT images, which are acquired for treatment planning purposes, to reconstruct 3D CBCT from a single-view 2D CBCT projection of any arbitrary angle during treatment, namely single-view reconstructed CBCT (svCBCT). The CG-DDPM framework encompasses a dual DDPM structure: the Projection-DDPM for synthesizing comprehensive full-view projections and the CBCT-DDPM for creating CBCT images. A key innovation is our Cycle-Domain Geometry-Integrated (CDGI) method, incorporating a Cone Beam X-ray Geometric Transformation Module (GTM) to ensure precise, synergistic operation between the dual DDPMs, thereby enhancing reconstruction accuracy and reducing artifacts. Evaluated in a study involving 37 lung cancer patients, the method demonstrated its ability to reconstruct CBCT not only from simulated X-ray projections but also from real-world data. The CG-DDPM significantly outperforms existing V-shape convolutional neural networks (V-nets), Generative Adversarial Networks (GANs), and DDPM methods in terms of reconstruction fidelity and artifact minimization. This was confirmed through extensive voxel-level, structural, visual, and clinical assessments. The capability of CG-DDPM to generate high-quality reconstructed CBCT from a single-view projection at any arbitrary angle using a single model opens the door for ultra-fast, in-treatment volumetric imaging. This is especially beneficial for radiotherapy at motion-associated cancer sites and image-guided interventional procedures.
Shaoyan Pan, Junbo Peng, Yuan Gao 0027, Shao-Yuan Lo, Tianyu Luan, Junyuan Li, Tonghe Wang, Chih-Wei Chang, Zhen Tian 0003, Xiaofeng Yang 0005
IEEE Trans. Medical Imaging10
2024 AnatoMask: Enhancing Medical Image Segmentation with Reconstruction-Guided Self-masking
Tianyu Luan, Yizhou Wu, Shaoyan Pan, Yenho Chen, Xiaofeng Yang 0005
ECCV (60)6
2024 Visual Attention Prompted Prediction and Learning
Yifei Zhang 0006, Bo Pan 0009, Siyi Gu, Guangji Bai, Meikang Qiu, Xiaofeng Yang 0005, Liang Zhao 0002
IJCAI6
2024 DUE: Dynamic Uncertainty-Aware Explanation Supervision via 3D Imputation
abstract
Explanation supervision aims to enhance deep learning models by integrating additional signals to guide the generation of model explanations, showcasing notable improvements in both the predictability and explainability of the model. However, the application of explanation supervision to higher-dimensional data, such as 3D medical images, remains an under-explored domain. Challenges associated with supervising visual explanations in the presence of an additional dimension include: 1) spatial correlation changed, 2) lack of direct 3D annotations, and 3) uncertainty varies across different parts of the explanation. To address these challenges, we propose a Dynamic Uncertainty-aware Explanation supervision (DUE) framework for 3D explanation supervision that ensures uncertainty-aware explanation guidance when dealing with sparsely annotated 3D data with diffusion-based 3D interpolation. Our proposed framework is validated through comprehensive experiments on diverse real-world medical imaging datasets. The results demonstrate the effectiveness of our framework in enhancing the predictability and explainability of deep learning models in the context of medical imaging diagnosis applications.
Qilong Zhao, Yifei Zhang 0006, Mengdan Zhu, Siyi Gu, Xiaofeng Yang 0005, Liang Zhao 0002
KDD6
2023 MAGI: Multi-Annotated Explanation-Guided Learning
abstract
Explanation supervision is a technique in which the model is guided by human-generated explanations during training. This technique aims to improve the predictability of the model by incorporating human understanding of the prediction process into the training phase. This is a challenging task since it relies on the accuracy of human annotation labels. To obtain high-quality explanation annotations, using multiple annotations to do explanation supervision is a reasonable method. However, how to use multiple annotations to improve accuracy is particularly challenging due to the following: 1) The noisiness of annotations from different annotators; 2) The lack of pre-given information about the corresponding relationship between annotations and annotators; 3) Missing annotations since some images are not labeled by all annotators. To solve these challenges, we propose a Multi-annotated explanation-guided learning (MAGI) framework to do explanation supervision with comprehensive and high-quality generated annotations. We first propose a novel generative model to generate annotations from all annotators and infer them using a newly proposed variational inference-based technique by learning the characteristics of each annotator. We also incorporate an alignment mechanism into the generative model to infer the correspondence between annotations and annotators in the training process. Extensive experiments on two datasets from the medical imaging domain demonstrate the effectiveness of our proposed framework in handling noisy annotations while obtaining superior prediction performance compared with previous SOTA.
Yifei Zhang 0006, Siyi Gu, Bo Pan 0009, Xiaofeng Yang 0005, Liang Zhao 0002
ICCV5
2023 ESSA: Explanation Iterative Supervision via Saliency-guided Data Augmentation
abstract
Explanation supervision is a technique in which the model is guided by human-generated explanations during training. This technique aims to improve both the interpretability and predictability of the model by incorporating human understanding into the training process. Since explanation supervision requires a large scale of training data, the data augmentation technique is necessary to be applied to increase the size and diversity of the original dataset. However, data augmentation on sophisticated data like medical images is particularly challenging due to the following: 1) scarcity of data in training the learning-based data augmenter, 2) difficulty in generating realistic and sophisticated images, and 3) difficulty in ensuring the augmented data indeed boosts the performance of explanation-guided learning. To solve these challenges, we propose an Explanation Iterative Supervision via Saliency-guided Data Augmentation (ESSA) framework for conducting explanation supervision and adversarial-trained image data augmentation via a synergized iterative loop that handles the translation from annotation to sophisticated images and the generation of synthetic image-annotation pairs with an alternating training strategy. Extensive experiments on two datasets from the medical imaging domain demonstrate the effectiveness of our proposed framework in improving both the predictability and explainability of the model.
Siyi Gu, Yifei Zhang 0006, Xiaofeng Yang 0005, Liang Zhao 0002
KDD4
2021 Biomechanically constrained non-rigid MR-TRUS prostate registration using deep learning based 3D point cloud matching
Yabo Fu, Yang Lei 0002, Tonghe Wang, Pretesh Patel, Ashesh B. Jani, Walter J. Curran, Tian Liu 0004, Xiaofeng Yang 0005
Medical Image Anal.9
2020 Multi-Needle Detection in 3D Ultrasound Images Using Unsupervised Order-Graph Regularized Sparse Dictionary Learning
abstract
Accurate and automatic multi-needle detection in three-dimensional (3D) ultrasound (US) is a key step of treatment planning for US-guided brachytherapy. However, most current studies are concentrated on single-needle detection by only using a small number of images with a needle, regardless of the massive database of US images without needles. In this paper, we propose a workflow for multi-needle detection by considering the images without needles as auxiliary. Concretely, we train position-specific dictionaries on 3D overlapping patches of auxiliary images, where we develop an enhanced sparse dictionary learning method by integrating spatial continuity of 3D US, dubbed order-graph regularized dictionary learning. Using the learned dictionaries, target images are reconstructed to obtain residual pixels which are then clustered in every slice to yield centers. With the obtained centers, regions of interest (ROIs) are constructed via seeking cylinders. Finally, we detect needles by using the random sample consensus algorithm per ROI and then locate the tips by finding the sharp intensity drops along the detected axis for every needle. Extensive experiments were conducted on a phantom dataset and a prostate dataset of 70/21 patients without/with needles. Visualization and quantitative results show the effectiveness of our proposed workflow. Specifically, our method can correctly detect 95% of needles with a tip location error of 1.01 mm on the prostate dataset. This technique provides accurate multi-needle detection for US-guided HDR prostate brachytherapy, facilitating the clinical workflow.
Xiuxiu He, Zhen Tian 0003, Jiwoong Jason Jeong, Yang Lei 0002, Tonghe Wang, Qiulan Zeng, Ashesh B. Jani, Walter J. Curran, Pretesh Patel, Tian Liu 0004, Xiaofeng Yang 0005
IEEE Trans. Medical Imaging12
2019 Multiscale Entropy Analysis of EEG Based on Non-uniform Time
Hongxia Deng, Jinxiu Guo, Xiaofeng Yang 0005, Jinxiu Hou, Haoqi Liu, Haifang Li 0001
PRCV (2)3
2013 Research and applications: Multiscale segmentation of the skull in MR images for MRI-based attenuation correction of combined MR/PET
abstract
BACKGROUND AND OBJECTIVE: Combined magnetic resonance/positron emission tomography (MR/PET) is a relatively new, hybrid imaging modality. MR-based attenuation correction often requires segmentation of the bone on MR images. In this study, we present an automatic segmentation method for the skull on MR images for attenuation correction in brain MR/PET applications. MATERIALS AND METHODS: Our method transforms T1-weighted MR images to the Radon domain and then detects the features of the skull image. In the Radon domain we use a bilateral filter to construct a multiscale image series. For the repeated convolution we increase the spatial smoothing in each scale and make the width of the spatial and range Gaussian function doubled in each scale. Two filters with different kernels along the vertical direction are applied along the scales from the coarse to fine levels. The results from a coarse scale give a mask for the next fine scale and supervise the segmentation in the next fine scale. The use of the multiscale bilateral filtering scheme is to improve the robustness of the method for noise MR images. After combining the two filtered sinograms, the reciprocal binary sinogram of the skull is obtained for the reconstruction of the skull image. RESULTS: This method has been tested with brain phantom data, simulated brain data, and real MRI data. For real MRI data the Dice overlap ratios are 92.2%±1.9% between our segmentation and manual segmentation. CONCLUSIONS: The multiscale segmentation method is robust and accurate and can be used for MRI-based attenuation correction in combined MR/PET.
Xiaofeng Yang 0005, Baowei Fei
J. Am. Medical Informatics Assoc.1