Yan Xia 0002

dblp:17/6518-2 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vision transformer Hook for dense predictions
abstract
Pre-trained vision transformers (ViTs) have demonstrated remarkable capability in learning semantically rich image representations. However, their underlying plain architectures yield low-resolution feature maps, lacking essential fine-grained spatial details required for dense prediction tasks. To better transfer the learned visual features, we present ViT-Hook, a novel hybrid backbone compatible with plain ViTs that effectively bridges the gap between global semantic understanding and local spatial encodings. Specifically, our method aims to broaden the scope and impact of ViT from the following perspectives: (1) We propose a simple transformer-decoder-inspired hook module that receives hierarchical CNN features as spatial queries and interacts with expressive ViT features from large-scale pre-training, therefore instantiating general-purpose representations into task-suited ones. (2) ViT-Hook is a plug-and-play solution for powerful vision foundation models, such as DINOv2 and RADIO. In this case, we find that only partially fine-tuning several intermediate ViT layers can outperform previous full fine-tuning methods, while substantially reducing compute and memory burdens with most parameters frozen. (3) We evaluate ViT-Hook with various pre-trained sources on multiple dense prediction tasks, including semantic segmentation, instance segmentation, and object detection. Notably, tested on the unified UperNet and Mask R-CNN frameworks, our ViT-Hook surpasses state-of-the-art by a large margin, achieving 59.7 (+4.7) mIoU on ADE20K val, 55.0 (+3.6) box AP and 48.5 (+3.3) mask AP on COCO val2017. • We propose ViT-Hook, a hybrid backbone that effectively enhances Vision Transformer performance on various dense prediction tasks. • The proposed spatial query and hook modules are lightweight yet powerful, achieving competitive results compared to SoTA on widely used benchmarks. • We introduce a novel partial fine-tuning strategy, which outperforms full fine-tuning while using only a fraction of compute and memory. • We validate the generalizability of ViT-Hook on multiple types of upstream pre-training methods, including the most recent vision foundation models.
Siyuan Mei, Mareike Thies, Yan Xia 0002, Yipeng Sun, Fei Wu 0025, Fuxin Fan, Mingxuan Gu, Chengze Ye, Yixing Huang, Vincent Christlein, Andreas K. Maier
Pattern Recognit.3
2025 Multi-view hybrid graph convolutional network for volume-to-mesh reconstruction in cardiovascular MRI
Nicolás Gaggion, Benjamin A. Matheson, Yan Xia 0002, Rodrigo Bonazzola, Nishant Ravikumar, Zeike A. Taylor, Diego H. Milone, Alejandro F. Frangi, Enzo Ferrante
Medical Image Anal.3
2025 SegMorph: Concurrent Motion Estimation and Segmentation for Cardiac MRI Sequences
abstract
We propose a novel recurrent variational netwo8=k]irk, SegMorph, to perform concurrent segmentation and motion estimation on cardiac cine magnetic resonance image (CMR) sequences. Our model establishes a recurrent latent space that captures spatiotemporal features from cine-MRI sequences for multitask inference and synthesis. The proposed model follows a recurrent variational auto-encoder framework and adopts a learnt prior from the temporal inputs. We utilise a multi-branch decoder to handle bi-ventricular segmentation and motion estimation simultaneously. In addition to the spatiotemporal features from the latent space, motion estimation enriches the supervision of sequential segmentation tasks by providing pseudo-ground truth. On the other hand, the segmentation branch helps with motion estimation by predicting deformation vector fields (DVFs) based on anatomical information. Experimental results demonstrate that the proposed method performs better than state-of-the-art approaches qualitatively and quantitatively for both segmentation and motion estimation tasks. We achieved an 81% average Dice Similarity Coefficient (DSC) and a less than 3.5 mm average Hausdorff distance on segmentation. Meanwhile, we achieved a motion estimation Dice Similarity Coefficient of over 79%, with approximately 0.14% of pixels displaying a negative Jacobian determinant in the estimated DVFs.
Ning Bi, Arezoo Zakeri, Yan Xia 0002, Nina Cheng, Alejandro F. Frangi, Ali Gooya
IEEE Trans. Medical Imaging3
2023 Virtual high-resolution MR angiography from non-angiographic multi-contrast MRIs: synthetic vascular model populations for in-silico trials
abstract
Despite success on multi-contrast MR image synthesis, generating specific modalities remains challenging. Those include Magnetic Resonance Angiography (MRA) that highlights details of vascular anatomy using specialised imaging sequences for emphasising inflow effect. This work proposes an end-to-end generative adversarial network that can synthesise anatomically plausible, high-resolution 3D MRA images using commonly acquired multi-contrast MR images (e.g. T1/T2/PD-weighted MR images) for the same subject whilst preserving the continuity of vascular anatomy. A reliable technique for MRA synthesis would unleash the research potential of very few population databases with imaging modalities (such as MRA) that enable quantitative characterisation of whole-brain vasculature. Our work is motivated by the need to generate digital twins and virtual patients of cerebrovascular anatomy for in-silico studies and/or in-silico trials. We propose a dedicated generator and discriminator that leverage the shared and complementary features of multi-source images. We design a composite loss function for emphasising vascular properties by minimising the statistical difference between the feature representations of the target images and the synthesised outputs in both 3D volumetric and 2D projection domains. Experimental results show that the proposed method can synthesise high-quality MRA images and outperform the state-of-the-art generative models both qualitatively and quantitatively. The importance assessment reveals that T2 and PD-weighted images are better predictors of MRA images than T1; and PD-weighted images contribute to better visibility of small vessel branches towards the peripheral regions. In addition, the proposed approach can generalise to unseen data acquired at different imaging centres with different scanners, whilst synthesising MRAs and vascular geometries that maintain vessel continuity. The results show the potential for use of the proposed approach to generating digital twin cohorts of cerebrovascular anatomy at scale from structural MR images typically acquired in population imaging initiatives.
Yan Xia 0002, Nishant Ravikumar, Toni Lassila, Alejandro F. Frangi
Medical Image Anal.1
2022 Automatic 3D+t four-chamber CMR quantification of the UK biobank: integrating imaging and non-imaging data priors at scale
abstract
Accurate 3D modelling of cardiac chambers is essential for clinical assessment of cardiac volume and function, including structural, and motion analysis. Furthermore, to study the correlation between cardiac morphology and other patient information within a large population, it is necessary to automatically generate cardiac mesh models of each subject within the population. In this study, we introduce MCSI-Net (Multi-Cue Shape Inference Network), where we embed a statistical shape model inside a convolutional neural network and leverage both phenotypic and demographic information from the cohort to infer subject-specific reconstructions of all four cardiac chambers in 3D. In this way, we leverage the ability of the network to learn the appearance of cardiac chambers in cine cardiac magnetic resonance (CMR) images, and generate plausible 3D cardiac shapes, by constraining the prediction using a shape prior, in the form of the statistical modes of shape variation learned a priori from a subset of the population. This, in turn, enables the network to generalise to samples across the entire population. To the best of our knowledge, this is the first work that uses such an approach for patient-specific cardiac shape generation. MCSI-Net is capable of producing accurate 3D shapes using just a fraction (about 23% to 46%) of the available image data, which is of significant importance to the community as it supports the acceleration of CMR scan acquisitions. Cardiac MR images from the UK Biobank were used to train and validate the proposed method. We also present the results from analysing 40,000 subjects of the UK Biobank at 50 time-frames, totalling two million image volumes. Our model can generate more globally consistent heart shape than that of manual annotations in the presence of inter-slice motion and shows strong agreement with the reference ranges for cardiac structure and function across cardiac ventricles and atria.
Yan Xia 0002, Xiang Chen 0008, Nishant Ravikumar, Christopher Kelly, Rahman Attar, Nay Aung, Stefan Neubauer, Steffen E. Petersen, Alejandro F. Frangi
Medical Image Anal.1
2022 Learning to complete incomplete hearts for population analysis of cardiac MR images
abstract
Cardiac MR acquisition with complete coverage from base to apex is required to ensure accurate subsequent analyses, such as volumetric and functional measurements. However, this requirement cannot be guaranteed when acquiring images in the presence of motion induced by cardiac muscle contraction and respiration. To address this problem, we propose an effective two-stage pipeline for detecting and synthesising absent slices in both the apical and basal region. The detection model comprises several dense blocks containing convolutional long short-term memory (ConvLSTM) layers, to leverage through-plane contextual and sequential ordering information of slices in cine MR data and achieve reliable classification results. The imputation network is based on a dedicated conditional generative adversarial network (GAN) that helps retain key visual cues and fine structural details in the synthesised image slices. The proposed network can infer multiple missing slices that are anatomically plausible and lead to improved accuracy of subsequent analyses on cardiac MRIs, e.g., ventricle segmentation, cardiac quantification compared to those derived from incomplete cardiac MR datasets. For instance, the results obtained when compensating for the absence of two basal slices show that the mean differences to the reference of stroke volume and ejection fraction are only -1.3 mL and -1.0%, respectively, which are significantly smaller than those calculated from the incomplete data (-26.8 mL and -6.7%). The proposed approach can improve the reliability of high-throughput image analysis in large-scale population studies, minimising the need for re-scanning patients or discarding incomplete acquisitions.
Yan Xia 0002, Nishant Ravikumar, Alejandro F. Frangi
Medical Image Anal.1
2021 A Deep Discontinuity-Preserving Image Registration Network
Xiang Chen 0008, Yan Xia 0002, Nishant Ravikumar, Alejandro F. Frangi
MICCAI (4)2
2021 Shape registration with learned deformations for 3D shape reconstruction from sparse and incomplete point clouds
abstract
Shape reconstruction from sparse point clouds/images is a challenging and relevant task required for a variety of applications in computer vision and medical image analysis (e.g. surgical navigation, cardiac motion analysis, augmented/virtual reality systems). A subset of such methods, viz. 3D shape reconstruction from 2D contours, is especially relevant for computer-aided diagnosis and intervention applications involving meshes derived from multiple 2D image slices, views or projections. We propose a deep learning architecture, coined Mesh Reconstruction Network (MR-Net), which tackles this problem. MR-Net enables accurate 3D mesh reconstruction in real-time despite missing data and with sparse annotations. Using 3D cardiac shape reconstruction from 2D contours defined on short-axis cardiac magnetic resonance image slices as an exemplar, we demonstrate that our approach consistently outperforms state-of-the-art techniques for shape reconstruction from unstructured point clouds. Our approach can reconstruct 3D cardiac meshes to within 2.5-mm point-to-point error, concerning the ground-truth data (the original image spatial resolution is ∼1.8×1.8×10mm3). We further evaluate the robustness of the proposed approach to incomplete data, and contours estimated using an automatic segmentation algorithm. MR-Net is generic and could reconstruct shapes of other organs, making it compelling as a tool for various applications in medical image analysis.
Xiang Chen 0008, Nishant Ravikumar, Yan Xia 0002, Rahman Attar, Andres Diaz-Pinto, Stefan K. Piechnik, Stefan Neubauer, Steffen E. Petersen, Alejandro F. Frangi
Medical Image Anal.3
2021 Super-Resolution of Cardiac MR Cine Imaging using Conditional GANs and Unsupervised Transfer Learning
abstract
High-resolution (HR), isotropic cardiac Magnetic Resonance (MR) cine imaging is challenging since it requires long acquisition and patient breath-hold times. Instead, 2D balanced steady-state free precession (SSFP) sequence is widely used in clinical routine. However, it produces highly-anisotropic image stacks, with large through-plane spacing that can hinder subsequent image analysis. To resolve this, we propose a novel, robust adversarial learning super-resolution (SR) algorithm based on conditional generative adversarial nets (GANs), that incorporates a state-of-the-art optical flow component to generate an auxiliary image to guide image synthesis. The approach is designed for real-world clinical scenarios and requires neither multiple low-resolution (LR) scans with multiple views, nor the corresponding HR scans, and is trained in an end-to-end unsupervised transfer learning fashion. The designed framework effectively incorporates visual properties and relevant structures of input images and can synthesise 3D isotropic, anatomically plausible cardiac MR images, consistent with the acquired slices. Experimental results show that the proposed SR method outperforms several state-of-the-art methods both qualitatively and quantitatively. We show that subsequent image analyses including ventricle segmentation, cardiac quantification, and non-rigid registration can benefit from the super-resolved, isotropic cardiac MR images, to produce more accurate quantitative results, without increasing the acquisition time. The average Dice similarity coefficient (DSC) for the left ventricular (LV) cavity and myocardium are 0.95 and 0.81, respectively, between real and synthesised slice segmentation. For non-rigid registration and motion tracking through the cardiac cycle, the proposed method improves the average DSC from 0.75 to 0.86, compared to the original resolution images.
Yan Xia 0002, Nishant Ravikumar, John P. Greenwood, Stefan Neubauer, Steffen E. Petersen, Alejandro F. Frangi
Medical Image Anal.1
2021 Recovering from missing data in population imaging - Cardiac MR image imputation via conditional generative adversarial nets
Yan Xia 0002, Le Zhang 0005, Nishant Ravikumar, Rahman Attar, Stefan K. Piechnik, Stefan Neubauer, Steffen E. Petersen, Alejandro F. Frangi
Medical Image Anal.1
2021 PMS-GAN: Parallel Multi-Stream Generative Adversarial Network for Multi-Material Decomposition in Spectral Computed Tomography
abstract
Spectral computed tomography is able to provide quantitative information on the scanned object and enables material decomposition. Traditional projection-based material decomposition methods suffer from the nonlinearity of the imaging system, which limits the decomposition accuracy. Inspired by the generative adversarial network, we proposed a novel parallel multi-stream generative adversarial network (PMS-GAN) to perform projection-based multi-material decomposition in spectral computed tomography. By designing the differential map and incorporating the adversarial network into loss function, the decomposition accuracy was significantly improved with robust performance. The proposed network was quantitatively evaluated by both simulation and experimental study. The results show that PMS-GAN outperformed the reference methods with certain robustness. Compared with Pix2pix-GAN, PMS-GAN increased the structural similarity index by 172% on the contrast agent Ultravist370, 11% on bones, and 71% on bone marrow, respectively, in a simulated test scenario. In an experimental test scenario, 9% and 38% improvements of the structural similarity index on the biopsy needle and on a torso phantom were observed, respectively. The proposed network demonstrates its capability of multi-material decomposition and has certain potential toward clinical applications.
Mufeng Geng, Zifeng Tian, Yunfei You, Ximeng Feng, Yan Xia 0002, Qiushi Ren, Xiangxi Meng 0001, Andreas K. Maier, Yanye Lu
IEEE Trans. Medical Imaging6
2014 Towards Clinical Application of a Laplace Operator-Based Region of Interest Reconstruction Algorithm in C-Arm CT
abstract
It is known that a reduction of the field-of-view in 3-D X-ray imaging is proportional to a reduction in radiation dose. The resulting truncation, however, is incompatible with conventional reconstruction algorithms. Recently, a novel method for region of interest reconstruction that uses neither prior knowledge nor extrapolation has been published, named approximated truncation robust algorithm for computed tomography (ATRACT). It is based on a decomposition of the standard ramp filter into a 2-D Laplace filtering and a 2-D Radon-based residual filtering step. In this paper, we present two variants of the original ATRACT. One is based on expressing the residual filter as an efficient 2-D convolution with an analytically derived kernel. The second variant is to apply ATRACT in 1-D to further reduce computational complexity. The proposed algorithms were evaluated by using a reconstruction benchmark, as well as two clinical data sets. The results are encouraging since the proposed algorithms achieve a speed-up factor of up to 245 compared to the 2-D Radon-based ATRACT. Reconstructions of high accuracy are obtained, e.g., even real-data reconstruction in the presence of severe truncation achieve a relative root mean square error of as little as 0.92% with respect to nontruncated data.
Yan Xia 0002, Hannes G. Hofmann, Frank Dennerlein, Kerstin Müller 0002, Chris Schwemmer, Sebastian Bauer 0001, Gouthami Chintalapani, Ponraj Chinnadurai, Joachim Hornegger, Andreas K. Maier
IEEE Trans. Medical Imaging1