EDBT 2026 Demo / reviewers in the wild / expert
Xiangyi Yan
dblp:249/2733
· DBLP profile ↗
18ranked-venue papers
4as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Rib Fracture Instance Segmentation and Classification From CT on the RibFrac ChallengeabstractRib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website (https://ribfrac.grand-challenge.org/). In addition, we further analyzed the impact of two post-challenge advancements-large-scale pretraining and rib segmentation-based on our internal baseline for rib fracture detection. These findings lay a foundation for future research and development in AI-assisted rib fracture diagnosis. Jiancheng Yang, Kaiming Kuang, Donglai Wei 0001, Shixuan Gu, Jianying Liu, Zhizhong Chai, Yongjie Xiao, Hao Chen 0011, Liming Xu, Bang Du, Xiangyi Yan, Hao Tang 0010, Adam M. Alessio, Gregory Holste, Jianye He, Lixuan Che, Hanspeter Pfister, Ming Li 0005, Bingbing Ni |
IEEE Trans. Medical Imaging | 15 |
| 2024 | Hybrid-CSR: Coupling Explicit and Implicit Reconstruction of Cortical Surface
Shanlin Sun, Pooya Khosravi, Chenyu You, Deying Kong, Xiangyi Yan, Xiaohui Xie |
BMVC | 8 |
| 2024 | Hybrid Neural Diffeomorphic Flow for Shape Representation and Generation via TriplaneabstractDeep Implicit Functions (DIFs) have gained popularity in 3D computer vision due to their compactness and continuous representation capabilities. However, addressing dense correspondences and semantic relationships across DIF-encoded shapes remains a critical challenge, limiting their applications in texture transfer and shape analysis. Moreover, recent endeavors in 3D shape generation using DIFs often neglect correspondence and topology preservation. This paper presents HNDF (Hybrid Neural Diffeomorphic Flow), a method that implicitly learns the underlying representation and decomposes intricate dense correspondences into explicitly axis-aligned triplane features. To avoid suboptimal representations trapped in local minima, we propose hybrid supervision that captures both local and global correspondences. Unlike conventional approaches that directly generate new 3D shapes, we further explore the idea of shape generation with deformed template shape via diffeomorphic flows, where the deformation is encoded by the generated triplane features. Leveraging a pre-existing 2D diffusion model, we produce high-quality and diverse 3D diffeomorphic flows through generated triplanes features, ensuring topological consistency with the template shape. Extensive experiments on medical image organ segmentation datasets evaluate the effectiveness of HNDF in 3D shape representation and generation. Shanlin Sun, Thanh-Tung Le, Xiangyi Yan, Chenyu You, Xiaohui Xie |
WACV | 4 |
| 2024 | CVTHead: One-shot Controllable Head Avatar with Vertex-feature TransformerabstractReconstructing personalized animatable head avatars has significant implications in the fields of AR/VR. Existing methods for achieving explicit face control of 3D Morphable Models (3DMM) typically rely on multi-view images or videos of a single subject, making the reconstruction process complex. Additionally, the traditional rendering pipeline is time-consuming, limiting real-time animation possibilities. In this paper, we introduce CVTHead, a novel approach that generates controllable neural head avatars from a single reference image using point-based neural rendering. CVT-Head considers the sparse vertices of mesh as the point set and employs the proposed Vertex-feature Transformer to learn local feature descriptors for each vertex. This enables the modeling of long-range dependencies among all the vertices. Experimental results on the VoxCeleb dataset demonstrate that CVTHead achieves comparable performance to state-of-the-art graphics-based methods. Moreover, it enables efficient rendering of novel human heads with various expressions, head poses, and camera views. These attributes can be explicitly controlled using the coefficients of 3DMMs, facilitating versatile and realistic animation in real-time scenarios. Codes and pre-trained model can be found at https://github.com/HowieMa/CVTHead. Shanlin Sun, Xiangyi Yan, Xiaohui Xie |
WACV | 4 |
| 2024 | AFTer-SAM: Adapting SAM with Axial Fusion Transformer for Medical Imaging SegmentationabstractThe Segmentation Anything Model (SAM) has demonstrated effectiveness in various segmentation tasks. However, its application to 3D medical data has posed challenges due to its inherent design for both 2D and natural images. While there have been attempts to apply SAM to medical images on a slice-by-slice basis, the outcomes have been less than optimal. In this study, we introduce AFTer-SAM, an adaptation of SAM designed for volumetric medical image segmentation. By incorporating an Axial Fusion Transformer, AFTer-SAM is capable of capturing both intra-slice details and inter-slice contextual information, essential for accurate medical image segmentation. Given the potential computational challenges of training this enhanced model, we utilize Low-Rank Adaptation (LoRA) to efficiently finetune the weights of the Axial Fusion Transformer. This ensures a streamlined training process without compromising on performance. Our results indicate that AFTer-SAM offers significant improvements in volumetric medical image segmentation, suggesting a promising direction for the application of large pre-trained models in medical imaging. Xiangyi Yan, Shanlin Sun, Thanh-Tung Le, Chenyu You, Xiaohui Xie |
WACV | 1 |
| 2024 | Medical image registration via neural fieldsabstractImage registration is an essential step in many medical image analysis tasks. Traditional methods for image registration are primarily optimization-driven, finding the optimal deformations that maximize the similarity between two images. Recent learning-based methods, trained to directly predict transformations between two images, run much faster, but suffer from performance deficiencies due to domain shift. Here we present a new neural network based image registration framework, called NIR (Neural Image Registration), which is based on optimization but utilizes deep neural networks to model deformations between image pairs. NIR represents the transformation between two images with a continuous function implemented via neural fields, receiving a 3D coordinate as input and outputting the corresponding deformation vector. NIR provides two ways of generating deformation field: directly output a displacement vector field for general deformable registration, or output a velocity vector field and integrate the velocity field to derive the deformation field for diffeomorphic image registration. The optimal registration is discovered by updating the parameters of the neural field via stochastic mini-batch gradient descent. We describe several design choices that facilitate model optimization, including coordinate encoding, sinusoidal activation, coordinate sampling, and intensity sampling. NIR is evaluated on two 3D MR brain scan datasets, demonstrating highly competitive performance in terms of both registration accuracy and regularity. Compared to traditional optimization-based methods, our approach achieves better results in shorter computation times. In addition, our methods exhibit performance on a cross-dataset registration task, compared to the pre-trained learning-based methods. Shanlin Sun, Chenyu You, Hao Tang 0010, Deying Kong, Junayed Naushad, Xiangyi Yan, Pooya Khosravi, James S. Duncan, Xiaohui Xie |
Medical Image Anal. | 7 |
| 2023 | MedGen3D: A Deep Generative Framework for Paired 3D Image and Mask Generation
Yifeng Xiong, Chenyu You, Pooya Khosravi, Shanlin Sun, Xiangyi Yan, James S. Duncan, Xiaohui Xie |
MICCAI (1) | 6 |
| 2023 | Localized Region Contrast for Enhancing Self-supervised Learning in Medical Image Segmentation
Xiangyi Yan, Junayed Naushad, Chenyu You, Hao Tang 0010, Shanlin Sun, James S. Duncan, Xiaohui Xie |
MICCAI (2) | 1 |
| 2023 | Diffeomorphic Image Registration with Neural Velocity FieldabstractDiffeomorphic image registration, offering smooth transformation and topology preservation, is required in many medical image analysis tasks. Traditional methods impose certain modeling constraints on the space of admissible transformations and use optimization to find the optimal transformation between two images. Specifying the right space of admissible transformations is challenging: the registration quality can be poor if the space is too restrictive, while the optimization can be hard to solve if the space is too general. Recent learning-based methods, utilizing deep neural networks to learn the transformation directly, achieve fast inference, but face challenges in accuracy due to the difficulties in capturing the small local deformations and generalization ability. Here we propose a new optimization-based method named DNVF (Diffeomorphic Image Registration with Neural Velocity Field) which utilizes deep neural network to model the space of admissible transformations. A multilayer perceptron (MLP) with sinusoidal activation function is used to represent the continuous velocity field and assigns a velocity vector to every point in space, providing the flexibility of modeling complex deformations as well as the convenience of optimization. Moreover, we propose a cascaded image registration framework (Cas-DNVF) by combining the benefits of both optimization and learning based methods, where a fully convolutional neural network (FCN) is trained to predict the initial deformation, followed by DNVF for further refinement. Experiments on two large-scale 3D MR brain scan datasets demonstrate that our proposed methods significantly outperform the state-of-the-art registration methods. Shanlin Sun, Xiangyi Yan, Chenyu You, Hao Tang 0010, Junayed Naushad, Deying Kong, Xiaohui Xie |
WACV | 3 |
| 2023 | Representation Recovering for Self-Supervised Pre-training on Medical ImagesabstractAdvances in self-supervised learning have drawn attention to developing techniques to extract effective visual representations from unlabeled images. Contrastive learning (CL) trains a model to extract consistent features by generating different views. Recent success of Masked Autoencoders (MAE) highlights the benefit of generative modeling in self-supervised learning. The generative approaches encode the input into a compact embedding and empower the model’s ability of recovering the original input. However, in our experiments, we found vanilla MAE mainly recovers coarse high level semantic information and is inadequate in recovering detailed low level information. We show that in dense downstream prediction tasks like multi-organ segmentation, directly applying MAE is not ideal. Here, we propose RepRec, a hybrid visual representation learning framework for self-supervised pre-training on large-scale unlabelled medical datasets, which takes advantage of both contrastive and generative modeling. To solve the aforementioned dilemma that MAE encounters, a convolutional encoder is pre-trained to provide low-level feature information, in a contrastive way; and a transformer encoder is pre-trained to produce high level semantic dependency, in a generative way – by recovering masked representations from the convolutional encoder. Extensive experiments on three multi-organ segmentation datasets demonstrate that our method outperforms current state-of-the-art methods. Xiangyi Yan, Junayed Naushad, Shanlin Sun, Hao Tang 0010, Deying Kong, Chenyu You, Xiaohui Xie |
WACV | 1 |
| 2022 | Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic FlowabstractDeep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences simultaneously, capturing semantic relationships across shapes of the same class by learning a DIFs-modeled shape template. These methods provide great flexibility and accuracy in reconstructing 3D shapes and inferring correspondences. However, the point correspondences built from these methods do not intrinsically preserve the topology of the shapes, unlike mesh-based template matching methods. This limits their applications on 3D geometries where underlying topological structures exist and matter, such as anatomical structures in medical images. In this paper, we propose a new model called Neural Diffeomorphic Flow (NDF) to learn deep implicit shape templates, representing shapes as conditional diffeomorphic deformations of templates, intrinsically preserving shape topologies. The diffeomorphic deformation is realized by an autodecoder consisting of Neural Ordinary Differential Equation (NODE) blocks that progressively map shapes to implicit templates. We conduct extensive experiments on several medical image organ segmentation datasets to evaluate the effectiveness of NDF on reconstructing and aligning shapes. NDF achieves consistently state-of-the-art organ shape reconstruction and registration results in both accuracy and quality. The source code is publicly available at https://github.com/Siwensun/Neural_Diffeomorphic_Flow-NDF. Shanlin Sun, Deying Kong, Hao Tang 0010, Xiangyi Yan, Xiaohui Xie |
CVPR | 5 |
| 2022 | Identity-Aware Hand Mesh Estimation and Personalization from RGB Images
Deying Kong, Linguang Zhang, Liangjian Chen, Xiangyi Yan, Shanlin Sun, Xingwei Liu, Xiaohui Xie |
ECCV (5) | 5 |
| 2022 | PPT: Token-Pruned Pose Transformer for Monocular and Multi-view Human Pose Estimation
Yifei Chen 0021, Deying Kong, Liangjian Chen, Xingwei Liu, Xiangyi Yan, Hao Tang 0010, Xiaohui Xie |
ECCV (5) | 7 |
| 2022 | AFTer-UNet: Axial Fusion Transformer UNet for Medical Image SegmentationabstractRecent advances in transformer-based models have drawn attention to exploring these techniques in medical image segmentation, especially in conjunction with the UNet model (or its variants), which has shown great success in medical image segmentation, under both 2D and 3D settings. Current 2D based methods either directly replace convolutional layers with pure transformers or consider a transformer as an additional intermediate encoder between the encoder and decoder of U-Net. However, these approaches only consider the attention encoding within one single slice and do not utilize the axial-axis information naturally provided by a 3D volume. In the 3D setting, convolution on volumetric data and transformers both consume large GPU memory. One has to either downsample the image or use cropped local patches to reduce GPU memory usage, which limits its performance. In this paper, we propose Axial Fusion Transformer UNet (AFTer-UNet), which takes both advantages of convolutional layers’ capability of extracting detailed features and transformers’ strength on long sequence modeling. It considers both intra-slice and inter-slice long-range cues to guide the segmentation. Meanwhile, it has fewer parameters and takes less GPU memory to train than the previous transformer-based models. Extensive experiments on three multi-organ segmentation datasets demonstrate that our method outperforms current state-of-the-art methods. Xiangyi Yan, Hao Tang 0010, Shanlin Sun, Deying Kong, Xiaohui Xie |
WACV | 1 |
| 2021 | TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
Liangjian Chen, Deying Kong, Xingwei Liu, Hao Tang 0010, Xiangyi Yan, Yusheng Xie, Shih-Yao Lin 0001, Xiaohui Xie |
BMVC | 7 |
| 2021 | Recurrent Mask Refinement for Few-Shot Medical Image SegmentationabstractAlthough having achieved great success in medical image segmentation, deep convolutional neural networks usually require a large dataset with manual annotations for training and are difficult to generalize to unseen classes. Few-shot learning has the potential to address these challenges by learning new classes from only a few labeled examples. In this work, we propose a new framework for few-shot medical image segmentation based on prototypical networks. Our innovation lies in the design of two key modules: 1) a context relation encoder (CRE) that uses correlation to capture local relation features between foreground and background regions; and 2) a recurrent mask refinement module that repeatedly uses the CRE and a prototypical network to recapture the change of context relationship and refine the segmentation mask iteratively. Experiments on two abdomen CT datasets and an abdomen MRI dataset show the proposed method obtains substantial improvement over the state-of-the-art methods by an average of 16.32%, 8.45% and 6.24% in terms of DSC, respectively. Code is publicly available1. Hao Tang 0010, Xingwei Liu, Shanlin Sun, Xiangyi Yan, Xiaohui Xie |
ICCV | 4 |
| 2020 | Nonparametric Structure Regularization Machine for 2D Hand Pose EstimationabstractHand pose estimation is more challenging than body pose estimation due to severe articulation, self-occlusion and high dexterity of the hand. Current approaches often rely on a popular body pose algorithm, such as the Convolutional Pose Machine (CPM), to learn 2D keypoint features. These algorithms cannot adequately address the unique challenges of hand pose estimation, because they are trained solely based on keypoint positions without seeking to explicitly model structural relationship between them. We propose a novel Nonparametric Structure Regularization Machine (NSRM) for 2D hand pose estimation, adopting a cascade multi-task architecture to learn hand structure and keypoint representations jointly. The structure learning is guided by synthetic hand mask representations, which are directly computed from keypoint positions, and is further strengthened by a novel probabilistic representation of hand limbs and an anatomically inspired composition strategy of mask synthesis. We conduct extensive studies on two public datasets - OneHand 10k and CMU Panoptic Hand. Experimental results demonstrate that explicitly enforcing structure learning consistently improves pose estimation accuracy of CPM baseline models, by 1.17% on the first dataset and 4.01% on the second one. The implementation and experiment code is freely available online1. Our proposal of incorporating structural learning to hand pose estimation requires no additional training information, and can be a generic add-on module to other pose estimation models. Yifei Chen 0021, Deying Kong, Xiangyi Yan, Jianbao Wu, Xiaohui Xie |
WACV | 4 |
| 2019 | Adaptive Graphical Model Network for 2D Handpose Estimation
Deying Kong, Yifei Chen 0021, Xiangyi Yan, Xiaohui Xie |
BMVC | 4 |