EDBT 2026 Demo / reviewers in the wild / expert
Shanlin Sun
dblp:141/8363
· DBLP profile ↗
23ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-4519-7287ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoMA: Compositional Human Motion Generation with Multi-modal Agentsabstract3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely due to the scarcity of motion datasets and the prohibitive cost of generating new training examples. To address these challenges, we introduce CoMA, an agent-based solution for complex human motion generation, editing, and comprehension. CoMA leverages multiple collaborative agents powered by large language and vision models, alongside a mask transformer-based motion generator featuring body part-specific encoders and codebooks for fine-grained control. Our framework enables generation of both short and long motion sequences with detailed instructions, text-guided motion editing, and self-correction for improved quality. Evaluations on the HumanML3D dataset demonstrate competitive performance against state-of-the-art methods. Additionally, we create a set of context-rich, compositional, and long text prompts, where user studies show our method significantly outperforms existing approaches. Shanlin Sun, Gabriel de Araujo, Shenghan Zhou, Ziheng Huang 0001, Chenyu You, Xiaohui Xie |
AAAI | 1 |
| 2025 | Ouroboros: Single-Step Diffusion Models for Cycle-Consistent Forward and Inverse RenderingabstractWhile multi-step diffusion models have advanced both forward and inverse rendering, existing approaches often treat these problems independently, leading to cycle inconsistency and slow inference speed. In this work, we present Ouroboros, a framework composed of two single-step diffusion models that handle forward and inverse rendering with mutual reinforcement. Our approach extends intrinsic decomposition to both indoor and outdoor scenes and introduces a cycle consistency mechanism that ensures coherence between forward and inverse rendering outputs. Experimental results demonstrate state-of-the-art performance across diverse scenes while achieving substantially faster inference speed compared to other diffusion-based methods. We also demonstrate that Ouroboros can transfer to video decomposition in a training-free manner, reducing temporal inconsistency in video sequences while maintaining high-quality per-frame inverse rendering. Shanlin Sun, Yifeng Xiong, Ruogu Fang, Xiaohui Xie, Chenyu You |
ICCV | 1 |
| 2024 | Hybrid-CSR: Coupling Explicit and Implicit Reconstruction of Cortical Surface
Shanlin Sun, Pooya Khosravi, Chenyu You, Deying Kong, Xiangyi Yan, Xiaohui Xie |
BMVC | 1 |
| 2024 | Integrating Efficient Optimal Transport and Functional Maps for Unsupervised Shape Correspondence LearningabstractIn the realm of computer vision and graphics, accurately establishing correspondences between geometric 3D shapes is pivotal for applications like object tracking, registration, texture transfer, and statistical shape analysis. Moving beyond traditional hand-crafted and data-driven feature learning methods, we incorporate spectral methods with deep learning, focusing on functional maps (FMs) and optimal transport (OT). Traditional OT-based approaches, often reliant on entropy regularization OT in learning-based framework, face computational challenges due to their quadratic cost. Our key contribution is to employ the sliced Wasserstein distance (SWD)for OT, which is a valid fast optimal transport metric in an unsupervised shape matching framework. This unsupervised framework integrates functional map regularizers with a novel OT-based loss derived from SWD, enhancing feature alignment between shapes treated as discrete probability measures. We also introduce an adaptive refinement process utilizing entropy regularized OT, further refining feature alignments for accurate point-to-point correspondences. Our method demonstrates superior performance in non-rigid shape matching, including near-isometric and non-isometric scenarios, and excels in downstream tasks like segmentation transfer. The empirical results on diverse datasets highlight our framework's effectiveness and generalization capabilities, setting new standards in non-rigid shape matching with efficient OT metrics and an adaptive refinement module. Code is available at11https://github.com/Tungthanhlee/EOT-Correspondence. Shanlin Sun, Nhat Ho, Xiaohui Xie |
CVPR | 3 |
| 2024 | LidaRF: Delving into Lidar for Neural Radiance Field on Street ScenesabstractPhotorealistic simulation plays a crucial role in applications such as autonomous driving, where advances in neural radiance fields (NeRFs) may allow better scalability through the automatic creation of digital 3D assets. However, reconstruction quality suffers on street scenes due to largely collinear camera motions and sparser samplings at higher speeds. On the other hand, the application often demands rendering from camera views that deviate from the inputs to accurately simulate behaviors like lane changes. In this paper, we propose several insights that allow a better utilization of Lidar data to improve NeRF quality on street scenes. First, our framework learns a geometric scene representation from Lidar, which are fused with the implicit grid-based representation for radiance decoding, thereby supplying stronger geometric information offered by explicit point cloud. Sec-ond, we put forth a robust occlusion-aware depth supervision scheme, which allows utilizing densified Lidar points by accumulation. Third, we generate augmented training views from Lidar points for further improvement. Our insights translate to largely improved novel view synthesis under real driving scenes. Shanlin Sun, Bingbing Zhuang, Ziyu Jiang, Buyu Liu, Xiaohui Xie, Manmohan Krishna Chandraker |
CVPR | 1 |
| 2024 | Diffeomorphic Mesh Deformation via Efficient Optimal Transport for Cortical Surface ReconstructionabstractMesh deformation plays a pivotal role in many 3D vision tasks including dynamic simulations, rendering, and reconstruction. However, defining an efficient discrepancy between predicted and target meshes remains an open problem. A prevalent approach in current deep learning is the set-based approach which measures the discrepancy between two surfaces by comparing two randomly sampled point-clouds from the two meshes with Chamfer pseudo-distance. Nevertheless, the set-based approach still has limitations such as lacking a theoretical guarantee for choosing the number of points in sampled point-clouds, and the pseudo-metricity and the quadratic complexity of the Chamfer divergence. To address these issues, we propose a novel metric for learning mesh deformation. The metric is defined by sliced Wasserstein distance on meshes represented as probability measures that generalize the set-based approach. By leveraging probability measure space, we gain flexibility in encoding meshes using diverse forms of probability measures, such as continuous, empirical, and discrete measures via \textit{varifold} representation. After having encoded probability measures, we can compare meshes by using the sliced Wasserstein distance which is an effective optimal transport distance with linear computational complexity and can provide a fast statistical rate for approximating the surface of meshes. To the end, we employ a neural ordinary differential equation (ODE) to deform the input surface into the target shape by modeling the trajectories of the points on the surface. Our experiments on cortical surface reconstruction demonstrate that our approach surpasses other competing methods in multiple datasets and metrics. Thanh-Tung Le, Shanlin Sun, Nhat Ho, Xiaohui Xie |
ICLR | 3 |
| 2024 | Hybrid Neural Diffeomorphic Flow for Shape Representation and Generation via TriplaneabstractDeep Implicit Functions (DIFs) have gained popularity in 3D computer vision due to their compactness and continuous representation capabilities. However, addressing dense correspondences and semantic relationships across DIF-encoded shapes remains a critical challenge, limiting their applications in texture transfer and shape analysis. Moreover, recent endeavors in 3D shape generation using DIFs often neglect correspondence and topology preservation. This paper presents HNDF (Hybrid Neural Diffeomorphic Flow), a method that implicitly learns the underlying representation and decomposes intricate dense correspondences into explicitly axis-aligned triplane features. To avoid suboptimal representations trapped in local minima, we propose hybrid supervision that captures both local and global correspondences. Unlike conventional approaches that directly generate new 3D shapes, we further explore the idea of shape generation with deformed template shape via diffeomorphic flows, where the deformation is encoded by the generated triplane features. Leveraging a pre-existing 2D diffusion model, we produce high-quality and diverse 3D diffeomorphic flows through generated triplanes features, ensuring topological consistency with the template shape. Extensive experiments on medical image organ segmentation datasets evaluate the effectiveness of HNDF in 3D shape representation and generation. Shanlin Sun, Thanh-Tung Le, Xiangyi Yan, Chenyu You, Xiaohui Xie |
WACV | 2 |
| 2024 | CVTHead: One-shot Controllable Head Avatar with Vertex-feature TransformerabstractReconstructing personalized animatable head avatars has significant implications in the fields of AR/VR. Existing methods for achieving explicit face control of 3D Morphable Models (3DMM) typically rely on multi-view images or videos of a single subject, making the reconstruction process complex. Additionally, the traditional rendering pipeline is time-consuming, limiting real-time animation possibilities. In this paper, we introduce CVTHead, a novel approach that generates controllable neural head avatars from a single reference image using point-based neural rendering. CVT-Head considers the sparse vertices of mesh as the point set and employs the proposed Vertex-feature Transformer to learn local feature descriptors for each vertex. This enables the modeling of long-range dependencies among all the vertices. Experimental results on the VoxCeleb dataset demonstrate that CVTHead achieves comparable performance to state-of-the-art graphics-based methods. Moreover, it enables efficient rendering of novel human heads with various expressions, head poses, and camera views. These attributes can be explicitly controlled using the coefficients of 3DMMs, facilitating versatile and realistic animation in real-time scenarios. Codes and pre-trained model can be found at https://github.com/HowieMa/CVTHead. Shanlin Sun, Xiangyi Yan, Xiaohui Xie |
WACV | 3 |
| 2024 | AFTer-SAM: Adapting SAM with Axial Fusion Transformer for Medical Imaging SegmentationabstractThe Segmentation Anything Model (SAM) has demonstrated effectiveness in various segmentation tasks. However, its application to 3D medical data has posed challenges due to its inherent design for both 2D and natural images. While there have been attempts to apply SAM to medical images on a slice-by-slice basis, the outcomes have been less than optimal. In this study, we introduce AFTer-SAM, an adaptation of SAM designed for volumetric medical image segmentation. By incorporating an Axial Fusion Transformer, AFTer-SAM is capable of capturing both intra-slice details and inter-slice contextual information, essential for accurate medical image segmentation. Given the potential computational challenges of training this enhanced model, we utilize Low-Rank Adaptation (LoRA) to efficiently finetune the weights of the Axial Fusion Transformer. This ensures a streamlined training process without compromising on performance. Our results indicate that AFTer-SAM offers significant improvements in volumetric medical image segmentation, suggesting a promising direction for the application of large pre-trained models in medical imaging. Xiangyi Yan, Shanlin Sun, Thanh-Tung Le, Chenyu You, Xiaohui Xie |
WACV | 2 |
| 2024 | Medical image registration via neural fieldsabstractImage registration is an essential step in many medical image analysis tasks. Traditional methods for image registration are primarily optimization-driven, finding the optimal deformations that maximize the similarity between two images. Recent learning-based methods, trained to directly predict transformations between two images, run much faster, but suffer from performance deficiencies due to domain shift. Here we present a new neural network based image registration framework, called NIR (Neural Image Registration), which is based on optimization but utilizes deep neural networks to model deformations between image pairs. NIR represents the transformation between two images with a continuous function implemented via neural fields, receiving a 3D coordinate as input and outputting the corresponding deformation vector. NIR provides two ways of generating deformation field: directly output a displacement vector field for general deformable registration, or output a velocity vector field and integrate the velocity field to derive the deformation field for diffeomorphic image registration. The optimal registration is discovered by updating the parameters of the neural field via stochastic mini-batch gradient descent. We describe several design choices that facilitate model optimization, including coordinate encoding, sinusoidal activation, coordinate sampling, and intensity sampling. NIR is evaluated on two 3D MR brain scan datasets, demonstrating highly competitive performance in terms of both registration accuracy and regularity. Compared to traditional optimization-based methods, our approach achieves better results in shorter computation times. In addition, our methods exhibit performance on a cross-dataset registration task, compared to the pre-trained learning-based methods. Shanlin Sun, Chenyu You, Hao Tang 0010, Deying Kong, Junayed Naushad, Xiangyi Yan, Pooya Khosravi, James S. Duncan, Xiaohui Xie |
Medical Image Anal. | 1 |
| 2023 | MedGen3D: A Deep Generative Framework for Paired 3D Image and Mask Generation
Yifeng Xiong, Chenyu You, Pooya Khosravi, Shanlin Sun, Xiangyi Yan, James S. Duncan, Xiaohui Xie |
MICCAI (1) | 5 |
| 2023 | Localized Region Contrast for Enhancing Self-supervised Learning in Medical Image Segmentation
Xiangyi Yan, Junayed Naushad, Chenyu You, Hao Tang 0010, Shanlin Sun, James S. Duncan, Xiaohui Xie |
MICCAI (2) | 5 |
| 2023 | Diffeomorphic Image Registration with Neural Velocity FieldabstractDiffeomorphic image registration, offering smooth transformation and topology preservation, is required in many medical image analysis tasks. Traditional methods impose certain modeling constraints on the space of admissible transformations and use optimization to find the optimal transformation between two images. Specifying the right space of admissible transformations is challenging: the registration quality can be poor if the space is too restrictive, while the optimization can be hard to solve if the space is too general. Recent learning-based methods, utilizing deep neural networks to learn the transformation directly, achieve fast inference, but face challenges in accuracy due to the difficulties in capturing the small local deformations and generalization ability. Here we propose a new optimization-based method named DNVF (Diffeomorphic Image Registration with Neural Velocity Field) which utilizes deep neural network to model the space of admissible transformations. A multilayer perceptron (MLP) with sinusoidal activation function is used to represent the continuous velocity field and assigns a velocity vector to every point in space, providing the flexibility of modeling complex deformations as well as the convenience of optimization. Moreover, we propose a cascaded image registration framework (Cas-DNVF) by combining the benefits of both optimization and learning based methods, where a fully convolutional neural network (FCN) is trained to predict the initial deformation, followed by DNVF for further refinement. Experiments on two large-scale 3D MR brain scan datasets demonstrate that our proposed methods significantly outperform the state-of-the-art registration methods. Shanlin Sun, Xiangyi Yan, Chenyu You, Hao Tang 0010, Junayed Naushad, Deying Kong, Xiaohui Xie |
WACV | 2 |
| 2023 | Representation Recovering for Self-Supervised Pre-training on Medical ImagesabstractAdvances in self-supervised learning have drawn attention to developing techniques to extract effective visual representations from unlabeled images. Contrastive learning (CL) trains a model to extract consistent features by generating different views. Recent success of Masked Autoencoders (MAE) highlights the benefit of generative modeling in self-supervised learning. The generative approaches encode the input into a compact embedding and empower the model’s ability of recovering the original input. However, in our experiments, we found vanilla MAE mainly recovers coarse high level semantic information and is inadequate in recovering detailed low level information. We show that in dense downstream prediction tasks like multi-organ segmentation, directly applying MAE is not ideal. Here, we propose RepRec, a hybrid visual representation learning framework for self-supervised pre-training on large-scale unlabelled medical datasets, which takes advantage of both contrastive and generative modeling. To solve the aforementioned dilemma that MAE encounters, a convolutional encoder is pre-trained to provide low-level feature information, in a contrastive way; and a transformer encoder is pre-trained to produce high level semantic dependency, in a generative way – by recovering masked representations from the convolutional encoder. Extensive experiments on three multi-organ segmentation datasets demonstrate that our method outperforms current state-of-the-art methods. Xiangyi Yan, Junayed Naushad, Shanlin Sun, Hao Tang 0010, Deying Kong, Chenyu You, Xiaohui Xie |
WACV | 3 |
| 2022 | Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic FlowabstractDeep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences simultaneously, capturing semantic relationships across shapes of the same class by learning a DIFs-modeled shape template. These methods provide great flexibility and accuracy in reconstructing 3D shapes and inferring correspondences. However, the point correspondences built from these methods do not intrinsically preserve the topology of the shapes, unlike mesh-based template matching methods. This limits their applications on 3D geometries where underlying topological structures exist and matter, such as anatomical structures in medical images. In this paper, we propose a new model called Neural Diffeomorphic Flow (NDF) to learn deep implicit shape templates, representing shapes as conditional diffeomorphic deformations of templates, intrinsically preserving shape topologies. The diffeomorphic deformation is realized by an autodecoder consisting of Neural Ordinary Differential Equation (NODE) blocks that progressively map shapes to implicit templates. We conduct extensive experiments on several medical image organ segmentation datasets to evaluate the effectiveness of NDF on reconstructing and aligning shapes. NDF achieves consistently state-of-the-art organ shape reconstruction and registration results in both accuracy and quality. The source code is publicly available at https://github.com/Siwensun/Neural_Diffeomorphic_Flow-NDF. Shanlin Sun, Deying Kong, Hao Tang 0010, Xiangyi Yan, Xiaohui Xie |
CVPR | 1 |
| 2022 | Identity-Aware Hand Mesh Estimation and Personalization from RGB Images
Deying Kong, Linguang Zhang, Liangjian Chen, Xiangyi Yan, Shanlin Sun, Xingwei Liu, Xiaohui Xie |
ECCV (5) | 6 |
| 2022 | AFTer-UNet: Axial Fusion Transformer UNet for Medical Image SegmentationabstractRecent advances in transformer-based models have drawn attention to exploring these techniques in medical image segmentation, especially in conjunction with the UNet model (or its variants), which has shown great success in medical image segmentation, under both 2D and 3D settings. Current 2D based methods either directly replace convolutional layers with pure transformers or consider a transformer as an additional intermediate encoder between the encoder and decoder of U-Net. However, these approaches only consider the attention encoding within one single slice and do not utilize the axial-axis information naturally provided by a 3D volume. In the 3D setting, convolution on volumetric data and transformers both consume large GPU memory. One has to either downsample the image or use cropped local patches to reduce GPU memory usage, which limits its performance. In this paper, we propose Axial Fusion Transformer UNet (AFTer-UNet), which takes both advantages of convolutional layers’ capability of extracting detailed features and transformers’ strength on long sequence modeling. It considers both intra-slice and inter-slice long-range cues to guide the segmentation. Meanwhile, it has fewer parameters and takes less GPU memory to train than the previous transformer-based models. Extensive experiments on three multi-organ segmentation datasets demonstrate that our method outperforms current state-of-the-art methods. Xiangyi Yan, Hao Tang 0010, Shanlin Sun, Deying Kong, Xiaohui Xie |
WACV | 3 |
| 2021 | Recurrent Mask Refinement for Few-Shot Medical Image SegmentationabstractAlthough having achieved great success in medical image segmentation, deep convolutional neural networks usually require a large dataset with manual annotations for training and are difficult to generalize to unseen classes. Few-shot learning has the potential to address these challenges by learning new classes from only a few labeled examples. In this work, we propose a new framework for few-shot medical image segmentation based on prototypical networks. Our innovation lies in the design of two key modules: 1) a context relation encoder (CRE) that uses correlation to capture local relation features between foreground and background regions; and 2) a recurrent mask refinement module that repeatedly uses the CRE and a prototypical network to recapture the change of context relationship and refine the segmentation mask iteratively. Experiments on two abdomen CT datasets and an abdomen MRI dataset show the proposed method obtains substantial improvement over the state-of-the-art methods by an average of 16.32%, 8.45% and 6.24% in terms of DSC, respectively. Code is publicly available1. Hao Tang 0010, Xingwei Liu, Shanlin Sun, Xiangyi Yan, Xiaohui Xie |
ICCV | 3 |
| 2021 | Spatial Context-Aware Self-Attention Model For Multi-Organ SegmentationabstractMulti-organ segmentation is one of most successful applications of deep learning in medical image analysis. Deep convolutional neural nets (CNNs) have shown great promise in achieving clinically applicable image segmentation performance on CT or MRI images. State-of-the-art CNN segmentation models apply either 2D or 3D convolutions on input images, with pros and cons associated with each method: 2D convolution is fast, less memory-intensive but inadequate for extracting 3D contextual information from volumetric images, while the opposite is true for 3D convolution. To fit a 3D CNN model on CT or MRI images on commodity GPUs, one usually has to either downsample input images or use cropped local regions as inputs, which limits the utility of 3D models for multi-organ segmentation. In this work, we propose a new framework for combining 3D and 2D models, in which the segmentation is realized through high-resolution 2D convolutions, but guided by spatial contextual information extracted from a low-resolution 3D model. We implement a self-attention mechanism to control which 3D features should be used to guide 2D segmentation. Our model is light on memory usage but fully equipped to take 3D contextual information into account. Experiments on multiple organ segmentation datasets demonstrate that by taking advantage of both 2D and 3D models, our method is consistently outperforms existing 2D and 3D models in organ segmentation accuracy, while being able to directly take raw whole-volume image data as inputs. Hao Tang 0010, Xingwei Liu, Xiaohui Xie, Xuming Chen, Shanlin Sun, Narisu Bai |
WACV | 8 |
| 2020 | 3D Model Retrieval Using Bipartite Graph Matching Based on Attention
Shanlin Sun, Yun Li 0006, Yunfeng Xie, Zhicheng Tan, Xing Yao, Rongyao Zhang |
Neural Process. Lett. | 1 |
| 2017 | Traffic lights detection and recognition based on multi-feature fusion
Shanlin Sun, Ming-Xin Jiang, Yunyang Yan, Xiaobing Chen |
Multim. Tools Appl. | 2 |
| 2016 | Resilient Control of Networked Control System Under DoS Attacks: A Unified Game ApproachabstractWe consider the problem of resilient control of networked control system (NCS) under denial-of-service (DoS) attack via a unified game approach. The DoS attacks lead to extra constraints in the NCS, where the packets may be jammed by a malicious adversary. Considering the attack-induced packet dropout, optimal control strategies with multitasking and central-tasking structures are developed using game theory in the delta domain, respectively. Based on the optimal control structures, we propose optimality criteria and algorithms for both cyber defenders and DoS attackers. Both simulation and experimental results are provided to illustrate the effectiveness of the proposed design procedure. Yuan Yuan 0006, Huanhuan Yuan, Lei Guo 0003, Hongjiu Yang, Shanlin Sun |
IEEE Trans. Ind. Informatics | 5 |
| 2013 | Asymptotic throughput analysis of random beamforming for multi-antenna cognitive broadcast networksabstractRandom beamforming (RBF) has received much attention recently in downlink beamforming because of its simple structure, low‐feedback load and same throughput scaling as that obtained using dirty paper coding at the transmitter. In this study, the authors analyse the performance of RBF for cognitive downlink multi‐antenna system in terms of the throughput of the secondary network. The authors consider a secondary broadcast station with multiple antennas utilising the licensed spectrum of each primary receiver to broadcast information to multiple secondary users (SUs) simultaneously, as long as the interference power inflicted at each primary receiver is less than a predefined threshold. Firstly, they derived the secondary network closed‐form throughput approximation of a single‐beam RBF by exploiting extreme value order statistics. Then, closed‐form approximation for secondary network throughput on multiple‐beam RBF is presented. Simulation results verify the validity of the authors approximation results analysis even with fewer SUs. Jianbo Ji, Wen Chen 0001, Shanlin Sun |
IET Commun. | 3 |