Weixing Xie

dblp:226/5771 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0009-0008-5558-0387ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MGD: Mesh-guided Gaussians with Diffusion Priors for Dynamic Objects Reconstruction from Monocular RGB-D Video
abstract
Reconstructing dynamic objects from monocular RGB-D video is critical for advancing 3D vision applications and enhancing user experience. However, monocular RGB-D video provides limited 3D observations, making the reconstruction of unobserved regions highly under-constrained. Despite recent advances that combine neural implicit surfaces with diffusion models, the inherent limitations of implicit representations and the lack of effective guidance in diffusion priors lead to blurry appearance and inaccurate geometry in dynamic object reconstruction. To address the issue, we present MGD, which leverages scene-adaptive diffusion priors and Mesh-guided Gaussians for realistic rendering and geometrically accurate reconstruction of dynamic objects, including unobserved regions. The dynamic 3D objects reconstructed by MGD are represented using our proposed Mesh-guided Gaussians, which leverage global and local Gaussians to capture large-scale deformations and fine-grained appearance details, respectively. Additionally, in order to utilize depth information, we integrate a depth ControlNet into the diffusion model and conduct scene-adaptive fine-tuning. We design a self-generated image-pair strategy to produce image pairs used for fine-tuning. Extensive experiments demonstrate that MGD achieves state-of-the-art performance in both high-fidelity reconstruction and structural completeness, while maintaining real-time efficiency during training and rendering.
Weixing Xie, Jintian Li, Bingchuan Li, Yanchen Lin, Junfeng Yao
AAAI1
2025 PD-SDF: Dynamic Surface Reconstruction Based on Plane Decomposition for Single View RGB-D Videos
abstract
Surface reconstruction of dynamic scenes from single view videos is a challenging task due to the highly ill-posed and under-constrained nature. Existing single view reconstruction methods suffer from severe quality issues, such as surface distortion and mesh adherison. In this paper, we propose an efficient dynamic representation network, PD-SDF, which consists of a 4D motion field and a 3D geometry field. The explicit disentanglement of motion and geometry based on planar factorization guarantees the mesh consistency across frames. Specifically, to address the mesh adhesion problem, we design a depth-guided sampling strategy to focus on optimizing the SDF field near the object surface. Due to insufficient geometric cues, we design various regularization strategies to constrain smoothness and topological correcness of the scene geometry. Extensive experiments show that our method outperforms existing methods in both appearance and geometry reconstruction. The project page: https://pd-sdf.github.io/
Weixing Xie, Junfeng Yao, Shaoqi Wu, Youhong Peng, Mengyuan Ge
ICASSP2
2025 LDG: Lightweight Deformable 3D Gaussians for Single-View Dynamic Scene Reconstruction
abstract
Recent deformable 3D Gaussians methods achieve high-quality reconstruction and real-time rendering. However, they require multi-view information and are not applicable to single-view dynamic scenes captured from mobile phones. Additionally, the high-dimensional hidden layer of deformation MLP and the excessive number of Gaussian primitives and attributes impose significant storage pressure, greatly limiting their practical application. To address the issues, we propose a novel Lightweight Deformable 3D Gaussians teacher-student framework. Specifically, we initialize Gaussian primitives with an initialization strategy designed for single-view scenes, and then optimize the teacher model using color and depth information. For the trained teacher model, we distill deformation MLP, prune Gaussian primitives and Gaussian attributes, and finally obtain a student model with low storage and high efficiency. Public benchmark experiments demonstrate the effectiveness of our framework, showing a compression rate exceeding 4× while maintaining satisfactory rendering quality. Project page: https://poyoki.github.io/ldg/.
Youhong Peng, Weixing Xie, Shaoqi Wu, Junfeng Yao
ICASSP2
2025 LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View Scenes
abstract
With the popularity of short video platforms, the number of single-view videos has increased significantly. Existing NeRF-based methods can reconstruct dynamic scenes in a single-view setting, but slow rendering speed and low rendering quality limit their practical applications. To address these challenges, we propose a fast single-view scenes reconstruction framework based on 3D Gaussian Splatting. Our method uses point clouds obtained with depth priors as the Gaussian initialization and introduces learnable parametric functions to model the time-dependent deformation of Gaussians. The explicit deformation modeling for Gaussians significantly reduces training and rendering time. Furthermore, to improve the rendering quality of challenging areas, we adopt an adaptive sampling strategy to densify Gaussians. For occlusion problems from single-view videos, we design a smooth loss function to restore the color of the occluded areas. Experimental results demonstrate that our method significantly reduces training time, enhances rendering quality, and accelerates rendering speed. Project page: https://github.com/LPGaussians.
Shaoqi Wu, Weixing Xie, Youhong Peng, Jiawei Yao, Junfeng Yao
ICASSP2
2025 RMAvatar: Photorealistic human avatar reconstruction from monocular video based on rectified mesh-embedded Gaussians
abstract
We introduce RMAvatar, a novel human avatar representation with Gaussian splatting embedded on mesh to learn clothed avatar from a monocular video. We utilize the explicit mesh geometry to represent motion and shape of a virtual human and implicit appearance rendering with Gaussian Splatting. Our method consists of two main modules: Gaussian initialization module and Gaussian rectification module. We embed Gaussians into triangular faces and control their motion through the mesh, which ensures low-frequency motion and surface deformation of the avatar. Due to the limitations of LBS formula, the human skeleton is hard to control complex non-rigid transformations. We then design a pose-related Gaussian rectification module to learn fine-detailed non-rigid deformations, further improving the realism and expressiveness of the avatar. We conduct extensive experiments on public datasets, and RMAvatar shows state-of-the-art performance on both rendering quality and quantitative evaluations. Please see our project page at https://rm-avatar.github.io .
Sen Peng, Weixing Xie, Xiaohu Guo, Zhonggui Chen, Baorong Yang
Graph. Model.2
2025 4D Gaussian Splatting for high-fidelity dynamic reconstruction of single-view scenes
Weixing Xie, Sen Peng, Yihang Fu, Wentao Fan 0001, Baorong Yang
Neurocomputing2
2025 Endo-HDR: Dynamic endoscopic reconstruction with deformable 3D Gaussians and hierarchical depth regularization
Weixing Xie, Qingqi Hong, Junfeng Yao, Shaoqi Wu, Rongzhou Zhou, Xiaohu Guo
Knowl. Based Syst.1
2024 S2CCT: Self-Supervised Collaborative CNN-Transformer for Few-shot Medical Image Segmentation
abstract
Self-supervised pre-training followed by fine-tuning is a potent paradigm for few-shot learning, leveraging extensive unlabeled data with remarkable efficacy. Current self-supervised methods often lean towards Vision Transformers (ViTs) rather than CNN-Transformer hybrid architectures, which generally demonstrate superior performance. However, this reliance on ViTs can lead to poor perception of local features by the model. The challenge lies in designing a suitable proxy task for hybrid architectures like CNN-Transformers, which have significant structural differences. Additionally, the current organization of CNN-Transformer hybrid backbones is often sequential, hindering collaboration during pre-training and the acquisition of robust representations. To address these issues, we propose Self-Supervised Collaborative CNN-Transformer (S2CCT) for few-shot medical image segmentation. This framework introduces three innovative designs: (1) a composite proxy task based on image masking and image super-resolution tailored for CNN-Transformer hybrid architectures, enabling the backbone to acquire robust representations during pre-training that can be transferred to downstream tasks; (2) a parallel CNN-Transformer architecture that better attends to multi-scale features in images, making it more suitable for dense prediction tasks like image segmentation; (3) a sparse and dense feature fusion module to enhance collaboration between the two encoders. Experiments demonstrate that S2CCT outperforms previous state-of-the-art methods on two public medical image segmentation benchmarks, i.e., ACDC and KiTs19. The code and pretrained models will be released soon.
Rongzhou Zhou, Ziqi Shu, Weixing Xie, Junfeng Yao, Qingqi Hong
BIBM3
2024 DRSM: Efficient Neural 4D Decomposition for Dynamic Reconstruction in Stationary Monocular Cameras
abstract
With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to scene reconstructions that exploit multi-view observations, the problem of modeling a dynamic scene from a single view is significantly more under-constrained and ill-posed. Inspired by recent progress in neural rendering, we present a novel framework to tackle 4D decomposition problem for dynamic scenes in monocular cameras. Our framework utilizes decomposed static and dynamic feature planes to represent 4D scenes and emphasizes the learning of dynamic regions through dense ray casting. Inadequate 3D clues from a single-view and occlusion are also particular challenges in scene reconstruction. To overcome these difficulties, we propose deep supervised optimization and ray casting strategies. With experiments on various videos, our method generates higher-fidelity results than existing methods for single-view dynamic scene representation.
Weixing Xie, Qiqin Lin, Jingze Chen, Junfeng Yao, Xiaohu Guo
ICASSP1
2024 ICR-Net: Semi-Supervised Medical Image Segmentation Guided By Intra-Sample Cross Reconstruction
abstract
Semi-supervised learning is becoming increasingly popular in medical image segmentation because of its ability to exploit large amounts of unlabelled data to extract additional information. However, most existing semi-supervised segmentation methods focus only on extracting information from unlabelled data, ignoring the potential of labelled data to further improve model performance. In this paper, we propose a new framework for Intra-Sample Cross Reconstruction Networks (ICR-Net) that utilises labelled data to help the network extract information from unlabelled data, thereby guiding the network’s regularisation learning. Our method contains two modules: Intra-Sample Cross Reconstruction (ICR) module and Synergistic Consistency Constraints (SCC) module. The ICR module processes the labelled data features in a more fine-grained manner, thus enabling the network to learn and capture the key patterns and features in the inputs more efficiently, and the SCC guides the network’s regularised learning by formulating additional model regularisations. Experiments on the LA dataset and the pancreas dataset show that our proposed framework is more effective than current state-of-the-art methods in medical image segmentation tasks.
Xianpeng Cao, Weixing Xie, Xianxing Cao, Qiqin Lin, Rongzhou Zhou, Junfeng Yao, Qingqi Hong
ICME2
2024 DPP-Net: Difficulty Perception-Processing Heterogeneous Network for Semi-supervised Medical Image Segmentation
abstract
In semi-supervised medical image segmentation, the scarcity of labeled data makes models prone to learning bias, causing persistent errors in certain regions and eventual over-fitting, significantly impacting segmentation performance. These problematic regions, termed difficult areas, are inadequately addressed by existing methods. To address this, We propose the Difficulty Perception-Processing Heterogeneous Network (DPP-Net). It guides the model in accurately perceiving and rectifying difficult areas, overcoming learning bias. Specifically, we introduce the Global Mutual Perception (GMP) to establish a comprehensive information perception channel between sample data, enabling a more holistic and accurate perception of difficult areas. The Difficulty-Aware Rectification (DAR) structure ensures continuous monitoring of difficult areas during training, allowing for timely adjustments to errors. Additionally, the Adaptive Competitive Pseudo-Label (ACP) Augmentation strategy enhances pseudo-labels through adaptive confidence competition. Experimental results on two different medical image databases (CT and MRI) demonstrate that our approach outperforms several state-of-the-art methods.
Qiqin Lin, Weixing Xie, Rongzhou Zhou, Xianpeng Cao, Jingze Chen, Junfeng Yao, Qingqi Hong
ICME2
2024 SurgicalGaussian: Deformable 3D Gaussians for High-Fidelity Surgical Scene Reconstruction
Weixing Xie, Junfeng Yao, Xianpeng Cao, Qiqin Lin, Zerui Tang, Xiaohu Guo
MICCAI (6)1
2024 A detail-preserving method for medial mesh computation in triangular meshes
abstract
The medial axis transform (MAT) of an object is the set of all points inside the object that have more than one closest point on the object’s boundary. Representing sharp edges and corners of triangular meshes using MAT poses a complex challenge. While some researchers have proposed using zero-radius medial spheres to depict these features, they have not clearly articulated how to establish proper connections among them. In this paper, we propose a novel framework for computing MAT of a triangular mesh while preserving its features. The initial medial axis mesh obtained may contain erroneous edges, which are discussed and addressed in Section 3.3. Furthermore, during the simplification process, it is crucial to ensure that the medial spheres remain within the confines of the triangular mesh. Our algorithm excels in preserving critical features throughout the simplification procedure, consistently ensuring that the spheres remain enclosed within the triangular mesh. Experiments on various types of 3D models demonstrate the robustness, shape fidelity, and efficiency in representation achieved by our algorithm.
Bingchuan Li, Yuping Ye, Junfeng Yao, Weixing Xie, Mengyuan Ge
Graph. Model.5
2023 LATrans-Unet: Improving CNN-Transformer with Location Adaptive for Medical Image Segmentation
Qiqin Lin, Junfeng Yao, Qingqi Hong, Xianpeng Cao, Rongzhou Zhou, Weixing Xie
PRCV (13)6