Zhixin Xu

dblp:173/2964 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic Gaussian Scene Reconstruction from Unsynchronized Videos
abstract
Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 3D Gaussian Splatting have demonstrated impressive capabilities in dynamic scene reconstruction, they typically rely on the assumption that input video streams are temporally synchronized. However, in real-world scenarios, this assumption often fails due to factors like camera trigger delays, frame rate discrepancies, or independent recording setups, leading to temporal misalignment across views and reduced reconstruction quality. To address this challenge, a novel temporal alignment strategy is proposed for high-quality 4DGS reconstruction from unsynchronized multi-view videos. Our method features a coarse-to-fine alignment module that estimates and compensates for each camera's time shift. The method first determines a coarse, frame-level offset and then refines it to achieve sub-frame accuracy. This strategy can be integrated as a plug-and-play module into existing 4DGS frameworks, enhancing their robustness when handling asynchronous data. Experiments show that this approach effectively processes temporally misaligned videos and significantly enhances baseline methods.
Zhixin Xu, Hengyu Zhou, Yuan Liu 0025, Wenhan Xue, Hao Pan 0001, Wenping Wang 0001, Bin Wang 0021
AAAI1
2026 D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
abstract
Free-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent compression and customized motion formats, limiting generalization and standardization. To address this, we propose D-FCGS, a novel Feedforward Compression framework for Dynamic Gaussian Splatting. Key innovations include: (1) a standardized Group-of-Frames (GoF) structure with I-P coding, leveraging sparse control points to extract inter-frame motion tensors; (2) a dual prior-aware entropy model that fuses hyperprior and spatial-temporal priors for accurate rate estimation; (3) a control-point-guided motion compensation mechanism and refinement network to enhance view-consistent fidelity. Trained on Gaussian frames derived from multi-view videos, D-FCGS generalizes across diverse scenes in a zero-shot fashion. Experiments show that it matches the rate-distortion performance of optimization-based methods, achieving over 40 times compression compared to the baseline while preserving visual quality across viewpoints. This work advances feedforward compression of dynamic 3DGS, facilitating scalable FVV transmission and storage for immersive applications.
Yan Zhao 0041, Qiang Wang 0061, Zhixin Xu, Li Song 0001, Zhengxue Cheng
AAAI4
2026 Delay-Guaranteed Multi-Satellite Communication System
Qi Duan, Changxin Shi, Zhixin Xu, Feng Yang 0006
ICC3
2025 Bio-IL: A Robust Decentralized Biometric Recognition System Using Isomerism Learning with Heterogeneous Models on Private Blockchain
abstract
With the rapid adoption of biometric authentication technologies, safeguarding users’ sensitive biometric data has become increasingly critical. Traditional centralized training methods pose significant risks due to exposure of raw biometric data. Federated learning offers a decentralized alternative but faces challenges from data and model heterogeneity, which can degrade performance and compromise privacy. To overcome these challenges, this work proposes Bio-IL, a decentralized and robust Biometric recognition system that integrates Isomerism Learning with heterogeneous models over a private blockchain infrastructure. The use of private blockchain ensures secure and tamper-resistant coordination of distributed training while safe-guarding data privacy. At the core of Bio-IL lies the novel IsoFus aggregation algorithm, an Isomerism Learning-based Fusion method that effectively combines heterogeneous local models and significantly improves recognition accuracy in realistic decentralized settings. Extensive experiments, with face recognition as a representative application, demonstrate that the proposed model outperforms baseline methods in both accuracy and robustness. This framework offers a scalable, secure, and adaptable solution for decentralized biometric authentication, making it well-suited for practical identity verification tasks.
Zhihao Hao, Zhixin Xu, Junping Du 0001
IJCB4
2025 Mask-guided Cross Palm Attention Network for Palmprint Image Super-Resolution
abstract
Palmprint has shown great potential for biometric recognition due to its high user-friendliness and low invasiveness. However, most existing palmprint studies focus solely on feature learning and matching without considering the quality of the images, while the palmprint images collected in contactless scenarios are usually low-quality with complex backgrounds, significantly degrading recognition performance. In this paper, we propose a mask-guided cross palm attention network (MCPAnet) for complex palm-print image super-resolution. It first applies multi-scale pyramid feature aggregation to detect the palm-specific regions from complex palmprint images containing various non-palm backgrounds. With the guidance of the palm-specific region masks, we develop multiple residual-based learning groups to exploit the palmprint intrinsic features through channel-wise and spatial-wise attention interaction for high-resolution palmprint image reconstruction. Moreover, we design palmpix and perceptual losses to make the HR palmprint images realistic at the visual level while simultaneously preserving the identity-aware features as the ground-truth ones. Experimental results on four public palmprint benchmarks clearly show the effectiveness of the proposed method for palmprint image super-resolution.
Kaiting Huang, Zhixin Xu, Yao Wang 0012, Lunke Fei, Jinrong Cui
IJCB2
2025 CCNet: A Cross-Channel Enhanced CNN for Blind Image Denoising
abstract
ABSTRACT Nowadays, blind image denoising with deep convolutional neural network (CNN) is one of the research hotspots in the field of image denoising. Relying on the convolutional operation and respective field, CNN is excellent in processing local information. However, this also brings the problems of lack of cross‐domain interaction and process of global feature information. We incorporate it with transformer's self‐attention mechanism and propose a cross‐channel enhanced CNN, namely CCNet, for image denoising. CCNet consists of three parts: the backbone encoder (BE), the cross‐channel enhancer (CCE), and the backbone decoder (BD). The BE and BD are constructed using the multiscale symmetric network U‐Net and use the residual connection block (RCB) as the basic block for image feature extraction and reconstruction. CCE introduces transposed attention, serving as a complement of BE for cross‐channel modeling. Meanwhile, we propose a unique gated fusion block (GFB) to fuse the information of these two modules and further feature learning. To improve training, we use random cropping, shuffling, and mixed noise strategies to expand the noise distribution learned by the model, increasing its noise adaptability. Extensive experiments on grayscale images, color images, and real noisy images demonstrate CCNet's strong performance in these tasks.
Minling Zhu, Zhixin Xu
Comput. Intell.2
2025 GCSTormer: Gated swin transformer with channel weights for image denoising
Minling Zhu, Zhixin Xu, Yonglin Liu, Dongbing Gu, Sendren Sheng-Dong Xu
Expert Syst. Appl.2
2025 Piecewise Ruled Approximation for Freeform Mesh Surfaces
abstract
A ruled surface is a shape swept out by moving a line in 3D space. Due to their simple geometric forms, ruled surfaces have applications in various domains such as architecture and engineering. In the past, various approaches have been proposed to approximate a target shape using developable surfaces, which are special ruled surfaces with zero Gaussian curvature. However, methods for shape approximation using general ruled surfaces remain limited and often require the target shape to be either represented as parametric surfaces or have non-positive Gaussian curvature. In this paper, we propose a method to compute a piecewise ruled surface that approximates an arbitrary freeform mesh surface. We first use a group-sparsity formulation to optimize the given mesh shape into an approximately piecewise ruled form, in conjunction with a tangent vector field that indicates the ruling directions. Afterward, we utilize the optimization result to extract seams that separate smooth families of rulings, and use the seams to construct the initial rulings. Finally, we further optimize the positions and orientations of the rulings to improve the alignment with the input target shape. We apply our method to a variety of freeform shapes with different topologies and complexity, demonstrating its effectiveness in approximating arbitrary shapes.
Yiling Pan, Zhixin Xu, Bin Wang 0021, Bailin Deng
ACM Trans. Graph.2
2024 CAWM: Class-Aware Weight Map for Improved Semi-Supervised Nuclei Segmentation
abstract
Due to the rich histopathological information of nuclei in whole slide images, nuclei segmentation becomes essential for medical analysis. Since collecting sufficient pixel-wise annotations for supervised training of nuclei segmentation networks is challenging, semi-supervised nuclei segmentation methods have been extensively studied. In particular, many of them use pseudo-labels generated from unlabeled images for training the segmentation model. In this Letter, we propose a new pseudo-label handling method for semi-supervised nuclei segmentation. Specifically, based on our observation that nuclear features within the same image share high similarities, we define confidence maps for pseudo-labels and use them to adapt consistency regularization and contrastive loss measures. From extensive experiments on three public datasets, we demonstrate the effectiveness of the proposed method compared with other semi-supervised training methods.
Seohoon Lim, Zhixin Xu, Yosep Chong, Seung-Won Jung
IEEE Signal Process. Lett.2
2024 Cross-Domain Denoising for Low-Dose Multi-Frame Spiral Computed Tomography
abstract
Computed tomography (CT) has been used worldwide as a non-invasive test to assist in diagnosis. However, the ionizing nature of X-ray exposure raises concerns about potential health risks such as cancer. The desire for lower radiation doses has driven researchers to improve reconstruction quality. Although previous studies on low-dose computed tomography (LDCT) denoising have demonstrated the effectiveness of learning-based methods, most were developed on the simulated data. However, the real-world scenario differs significantly from the simulation domain, especially when using the multi-slice spiral scanner geometry. This paper proposes a two-stage method for the commercially available multi-slice spiral CT scanners that better exploits the complete reconstruction pipeline for LDCT denoising across different domains. Our approach makes good use of the high redundancy of multi-slice projections and the volumetric reconstructions while leveraging the over-smoothing issue in conventional cascaded frameworks caused by aggressive denoising. The dedicated design also provides a more explicit interpretation of the data flow. Extensive experiments on various datasets showed that the proposed method could remove up to 70% of noise without compromised spatial resolution, while subjective evaluations by two experienced radiologists further supported its superior performance against state-of-the-art methods in clinical practice. Code is available at https://github.com/YCL92/TMD-LDCT.
Yucheng Lu 0001, Zhixin Xu, Moon Hyung Choi, Seung-Won Jung
IEEE Trans. Medical Imaging2
2020 Filter Grafting for Deep Neural Networks
abstract
This paper proposes a new learning paradigm called filter grafting, which aims to improve the representation capability of Deep Neural Networks (DNNs). The motivation is that DNNs have unimportant (invalid) filters (e.g., l1norm close to 0). These filters limit the potential of DNNs since they are identified as having little effect on the network. While filter pruning removes these invalid filters for efficiency consideration, filter grafting re-activates them from an accuracy boosting perspective. The activation is processed by grafting external information (weights) into invalid filters. To better perform the grafting process, we develop an entropy-based criterion to measure the information of filters and an adaptive weighting strategy for balancing the grafted information among networks. After the grafting operation, the network has very few invalid filters compared with its untouched state, empowering the model with more representation capacity. We also perform extensive experiments on the classification and recognition tasks to show the superiority of our method. For example, the grafted MobileNetV2 outperforms the non-grafted MobileNetV2 by about 7 percent on CIFAR-100 dataset.
Fanxu Meng 0003, Hao Cheng 0012, Ke Li 0015, Zhixin Xu, Rongrong Ji, Xing Sun 0001, Guangming Lu 0002
CVPR4
2019 Single Image Reflection Removal Based on Deep Residual Learning
Zhixin Xu, Xiaobao Guo, Guangming Lu 0002
PRCV (2)1
2016 Incremental regularized extreme learning machine and it's enhancement
Zhixin Xu, Zhaohui Wu 0001, Weihui Dai
Neurocomputing1