Xitie Zhang

dblp:353/9011 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-8114-9119ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Credible and Detailed 3D Face Reconstruction in Large Pose
abstract
The existing monocular methods face huge challenges in reconstructing credible details of non-visible areas in large pose images. Due to the fact that facial details are lost in non-visible areas of large pose images, existing methods lose basis when reconstructing details, resulting in unreliable results. Even if the generative model is used to repair the image first, the reconstruction process is very complex and costly. To this end, we propose an end-to-end and self-supervised RGB to depth method that equates small pose to large pose in UV space to obtain labels for non-visible areas. Then, we infer the depth values of non-visible areas and reconstruct a detailed 3D face. Finally, we render it into a face image and use it alongside the labels for self-supervised training of the network. During inference for large pose image, our method could reconstruct credible details of non-visible areas with a basis, rather than blindly. In addition, coarse and detailed reconstruction have mutually exclusive requirements for training images while coarse reconstruction is often not met, resulting in limited reconstruction accuracy. We propose a mutual exclusion elimination method to solve it, improving its accuracy. Extensive experiments demonstrate that our method could reconstruct precise 3D faces with credible details from large pose images. Our supplementary material is published on: https://github.com/lxy-nxu/ICASSP2025/tree/main
Xinyu Li 0014, Xitie Zhang, Suping Wu, Ruijie Peng, Kehua Ma
ICASSP2
2025 MSANet: Mixed Spectral and Attention Network for Robust 3D Human Pose Estimation
abstract
Despite significant advances in 3D human pose estimation from a single-view video, existing methods often struggle to produce reasonable human poses when the human is heavily occluded or blurred. To address this issue, we propose a Mixed Spectral and Attention Network (MSANet) that stacks spectral and attention blocks alternately. The attention block captures visual cues before and after occlusions or blurs, while the spectral block perceives subtle localized occlusions or blurs for robust 3D human pose estimation. Specifically, our attention block captures the global information of intra-frame joints and enhances the coherent representation of inter-frame joints, the spectral block complements the intra- and inter-frame local occlusions information which is difficult to capture by the attention block. In addition, we improve the regression head (IRH) to narrow the grained gap between joint-level feature extraction and frame-level pose regression for smooth regression. With better temporal consistency and subtle localized occlusion awareness, our MSANet outperforms previous state-of-the-art methods on the commonly used benchmarks Human3.6M and MPI-INF-3DHP. Moreover, MSANet demonstrates broad real world applicability, realizing occlusions and blurs robust and accurate 3D pose estimation. The Code will be made public.
Suping Wu, Xitie Zhang, Liyuan Shi, Zhijian Duan 0003, Tuo Xiong
ICASSP3
2025 Complementary Multi-dimensional Variance Attention Learning for 3D Human Mesh Reconstruction from Videos
abstract
Recently, great progress has been made in the field of video 3D human mesh reconstruction. Existing methods usually consider the correlation between features while ignoring essential feature learning, i.e. ignoring considering non-correlation between features, which results in learning a large amount of redundant and pseudo-correlated information, especially in complex scenarios. To address the above problem, we propose a multidimensional variance attention method that could learn non-correlation between features from multiple dimensions effectively. Specifically, we first design a variance attention network that filters out and weights non-correlation features by variance, the variance attention network can extract non-correlation features to approximate essential features and reduce redundancy. Furthermore, we design a multi-dimensional variance attention network that regards different dimensions as another kind of non-correlation, weights and fuses the selected features from time, channel and frequency domain dimensions to extract essential features. At the same time, the network extracts high-frequency information from the frequency domain by using a high-pass filter. These essential features and high-frequency information are fused to acquire the output that considers both correlation and non-correlation, which can conduct effective complementary learning. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method.
Tuo Xiong, Suping Wu, Ruijie Peng, Xitie Zhang, Zhijian Duan 0003
ICME6
2024 Towards Accurate 3D Face Alignment Under Extreme Scenarios Via Multi-Granularity Perturbation Relearning
abstract
3D face alignment from monocular images in challenging scenarios such as large poses and occlusions presents a huge challenge. To overcome this challenge, we propose a Multi-granularity Perturbation Relearning Network (MPRN), utilizing relearning attention to capture crucial features. Specifically, MPRN employs an attention mechanism to highlight effective features and further conducts relearning for attention to refine its accuracy. However, in extreme scenarios, the loss of key 3D facial information hampers the effective functioning of relearning attention. To this end, we construct multi-granularity perturbation graphs to infer the missing key 3D facial information and correspondingly guide the multiple times learning of attention module using perturbation graphs at various granularities. By doing this, our MPRN could effectively capture crucial 3D facial features in extreme scenarios, thereby achieving precise 3D face alignment. Experiments on the AFLW 2000-3D and AFLW datasets demonstrate the effectiveness of our MPRN.
Xinyu Li 0014, Xiaoxiao Yang, Suping Wu, Xiangzheng Li, Xitie Zhang
ICME6
2024 Correlation Disentangling and Spatio-Temporal Cooperative Optimizing Network for Temperature Prediction Revision
Aoao Wei, Xitie Zhang, Suping Wu, Kehua Ma
ICONIP (1)2
2024 CLTalk: Speech-Driven 3D Facial Animation with Contrastive Learning
abstract
Speech-driven 3D facial animation aims to generate realistic and vivid 3D facial animations from speech.However, the scarcity of labeled data and the tendency of existing methods to treat this cross-modal mapping problem as a regression task can result in inadequate learning of discriminative features from the speech.This deficiency often leads to excessively smooth facial movements, particularly in lip movements.To address these issues and enhance the accuracy of lip generation while reducing reliance on labeled data, we propose CLTalk, a framework based on a contrastive learning strategy.This framework comprises three main parts: a temporal domain contrastive learning strategy that facilitates the learning of discriminative features from different audio frames, a correlation learning method that ensures consistency between the distribution of audio features and Mesh labels, and a mouth opening angle constraint method to further improve the accuracy of lip generation.Extensive experimental results on the challenging, widely evaluated datasets indicate the effectiveness of our method compared with the state of the arts.
Xitie Zhang, Suping Wu
ICMR1
2024 Self-supervised Edge Structure Learning for Multi-view Stereo and Parallel Optimization
Suping Wu, Xitie Zhang, Yuxin Peng 0006
MMM (3)3
2024 Unsupervised Multi-collaborative Learning Network for 3D Face Reconstruction
Suping Wu, Xitie Zhang, Shengjia Zhang
MMM (3)3
2023 A Detail Geometry Learning Network for High-Fidelity Face Reconstruction
Kehua Ma, Xitie Zhang, Suping Wu, Leyang Yang, Zhixiang Yuan
ICANN (2)2
2023 CLN: Complementary Learning Network for 3D Face Reconstruction and Alignment
Kangbo Wu, Xitie Zhang, Xing Zheng, Suping Wu, Yongrong Cao, Kehua Ma
ICANN (2)2
2023 Multi Hybrid Extractor Network for 3D Human Pose Estimation
abstract
Monocular image or video based 3D human pose estimation remains a very challenging task because of depth ambiguity and occluded joints. To relieve this limitation, we propose a Multiple Hybrid Extraction Network (MHENet), which obtains three different representations of pose hypotheses features by multiple hybrid extractors with different structures, and uses pose interaction and fusion to obtain accurate 3D pose. The Hybrid Extraction Module obtains three hypotheses features: base features correspond to structural information, diverse features correspond to detail information, and condensed features correspond to action information. Hypotheses Interaction Fusion Modul builds relationships across hypotheses feature to generate more accurate 3D poses. Extensive qualitative and quantitative experimental results on a large-scale publicly available dataset demonstrate that our approach achieves competitive performance compared to state-of-the-art methods. The code will be made publicly.
Zhixiang Yuan, Xitie Zhang, Suping Wu, Yuxin Peng 0006
ICIP2
2023 A Bi-directional Optimization Network for De-obscured 3D High-Fidelity Face Reconstruction
Xitie Zhang, Suping Wu, Zhixiang Yuan, Kehua Ma, Leyang Yang
ICONIP (14)1