EDBT 2026 Demo / reviewers in the wild / expert
Xitie Zhang
dblp:353/9011
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-8114-9119ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Credible and Detailed 3D Face Reconstruction in Large PoseabstractThe existing monocular methods face huge challenges in reconstructing credible details of non-visible areas in large pose images. Due to the fact that facial details are lost in non-visible areas of large pose images, existing methods lose basis when reconstructing details, resulting in unreliable results. Even if the generative model is used to repair the image first, the reconstruction process is very complex and costly. To this end, we propose an end-to-end and self-supervised RGB to depth method that equates small pose to large pose in UV space to obtain labels for non-visible areas. Then, we infer the depth values of non-visible areas and reconstruct a detailed 3D face. Finally, we render it into a face image and use it alongside the labels for self-supervised training of the network. During inference for large pose image, our method could reconstruct credible details of non-visible areas with a basis, rather than blindly. In addition, coarse and detailed reconstruction have mutually exclusive requirements for training images while coarse reconstruction is often not met, resulting in limited reconstruction accuracy. We propose a mutual exclusion elimination method to solve it, improving its accuracy. Extensive experiments demonstrate that our method could reconstruct precise 3D faces with credible details from large pose images. Our supplementary material is published on: https://github.com/lxy-nxu/ICASSP2025/tree/main Xinyu Li 0014, Xitie Zhang, Suping Wu, Ruijie Peng, Kehua Ma |
ICASSP | 2 |
| 2025 | MSANet: Mixed Spectral and Attention Network for Robust 3D Human Pose EstimationabstractDespite significant advances in 3D human pose estimation from a single-view video, existing methods often struggle to produce reasonable human poses when the human is heavily occluded or blurred. To address this issue, we propose a Mixed Spectral and Attention Network (MSANet) that stacks spectral and attention blocks alternately. The attention block captures visual cues before and after occlusions or blurs, while the spectral block perceives subtle localized occlusions or blurs for robust 3D human pose estimation. Specifically, our attention block captures the global information of intra-frame joints and enhances the coherent representation of inter-frame joints, the spectral block complements the intra- and inter-frame local occlusions information which is difficult to capture by the attention block. In addition, we improve the regression head (IRH) to narrow the grained gap between joint-level feature extraction and frame-level pose regression for smooth regression. With better temporal consistency and subtle localized occlusion awareness, our MSANet outperforms previous state-of-the-art methods on the commonly used benchmarks Human3.6M and MPI-INF-3DHP. Moreover, MSANet demonstrates broad real world applicability, realizing occlusions and blurs robust and accurate 3D pose estimation. The Code will be made public. Suping Wu, Xitie Zhang, Liyuan Shi, Zhijian Duan 0003, Tuo Xiong |
ICASSP | 3 |
| 2025 | Complementary Multi-dimensional Variance Attention Learning for 3D Human Mesh Reconstruction from VideosabstractRecently, great progress has been made in the field of video 3D human mesh reconstruction. Existing methods usually consider the correlation between features while ignoring essential feature learning, i.e. ignoring considering non-correlation between features, which results in learning a large amount of redundant and pseudo-correlated information, especially in complex scenarios. To address the above problem, we propose a multidimensional variance attention method that could learn non-correlation between features from multiple dimensions effectively. Specifically, we first design a variance attention network that filters out and weights non-correlation features by variance, the variance attention network can extract non-correlation features to approximate essential features and reduce redundancy. Furthermore, we design a multi-dimensional variance attention network that regards different dimensions as another kind of non-correlation, weights and fuses the selected features from time, channel and frequency domain dimensions to extract essential features. At the same time, the network extracts high-frequency information from the frequency domain by using a high-pass filter. These essential features and high-frequency information are fused to acquire the output that considers both correlation and non-correlation, which can conduct effective complementary learning. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method. Tuo Xiong, Suping Wu, Ruijie Peng, Xitie Zhang, Zhijian Duan 0003 |
ICME | 6 |
| 2024 | Towards Accurate 3D Face Alignment Under Extreme Scenarios Via Multi-Granularity Perturbation Relearningabstract3D face alignment from monocular images in challenging scenarios such as large poses and occlusions presents a huge challenge. To overcome this challenge, we propose a Multi-granularity Perturbation Relearning Network (MPRN), utilizing relearning attention to capture crucial features. Specifically, MPRN employs an attention mechanism to highlight effective features and further conducts relearning for attention to refine its accuracy. However, in extreme scenarios, the loss of key 3D facial information hampers the effective functioning of relearning attention. To this end, we construct multi-granularity perturbation graphs to infer the missing key 3D facial information and correspondingly guide the multiple times learning of attention module using perturbation graphs at various granularities. By doing this, our MPRN could effectively capture crucial 3D facial features in extreme scenarios, thereby achieving precise 3D face alignment. Experiments on the AFLW 2000-3D and AFLW datasets demonstrate the effectiveness of our MPRN. Xinyu Li 0014, Xiaoxiao Yang, Suping Wu, Xiangzheng Li, Xitie Zhang |
ICME | 6 |
| 2024 | Correlation Disentangling and Spatio-Temporal Cooperative Optimizing Network for Temperature Prediction Revision
Aoao Wei, Xitie Zhang, Suping Wu, Kehua Ma |
ICONIP (1) | 2 |
| 2024 | CLTalk: Speech-Driven 3D Facial Animation with Contrastive LearningabstractSpeech-driven 3D facial animation aims to generate realistic and vivid 3D facial animations from speech.However, the scarcity of labeled data and the tendency of existing methods to treat this cross-modal mapping problem as a regression task can result in inadequate learning of discriminative features from the speech.This deficiency often leads to excessively smooth facial movements, particularly in lip movements.To address these issues and enhance the accuracy of lip generation while reducing reliance on labeled data, we propose CLTalk, a framework based on a contrastive learning strategy.This framework comprises three main parts: a temporal domain contrastive learning strategy that facilitates the learning of discriminative features from different audio frames, a correlation learning method that ensures consistency between the distribution of audio features and Mesh labels, and a mouth opening angle constraint method to further improve the accuracy of lip generation.Extensive experimental results on the challenging, widely evaluated datasets indicate the effectiveness of our method compared with the state of the arts. Xitie Zhang, Suping Wu |
ICMR | 1 |
| 2024 | Self-supervised Edge Structure Learning for Multi-view Stereo and Parallel Optimization
Suping Wu, Xitie Zhang, Yuxin Peng 0006 |
MMM (3) | 3 |
| 2024 | Unsupervised Multi-collaborative Learning Network for 3D Face Reconstruction
Suping Wu, Xitie Zhang, Shengjia Zhang |
MMM (3) | 3 |
| 2023 | A Detail Geometry Learning Network for High-Fidelity Face Reconstruction
Kehua Ma, Xitie Zhang, Suping Wu, Leyang Yang, Zhixiang Yuan |
ICANN (2) | 2 |
| 2023 | CLN: Complementary Learning Network for 3D Face Reconstruction and Alignment
Kangbo Wu, Xitie Zhang, Xing Zheng, Suping Wu, Yongrong Cao, Kehua Ma |
ICANN (2) | 2 |
| 2023 | Multi Hybrid Extractor Network for 3D Human Pose EstimationabstractMonocular image or video based 3D human pose estimation remains a very challenging task because of depth ambiguity and occluded joints. To relieve this limitation, we propose a Multiple Hybrid Extraction Network (MHENet), which obtains three different representations of pose hypotheses features by multiple hybrid extractors with different structures, and uses pose interaction and fusion to obtain accurate 3D pose. The Hybrid Extraction Module obtains three hypotheses features: base features correspond to structural information, diverse features correspond to detail information, and condensed features correspond to action information. Hypotheses Interaction Fusion Modul builds relationships across hypotheses feature to generate more accurate 3D poses. Extensive qualitative and quantitative experimental results on a large-scale publicly available dataset demonstrate that our approach achieves competitive performance compared to state-of-the-art methods. The code will be made publicly. Zhixiang Yuan, Xitie Zhang, Suping Wu, Yuxin Peng 0006 |
ICIP | 2 |
| 2023 | A Bi-directional Optimization Network for De-obscured 3D High-Fidelity Face Reconstruction
Xitie Zhang, Suping Wu, Zhixiang Yuan, Kehua Ma, Leyang Yang |
ICONIP (14) | 1 |