VLDB 2026 Research / reviewers in the wild / expert
Suping Wu
dblp:172/8665
· DBLP profile ↗
42ranked-venue papers
0as first author
32since 2021 · last 2025
0000-0001-5207-1802ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 26 since 2021Artificial intelligence and machine learning · 16 · 11 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Credible and Detailed 3D Face Reconstruction in Large PoseabstractThe existing monocular methods face huge challenges in reconstructing credible details of non-visible areas in large pose images. Due to the fact that facial details are lost in non-visible areas of large pose images, existing methods lose basis when reconstructing details, resulting in unreliable results. Even if the generative model is used to repair the image first, the reconstruction process is very complex and costly. To this end, we propose an end-to-end and self-supervised RGB to depth method that equates small pose to large pose in UV space to obtain labels for non-visible areas. Then, we infer the depth values of non-visible areas and reconstruct a detailed 3D face. Finally, we render it into a face image and use it alongside the labels for self-supervised training of the network. During inference for large pose image, our method could reconstruct credible details of non-visible areas with a basis, rather than blindly. In addition, coarse and detailed reconstruction have mutually exclusive requirements for training images while coarse reconstruction is often not met, resulting in limited reconstruction accuracy. We propose a mutual exclusion elimination method to solve it, improving its accuracy. Extensive experiments demonstrate that our method could reconstruct precise 3D faces with credible details from large pose images. Our supplementary material is published on: https://github.com/lxy-nxu/ICASSP2025/tree/main Xinyu Li 0014, Xitie Zhang, Suping Wu, Ruijie Peng, Kehua Ma |
ICASSP | 3 |
| 2025 | MSANet: Mixed Spectral and Attention Network for Robust 3D Human Pose EstimationabstractDespite significant advances in 3D human pose estimation from a single-view video, existing methods often struggle to produce reasonable human poses when the human is heavily occluded or blurred. To address this issue, we propose a Mixed Spectral and Attention Network (MSANet) that stacks spectral and attention blocks alternately. The attention block captures visual cues before and after occlusions or blurs, while the spectral block perceives subtle localized occlusions or blurs for robust 3D human pose estimation. Specifically, our attention block captures the global information of intra-frame joints and enhances the coherent representation of inter-frame joints, the spectral block complements the intra- and inter-frame local occlusions information which is difficult to capture by the attention block. In addition, we improve the regression head (IRH) to narrow the grained gap between joint-level feature extraction and frame-level pose regression for smooth regression. With better temporal consistency and subtle localized occlusion awareness, our MSANet outperforms previous state-of-the-art methods on the commonly used benchmarks Human3.6M and MPI-INF-3DHP. Moreover, MSANet demonstrates broad real world applicability, realizing occlusions and blurs robust and accurate 3D pose estimation. The Code will be made public. Suping Wu, Xitie Zhang, Liyuan Shi, Zhijian Duan 0003, Tuo Xiong |
ICASSP | 2 |
| 2025 | WL-MVSNet: Frequency-Aware and Regularized Learning for Multi-View StereoabstractIn recent years, learning-based Multi-View Stereo (MVS) methods have gained increasing attention due to their superior feature representation capabilities and end-to-end optimization frameworks. However, most existing approaches encounter significant challenges in capturing high-frequency details and preserving intricate edge structures, which often leads to suboptimal reconstruction results. To address these limitations, we propose WL-MVSNet, a novel framework designed to enhance matching precision and improve depth estimation performance. Specifically, we propose a wavelet-domain feature refinement strategy that extracts more salient features and incorporate a high-frequency clipping approach to mitigate noise interference, thereby enhancing reconstruction accuracy. Moreover, we introduce a Laplacian operator-based loss function to explicitly constrain edge details in depth maps, thus enhancing the accuracy of depth estimation along object boundaries. Experimental results demonstrate that WL-MVSNet significantly improves the robustness of MVS methods in complex scenarios, especially excelling in weakly-textured regions and along edges. Ruijie Peng, Suping Wu |
ICME | 3 |
| 2025 | Complementary Multi-dimensional Variance Attention Learning for 3D Human Mesh Reconstruction from VideosabstractRecently, great progress has been made in the field of video 3D human mesh reconstruction. Existing methods usually consider the correlation between features while ignoring essential feature learning, i.e. ignoring considering non-correlation between features, which results in learning a large amount of redundant and pseudo-correlated information, especially in complex scenarios. To address the above problem, we propose a multidimensional variance attention method that could learn non-correlation between features from multiple dimensions effectively. Specifically, we first design a variance attention network that filters out and weights non-correlation features by variance, the variance attention network can extract non-correlation features to approximate essential features and reduce redundancy. Furthermore, we design a multi-dimensional variance attention network that regards different dimensions as another kind of non-correlation, weights and fuses the selected features from time, channel and frequency domain dimensions to extract essential features. At the same time, the network extracts high-frequency information from the frequency domain by using a high-pass filter. These essential features and high-frequency information are fused to acquire the output that considers both correlation and non-correlation, which can conduct effective complementary learning. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method. Tuo Xiong, Suping Wu, Ruijie Peng, Xitie Zhang, Zhijian Duan 0003 |
ICME | 2 |
| 2025 | Hyperbolic Space Learning Method Leveraging Temporal Motion Priors for Human Mesh Recoveryabstract3D human meshes show a natural hierarchical structure (like torso-limbs-fingers). But existing video-based 3D human mesh recovery methods usually learn mesh features in Euclidean space. It’s hard to catch this hierarchical structure accurately. So wrong human meshes are reconstructed. To solve this problem, we propose a hyperbolic space learning method leveraging temporal motion prior for recovering 3D human meshes from videos. First, we design a temporal motion prior extraction module. This module extracts the temporal motion features from the input 3D pose sequences and image feature sequences respectively. Then it combines them into the temporal motion prior. In this way, it can strengthen the ability to express features in the temporal motion dimension. Since data representation in non-Euclidean space has been proved to effectively capture hierarchical relationships in real-world datasets (especially in hyperbolic space), we further design a hyperbolic space optimization learning strategy. This strategy uses the temporal motion prior information to assist learning, and uses 3D pose and pose motion information respectively in the hyperbolic space to optimize and learn the mesh features. Then, we combine the optimized results to get an accurate and smooth human mesh. Besides, to make the optimization learning process of human meshes in hyperbolic space stable and effective, we propose a hyperbolic mesh optimization loss. Extensive experimental results on large publicly available datasets indicate superiority in comparison with most state-of-the-art. Suping Wu, Weibin Qiu, Zhaocheng Jin |
ICME | 2 |
| 2025 | Latent-Info and Low-Dimensional Learning for Human Mesh Recovery and Parallel OptimizationabstractExisting 3D human mesh recovery methods often fail to fully exploit the latent information (e.g., human motion, shape alignment), leading to issues with limb misalignment and insufficient local details in the reconstructed human mesh (especially in complex scenes). Furthermore, the performance improvement gained by modelling mesh vertices and pose node interactions using attention mechanisms comes at a high computational cost. To address these issues, we propose a two-stage network for human mesh recovery based on latent information and low dimensional learning. Specifically, the first stage of the network fully excavates global (e.g., the overall shape alignment) and local (e.g., textures, detail) information from the low and high-frequency components of image features and aggregates this information into a hybrid latent frequency domain feature. This strategy effectively extracts latent information. Subsequently, utilizing extracted hybrid latent frequency domain features collaborates to enhance 2D poses to 3D learning. In the second stage, with the assistance of hybrid latent features, we model the interaction learning between the rough 3D human mesh template and the 3D pose, optimizing the pose and shape of the human mesh. Unlike existing mesh pose interaction methods, we design a low-dimensional mesh pose interaction method through dimensionality reduction and parallel optimization that significantly reduces computational costs without sacrificing reconstruction accuracy. Extensive experimental results on large publicly available datasets indicate superiority compared to the most state-of-the-art. Suping Wu |
ICME | 2 |
| 2025 | Text-Guided Diffusion with Spectral Convolution for 3D Human Pose EstimationabstractAbstract Although significant progress has been made in monocular video‐based 3D human pose estimation, existing methods lack guidance from fine‐grained high‐level prior knowledge such as action semantics and camera viewpoints, leading to significant challenges for pose reconstruction accuracy under scenarios with severely missing visual features, i.e., complex occlusion situations. We identify that the 3D human pose estimation task fundamentally constitutes a canonical inverse problem, and propose a motion‐semantics‐based diffusion(MS‐Diff) framework to address this issue by incorporating high‐level motion semantics with spectral feature regularization to eliminate interference noise in complex scenes and improve estimation accuracy. Specifically, we design a Multimodal Diffusion Interaction (MDI) module that incorporates motion semantics including action categories and camera viewpoints into the diffusion process, establishing semantic‐visual feature alignment through a cross‐modal mechanism to resolve pose ambiguities and effectively handle occlusions. Additionally, we leverage a Spectral Convolutional Regularization (SCR) module that implements adaptive filtering in the frequency domain to selectively suppress noise components. Extensive experiments on large‐scale public datasets Human3.6M and MPI‐INF‐3DHP demonstrate that our method achieves state‐of‐the‐art performance. Liyuan Shi, Suping Wu, Weibin Qiu, Dong Qiang, Jiarui Zhao |
Comput. Graph. Forum | 2 |
| 2024 | Towards Accurate 3D Face Alignment Under Extreme Scenarios Via Multi-Granularity Perturbation Relearningabstract3D face alignment from monocular images in challenging scenarios such as large poses and occlusions presents a huge challenge. To overcome this challenge, we propose a Multi-granularity Perturbation Relearning Network (MPRN), utilizing relearning attention to capture crucial features. Specifically, MPRN employs an attention mechanism to highlight effective features and further conducts relearning for attention to refine its accuracy. However, in extreme scenarios, the loss of key 3D facial information hampers the effective functioning of relearning attention. To this end, we construct multi-granularity perturbation graphs to infer the missing key 3D facial information and correspondingly guide the multiple times learning of attention module using perturbation graphs at various granularities. By doing this, our MPRN could effectively capture crucial 3D facial features in extreme scenarios, thereby achieving precise 3D face alignment. Experiments on the AFLW 2000-3D and AFLW datasets demonstrate the effectiveness of our MPRN. Xinyu Li 0014, Xiaoxiao Yang, Suping Wu, Xiangzheng Li, Xitie Zhang |
ICME | 4 |
| 2024 | GFAvatar: A High-Quality Facial Avatar Reconstruction MethodabstractDigitally modeling and reconstructing talking humans is important in telepresence applications of AR or VR environments. However, current methods often fail to effectively address the inability to capture local details of avatars due to resolution or image quality limitations or effectively generate realistic and natural 3D color representations. Meanwhile, invisible regions may cause the reconstruction results to appear hollow or missing. To alleviate these problems, in this paper, we propose a novel approach, called GFAvatar. Compared to existing methods, GFAvatar improves the quality of point cloud texture features by designing the fusion of image texture and 3D texture information. We achieve end-to-end learning by using an image super-resolution approach combined with our designed PointNet variant to extract detailed features from head avatars and enhance the representation of point cloud features. Our multimodal color fusion network combines image and point cloud color data, generating more precise and expressive 3D color representations for better avatar quality. We also design a texture consistency loss function to address the problem of abnormal local color in the fusion network. Further, to efficiently address the challenges posed by disordered point clouds, we carefully elaborate a 3D grids optimization to improve the integrity of facial reconstruction. Extensive experimental results on available datasets indicate the superiority in comparison with most state-of-the-arts. Shengjia Zhang, Suping Wu |
ICME | 2 |
| 2024 | Multi-scale Feature Edge Enhancement for Multi-view Stereo
Zhijian Duan 0003, Suping Wu, Ruijie Peng, Weibin Qiu, Tuo Xiong |
ICONIP (8) | 2 |
| 2024 | Correlation Disentangling and Spatio-Temporal Cooperative Optimizing Network for Temperature Prediction Revision
Aoao Wei, Xitie Zhang, Suping Wu, Kehua Ma |
ICONIP (1) | 3 |
| 2024 | CLTalk: Speech-Driven 3D Facial Animation with Contrastive LearningabstractSpeech-driven 3D facial animation aims to generate realistic and vivid 3D facial animations from speech.However, the scarcity of labeled data and the tendency of existing methods to treat this cross-modal mapping problem as a regression task can result in inadequate learning of discriminative features from the speech.This deficiency often leads to excessively smooth facial movements, particularly in lip movements.To address these issues and enhance the accuracy of lip generation while reducing reliance on labeled data, we propose CLTalk, a framework based on a contrastive learning strategy.This framework comprises three main parts: a temporal domain contrastive learning strategy that facilitates the learning of discriminative features from different audio frames, a correlation learning method that ensures consistency between the distribution of audio features and Mesh labels, and a mouth opening angle constraint method to further improve the accuracy of lip generation.Extensive experimental results on the challenging, widely evaluated datasets indicate the effectiveness of our method compared with the state of the arts. Xitie Zhang, Suping Wu |
ICMR | 2 |
| 2024 | Self-supervised Edge Structure Learning for Multi-view Stereo and Parallel Optimization
Suping Wu, Xitie Zhang, Yuxin Peng 0006 |
MMM (3) | 2 |
| 2024 | Unsupervised Multi-collaborative Learning Network for 3D Face Reconstruction
Suping Wu, Xitie Zhang, Shengjia Zhang |
MMM (3) | 2 |
| 2024 | Multi-granularity relationship reasoning network for high-fidelity 3D shape reconstruction
Suping Wu |
Pattern Recognit. | 3 |
| 2023 | Two-stage Co-segmentation Network Based on Discriminative Representation for Recovering Human Mesh from VideosabstractRecovering 3D human mesh from videos has recently made significant progress. However, most of the existing methods focus on the temporal consistency of videos, while ignoring the spatial representation in complex scenes, thus failing to recover a reasonable and smooth human mesh sequence under extreme illumination and chaotic backgrounds. To alleviate this problem, we propose a two-stage co-segmentation network based on discriminative representationfor recovering human body meshes from videos. Specifically, the first stage of the network segments the video spatial domain to spotlight spatially fine-grained information, and then learns and enhances the intra-frame discriminative representation through a dual-excitation mechanism and a frequency domain enhancement module, while sup-pressing irrelevant information (e.g., background). The second stage focuses on temporal context by segmenting the video temporal domain, and models inter-frame discriminative representation via a dynamic integration strategy. Further, to efficiently generate reasonable human discriminative actions, we carefully elaborate a landmark anchor area loss to constrain the variation of the human motion area. Extensive experimental results on large publicly available datasets indicate superiority in comparison with most state-of-the-art. The Code will be made public. Kehua Ma, Suping Wu, Zhixiang Yuan |
CVPR | 3 |
| 2023 | A Detail Geometry Learning Network for High-Fidelity Face Reconstruction
Kehua Ma, Xitie Zhang, Suping Wu, Leyang Yang, Zhixiang Yuan |
ICANN (2) | 3 |
| 2023 | CLN: Complementary Learning Network for 3D Face Reconstruction and Alignment
Kangbo Wu, Xitie Zhang, Xing Zheng, Suping Wu, Yongrong Cao, Kehua Ma |
ICANN (2) | 4 |
| 2023 | Time-Frequency Awareness Network For Human Mesh Recovery From VideosabstractThis paper focuses on the problem of 3D human mesh recovery from videos. Most recent works mainly centered on human spatio-temporal modeling in the time domain, but these time-domain methods often concentrate on the short-range spatio-temporal receptive field and information transfer of the video, thus cannot adaptively sense effective spatiotemporal dependencies in a long-range, furthermore lacking the ability to perceive local motion with a small-scale movement. In this work, we propose a Time-Frequency Awareness Network for human mesh recovery. We present a novel paradigm that learns human feature representations by introducing frequency domain. Specifically, we first design a time-frequency aware attention module that uses frequency domain information as a guide to model temporal long-range dependence and spatial long-range dependence in a unified manner. Secondly, we carefully develop a time-frequency-aware recurrent module that treats moving humans as discrete signals over time in the frequency domain to capture the spatio-temporal information accumulated by human movement in videos. In addition, we also elaborate design of a local awareness loss constraint on human motion of the small scale, which helps to mitigate the interference of global motion on the prediction results. Extensive experimental results on large publicly available datasets show advantages over most state-of-the-art methods. Our code are available at https://github.com/Changboyang/TFNet.git. Suping Wu, Meining Jia |
ICASSP | 2 |
| 2023 | Multi Hybrid Extractor Network for 3D Human Pose EstimationabstractMonocular image or video based 3D human pose estimation remains a very challenging task because of depth ambiguity and occluded joints. To relieve this limitation, we propose a Multiple Hybrid Extraction Network (MHENet), which obtains three different representations of pose hypotheses features by multiple hybrid extractors with different structures, and uses pose interaction and fusion to obtain accurate 3D pose. The Hybrid Extraction Module obtains three hypotheses features: base features correspond to structural information, diverse features correspond to detail information, and condensed features correspond to action information. Hypotheses Interaction Fusion Modul builds relationships across hypotheses feature to generate more accurate 3D poses. Extensive qualitative and quantitative experimental results on a large-scale publicly available dataset demonstrate that our approach achieves competitive performance compared to state-of-the-art methods. The code will be made publicly. Zhixiang Yuan, Xitie Zhang, Suping Wu, Yuxin Peng 0006 |
ICIP | 3 |
| 2023 | A Lightweight Grouped Low-rank Tensor Approximation Network for 3D Mesh Reconstruction From VideosabstractExisting methods for 3D mesh reconstruction from videos suffer from increasingly large parameter counts and model sizes due to encoders such as multi-hidden state recurrence. Therefore many models become complex and more difficult to be applied in practice. Based on this problem, we propose a lightweight grouped low-rank tensor approximation network for 3D mesh reconstruction from videos. Specifically, firstly we propose a generalized grouped low-rank tensor approximation algorithm, which decomposes the original high-rank tensor into multiple weighted groups with different low-rank tensors to maximize the approximation of the high-rank tensor. Then we also design a lightweight selection rearranging strategy to reduce feature redundancy and focus on local fragment features. Notably, our method can be flexibly plugged into other 3D reconstruction tasks. Experiments show that our method not only improves the performance but also reduces the parameters of the entire network by about 90% compared with the existing SOTA method. Our grouped low-rank tensor approximation method reduces the parameters by about 99.8% in a single GRU. We demonstrate the outstanding generalization of our method in other 3D reconstruction tasks, eg. 3D face reconstruction and multi-view stereo. Suping Wu, Leyang Yang |
ICME | 2 |
| 2023 | A Bi-directional Optimization Network for De-obscured 3D High-Fidelity Face Reconstruction
Xitie Zhang, Suping Wu, Zhixiang Yuan, Kehua Ma, Leyang Yang |
ICONIP (14) | 2 |
| 2023 | Multi-scale Edge-guided Learning for 3D ReconstructionabstractSingle-view three-dimensional (3D) object reconstruction has always been a long-term challenging task. Objects with complex topologies are hard to accurately reconstruct, which makes existing methods suffer from blurring of shape boundaries between multiple components in the object. Moreover, most of them cannot balance learning between global geometric structure information and local detail information. In this article, we propose a multi-scale edge-guided learning network (MEGLN) to utilize the global edge information guiding the network to better capture and recover local details. The goal is to exploit the multi-scale learning strategy to learn global edge information and local details, thus achieving robust 3D object reconstruction. We first design a multi-scale Gaussian difference block (MGDB) to extract global edge geometry features for input images of different scales and adopt the attention mechanism to aggregate the extracted global edge geometry features of different scales. Second, we design a multi-scale feature interaction block (MFIB) to learn local details, which utilizes the multi-scale feature interaction to capture the features of multiple objects or components at multiple scales. The MFIB can learn and capture better as much local detail information as possible under the guidance of global edge information. Finally, we dynamically fuse the predicted probabilities of the MGDB and MFIB to obtain the final predicted result, which makes our MEGLN able to recover 3D shapes with global complex topological structures and rich local details via the multi-scale learning strategy. Extensive qualitative and quantitative experimental results on the ShapeNet dataset demonstrate that our approach achieves competitive performance compared with state-of-the-art methods. Code is available at https://github.com/Ray-tju/MEGLN . Lei Li 0044, Suping Wu, Yongrong Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Global Contextual Complementary Network for Multi-View Stereo
Yongrong Cao, Suping Wu, Xing Zheng, Zhixiang Yuan, Yuxin Peng 0006 |
BMVC | 2 |
| 2022 | Spatio-temporal tendency reasoning for human body pose and shape estimation from videos
Suping Wu, Hu Cao, Kehua Ma |
BMVC | 2 |
| 2022 | GNC: Geometry Normal Consistency Loss for 3D Face Reconstruction and Dense AlignmentabstractIn this work, we propose Geometry Normal Consistency Loss (GNC) for 3D face reconstruction and dense alignment. The existing methods based on the strong constraints of the 3DMM parameter regression only consider reducing the error between 68 landmarks, while they rarely consider the geometric contour structure relation of the face. Instead, we take into account the discrete 68 landmarks as loss constrain by introducing geometry area and normal consistency loss, which naturally defines the holistic and local geometric contour structure of the face. In detail, we select the inverted triangle formed by the leftmost and rightmost landmarks on the cheek and the lowest point of the chin to globally constrain the entire facial feature. In addition, we triangulate the facial landmarks to construct triangular patches, and calculate the normals of the patches, which aims to make use of normal consistency between the local patch and the corresponding ground truth to reconstruct rich local details. Extensive experimental results on AFLW2000-3D and AFLW datasets demonstrate that our GNC achieves compelling performance compared to the state of the arts. Xing Zheng, Yongrong Cao, Lei Li 0044, Meining Jia, Suping Wu |
ICME | 6 |
| 2021 | Frame-level Feature Tokenization Learning for Human Body Pose and Shape EstimationabstractIn this paper, we propose a frame-level feature tokenization method for human body pose and shape es-timation(FTHE). Despite conventional 3D human pose and shape estimation methods have achieved success based on a single image, recovering accurate and smooth 3D human motion from a video is still challenging. Different from existing methods, our FTHE aims to pay attention to the meaningful detailed temporal feature between different granular tokens of video objects, and reduce the dominance of the current static frame. To this end, we carefully design an accurate and interpretable temporal encoding module for feature extraction and motion reconstruction. More specifically, our model captures temporal features and static features of different granular tokens, and simultaneously enhances their correlation and multi-granularity consistency. Extensive experimental results on large-scale publicly available datasets demonstrate that our FTHE achieves compelling performance compared to the state-of-the-art. Code has been made available at: https://githuh.com/chriful/FG_2020_FTHE. Hu Cao, Meining Jia, Suping Wu |
FG | 3 |
| 2021 | Replay Attention and Data Augmentation Network for 3D Dense Alignment and Face Reconstructionabstract3D face reconstruction from a single-view image in the wild is a long-standing challenging problem. Traditional 3DMM-based methods directly regressed parameters, which probably caused that the network learned the discriminative informative features in the face insufficiently. In this paper, we propose a replay attention and data augmentation network (RADAN) for 3D dense alignment and face reconstruction. Instead of the traditional attention mechanism, our replay attention module aims to increase the sensitivity of the network to informative features by adaptively recalibrating the weight response in the attention mechanism, which typically reinforces the distinguishability of the learned feature representation. In this way, the network is able to further improve the accuracy of face reconstruction and dense alignment in an unconstrained environment. Moreover, to improve the generalization performance of the model and the ability of the network to capture local details, we present a data augmentation strategy to preprocess the sample data, which generates the images that contain more local details and occluded face in a cropping and pasting manner. Extensive qualitative and quantitative experimental results on widely-evaluated benchmarking datasets demonstrate that our approach achieves competitive performance compared to state-of-the-art methods. Code is available at https://github.com/zhouzhiyuanl/RADANet. Lei Li 0044, Suping Wu |
FG | 3 |
| 2021 | Multi-Granularity Feature Interaction and Relation Reasoning for 3D Dense Alignment and Face ReconstructionabstractIn this paper, we propose a multi-granularity feature interaction and relation reasoning network (MFIRRN) which can recover a detail-rich 3D face and perform more accurate dense alignment in an unconstrained environment. Traditional 3DMM-based methods directly regress parameters, resulting in the lack of fine-grained details in the reconstruction 3D face. To this end, we use different branches to capture discriminative features at different granularities, especially local features at medium and fine granularities. Meanwhile, the finer-grained branch network shares its information with the adjacent coarser-grained branch network to achieve feature interaction. Our model performs cross-granular information integration and inter-granular relationship reasoning to obtain prediction results. Extensive experiments on AFLW2000-3D and AFLW datasets demonstrate the validity of our method. The code is publicly available at https://github.com/leilimaster/MFIRRN. Lei Li 0044, Xiangzheng Li, Kangbo Wu, Kui Lin, Suping Wu |
ICASSP | 5 |
| 2021 | High-Resolution Multi-View Stereo with Dynamic Depth Edge FlowabstractMulti-view stereo based on deep learning is mostly dedicated to improving the accuracy of point clouds, whlile complex scenes, occlusion, and other factors limit their reconstruction completeness, especially in the area with drastic changes in depth direction. In this paper, we propose a multi-view stereo network based on depth edge flow (DEF-MVSNet), using the reference image as a guide to dynamically infer the edge coordinates to improve reconstruction completeness. First, we ignore the boundaries in the depth prediction stage to generate better initial depth inference results. Then, we use a EdgeDetect module to extract the obviously features of the reference image and predict the pixel offset of the depth map. Finally, the EdgeFlow module modifies the initial depth map coordinates according to the offset and uses multiple iterations to dynamically update the depth map. The experimental results prove that our method has a great improvement in the completeness of reconstruction compared with MVSNet and R-MVSNet without increasing the memory and time overhead. Code and models are publicly available at https://github.com/linkuizzZ/EF-MVSNet. Kui Lin, Lei Li 0044, Xing Zheng, Suping Wu |
ICME | 5 |
| 2021 | Towards Rich-Detail 3D Face Reconstruction and Dense Alignment via Multi-Scale Detail Augmentationabstract3D face reconstruction based on a single image is a longstanding challenging problem in computer vision. Existing end-to-end methods are difficult to reconstruct rich 3D face details. To solve this problem, we propose a two-stream convolutional neural network combined with a face super-resolution method, which can effectively restore the image’s 3D position information. Our method combines an attention fusion mechanism, which can learn the individual attention mapping of each feature subspace, and effectively learn cross-channel information while learning multi-scale and multi-frequency features. Meanwhile, our module obtains the most discriminative features in different local areas, and enhances the consistency and correlation between the attention areas. Experimental results show that our SRCNet has made significant improvements in the 3D face reconstruction and face alignment of the AFLW2000-3D and AFLW datasets. Suping Wu, Lei Li 0044, Kui Lin, Xing Zheng, Hu Cao |
ICME | 2 |
| 2021 | Graph Structure Reasoning Network for Face Alignment and Reconstruction
Suping Wu |
MMM (1) | 3 |
| 2020 | DmifNet: 3D Shape Reconstruction based on Dynamic Multi-Branch Information Fusionabstract3D object reconstruction from a single-view image is a long-standing challenging problem. Previous work was difficult to accurately reconstruct 3D shapes with a complex topology which has rich details at the edges and corners. Moreover, previous works used synthetic data to train their network, but domain adaptation problems occurred when tested on real data. In this paper, we propose a Dynamic Multi-branch Information Fusion Network (DmifNet) which can recover a high-fidelity 3D shape of arbitrary topology from a 2D image. Specifically, we design several side branches from the intermediate layers to make the network produce more diverse representations to improve the generalization ability of network. In addition, we utilize DoG (Difference of Gaussians) to extract edge geometry and corners information from input images. Then, we use a separate side branch network to process the extracted data to better capture edge geometry and corners feature information. Finally, we dynamically fuse the information of all branches to gain final predicted probability. Extensive qualitative and quantitative experiments on a large-scale publicly available dataset demonstrate the validity and efficiency of our method. Code and models are publicly available at https://github.com/leilimaster/DmifNet. Lei Li 0044, Suping Wu |
ICPR | 2 |
| 2020 | Multi-Attribute Regression Network for Face ReconstructionabstractIn this paper, we propose a multi-attribute regression network (MARN) to investigate the problem of face reconstruction, especially in challenging cases when faces undergo large variations including severe poses, extreme expressions, and partial occlusions in unconstrained environments. The traditional 3DMM parametric regression method does not distinguish the learning of identity, expression, and attitude attributes, resulting in lacking geometric details in the reconstructed face. We propose to learn a face multi-attribute features during 3D face reconstruction from single 2D images. Our MARN enables the network to better extract the feature information of face identity, expression, and pose attributes. We introduce three loss functions to constrain the above three face attributes respectively. At the same time, we carefully design the geometric contour constraint loss function, using the constraints of sparse 2D face landmarks to improve the reconstructed geometric contour information. The experimental results show that our MARN has achieved significant improvements in 3D face reconstruction and face alignment on the AFLW2000-3D and AFLW datasets. Xiangzheng Li, Suping Wu |
ICPR | 2 |
| 2020 | 3D Human Pose Estimation based on Center of GravityabstractIn this paper, we propose a method about 3D human pose estimation with only 2D joints as input. Previous methods generally lift 2D poses to 3D space through a single mapping function, in which case some large-pose samples far away from the majority distribution may not be well concerned. To address the issue above, we design a multi-branch network based on the human center of gravity (COG) to enhance the robustness of the model to large-pose samples. Specifically, noticing the correspondence between the COG and human pose, by clustering the COG, we separate the large-pose samples from the normal ones in an unsupervised pattern, and lift them with separate branch network. In addition, we introduce a global loss function to regularize the integrality of 3D joints. Extensive experiments on the largest publicly available dataset demonstrate the validity and efficiency of our method. Suping Wu |
IJCNN | 2 |
| 2020 | Learning Reasoning-Decision Networks for Robust Face AlignmentabstractIn this paper, we propose an end-to-end reasoning-decision networks (RDN) approach for robust face alignment via policy gradient. Unlike the conventional coarse-to-fine approaches which likely lead to bias prediction due to poor initialization, our approach aims to learn a policy by leveraging raw pixels to reason a subset of shape candidates, sequentially making plausible decisions to remove outliers for robust initialization. To achieve this, we formulate face alignment as a Markov decision process by defining an agent, which typically interacts with a trajectory of states, actions, state transitions and rewards. The agent seeks an optimal shape searching policy over the whole shape space by maximizing a discounted sum of the received values. To further improve the alignment performance, we develop an LSTM-based value function to evaluate the shape quality. During the training procedure, we adjust the gradient of our value function in directions of the policy gradient. This prevents our training goal from being trapped into local optima entangled by both the pose deformations and appearance variations especially in unconstrained environments. Experimental results show that our proposed RDN consistently outperforms most state-of-the-art approaches on four widely-evaluated challenging datasets. Hao Liu 0019, Jiwen Lu, Suping Wu, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Similarity-Aware and Variational Deep Adversarial Learning for Robust Facial Age EstimationabstractIn this paper, we propose a similarity-aware deep adversarial learning (SADAL) approach for facial age estimation. Instead of making full access to the limited training samples which likely leads to bias age prediction, our SADAL aims to seek batches of unobserved hard-negative samples based on existing training samples, which typically reinforces the discriminativeness of the learned feature representation for facial ages. Motivated by the fact that age labels are usually correlated in real-world scenarios, we carefully develop a similarity-aware function to well measure the distance of each face pair based on the age value gaps. Consequently, the age-difference information is exploited in the synthetic feature space for robust age estimation. During the learning process, we jointly optimize both procedures of generating hard negatives and learning discriminative age ranker via a sequence of adversarial-game iterations. Another major issue lies on that existing methods only enforce the indiscriminativeness within each class, which is probably trapped into model overfitting and thus the generation capacity is limited particularly on unseen age classes with many individuals. To circumvent this problem, we propose a variational deep adversarial learning (VDAL) paradigm, which learns to encode each face sample in two factorized parts, i.e., the intra-class variance distribution and the intra-class invariant class center. Moreover, our VDAL principally optimizes the variational confidence lower bound on the variational factorized feature representation. To better enhance the discriminativeness of the age representation, our VDAL further learns to encode the ordinal relationship among age labels in the reconstructed subspace. Experimental results on folds of widely-evaluated benchmarking datasets demonstrate that our approach achieves promising performance in contrast to most state-of-the-art age estimation methods. Hao Liu 0019, Penghui Sun, Suping Wu, Zhenhua Yu 0002, Xuehong Sun |
IEEE Trans. Multim. | 4 |
| 2019 | Learning Relational-Structural Networks for Robust Face Alignment
Congcong Zhu, Suping Wu, Zhenhua Yu 0002 |
ICANN (3) | 3 |
| 2019 | Disentangled Representation Learning for Leaf Diseases Recognition
Congcong Zhu, Suping Wu |
ICIG (1) | 3 |
| 2019 | Learning Deformable Hourglass Networks (DHGN) for Unconstrained Face AlignmentabstractIn this paper, we propose a deformable hourglass networks (DHGN) approach to investigate the problem of face alignment, especially in such challenging cases when faces undergo large variations including severe poses, diverse expressions and partial occlusions in unconstrained environments. Unlike conventional feature extractions which cannot explicitly exploit irregular geometric structures for facial shapes, our DHGN learns a deformable mask to reduce the variances of facial deformation and extract attentional facial regions for robust feature representation. To achieve this, we carefully design a differential module, dubbed the deformable transformer, which typically incorporates with a regression sub-net to predict a set of offsets and a masking operator to filter the semantic facial parts for feature representation learning. To further reinforce the alignment performance, we integrate our designed modules in the paradigm of stacked hourglass networks and jointly optimize the network parameters in an end-to-end manner. Extensive experimental results demonstrate very compelling performance in comparisons to most state-of-the-art methods. Congcong Zhu, Suping Wu, Zhenhua Yu 0002, Xuehong Sun, Hao Liu 0019 |
ICIP | 3 |
| 2019 | Multi-Agent Deep Collaboration Learning for Face Alignment Under Different PerspectivesabstractIn this paper, we propose a multi-agent deep collaboration learning method (MADCL) for simultaneously detecting 2D facial landmarks and 3D facial landmarks projected from 3D to 2D, which aims at distinguishing the ambiguity caused by different perspectives. Above two facial annotations, there are a large number of public semantic areas and some very important private semantic areas. Our single agent captures and memorizes private features for iterations and multiple agents collaborate to learn public features. To achieve this, we design a collaboration learning mechanism to capture, memorize and share semantic information for enhancing the feature representation. Moreover, the input of traditional cascade regression methods is cropped directly from the raw facial image via the shape-indexed manner, which leads that the poor initial shapes likely bring about the predicted results getting worse and worse. We introduce the Markov decision process (MDP) to reason a better position of the initial shape by a reward function that reflects the shape quality. Authentic experimental results indicate that our MADCL consistently outperforms most state-of-the-art methods on two widely-evaluated challenging datasets. Congcong Zhu, Suping Wu, Zhenhua Yu 0002, Hao Liu 0019 |
ICIP | 2 |
| 2019 | Similarity-Aware Deep Adversarial Learning for Facial Age EstimationabstractIn this paper, we propose a similarity-aware deep adversarial learning (SADAL) approach for facial age estimation. Instead of making access to limited training samples which likely leads to sub-optima, our SADAL seeks sets of unobserved and plausible hard-examples based on existing training samples, which typically reinforces the discriminativeness of the learned feature descriptor for ages. Motivated by the fact that age labels are usually correlated in the real-world applications, we carefully develop a similarity-aware function in our approach, which dynamically measures each face pair with different weights based on different age value gaps. During the learning process, we jointly optimize both procedures of generating hard-examples and learning age estimator via a sequence of adversarial-game iterations. As a result, the smoothing aging pattern is exploited in the reconstructed hard-example space for robust age estimation. Experimental results on two standard benchmarking datasets show that our approach achieves superior performance compared with most state-of-the-art age estimation methods. Penghui Sun, Hao Liu 0019, Zhenhua Yu 0002, Suping Wu |
ICME | 5 |