EDBT 2026 Demo / reviewers in the wild / expert
Wanjuan Su
dblp:209/0023
· DBLP profile ↗
21ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-5497-4682ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One-Stage Absolute Human Mesh RecoveryabstractThe reconstruction of realistic and precise human meshes in world coordinates is facilitated by considering scene information. Challenges related to accuracy, robustness, and computation time are faced by existing absolute human mesh recovery methods. In this paper, a one-stage model for absolute human mesh recovery with superior reconstruction precision and inference speed is presented. The proposed one-stage model is composed of two parallel branches to achieve root position estimation and human mesh regression. To effectively connect the two branches, a scene-image information aggregation module is designed. The accuracy of the estimated human meshes is improved and the end-to-end training of the whole model is facilitated by this module. Experiments are conducted on three diverse datasets, and a GMPJPE decrease of 72.3 mm/27.32% and an MPJPE reduction of 25.6 mm/27.26% are achieved by the proposed method with the lowest inference time compared to previous SOTA methods. Xinyao Liao, Wanjuan Su, Chen Zhang 0043, Ximeng Li 0007, Wenbing Tao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | High-Fidelity Lightweight Mesh Reconstruction from Point CloudsabstractRecently, learning signed distance functions (SDFs) from point clouds has become popular for reconstruction. To ensure accuracy, most methods require using high-resolution Marching Cubes for surface extraction. However, this results in redundant mesh elements, making the mesh inconvenient to use. To solve the problem, we propose an adaptive meshing method to extract resolution-adaptive meshes based on surface curvature, enabling the recovery of high-fidelity lightweight meshes. Specifically, we first use point-based representation to perceive implicit surfaces and calculate surface curvature. A vertex generator is designed to produce curvature-adaptive vertices with any specified number on the implicit surface, preserving the overall structure and high-curvature features. Then we develop a Delaunay meshing algorithm to generate meshes from vertices, ensuring geometric fidelity and correct topology. In addition, to obtain accurate SDFs for adaptive meshing and achieve better lightweight reconstruction, we design a hybrid representation combining feature grid and feature tri-plane for better detail capture. Experiments demonstrate that our method can generate high-quality lightweight meshes from point clouds. Compared with methods from various categories, our approach achieves superior results, especially in capturing more details with fewer elements. Chen Zhang 0043, Ximeng Li 0007, Xinyao Liao, Wanjuan Su, Wenbing Tao |
CVPR | 5 |
| 2025 | Context-Aware Multi-view Stereo Network for Efficient Edge-Preserving Depth Estimation
Wanjuan Su, Wenbing Tao |
Int. J. Comput. Vis. | 1 |
| 2025 | FE-GS: 3D feature-embedded Gaussian splatting with geometric regularizations for high-fidelity rendering
Yining Peng, Chen Zhang 0043, Wanjuan Su, Wenbing Tao |
Knowl. Based Syst. | 3 |
| 2025 | InstaHMR: Instance-Aware One-Stage Multi-Person Human Mesh RecoveryabstractHuman mesh recovery aims to estimate all human meshes within a given image. In this article, we propose an Instance-aware Multi-person 3D Human Mesh Recovery (InstaHMR) network based on the one-stage framework. Compared to former one-stage methods, instance-aware single person feature is exploited to represent more accurate human mesh. Specifically, we propose the Contextual Instance Guidance (CIG) module which generates instance-aware single person feature by leveraging spatial and channel attention operations. In this way, it preserves more instance-specific information compared to the pixel-level feature used in some existing one-stage methods. Besides, we further introduce two auxiliary losses for better mesh recovery, namely the Human Triplet Planes (HTP) loss and the T-pose Shape (TS) loss. The HTP loss encourages the model to capture subtle differences in human joint positions, while the TS loss facilitates the learning of abstract shape parameters. By incorporating these advancements, our model achieves state-of-the-art results on four multi-person datasets. Xinyao Liao, Chen Zhang 0043, Jianyao Xu, Wanjuan Su, Zhi Chen 0011, Wenbing Tao |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | PSDF: Prior-Driven Neural Implicit Surface Learning for Multi-View ReconstructionabstractSurface reconstruction has traditionally relied on the Multi-View Stereo (MVS)-based pipeline, which often suffers from noisy and incomplete geometry. This is due to that although MVS has been proven to be an effective way to recover the geometry of the scenes, especially for locally detailed areas with rich textures, it struggles to deal with areas with low texture and large variations of illumination where the photometric consistency is unreliable. Recently, Neural Implicit Surface Reconstruction (NISR) combines surface rendering and volume rendering techniques and bypasses the MVS as an intermediate step, which has emerged as a promising alternative to overcome the limitations of traditional pipelines. While NISR has shown impressive results on simple scenes, it remains challenging to recover delicate geometry from uncontrolled real-world scenes which is caused by its underconstrained optimization. To this end, the framework PSDF is proposed which resorts to external geometric priors from a pretrained MVS network and internal geometric priors inherent in the NISR model to facilitate high-quality neural implicit surface learning. Specifically, the visibility-aware feature consistency loss and depth prior-assisted sampling based on external geometric priors are introduced. These proposals provide powerfully geometric consistency constraints and aid in locating surface intersection points, thereby significantly improving the accuracy and delicate reconstruction of NISR. Meanwhile, the internal prior-guided importance rendering is presented to enhance the fidelity of the reconstructed surface mesh by mitigating the biased rendering issue in NISR. Extensive experiments on Tanks and Temples datasets show that PSDF achieves state-of-the-art performance on complex uncontrolled scenes. Wanjuan Su, Chen Zhang 0043, Qingshan Xu 0001, Wenbing Tao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | PG-NeuS: Robust and Efficient Point Guidance for Multi-View Neural Surface ReconstructionabstractRecently, learning multi-view neural surface reconstruction with the supervision of point clouds or depth maps has been a promising way. However, due to weak perception and underutilization of prior information, current methods still struggle with the challenges of limited accuracy and excessive time complexity. In addition, prior data perturbation is also an important yet rarely considered issue, often resulting in distorted geometry. To address these challenges, we propose a novel point-guided method named PG-NeuS, which achieves accurate and efficient reconstruction while robustly coping with point noise. Specifically, the aleatoric uncertainty of the point cloud is modeled to capture the noise distribution, estimating the reliability of each point and enhancing robustness against noise. Moreover, a Neural Projection module is proposed to connect points and images, adding geometric constraints to the implicit surface and achieving more precise point guidance. To better compensate for geometric bias between volume rendering and point modeling, we additionally design a Bias network that leverages the geometric information in high-fidelity points to enhance detail representation. Benefiting from the effective point guidance, the proposed PG-NeuS achieves an 11x speed increase and a 33.3% accuracy improvement compared to NeuS on DTU, even with a lightweight network. Extensive experiments show that our method yields high-quality surfaces with high efficiency, especially for fine-grained details and smooth regions, outperforming the state-of-the-art methods. Moreover, it exhibits strong robustness to noisy data and sparse data. Chen Zhang 0043, Wanjuan Su, Qingshan Xu 0001, Xinyao Liao, Wenbing Tao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | IINet: Implicit Intra-inter Information Fusion for Real-Time Stereo MatchingabstractRecently, there has been a growing interest in 3D CNN-based stereo matching methods due to their remarkable accuracy. However, the high complexity of 3D convolution makes it challenging to strike a balance between accuracy and speed. Notably, explicit 3D volumes contain considerable redundancy. In this study, we delve into more compact 2D implicit network to eliminate redundancy and boost real-time performance. However, simply replacing explicit 3D networks with 2D implicit networks causes issues that can lead to performance degradation, including the loss of structural information, the quality decline of inter-image information, as well as the inaccurate regression caused by low-level features. To address these issues, we first integrate intra-image information to fuse with inter-image information, facilitating propagation guided by structural cues. Subsequently, we introduce the Fast Multi-scale Score Volume (FMSV) and Confidence Based Filtering (CBF) to efficiently acquire accurate multi-scale, noise-free inter-image information. Furthermore, combined with the Residual Context-aware Upsampler (RCU), our Intra-Inter Fusing network is meticulously designed to enhance information transmission on both feature-level and disparity-level, thereby enabling accurate and robust regression. Experimental results affirm the superiority of our network in terms of both speed and accuracy compared to all other fast methods. Ximeng Li 0007, Chen Zhang 0043, Wanjuan Su, Wenbing Tao |
AAAI | 3 |
| 2023 | Efficient Edge-Preserving Multi-View Stereo Network for Depth EstimationabstractOver the years, learning-based multi-view stereo methods have achieved great success based on their coarse-to-fine depth estimation frameworks. However, 3D CNN-based cost volume regularization inevitably leads to over-smoothing problems at object boundaries due to its smooth properties. Moreover, discrete and sparse depth hypothesis sampling exacerbates the difficulty in recovering the depth of thin structures and object boundaries. To this end, we present an Efficient edge-Preserving multi-view stereo Network (EPNet) for practical depth estimation. To keep delicate estimation at details, a Hierarchical Edge-Preserving Residual learning (HEPR) module is proposed to progressively rectify the upsampling errors and help refine multi-scale depth estimation. After that, a Cross-view Photometric Consistency (CPC) is proposed to enhance the gradient flow for detailed structures, which further boosts the estimation accuracy. Last, we design a lightweight cascade framework and inject the above two strategies into it to achieve better efficiency and performance trade-offs. Extensive experiments show that our method achieves state-of-the-art performance with fast inference speed and low memory usage. Notably, our method tops the first place on challenging Tanks and Temples advanced dataset and ETH3D high-res benchmark among all published learning-based methods. Code will be available at https://github.com/susuwj/EPNet. Wanjuan Su, Wenbing Tao |
AAAI | 1 |
| 2023 | 3D hand pose and shape estimation from monocular RGB via efficient 2D cuesabstractEstimating 3D hand shape from a single-view RGB image is important for many applications. However, the diversity of hand shapes and postures, depth ambiguity, and occlusion may result in pose errors and noisy hand meshes. Making full use of 2D cues such as 2D pose can effectively improve the quality of 3D human hand shape estimation. In this paper, we use 2D joint heatmaps to obtain spatial details for robust pose estimation. We also introduce a depth-independent 2D mesh to avoid depth ambiguity in mesh regression for efficient hand-image alignment. Our method has four cascaded stages: 2D cue extraction, pose feature encoding, initial reconstruction, and reconstruction refinement. Specifically, we first encode the image to determine semantic features during 2D cue extraction; this is also used to predict hand joints and for segmentation. Then, during the pose feature encoding stage, we use a hand joints encoder to learn spatial information from the joint heatmaps. Next, a coarse 3D hand mesh and 2D mesh are obtained in the initial reconstruction step; a mesh squeeze-and-excitation block is used to fuse different hand features to enhance perception of 3D hand structures. Finally, a global mesh refinement stage learns non-local relations between vertices of the hand mesh from the predicted 2D mesh, to predict an offset hand mesh to fine-tune the reconstruction results. Quantitative and qualitative results on the FreiHAND benchmark dataset demonstrate that our approach achieves state-of-the-art performance. Fenghao Zhang, Lin Zhao 0012, Shengling Li, Wanjuan Su, Liman Liu, Wenbing Tao |
Comput. Vis. Media | 4 |
| 2023 | Edge-Aware Spatial Propagation Network for Multi-view Depth Estimation
Qingshan Xu 0001, Wanjuan Su, Wenbing Tao |
Neural Process. Lett. | 3 |
| 2023 | LGP-MVS: combined local and global planar priors guidance for indoor multi-view stereo
Weihang Kong, Qingshan Xu 0001, Wanjuan Su, Wenbing Tao |
Vis. Comput. | 3 |
| 2022 | Sparse prior guided deep multi-view stereo
Yuhang Qi, Wanjuan Su, Qingshan Xu 0001, Wenbing Tao |
Comput. Graph. | 2 |
| 2022 | Learning Inverse Depth Regression for Pixelwise Visibility-Aware Multi-View Stereo Networks
Qingshan Xu 0001, Wanjuan Su, Yuhang Qi, Wenbing Tao, Marc Pollefeys |
Int. J. Comput. Vis. | 2 |
| 2022 | Two-Stage Fuzzy Fusion Based-Convolution Neural Network for Dynamic Emotion RecognitionabstractThe two-stage fuzzy fusion based-convolution neural network is proposed for dynamic emotion recognition by using both facial expression and speech modalities, which not only can extract discriminative emotion features which contain spatio-temporal information, but also can effectively fuse facial expression and speech modalities. Moreover, the proposal is able to handle situations where the contributions of each modality data to emotion recognition are very imbalanced. The local binary patterns coming from three orthogonal planes and spectrogram are considered first to extract low-level dynamic emotion, so that the spatio-temporal information of these modalities can be obtained. To reveal more discriminative features, two deep convolution neural networks are constructed to extract high-level emotion semantic features. Moreover, the two stage fuzzy fusion strategy is developed by integrating canonical correlation analysis and fuzzy broad learning system, so as to take into account the correlation and difference between different modal features, as well as handle the ambiguity of emotional state information. The experimental results obtained on benchmark databases show that the accuracies of the proposed method are higher than those of existing methods (such as the hybrid deep model, and the rule-based and machine learning method) on SAVEE, eNTERFACE’05, and AFEW databases. Min Wu 0002, Wanjuan Su, Luefeng Chen, Witold Pedrycz, Kaoru Hirota |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Uncertainty Guided Multi-View Stereo Network for Depth EstimationabstractDeep learning has greatly promoted the development of multi-view stereo in recent years. However, how to measure the reliability of the estimated depth map for practical applications and make reasonable depth hypothesis sampling for the cost volume building in the coarse-to-fine architecture are still unresolved crucial problems. To this end, an Uncertainty Guided multi-view Network (UGNet) is proposed in this paper. In order to enable the network to perceive the uncertainty, an uncertainty-aware loss function is introduced, which not only can infer uncertainty implicitly in an unsupervised manner but also can reduce the bad impact of high uncertainty regions and the erroneous labels in the training set during training. Moreover, an uncertainty-based depth hypothesis sampling strategy is further proposed to adaptively determine the depth search range of each pixel for finer stages, which helps to generate more rational depth intervals compared with other methods and build more compact cost volumes without redundancy. Experimental results on DTU dataset, BlendedMVS dataset, Tanks and Temples dataset and ETH3D high-res benchmark show that our method achieves promising reconstruction results compared with other state-of-the-art methods. Wanjuan Su, Qingshan Xu 0001, Wenbing Tao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | A population randomization-based multi-objective genetic algorithm for gesture adaptation in human-robot interaction
Luefeng Chen, Wanjuan Su, Min Li 0087, Min Wu 0002, Witold Pedrycz, Kaoru Hirota |
Sci. China Inf. Sci. | 2 |
| 2021 | Weight-Adapted Convolution Neural Network for Facial Expression Recognition in Human-Robot InteractionabstractThe weight-adapted convolution neural network (WACNN) is proposed to extract discriminative expression representations for recognizing facial expression. It aims to make good use of the convolution neural network's (CNN's) potential performance in avoiding local optima and speeding up convergence by the hybrid genetic algorithm (HGA) with optimal initial population, in such a way that it realizes deep and global emotion understanding in human-robot interaction. Moreover, the idea of novelty search is introduced to solve the deception problem in the HGA, which can expend the search space to help genetic algorithm jump out of local optimum and optimize large-scale parameters. In the proposal, the facial expression image preprocessing is conducted first, then the low-level expression features are extracted by using a principal component analysis. Finally, the high-level expression semantic features are extracted and recognized by WACNN which is optimized by HGA. In order to evaluate the effectiveness of WACNN, experiments on JAFFE, CK+, and static facial expressions in the wild 2.0 databases are carried out by using k -fold cross validation, and experimental results show the recognition accuracies of the proposal are superior to that of the state-of-the-art, such as local directional ternary pattern and weighted mixture deep neural network (DNN), which aim to extract discriminative and are the DNN-based methods. Moreover, recognition accuracies of the proposal are also higher than the deep CNN without HGA, which indicates that the proposal has better global optimization ability. Meanwhile, preliminary application experiments are also carried out by using the proposed algorithm on the emotional social robot system, where nine volunteers and two-wheeled robots experience the scenario of emotion understanding. Application results indicate that the wheeled robots can recognize basic expressions, such as happy, surprise, and so on. Min Wu 0002, Wanjuan Su, Luefeng Chen, Zhentao Liu 0001, Kaoru Hirota |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Two-layer fuzzy multiple random forest for speech emotion recognition in human-robot interaction
Luefeng Chen, Wanjuan Su, Min Wu 0002, Jinhua She, Kaoru Hirota |
Inf. Sci. | 2 |
| 2020 | A Fuzzy Deep Neural Network With Sparse Autoencoder for Emotional Intention Understanding in Human-Robot InteractionabstractA fuzzy deep neural network with sparse autoencoder (FDNNSA) is proposed for intention understanding based on human emotions and identification information (i.e., age, gender, and region), in which the fuzzy C-means (FCM) is used to cluster the input data, and deep neural network with sparse autoencoder (DNNSA) is designed for emotional intention understanding in human-robot interaction. It aims to make robots capable of recognizing human emotions and understanding related emotional intention, the FCM is suitable for gathering similar information so that the calculations of dimensionality of DNNSA will be reduced, and the sparse autoencoder of DNNSA can make the neuron of DNNSA sparse to reduce the complexity of the network in such a way human-robot interaction is running smoothly. To validate the proposal, simulation experiments based on benchmark databases such as facial expression database of CK+, and speech emotion corpus of CASIA were completed. The experimental results show that the proposal outperforms the baseline algorithms of Softmax regression (SR), DNNSA, FCM-based SR (FSR), Softplus, Gath Geva-based DNNSA (GDNNSA), and ensemble DNNSA (EDNNSA). Preliminary application experiments are performed in the development of emotional social robot system, where volunteers experience the scenario of “drinking at the bar”. The obtained results indicate that the proposed FDNNSA can promote robot understanding of emotional intention of human. Luefeng Chen, Wanjuan Su, Min Wu 0002, Witold Pedrycz, Kaoru Hirota |
IEEE Trans. Fuzzy Syst. | 2 |
| 2018 | Softmax regression based deep sparse autoencoder network for facial emotion recognition in human-robot interaction
Luefeng Chen, Mengtian Zhou, Wanjuan Su, Min Wu 0002, Jinhua She, Kaoru Hirota |
Inf. Sci. | 3 |