EDBT 2026 Demo / reviewers in the wild / expert
Yongtang Bao
dblp:21/7777
· DBLP profile ↗
21ranked-venue papers
16as first author
18since 2021 · last 2026
0000-0002-1010-7229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FaceCLIP: CLIP-Driven Accurate and Detailed 3D Face Reconstruction from a Single ImageabstractIn recent years, 3D face reconstruction has become a research hotspot in computer graphics and computer vision. Most current 3DMM-based methods focus on learning displacement maps to recover high-frequency facial details. However, they focus less on learning mid-frequency facial details and introduce displacement maps with noise, decreasing face reconstruction accuracy. Thus, this work presents a novel approach to regressing accurate and detailed 3D face shapes. First, we design a novel feature consistency loss to recover mid-frequency facial details. Specifically, we exploit the powerful CLIP as prior knowledge of faces to extract geometric and semantic features, which helps guide the reconstructed 3D geometric details to match local details in the input image. Furthermore, we propose a parameter refinement module to learn fine-grained features. It helps to obtain accurate model parameters and improve the accuracy of facial reconstruction. Extensive experiments on a FaceScape and a REALY benchmark demonstrate that our method outperforms several state-of-the-art methods in reconstruction accuracy. Furthermore, comprehensive qualitative results show that our approach achieves better visual performance than existing methods. Yongtang Bao |
Comput. Vis. Media | 1 |
| 2026 | C&D-CLIP: Cascaded decoder and deep cross visual prompt tuning for zero-shot semantic segmentation
Linchuan Li, Zhihui Wang 0003, Yongtang Bao, Shansong Yang |
Pattern Recognit. | 4 |
| 2025 | Seg-Wild: Interactive Segmentation based on 3D Gaussian Splatting for Unconstrained Image Collections
Yongtang Bao, Chengjie Tang, Yuze Wang 0006 |
ACM Multimedia | 1 |
| 2025 | Weighted cross-integrated fusion network based on knowledge distillation for multi-modal personality recognition
Yongtang Bao |
Appl. Intell. | 1 |
| 2025 | Efficient interactive segmentation of three-dimensional Gaussians with optimal view selection
Yongtang Bao, Chengjie Tang, Yuze Wang 0006, Yutong Qi |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Emotion-Assisted multi-modal Personality Recognition using adversarial Contrastive learning
Yongtang Bao, Yutong Qi, Liping Feng |
Knowl. Based Syst. | 1 |
| 2025 | HIAN: A hybrid interactive attention network for multimodal sarcasm detection
Yongtang Bao, Peng Zhang 0057 |
Pattern Recognit. | 1 |
| 2024 | SSCCPC-Net: Simultaneously Learning 2D and 3D Features with CLIP for Semantic Scene Completion on Point Cloud
Wantong Duan, Yongtang Bao |
CGI (3) | 2 |
| 2024 | Identity Semantic Correspondence for Cloth-Changing Person Re-IdentificationabstractCloth-changing Person Re-Identification (CC-ReID) aims at retrieving the same person who might change clothes across different locations. Although remarkable progress has been achieved in recent studies, most of the recent methods still lack sufficient emphasis on identity-related regions. To address these issues, we propose a novel Identity Semantic Correspondence framework (ISC) to fully utilize human semantic information, which includes dual-stream identity semantic correspondance networks, i.e., a Clothing-invariant Identity (CI) stream and a Fine-grained Identity Semantic Highlighting (FISH) stream. The CI mitigates the interference of human dressing and enhances clothing-invariant identity information by erasing clothing information. And, the FISH exploits fine-grained identity information by highlighting contribution of different body parts to identity with a part-aware weighting module. Additionally, an identity semantic consistency module is further proposed to extract the most representative and discriminative semantic features for each identity. Besides, we employ a mutual loss to transfer identity-related knowledge between components, which enables the original appearance module to be deployed independently during the inference stage. Extensive experiments on two CC-ReID benchmarks, including PRCC and VC-Clothes, are conducted to demonstrate the effectiveness of the proposed ISC method. Yongtang Bao, Kun Zhan, Peng Zhang 0057 |
IJCNN | 1 |
| 2024 | SCARF: Scalable Continual Learning Framework for Memory-efficient Multiple Neural Radiance FieldsabstractAbstract This paper introduces a novel continual learning framework for synthesising novel views of multiple scenes, learning multiple 3D scenes incrementally, and updating the network parameters only with the training data of the upcoming new scene. We build on Neural Radiance Fields (NeRF), which uses multi‐layer perceptron to model the density and radiance field of a scene as the implicit function. While NeRF and its extensions have shown a powerful capability of rendering photo‐realistic novel views in a single 3D scene, managing these growing 3D NeRF assets efficiently is a new scientific problem. Very few works focus on the efficient representation or continuous learning capability of multiple scenes, which is crucial for the practical applications of NeRF. To achieve these goals, our key idea is to represent multiple scenes as the linear combination of a cross‐scene weight matrix and a set of scene‐specific weight matrices generated from a global parameter generator. Furthermore, we propose an uncertain surface knowledge distillation strategy to transfer the radiance field knowledge of previous scenes to the new model. Representing multiple 3D scenes with such weight matrices significantly reduces memory requirements. At the same time, the uncertain surface distillation strategy greatly overcomes the catastrophic forgetting problem and maintains the photo‐realistic rendering quality of previous scenes. Experiments show that the proposed approach achieves state‐of‐the‐art rendering quality of continual learning NeRF on NeRF‐Synthetic, LLFF, and TanksAndTemples datasets while preserving extra low storage cost. Yuze Wang 0006, Junyi Wang 0001, Chen Wang 0043, Wantong Duan, Yongtang Bao |
Comput. Graph. Forum | 5 |
| 2024 | Adaptive information fusion network for multi-modal personality recognitionabstractAbstract Personality recognition is of great significance in deepening the understanding of social relations. While personality recognition methods have made significant strides in recent years, the challenge of heterogeneity between modalities during feature fusion still needs to be solved. This paper introduces an adaptive multi‐modal information fusion network (AMIF‐Net) capable of concurrently processing video, audio, and text data. First, utilizing the AMIF‐Net encoder, we process the extracted audio and video features separately, effectively capturing long‐term data relationships. Then, adding adaptive elements in the fusion network can alleviate the problem of heterogeneity between modes. Lastly, we concatenate audio‐video and text features into a regression network to obtain Big Five personality trait scores. Furthermore, we introduce a novel loss function to address the problem of training inaccuracies, taking advantage of its unique property of exhibiting a peak at the critical mean. Our tests on the ChaLearn First Impressions V2 multi‐modal dataset show partial performance surpassing state‐of‐the‐art networks. Yongtang Bao |
Comput. Animat. Virtual Worlds | 1 |
| 2024 | Boosting Micro-Expression Recognition via Self-Expression Reconstruction and Memory Contrastive LearningabstractMicro-expression (ME) is an instinctive reaction that is not controlled by thoughts. It reveals one's inner feelings, which is significant in sentiment analysis and lie detection. Since micro-expression is expressed as subtle facial changes within particular facial action units, learning discriminative and generalized features for Micro-expression Recognition (MER) is challenging. To achieve the purpose, this paper proposes a novel MER framework that simultaneously integrates supervised Prototype-based Memory Contrastive Learning (PMCL) for discriminative feature mining and adds Self-expression Reconstruction (SER) as an auxiliary task and regularization for better generalization. In particular, the proposed SER module is forced as a regularization by reconstructing input ME from the randomly dropped patch- wise features in the bottleneck. And, the PMCL module globally compares historical and current cluster agents learned from training instances to enhance intra-class compactness and inter-class separability. Extensive experiments are conducted on three benchmarks, e.g., SMIC, CASME II, and SAMM, under evaluation criteria of both Composite Database Evaluation (CDE) and Single Database Evaluation (SDE) protocols. The results show our method surpasses other state-of-the-art approaches under various evaluation metrics, achieving overall 86.30% unweighed F1-score and 88.30% unweighed average recall on the composite dataset. Furthermore, the ablation studies verify the effectiveness of our SER for better generalization and PMCL for better discrimination in learning feature representation from limited micro-expression samples. Yongtang Bao, Peng Zhang 0057, Caifeng Shan, Xianye Ben |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D ImageabstractCategory-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics. Yongtang Bao, Chunjian Su, Yutong Qi, Yanbing Geng |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Multiple object tracking with adaptive multi-features fusion and improved learnable graph matching
Yongtang Bao, Zhihui Wang 0003 |
Vis. Comput. | 1 |
| 2023 | Local Context and Dimensional Relation Aware Transformer Network for Continuous Affect EstimationabstractIn recent years, video-based continuous affect estimation has received more attention in computer vision. Therefore, how to robustly and accurately model the temporal information during facial expression change is crucial. Hence, we propose a transformer network that incorporates both local context and dimensional correlation to model visual information in an efficient manner. Specifically, noise, such as instantaneous head poses and lighting changes, may affect the model’s performance due to the local context insensitivity of the transformer’s self-attention layer. Therefore, a local-wise transformer encoder is adopted to enhance the transformer’s ability to capture local contextual information. In addition, considering the prior knowledge of the correlation between valence and arousal,we design the va-relevance bootstrap module and the corresponding valence-arousal relevance loss (va loss). Experiments on Aff-Wild2 and AFEW-VA datasets show the superior performance of our method for continuous affect estimation. Yongtang Bao |
ICIP | 2 |
| 2023 | Learning Unoccluded Face Texture Completion from Single Image in the Wild
Yongtang Bao |
Neural Process. Lett. | 1 |
| 2022 | AE-GAN: Attention Embedded GAN for Irregular and Large-Area Mask Face Image Inpainting
Yongtang Bao, Xinfei Xiao |
CGI | 1 |
| 2021 | Effective multiple pedestrian tracking system in video surveillance with monocular stationary camera
Zhihui Wang 0003, Ming Li 0065, Yu Lu 0006, Yongtang Bao, Zhe Li 0015, Jianli Zhao 0002 |
Expert Syst. Appl. | 4 |
| 2018 | A motion-field-based data-driven method for hair animationabstractAbstract We propose a motion‐field‐based data‐driven approach that animates static hair models from a set of precomputed motion data. We first sample a motion sequence to construct a motion database from physics‐based hair simulation data or dynamic hair capture data. We also define the preliminary definitions of motion states and construct a motion field for this motion database. Finally, we generate a sequence of target hairstyle from the input motion field by strand correspondence and motion control. Experimental results show that our approach achieves comparable quality with physics‐based methods but in orders of magnitude faster performance. Yongtang Bao |
Comput. Animat. Virtual Worlds | 1 |
| 2017 | An adaptive floating tangents fitting with helices method for image-based hair modelingabstractCurrently, hair geometry is mostly represented as sequences of 3D points. It is difficult to simulate hair directly from this representation. This paper proposes a novel approach to convert hair geometry model into helices, which could be easily plugged into dynamic hair simulation. We construct a hair model from a hybrid orientation field. Then we use adaptive floating tangents fitting algorithm to convert this hair geometry model into a physics-based hair model. We simulate dynamic hair by Lagrange equations. Results show that this approach can preserve structural details of 3D hair models, and can be applied to simulate various hair geometries. Yongtang Bao |
CGI | 1 |
| 2016 | Realistic hair modeling from a hybrid orientation field
Yongtang Bao |
Vis. Comput. | 1 |