Guanwen Feng

dblp:283/3573 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-1190-676XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A two-stage sign language generation framework with self-supervised latent representation learning
Qiguang Miao, Guanwen Feng, Junwei Jing, Yilin Zhang 0007, Yunan Li 0001, Chi-Man Pun
Knowl. Based Syst.2
2026 LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
abstract
While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be considered fine-grained, it needs to provide clear definitions and sufficiently detailed differentiation. We present LES-Talker, a novel one-shot talking head generation model with high interpretability, to achieve fine-grained emotion editing across emotion types, emotion levels, and facial units. We propose a Linear Emotion Space (LES) definition based on Facial Action Units to characterize emotion transformations as vector transformations. We design the Cross-Dimension Attention Net (CDAN) to deeply mine the correlation between LES representation and 3D model representation. Through mining multiple relationships across different feature and structure dimensions, we enable LES representation to guide the controllable deformation of 3D model. In order to adapt the multimodal data with deviations to the LES and enhance visual quality, we utilize specialized network design and training strategies. Experiments show that our method provides high visual quality along with multilevel and inter pretable fine-grained emotion editing, outperforming mainstream methods. Project page: https://peterfanfan.github.io/LES-Talker/
Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao, Chi-Man Pun
IEEE Trans. Affect. Comput.1
2026 EmoSpeaker: One-Shot Fine-Grained Emotion-Controlled Talking Face Generation
abstract
Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page:https://peterfanfan.github.io/EmoSpeaker/
Guanwen Feng, Yunan Li 0001, Chaoneng Li, Zhihao Qian 0001, Qiguang Miao, Chi-Man Pun
IEEE Trans. Multim.1
2025 KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation
abstract
Despite significant progress in NeRF-based talking head generation, problems like poor lip synchronization and inefficient resource usage remain. To solve these, we propose KANFace, a lightweight framework. In preprocessing, we introduce a Lip-Sync Enhancement Module that uses Wav2Lip to extract high-resolution audio features and map them to an explicit intermediate representation, ensuring precise lip movement alignment with the speaker’s identity. These predicted lip-sync features are combined with fundamental audio-extracted lip features and injected into the rendering module to improve synchronization. For rendering, we introduce FastKAN to map spatial points to color values. As a variant of KAN, FastKAN’s sensitivity to 3D scenes and efficient structure enable precise, fast color prediction. Our framework reduces resource consumption while enhancing lip-sync accuracy and facial reconstruction, making it ideal for talking head generation tasks in resource-limited settings. Project: https://peterfanfan.github.io/KAN-Face/
Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao
ICASSP1
2025 Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian Splatting
abstract
In recent years, audio-driven neural radiance field (NeRF)-based talking head generation techniques have achieved impressive results. However, these methods still have some limitations, such as unsynchronized lip movements and visual jitter. Recently, 3D Gaussian splatting has gradually replaced NeRF. Compared to NeRF, 3D Gaussian offers notable advantages, including higher efficiency and better reconstruction quality. Based on this, we propose Gaussian-Face, an audio-driven Gaussian-based facial avatar. With just a few minutes of monocular video and audio, a high-fidelity, driveable facial avatar can be reconstructed within hours. To achieve this, we first use FLAME to obtain the 3D representation of the face and then design a Lip Motion Translator to map audio to 3D lip representations. To model higher-quality facial details, we propose a hybrid density modeling method that balances rendering speed and quality, enabling our approach to render high-fidelity facial avatars at more than 160 FPS. Project page: https://peterfanfan.github.io/Gaussian-Face/
Guanwen Feng, Yilin Zhang 0007, Yunan Li 0001, Qiguang Miao
ICASSP1
2025 Sign-Mamba: Advanced Mamba-Based Sign Language Generation
abstract
In the field of sign language generation, Transformer-based models have been widely studied, but their quadratic computational complexity poses challenges. State Space Models (SSMs), like Mamba, offer a promising alternative with efficient long-range interaction modeling and linear complexity. In this study, we propose a two-stage generative framework for sign language generation, called Sign-Mamba, based on the state selection mechanism of SSM. In the first stage, we designed a Mamba-based encoder-decoder architecture, where the encoder captures the latent space representation and a symmetric decoder reconstructs it back into sign language skeletal points. In the second stage, the sign language latent space predicted by the Mamba-based latent predictor is used as a condition to guide the reconstruction network in generating more precise skeletal point sequences. We conducted comprehensive experiments on the PHOENIX and How2Sign datasets, and the results indicate that Sign-Mamba demonstrates competitive performance in sign language generation tasks. Project page: https://peterfanfan.github.io/Sign-Mamba/
Guanwen Feng, Yilin Zhang 0007, Yunan Li 0001, Qiguang Miao
ICASSP1
2025 One-shot handwriting imitation via self-supervised cross spatial transformer networks
Bocheng Zhao, Guanwen Feng, Wenxing Zhang, Yunan Li 0001, Qiguang Miao, Xiangzeng Liu, Ruyi Liu 0001
Neurocomputing2
2024 DiffTAD: Denoising diffusion probabilistic models for vehicle trajectory anomaly detection
Chaoneng Li, Guanwen Feng, Yunan Li 0001, Ruyi Liu 0001, Qiguang Miao, Liang Chang 0003
Knowl. Based Syst.2
2024 Exploration of Class Center for Fine-Grained Visual Classification
abstract
Different from large-scale classification tasks, fine-grained visual classification is a challenging task due to two critical problems: 1) evident intra-class variances and subtle inter-class differences, and 2) overfitting owing to fewer training samples in datasets. Most existing methods extract key features to reduce intra-class variances, but pay no attention to subtle inter-class differences in fine-grained visual classification. To address this issue, we propose a loss function named exploration of class center, which consists of a multiple class-center constraint and a class-center label generation. This loss function fully utilizes the information of the class center from the perspective of features and labels. From the feature perspective, the multiple class-center constraint pulls samples closer to the target class center, and pushes samples away from the most similar nontarget class center. Thus, the constraint reduces intra-class variances and enlarges inter-class differences. From the label perspective, the class-center label generation utilizes class-center distributions to generate soft labels to alleviate overfitting. Our method can be easily integrated with existing fine-grained visual classification approaches as a loss function, to further boost excellent performance with only slight training costs. Extensive experiments are conducted to demonstrate consistent improvements achieved by our method on four widely-used fine-grained visual classification datasets. In particular, our method achieves state-of-the-art performance on the FGVC-Aircraft and CUB-200-2011 datasets.
Hang Yao 0001, Qiguang Miao, Chaoneng Li, Guanwen Feng, Ruyi Liu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture Recognition
abstract
Although previous RGB-D-based gesture recognition methods have shown promising performance, researchers often overlook the interference of task-irrelevant cues like illumination and background. These unnecessary factors are learned together with the predictive ones by the network and hinder accurate recognition. In this paper, we propose a convenient and analytical framework to learn a robust feature representation that is impervious to gesture-irrelevant factors. Based on the Information Bottleneck theory, two rules of Sufficiency and Compactness are derived to develop a new information-theoretic loss function, which cultivates a more sufficient and compact representation from the feature encoding and mitigates the impact of gesture-irrelevant information. To highlight the predictive information, we further integrate a memory network. Using our proposed content-based and contextual memory addressing scheme, we weaken the nuisances while preserving the task-relevant information, providing guidance for refining the feature representation. Experiments conducted on three public datasets demonstrate that our approach leads to a better feature representation and achieves better performance than state-of-the-art methods. The code of our method is available at: https://github.com/Carpumpkin/InBoMem.
Yunan Li 0001, Huizhou Chen, Guanwen Feng, Qiguang Miao
ICCV3
2022 Fidelity Evaluation of Virtual Traffic Based on Anomalous Trajectory Detection
abstract
Measuring the fidelity of synthesized virtual traffic has become an important and fundamental concern for evaluating the performance of different traffic simulation techniques and applications of autonomous vehicle testing. In this work, we propose a novel method to evaluate the fidelity of any trajectory data from the perspective of anomalous trajectory detection. First, given the trajectory data to be evaluated as input, the method learns spatio-temporal traffic features and reconstructs the input trajectory through a Long Short-Term Memory (LSTM)-based autoencoder architecture. Then, the anomalous trajectories are detected by comparing the reconstructed trajectories and the input ones using the reconstruction error as the benchmark. Our method can detect eight different kinds of anomalous trajectory in terms of changes in velocity and moving direction. In order to evaluate the fidelity of the input trajectory, we design a perceptual evaluation on virtual traffic fidelity and derive a mapping from the reconstruction error to the evaluation score. We demonstrated the effectiveness and robustness of our metric through many experiments on real-world and synthetic trajectory data containing different types of motion anomalies.
Chaoneng Li, Qianwen Chao, Guanwen Feng, Qiongyan Wang, Yunan Li 0001, Qiguang Miao
IROS3
2021 SuccSPred: Succinylation Sites Prediction Using Fused Feature Representation and Ranking Method
Ruiquan Ge, Yizhang Luo, Guanwen Feng, Gangyong Jia, Gang Xu 0001
ISBRA3
2020 ProFPred: a two-step protein function prediction model based on sequence and evolutionary information
abstract
In post-genomic era, the understanding of protein function has been seriously behind the development of sequencing technology. Experimental verification for protein function is difficult, time consuming and expensive. Meanwhile, the new proteins vary in function. It is difficult for traditional methods to fully and accurately understand its functions. In this work, we present a novel method ProFPred to predict protein function based on sequence and evolutionary information to deal with small samples data corresponding to new diseases or discoveries. Experimental results demonstrate that our method can achieve better or comparable performances compared with current state-of-the-art methods.
Ruiquan Ge, Guanwen Feng, Qiguang Miao
BIBM2