Suzhen Wang 0001

dblp:08/6077-1 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-7271-4481ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2025 EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting
abstract
Character customization, or ’face crafting,’ is a vital feature in role-playing games (RPGs), enhancing player engagement by enabling the creation of personalized avatars. Existing automated methods often struggle with generalizability across diverse game engines due to their reliance on the intermediate constraints of specific image domain and typically support only one type of input, either text or image. To overcome these challenges, we introduce EasyCraft, an innovative end-to-end feedforward framework that automates character crafting by uniquely supporting both text and image inputs. Our approach employs a translator capable of converting facial images of any style into crafting parameters. We first establish a unified feature distribution in the translator’s image encoder through self-supervised learning on a large-scale dataset, enabling photos of any style to be embedded into a unified feature representation.Subsequently, we map this unified feature distribution to crafting parameters specific to a game engine, a process that can be easily adapted to most game engines and thus enhances EasyCraft’s generalizability. By integrating text-to-image techniques with our translator, EasyCraft also facilitates precise, text-based character crafting. EasyCraft’s ability to integrate diverse inputs significantly enhances the versatility and accuracy of avatar creation. Extensive experiments on two RPG games demonstrate the effectiveness of our method, achieving state-of-the-art results and facilitating adaptability across various avatar engines.
Suzhen Wang 0001, Wei Zhang 0219, Minda Zhao, Lincheng Li, Zhipeng Hu, Xin Yu 0002
CVPR1
2025 TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
abstract
Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to find and hence limits their usage. In this work, we propose TalkCLIP, a framework that can generate talking heads where the expressions are specified by natural language, hence allowing for specifying expressions more conveniently. To model the mapping from text to expressions, we first construct a text-video paired talking head dataset where each video has diverse text descriptions that depict both coarse-grained emotions and fine-grained facial movements. Leveraging the proposed dataset, we introduce a CLIP-based style encoder that projects natural language-based descriptions to the representations of expressions. TalkCLIP can even infer expressions for descriptions unseen during training. TalkCLIP can also use text to modulate expression intensity and edit expressions. Extensive experiments demonstrate that TalkCLIP achieves the advanced capability of generating photo-realistic talking heads with vivid facial expressions guided by text descriptions.
Yifeng Ma 0001, Suzhen Wang 0001, Yu Ding 0001, Tangjie Lv, Changjie Fan, Zhipeng Hu, Zhidong Deng, Xin Yu 0002
IEEE Trans. Multim.2
2024 StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads
abstract
Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse speaking styles in the final videos. To address this challenge, we propose a one-shot style-controllable talking face generation method that can obtain speaking styles from reference speaking videos and drive the one-shot portrait to speak with the reference speaking styles and another piece of audio. Our method aims to synthesize the style-controllable coefficients of a 3D Morphable Model (3DMM), including facial expressions and head movements, in a unified framework. Specifically, the proposed framework first leverages a style encoder to extract the desired speaking styles from the reference videos and transform them into style codes. Then, the framework uses a style-aware decoder to synthesize the coefficients of 3DMM from the audio input and style codes. During decoding, our framework adopts a two-branch architecture, which generates the stylized facial expression coefficients and stylized head movement coefficients, respectively. After obtaining the coefficients of 3DMM, an image renderer renders the expression coefficients into a specific person's talking-head video. Extensive experiments demonstrate that our method generates visually authentic talking head videos with diverse speaking styles from only one portrait image and an audio clip.
Suzhen Wang 0001, Yifeng Ma 0001, Yu Ding 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Zhidong Deng, Xin Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 StyleTalk: One-Shot Talking Head Generation with Controllable Speaking Styles
abstract
Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate diverse speaking styles in the final talking head videos. To tackle this problem, we propose a one-shot style-controllable talking face generation framework. In a nutshell, we aim to attain a speaking style from an arbitrary reference speaking video and then drive the one-shot portrait to speak with the reference speaking style and another piece of audio. Specifically, we first develop a style encoder to extract dynamic facial motion patterns of a style reference video and then encode them into a style code. Afterward, we introduce a style-controllable decoder to synthesize stylized facial animations from the speech content and style code. In order to integrate the reference speaking style into generated videos, we design a style-aware adaptive transformer, which enables the encoded style code to adjust the weights of the feed-forward layers accordingly. Thanks to the style-aware adaptation mechanism, the reference speaking style can be better embedded into synthesized videos during decoding. Extensive experiments demonstrate that our method is capable of generating talking head videos with diverse speaking styles from only one portrait image and an audio clip while achieving authentic visual effects. Project Page: https://github.com/FuxiVirtualHuman/styletalk.
Yifeng Ma 0001, Suzhen Wang 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001, Zhidong Deng, Xin Yu 0002
AAAI2
2023 FlowFace: Semantic Flow-Guided Shape-Aware Face Swapping
abstract
In this work, we propose a semantic flow-guided two-stage framework for shape-aware face swapping, namely FlowFace. Unlike most previous methods that focus on transferring the source inner facial features but neglect facial contours, our FlowFace can transfer both of them to a target face, thus leading to more realistic face swapping. Concretely, our FlowFace consists of a face reshaping network and a face swapping network. The face reshaping network addresses the shape outline differences between the source and target faces. It first estimates a semantic flow (i.e. face shape differences) between the source and the target face, and then explicitly warps the target face shape with the estimated semantic flow. After reshaping, the face swapping network generates inner facial features that exhibit the identity of the source face. We employ a pre-trained face masked autoencoder (MAE) to extract facial features from both the source face and the target face. In contrast to previous methods that use identity embedding to preserve identity information, the features extracted by our encoder can better capture facial appearances and identity information. Then, we develop a cross-attention fusion module to adaptively fuse inner facial features from the source face with the target facial attributes, thus leading to better identity preservation. Extensive quantitative and qualitative experiments on in-the-wild faces demonstrate that our FlowFace outperforms the state-of-the-art significantly.
Hao Zeng 0001, Wei Zhang 0219, Changjie Fan, Tangjie Lv, Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Xin Yu 0002
AAAI5
2023 Exploring Complementary Features in Multi-Modal Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is of great importance in human-computer interaction. Recent research has demonstrated that self-supervised learned acoustic and linguistic features are helpful in this task. However, few works have fully exploit the advantages of the pre-trained features in SER. The primary challenge is how to effectively extract the complementary emotional information implied in the pre-trained features of the respective modality. To tackle this challenge, we propose a novel modality-sensitive multimodal speech emotion recognition framework. In a nutshell, we aim to exploit the typical emotion features in each modality and then fuse the complementary emotional information for classification. Specifically, we first utilize the parallel uni-modal encoders to refine the emotion-related information from the pre-trained features of each modality. For better fusion of the multimodal features, we develop a group of learnable emotion query tokens to gather the emotional information from the refined acoustic and linguistic features with the cross-attention mechanism in the transformer decoder. Observing the modality bias problem in multimodal methods, we introduce the random modality masking training strategy to maximize the utilization of the emotional information in each modality and mitigate this problem. We evaluate our method on the widely used IEMOCAP dataset and achieve 1.1% and 0.9% improvements on the unweighted accuracy and weighted accuracy, respectively. Extensive experiments demonstrate the effectiveness of the proposed method.
Suzhen Wang 0001, Yifeng Ma 0001, Yu Ding 0001
ICASSP1
2023 Deep learning applications in games: a survey from a data perspective
Zhipeng Hu, Yu Ding 0001, Runze Wu 0001, Lincheng Li, Yujing Hu, Kai Wang 0064, Yongqiang Zhang 0003, Ji Jiang, Yadong Xi, Jiashu Pu, Wei Zhang 0219, Suzhen Wang 0001, Ke Chen 0005, Tianze Zhou, Jiarui Chen, Tangjie Lv, Changjie Fan
Appl. Intell.16
2022 One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation Learning
abstract
Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those methods struggle to learn a consistent speech style from different speakers. We observe that it would be much easier to learn a consistent speech style from a specific speaker, which leads to authentic mouth movements. Hence, we propose a novel one-shot talking face generation framework by exploring consistent correlations between audio and visual motions from a specific speaker and then transferring audio-driven motion fields to a reference image. Specifically, we develop an Audio-Visual Correlation Transformer (AVCT) that aims to infer talking motions represented by keypoint based dense motion fields from an input audio. In particular, considering audio may come from different identities in deployment, we incorporate phonemes to represent audio signals. In this manner, our AVCT can inherently generalize to audio spoken by other identities. Moreover, as face keypoints are used to represent speakers, AVCT is agnostic against appearances of the training speaker, and thus allows us to manipulate face images of different identities readily. Considering different face shapes lead to different motions, a motion field transfer module is exploited to reduce the audio-driven dense motion field gap between the training identity and the one-shot reference. Once we obtained the dense motion field of the reference image, we employ an image renderer to generate its talking face videos from an audio clip. Thanks to our learned consistent speaking style, our method generates authentic mouth shapes and vivid movements. Extensive experiments demonstrate that our synthesized videos outperform the state-of-the-art in terms of visual quality and lip-sync.
Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Xin Yu 0002
AAAI1
2021 Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation
abstract
In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be specific, our framework consists of a speaker-independent stage and a speaker-specific stage. In the speaker-independent stage, we design three parallel networks to generate animation parameters of the mouth, upper face, and head from texts, separately. In the speaker-specific stage, we present a 3D face model guided attention network to synthesize videos tailored for different individuals. It takes the animation parameters as input and exploits an attention mask to manipulate facial expression changes for the input individuals. Furthermore, to better establish authentic correspondences between visual motions (i.e., facial expression changes and head movements) and audios, we leverage a high-accuracy motion capture dataset instead of relying on long videos of specific individuals. After attaining the visual and audio correspondences, we can effectively train our network in an end-to-end fashion. Extensive experiments on qualitative and quantitative results demonstrate that our algorithm achieves high-quality photo-realistic talking-head videos including various facial expressions and head motions according to speech rhythms and outperforms the state-of-the-art.
Lincheng Li, Suzhen Wang 0001, Yu Ding 0001, Yixing Zheng, Xin Yu 0002, Changjie Fan
AAAI2
2021 Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
abstract
We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)} maintaining the appearance of a speaker in a large head motion while stabilizing the non-face regions. We first design a head pose predictor by modeling rigid 6D head movements with a motion-aware recurrent neural network (RNN). In this way, the predicted head poses act as the low-frequency holistic movements of a talking head, thus allowing our latter network to focus on detailed facial movement generation. To depict the entire image motions arising from audio, we exploit a keypoint based dense motion field representation. Then, we develop a motion field generator to produce the dense motion fields from input audio, head poses, and a reference image. As this keypoint based representation models the motions of facial regions, head, and backgrounds integrally, our method can better constrain the spatial and temporal consistency of the generated videos. Finally, an image generation network is employed to render photo-realistic talking-head videos from the estimated keypoint based motion fields and the input reference image. Extensive experiments demonstrate that our method produces videos with plausible head motions, synchronized facial expressions, and stable backgrounds and outperforms the state-of-the-art.
Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Changjie Fan, Xin Yu 0002
IJCAI1
2021 Learning a deep motion interpolation network for human skeleton animations
abstract
Abstract Motion interpolation technology produces transition motion frames between two discrete movements. It is wildly used in video games, virtual reality and augmented reality. In the fields of computer graphics and animations, our data‐driven method generates transition motions of two arbitrary animations without additional control signals. In this work, we propose a novel carefully designed deep learning framework, named deep motion interpolation network (DMIN), to learn human movement habits from a real dataset and then to perform the interpolation function specific for human motions. It is a data‐driven approach to capture overall rhythm of two given discrete movements and generate natural in‐between motion frames. The sequence‐by‐sequence architecture allows completing all missing frames within single forward inference, which reduces computation time for interpolation. Experiments on human motion datasets show that our network achieves promising interpolation performance. The ablation study demonstrates the effectiveness of the carefully designed DMIN.1
Chi Zhou 0005, Zhangjiong Lai, Suzhen Wang 0001, Lincheng Li, Yu Ding 0001
Comput. Animat. Virtual Worlds3
2018 Fine-Grained Grocery Product Recognition by One-Shot Learning
abstract
Fine-grained grocery product recognition via camera is a challenging task to identify the visually similar products with subtle differences by using single-shot training examples. To address this issue? we present a novel hybrid classification approach that combines feature-based matching and one-shot deep learning with a coarse-to-fine strategy. The candidate regions of product instances are first detected and coarsely labeled by recurring features in product images without any training. Then, attention maps are generated to guide the classifier to focus on fine discriminative details by magnifying the influences of the features in the candidate regions of interest (ROI) and suppressing the interferences of the features outside, improving the accuracy of fine-grained grocery products recognition effectively. Our framework also performs a good adaptability which allows existing classifier to be refined without retraining for new coming product classes. As an additional contribution, we collect a new grocery product database with 102 classes from 2 stores. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods.
Weidong Geng, Feilin Han, Jiangke Lin, Liuyi Zhu, Jieming Bai, Suzhen Wang 0001, Zhangjiong Lai
ACM Multimedia6