EDBT 2026 Demo / reviewers in the wild / expert
Yu Ding 0001
dblp:77/6871-1
· DBLP profile ↗
54ranked-venue papers
7as first author
38since 2021 · last 2026
0000-0003-1834-4429ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 1 first-author · 30 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Multimodal Semantic Balancing Framework for Sentiment Analysis
Jiajia Tang, Feiwei Zhou, Xiping Wang, Qibin Zhao, Yu Ding 0001, Wanzeng Kong |
IEEE Trans. Multim. | 6 |
| 2025 | MECG: modality-enhanced convolutional graph for unbalanced multimodal representations
Jiajia Tang, Binbin Ni, Yutao Yang, Yu Ding 0001, Wanzeng Kong |
J. Supercomput. | 4 |
| 2025 | TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking StylesabstractAudio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to find and hence limits their usage. In this work, we propose TalkCLIP, a framework that can generate talking heads where the expressions are specified by natural language, hence allowing for specifying expressions more conveniently. To model the mapping from text to expressions, we first construct a text-video paired talking head dataset where each video has diverse text descriptions that depict both coarse-grained emotions and fine-grained facial movements. Leveraging the proposed dataset, we introduce a CLIP-based style encoder that projects natural language-based descriptions to the representations of expressions. TalkCLIP can even infer expressions for descriptions unseen during training. TalkCLIP can also use text to modulate expression intensity and edit expressions. Extensive experiments demonstrate that TalkCLIP achieves the advanced capability of generating photo-realistic talking heads with vivid facial expressions guided by text descriptions. Yifeng Ma 0001, Suzhen Wang 0001, Yu Ding 0001, Tangjie Lv, Changjie Fan, Zhipeng Hu, Zhidong Deng, Xin Yu 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | Fine-grained Semantic Disentanglement Network for Multimodal Sarcasm AnalysisabstractMultimodal sarcasm analysis is one of the most challenging research branch of the sentiment analysis area, due to the presence of cross-modality incongruity. However, existing works mainly attend to the coarse-grained incongruity analysis, and totally ignore the sentiment semantic coupling issue. This indeed limits the discriminate capability and robustness of the sarcasm analysis model. In order to address the above issue, we propose a novel Fine-grained Semantic Disentanglement Network (FSDN). Specifically, the intra-modality semantic disentanglement is performed to investigate the more intrinsic semantic cues of the same modality. Additionally, the inter-modality semantic disentanglement is leveraged to simultaneously facilitate the common and intrinsic semantic cues across modalities. Furthermore, the dual-spatial semantic interaction block is presented to explore the long-range cross-spatial semantic context between the obtained verbal and non-verbal semantic space with the global view. The above semantic disentanglement processes with both local and global views significantly unleash much more robustness even for the sarcasm case consisting of multiple semantic message. Various experiments indicate that the FSDN can receive state-of-the-art or competitive performance. Jiajia Tang, Binbin Ni, Feiwei Zhou, Dongjun Liu, Yu Ding 0001, Yong Peng 0001, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Say Anything with Any StyleabstractGenerating stylized talking head with diverse head motions is crucial for achieving natural-looking videos but still remains challenging. Previous works either adopt a regressive method to capture the speaking style, resulting in a coarse style that is averaged across all training data, or employ a universal network to synthesize videos with different styles which causes suboptimal performance. To address these, we propose a novel dynamic-weight method, namely Say Anything with Any Style (SAAS), which queries the discrete style representation via a generative model with a learned style codebook. Specifically, we develop a multi-task VQ-VAE that incorporates three closely related tasks to learn a style codebook as a prior for style extraction. This discrete prior, along with the generative model, enhances the precision and robustness when extracting the speaking styles of the given style clips. By utilizing the extracted style, a residual architecture comprising a canonical branch and style-specific branch is employed to predict the mouth shapes conditioned on any driving audio while transferring the speaking style from the source to any desired one. To adapt to different speaking styles, we steer clear of employing a universal network by exploring an elaborate HyperStyle to produce the style-specific weights offset for the style branch. Furthermore, we construct a pose generator and a pose codebook to store the quantized pose representation, allowing us to sample diverse head motions aligned with the audio and the extracted style. Experiments demonstrate that our approach surpasses state-of-the-art methods in terms of both lip-synchronization and stylized expression. Besides, we extend our SAAS to video-driven style editing field and achieve satisfactory performance as well. Shuai Tan 0002, Bin Ji 0004, Yu Ding 0001 |
AAAI | 3 |
| 2024 | Towards a Simultaneous and Granular Identity-Expression Control in Personalized Face GenerationabstractIn human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted por-trait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts towards personalized face generation. To this end, we propose a novel multi-modal face generation frame-work, capable of simultaneous identity-expression control and more fine-grained expression synthesis. Our expression control is so sophisticated that it can be specialized by the fine-grained emotional vocabulary. We devise a novel dif-fusion model that can undertake the task of simultaneously face swapping and reenactment. Due to the entanglement of identity and expression, separately and precisely control-ling them within one framework is a nontrivial task, thus has not been explored yet. To overcome this, we propose sev-eral innovative designs in the conditional diffusion model, including balancing identity and expression encoder, improved midpoint sampling, and explicitly background con-ditioning. Extensive experiments have demonstrated the controllability and scalability of the proposed framework, in comparison with state-of-the-art text-to-image, face swap-ping, and face reenactment methods. Renshuai Liu, Wei Zhang 0219, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001 |
CVPR | 7 |
| 2024 | Norface: Improving Facial Expression Analysis by Identity Normalization
Hanwei Liu, Rudong An, Wei Zhang 0219, Yujing Hu, Yu Ding 0001 |
ECCV (55) | 9 |
| 2024 | FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation ModelabstractVideo-driven 3D facial animation transfer aims to drive avatars to reproduce the expressions of actors. Existing methods have achieved remarkable results by constraining both geometric and perceptual consistency. However, geometric constraints (like those designed on facial landmarks) are insufficient to capture subtle emotions, while expression features trained on classification tasks lack fine granularity for complex emotions. To address this, we propose \textbf{FreeAvatar}, a robust facial animation transfer method that relies solely on our learned expression representation. Specifically, FreeAvatar consists of two main components: the expression foundation model and the facial animation transfer model. In the first component, we initially construct a facial feature space through a face reconstruction task and then optimize the expression feature space by exploring the similarities among different expressions. Benefiting from training on the amounts of unlabeled facial images and re-collected expression comparison dataset, our model adapts freely and effectively to any in-the-wild input facial images. In the facial animation transfer component, we propose a novel Expression-driven Multi-avatar Animator, which first maps expressive semantics to the facial control parameters of 3D avatars and then imposes perceptual constraints between the input and output images to maintain expression consistency. To make the entire process differentiable, we employ a trained neural renderer to translate rig parameters into corresponding images. Furthermore, unlike previous methods that require separate decoders for each avatar, we propose a dynamic identity injection module that allows for the joint training of multiple avatars within a single network. Wei Zhang 0219, Chen Liu 0028, Rudong An, Lincheng Li, Yu Ding 0001, Changjie Fan, Zhipeng Hu, Xin Yu 0002 |
SIGGRAPH Asia | 6 |
| 2024 | Learning facial expression-aware global-to-local representation for robust action unit detection
Rudong An, Aobo Jin, Wei Chen 0157, Wei Zhang 0219, Hao Zeng 0001, Zhigang Deng 0001, Yu Ding 0001 |
Appl. Intell. | 7 |
| 2024 | Emotion knowledge-based fine-grained facial expression recognition
Yu Ding 0001, Hanwei Liu, Zhanpeng Lin, Wenxing Hong |
Neurocomputing | 2 |
| 2024 | Learning a compact embedding for fine-grained few-shot static gesture recognition
Zhipeng Hu, Wei Zhang 0219, Yu Ding 0001, Tangjie Lv, Changjie Fan |
Multim. Tools Appl. | 5 |
| 2024 | StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking HeadsabstractIndividuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse speaking styles in the final videos. To address this challenge, we propose a one-shot style-controllable talking face generation method that can obtain speaking styles from reference speaking videos and drive the one-shot portrait to speak with the reference speaking styles and another piece of audio. Our method aims to synthesize the style-controllable coefficients of a 3D Morphable Model (3DMM), including facial expressions and head movements, in a unified framework. Specifically, the proposed framework first leverages a style encoder to extract the desired speaking styles from the reference videos and transform them into style codes. Then, the framework uses a style-aware decoder to synthesize the coefficients of 3DMM from the audio input and style codes. During decoding, our framework adopts a two-branch architecture, which generates the stylized facial expression coefficients and stylized head movement coefficients, respectively. After obtaining the coefficients of 3DMM, an image renderer renders the expression coefficients into a specific person's talking-head video. Extensive experiments demonstrate that our method generates visually authentic talking head videos with diverse speaking styles from only one portrait image and an audio clip. Suzhen Wang 0001, Yifeng Ma 0001, Yu Ding 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Zhidong Deng, Xin Yu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Facial Action Unit Detection and Intensity Estimation From Self-Supervised RepresentationabstractAs a fine-grained and local expression behavior measurement, facial action unit (FAU) analysis (e.g., detection and intensity estimation) has been documented for its time-consuming, labor-intensive, and error-prone annotation. Thus a long-standing challenge of FAU analysis arises from the data scarcity of manual annotations, limiting the generalization ability of trained models to a large extent. Amounts of previous works have made efforts to alleviate this issue via semi/weakly supervised methods and extra auxiliary information. However, these methods still require domain knowledge and have not yet avoided the high dependency on data annotation. This article introduces a robust facial representation model MAE-Face for AU analysis. Using masked autoencoding as the self-supervised pre-training approach, MAE-Face first learns a high-capacity model from a feasible collection of face images without additional data annotations. Then after being fine-tuned on AU datasets, MAE-Face exhibits convincing performance for both AU detection and AU intensity estimation, achieving a new state-of-the-art on nearly all the evaluation results. Further investigation shows that MAE-Face achieves decent performance even when fine-tuned on only 1% of the AU training set, strongly proving its robustness and generalization performance. The pre-trained model is available at our GitHub repository. Rudong An, Wei Zhang 0219, Yu Ding 0001, Zeng Zhao, Tangjie Lv, Changjie Fan, Zhipeng Hu |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Detecting Facial Action Units From Global-Local Fine-Grained ExpressionsabstractSince Facial Action Unit (AU) annotations require domain expertise, common AU datasets only contain a limited number of subjects. As a result, a crucial challenge for AU detection is addressing identity overfitting. We find that AUs and facial expressions are highly associated, and existing facial expression datasets often contain a large number of identities. In this paper, we aim to utilize the expression datasets without AU labels to facilitate AU detection. Specifically, we develop a novel AU detection framework aided by the Global-Local facial Expressions Embedding, dubbed GLEE-Net. Our GLEE-Net consists of three branches to extract identity-independent expression features for AU detection. We introduce a global branch for modeling the overall facial expression while eliminating the impacts of identities. We also design a local branch focusing on specific local face regions. The combined output of global and local branches is firstly pre-trained on an expression dataset as an identity-independent expression embedding, and then finetuned on AU datasets. Therefore, we significantly alleviate the issue of limited identities. Furthermore, we introduce a 3D global branch that extracts expression coefficients through 3D face reconstruction to consolidate 2D AU descriptions. Finally, a Transformer-based multi-label classifier is employed to fuse all the representations for AU detection. Extensive experiments demonstrate that our method significantly outperforms the state-of-the-art on the widely-used DISFA, BP4D and BP4D+ datasets. Wei Zhang 0219, Lincheng Li, Yu Ding 0001, Wei Chen 0157, Zhigang Deng 0001, Xin Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesabstractDifferent people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate diverse speaking styles in the final talking head videos. To tackle this problem, we propose a one-shot style-controllable talking face generation framework. In a nutshell, we aim to attain a speaking style from an arbitrary reference speaking video and then drive the one-shot portrait to speak with the reference speaking style and another piece of audio. Specifically, we first develop a style encoder to extract dynamic facial motion patterns of a style reference video and then encode them into a style code. Afterward, we introduce a style-controllable decoder to synthesize stylized facial animations from the speech content and style code. In order to integrate the reference speaking style into generated videos, we design a style-aware adaptive transformer, which enables the encoded style code to adjust the weights of the feed-forward layers accordingly. Thanks to the style-aware adaptation mechanism, the reference speaking style can be better embedded into synthesized videos during decoding. Extensive experiments demonstrate that our method is capable of generating talking head videos with diverse speaking styles from only one portrait image and an audio clip while achieving authentic visual effects. Project Page: https://github.com/FuxiVirtualHuman/styletalk. Yifeng Ma 0001, Suzhen Wang 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001, Zhidong Deng, Xin Yu 0002 |
AAAI | 6 |
| 2023 | FlowFace: Semantic Flow-Guided Shape-Aware Face SwappingabstractIn this work, we propose a semantic flow-guided two-stage framework for shape-aware face swapping, namely FlowFace. Unlike most previous methods that focus on transferring the source inner facial features but neglect facial contours, our FlowFace can transfer both of them to a target face, thus leading to more realistic face swapping. Concretely, our FlowFace consists of a face reshaping network and a face swapping network. The face reshaping network addresses the shape outline differences between the source and target faces. It first estimates a semantic flow (i.e. face shape differences) between the source and the target face, and then explicitly warps the target face shape with the estimated semantic flow. After reshaping, the face swapping network generates inner facial features that exhibit the identity of the source face. We employ a pre-trained face masked autoencoder (MAE) to extract facial features from both the source face and the target face. In contrast to previous methods that use identity embedding to preserve identity information, the features extracted by our encoder can better capture facial appearances and identity information. Then, we develop a cross-attention fusion module to adaptively fuse inner facial features from the source face with the target facial attributes, thus leading to better identity preservation. Extensive quantitative and qualitative experiments on in-the-wild faces demonstrate that our FlowFace outperforms the state-of-the-art significantly. Hao Zeng 0001, Wei Zhang 0219, Changjie Fan, Tangjie Lv, Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Xin Yu 0002 |
AAAI | 9 |
| 2023 | DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoabstractFor few-shot learning, it is still a critical challenge to realize photo-realistic face visually dubbing on high-resolution videos. Previous works fail to generate high-fidelity dubbing results. To address the above problem, this paper proposes a Deformation Inpainting Network (DINet) for high-resolution face visually dubbing. Different from previous works relying on multiple up-sample layers to directly generate pixels from latent embeddings, DINet performs spatial deformation on feature maps of reference images to better preserve high-frequency textural details. Specifically, DINet consists of one deformation part and one inpainting part. In the first part, five reference facial images adaptively perform spatial deformation to create deformed feature maps encoding mouth shapes at each frame, in order to align with input driving audio and also the head poses of input source images. In the second part, to produce face visually dubbing, a feature decoder is responsible for adaptively incorporating mouth movements from the deformed feature maps and other attributes (i.e., head pose and upper facial expression) from the source feature maps together. Finally, DINet achieves face visually dubbing with rich textural details. We conduct qualitative and quantitative comparisons to validate our DINet on high-resolution videos. The experimental results show that our method outperforms state-of-the-art works. Zhipeng Hu, Wenjin Deng, Changjie Fan, Tangjie Lv, Yu Ding 0001 |
AAAI | 6 |
| 2023 | Exploring Complementary Features in Multi-Modal Speech Emotion RecognitionabstractSpeech emotion recognition (SER) is of great importance in human-computer interaction. Recent research has demonstrated that self-supervised learned acoustic and linguistic features are helpful in this task. However, few works have fully exploit the advantages of the pre-trained features in SER. The primary challenge is how to effectively extract the complementary emotional information implied in the pre-trained features of the respective modality. To tackle this challenge, we propose a novel modality-sensitive multimodal speech emotion recognition framework. In a nutshell, we aim to exploit the typical emotion features in each modality and then fuse the complementary emotional information for classification. Specifically, we first utilize the parallel uni-modal encoders to refine the emotion-related information from the pre-trained features of each modality. For better fusion of the multimodal features, we develop a group of learnable emotion query tokens to gather the emotional information from the refined acoustic and linguistic features with the cross-attention mechanism in the transformer decoder. Observing the modality bias problem in multimodal methods, we introduce the random modality masking training strategy to maximize the utilization of the emotional information in each modality and mitigate this problem. We evaluate our method on the widely used IEMOCAP dataset and achieve 1.1% and 0.9% improvements on the unweighted accuracy and weighted accuracy, respectively. Extensive experiments demonstrate the effectiveness of the proposed method. Suzhen Wang 0001, Yifeng Ma 0001, Yu Ding 0001 |
ICASSP | 3 |
| 2023 | Real-time Facial Animation for 3D Stylized Character with Emotion DynamicsabstractOur aim is to improve animation production techniques' efficiency and effectiveness. We present two real-time solutions which drive character expressions in a geometrically consistent and perceptually valid way. Our first solution combines keyframe animation techniques with machine learning models. We propose a 3D emotion transfer network makes use of a 2D human image to generate a stylized 3D rig parameter. Our second solution combines blendshape-based motion capture animation techniques with machine learning models. We propose a blendshape adaption network which generates the character rig parameter motions with geometric consistency and temporally stability. We demonstrate the effectiveness of our system by comparing it to a commercial product Faceware. Results reveal that ratings of the recognition, intensity, and attractiveness of expressions depicted for animated characters via our systems are statistically higher than Faceware. Our results may be implemented into the animation pipeline, supporting animators to create expressions more rapidly and precisely. Ruisi Zhang, Yu Ding 0001, Kenny Mitchell |
ACM Multimedia | 4 |
| 2023 | Fully Automatic Blendshape Generation for Stylized CharactersabstractAvatars are one of the most important elements in virtual environments. Real-time facial retargeting technology is of vital importance in AR/VR interactions, the filmmaking, and the entertainment industry, and blendshapes for avatars are one of its important materials. Previous works either focused on the characters with the same topology, which cannot be generalized to universal avatars, or used optimization methods that have high demand on the dataset. In this paper, we adopt the essence of deep learning and feature transfer to realize deformation transfer, thereby generating blendshapes for target avatars based on the given sources. We proposed a Variational Autoencoder (VAE) to extract the latent space of the avatars and then use a Multilayer Perceptron (MLP) model to realize the translation between the latent spaces of the source avatar and target avatars. By decoding the latent code of different blendshapes, we can obtain the blendshapes for the target avatars with the same semantics as that of the source. We qualitatively and quantitatively compared our method with both classical and learning-based methods. The results revealed that the blendshapes generated by our method achieves higher similarity to the groundtruth blendshapes than the state-of-art methods. We also demonstrated that our method can be applied to expression transfer for stylized characters with different topologies. Yilin Qiu, Yu Ding 0001 |
VR | 4 |
| 2023 | Deep learning applications in games: a survey from a data perspective
Zhipeng Hu, Yu Ding 0001, Runze Wu 0001, Lincheng Li, Yujing Hu, Kai Wang 0064, Yongqiang Zhang 0003, Ji Jiang, Yadong Xi, Jiashu Pu, Wei Zhang 0219, Suzhen Wang 0001, Ke Chen 0005, Tianze Zhou, Jiarui Chen, Tangjie Lv, Changjie Fan |
Appl. Intell. | 2 |
| 2023 | Face identity and expression consistency for game character face swapping
Hao Zeng 0001, Wei Zhang 0219, Lincheng Li, Yu Ding 0001 |
Comput. Vis. Image Underst. | 6 |
| 2023 | BAFN: Bi-Direction Attention Based Fusion Network for Multimodal Sentiment AnalysisabstractAttention-based networks currently identify their effectiveness in multimodal sentiment analysis. However, existing methods ignore the redundancy of auxiliary modalities. More importantly, existing methods only attend to top-down attention (static process) or down-top attention (implicit process), leading to the coarse-grained multimodal sentiment context. In this paper, during the preprocessing period, we first propose the multimodal dynamic enhanced block to capture the intra-modality sentiment context. This can effectively decrease the intra-modality redundancy of auxiliary modalities. Furthermore, the bi-direction attention block is proposed to capture fine-grained multimodal sentiment context via the novel bi-direction multimodal dynamic routing mechanism. Specifically, the bi-direction attention block first highlights the explicit and low-level multimodal sentiment context. Then, the low-level multimodal context is transmitted to a carefully designed bi-direction multimodal dynamic routing procedure. This allows us to dynamically update and investigate high-level and much more fine-grained multimodal sentiment contexts. The experiments demonstrate that our fusion network can achieve state-of-the-art performance. Notably, our model outperforms the best baseline on the metric ‘Acc-7’ with an improvement of 6.9%. Jiajia Tang, Dongjun Liu, Xuanyu Jin, Yong Peng 0001, Qibin Zhao, Yu Ding 0001, Wanzeng Kong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | A Music-Driven Deep Generative Adversarial Model for Guzheng Playing AnimationabstractTo date relatively few efforts have been made on the automatic generation of musical instrument playing animations. This problem is challenging due to the intrinsically complex, temporal relationship between music and human motion as well as the lacking of high quality music-playing motion datasets. In this article, we propose a fully automatic, deep learning based framework to synthesize realistic upper body animations based on novel guzheng music input. Specifically, based on a recorded audiovisual motion capture dataset, we delicately design a generative adversarial network (GAN) based approach to capture the temporal relationship between the music and the human motion data. In this process, data augmentation is employed to improve the generalization of our approach to handle a variety of guzheng music inputs. Through extensive objective and subjective experiments, we show that our method can generate visually plausible guzheng-playing animations that are well synchronized with the input guzheng music, and it can significantly outperform the state-of-the-art methods. In addition, through an ablation study, we validate the contributions of the carefully-designed modules in our framework. Changjie Fan, Gongzheng Li, Zeng Zhao, Zhigang Deng 0001, Yu Ding 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | Emotional Voice PuppetryabstractThe paper presents emotional voice puppetry, an audio-based facial animation approach to portray characters with vivid emotional changes. The lips motion and the surrounding facial areas are controlled by the contents of the audio, and the facial dynamics are established by category of the emotion and the intensity. Our approach is exclusive because it takes account of perceptual validity and geometry instead of pure geometric processes. Another highlight of our approach is the generalizability to multiple characters. The findings showed that training new secondary characters when the rig parameters are categorized as eye, eyebrows, nose, mouth, and signature wrinkles is significant in achieving better generalization results compared to joint training. User studies demonstrate the effectiveness of our approach both qualitatively and quantitatively. Our approach can be applicable in AR/VR and 3DUI, namely, virtual reality avatars/self-avatars, teleconferencing and in-game dialogue. Ruisi Zhang, Shengran Cheng, Shuai Tan 0002, Yu Ding 0001, Kenny Mitchell, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation LearningabstractAudio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those methods struggle to learn a consistent speech style from different speakers. We observe that it would be much easier to learn a consistent speech style from a specific speaker, which leads to authentic mouth movements. Hence, we propose a novel one-shot talking face generation framework by exploring consistent correlations between audio and visual motions from a specific speaker and then transferring audio-driven motion fields to a reference image. Specifically, we develop an Audio-Visual Correlation Transformer (AVCT) that aims to infer talking motions represented by keypoint based dense motion fields from an input audio. In particular, considering audio may come from different identities in deployment, we incorporate phonemes to represent audio signals. In this manner, our AVCT can inherently generalize to audio spoken by other identities. Moreover, as face keypoints are used to represent speakers, AVCT is agnostic against appearances of the training speaker, and thus allows us to manipulate face images of different identities readily. Considering different face shapes lead to different motions, a motion field transfer module is exploited to reduce the audio-driven dense motion field gap between the training identity and the one-shot reference. Once we obtained the dense motion field of the reference image, we employ an image renderer to generate its talking face videos from an audio clip. Thanks to our learned consistent speaking style, our method generates authentic mouth shapes and vivid movements. Extensive experiments demonstrate that our synthesized videos outperform the state-of-the-art in terms of visual quality and lip-sync. Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Xin Yu 0002 |
AAAI | 3 |
| 2022 | Multi-Dimensional Prediction of Guild Health in Online Games: A Stability-Aware Multi-Task Learning ApproachabstractGuild is the most important long-term virtual community and emotional bond in massively multiplayer online role-playing games (MMORPGs). It matters a lot to the player retention and game ecology how the guilds are going, e.g., healthy or not. The main challenge now is to characterize and predict the guild health in a quantitative, dynamic, and multi-dimensional manner based on complicated multi-media data streams. To this end, we propose a novel framework, namely Stability-Aware Multi-task Learning Approach(SAMLA) to address these challenges. Specifically, different media-specific modules are designed to extract information from multiple media types of tabular data, time seriescharacteristics, and heterogeneous graphs. To capture the dynamics of guild health, we introduce a representation encoder to provide a time series view of multi-media data that is used for task prediction. Inspiredby well-received theories on organization management, we delicately define five specific and quantitative dimensions of guild health and make parallel predictions based on a multi-task approach. Besides, we devise a novel auxiliary task, i.e.,the guild stability, to boost the performance of the guild health prediction task. Extensive experiments on a real-world large-scale MMORPG dataset verify that our proposed method outperforms the state-of-the-art methods in the task of organizational health characterization and prediction. Moreover, our work has been practically deployed in online MMORPG, and case studies clearly illustrate the significant value. Chuang Zhao 0002, Hongke Zhao, Runze Wu 0001, Yu Ding 0001, Jianrong Tao, Changjie Fan |
AAAI | 5 |
| 2022 | Paste You Into Game: Towards Expression and Identity Consistency Face SwappingabstractCustomizing game characters for individual players has been a long-standing attractive feature in the game industry. However, traditional solutions like manual editing within a game engine are always time-consuming and unsatisfying. Our work proposes a novel automatic face swapping method for arbitrary users and game characters, addressing three challenges including style gap between human and game faces, identity preservation, and expression consistency. A game face dataset is collected to handle the cross-style gap; an identity compound embedding is proposed to ease the bias existing in the commonly-used ID identifiers and it provides a more robust identity representation; a novel expression embedding loss is proposed to enforce the expression consistency between the swapped and target faces and it achieves better expression consistency than the previous methods, especially when the expression is very subtle. The visualized results, as well as the qualitative and quantitative comparisons, reveal the significance and effectiveness of our proposed solutions. Hao Zeng 0001, Wei Zhang 0219, Lincheng Li, Yu Ding 0001 |
CoG | 6 |
| 2022 | MMT: Multi-way Multi-modal Transformer for Multimodal LearningabstractThe heart of multimodal learning research lies the challenge of effectively exploiting fusion representations among multiple modalities.However, existing two-way cross-modality unidirectional attention could only exploit the intermodal interactions from one source to one target modality. This indeed fails to unleash the complete expressive power of multimodal fusion with restricted number of modalities and fixed interactive direction.In this work, the multiway multimodal transformer (MMT) is proposed to simultaneously explore multiway multimodal intercorrelations for each modality via single block rather than multiple stacked cross-modality blocks. The core idea of MMT is the multiway multimodal attention, where the multiple modalities are leveraged to compute the multiway attention tensor. This naturally benefits us to exploit comprehensive many-to-many multimodal interactive paths. Specifically, the multiway tensor is comprised of multiple interconnected modality-aware core tensors that consist of the intramodal interactions. Additionally, the tensor contraction operation is utilized to investigate intermodal dependencies between distinct core tensors.Essentially, our tensor-based multiway structure allows for easily extending MMT to the case associated with an arbitrary number of modalities. Taking MMT as the basis, the hierarchical network is further established to recursively transmit the low-level multiway multimodal interactions to high-level ones. The experiments demonstrate that MMT can achieve state-of-the-art or comparable performance. Jiajia Tang, Xuanyu Jin, Wanzeng Kong, Yu Ding 0001, Qibin Zhao |
IJCAI | 6 |
| 2022 | Dynamically Adjust Word Representations Using Unaligned Multimodal InformationabstractMultimodal Sentiment Analysis is a promising research area for modeling multiple heterogeneous modalities. Two major challenges that exist in this area are a) multimodal data is unaligned in nature due to the different sampling rates of each modality, and b) long-range dependencies between elements across modalities. These challenges increase the difficulty of conducting efficient multimodal fusion. In this work, we propose a novel end-to-end network named Cross Hyper-modality Fusion Network (CHFN). The CHFN is an interpretable Transformer-based neural model that provides an efficient framework for fusing unaligned multimodal sequences. The heart of our model is to dynamically adjust word representations in different non-verbal contexts using unaligned multimodal sequences. It is concerned with the influence of non-verbal behavioral information at the scale of the entire utterances and then integrates this influence into verbal expression. We conducted experiments on both publicly available multimodal sentiment analysis datasets CMU-MOSI and CMU-MOSEI. The experiment results demonstrate that our model surpasses state-of-the-art models. In addition, we visualize the learned interactions between language modality and non-verbal behavior information and explore the underlying dynamics of multimodal language data. Jiwei Guo, Jiajia Tang, Weichen Dai 0001, Yu Ding 0001, Wanzeng Kong |
ACM Multimedia | 4 |
| 2022 | Adaptive Affine Transformation: A Simple and Effective Operation for Spatial Misaligned Image GenerationabstractOne challenging problem, named spatial misaligned image generation, describing a translation between two face/pose images with large spatial deformation, is widely faced in tasks of face/pose reenactment. Advanced researchers use the dense flow to solve this problem. However, under a complex spatial deformation, even using carefully designed networks, intrinsical complexities make it difficult to compute an accurate dense flow, leading to distorted results. Different from those dense flow based methods, we propose one simple but effective operator named AdaAT (Adaptive Affine Transformation) to realize misaligned image generation. AdaAT simulates spatial deformation by computing hundreds of affine transformations, resulting in less distortions. Without computing any dense flow, AdaAT directly carries out affine transformations in feature channel spaces. Furthermore, we package several AdaAT operators to one universal AdaAT module that is used for different face/pose generation tasks. To validate the effectiveness of our AdaAT, we conduct qualitative and quantitative experiments on four common datasets in the tasks of talking face generation, face reenactment, pose transfer and person image generation. We achieve state-of-the-art results on three of them. Yu Ding 0001 |
ACM Multimedia | 2 |
| 2022 | Semantic-Rich Facial Emotional Expression RecognitionabstractThe ability to perceive human facial emotions is an essential feature of various multi-modal applications, especially in the intelligent human-computer interaction (HCI) area. In recent decades, considerable efforts have been put into researching automatic facial emotion recognition (FER). However, most of the existing FER methods only focus on either basic emotions such as the seven/eight categories (e.g.,happiness, angerandsurprise) or abstract dimensions (valence, arousal, etc.), while neglecting the fruitful nature of emotion statements. In real-world scenarios, there is definitely a larger vocabulary for describing human's inner feelings as well as their reflection on facial expressions. In this work, we propose to address the semantic richness issue in the FER problem, with an emphasis on the granularity of the emotion concepts. Particularly, we take inspiration from former psycho-linguistic research, which conducted a prototypicality rating study and chose 135 emotion names from hundreds of English emotion terms. Based on the 135 emotion categories, we investigate the corresponding facial expressions by collecting a large-scale 135-class FER image dataset and propose a consequent facial emotion recognition framework. To demonstrate the accessibility of prompting FER research to a fine-grained level, we conduct extensive evaluations on the dataset credibility and the accompanying baseline classification model. The qualitative and quantitative results prove that the problem is meaningful and our solution is effective. To the best of our knowledge, this is the first work aimed at exploiting such a large semantic space for emotion representation in the FER problem. Changjie Fan, Wei Zhang 0219, Yu Ding 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2021 | Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationabstractIn this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be specific, our framework consists of a speaker-independent stage and a speaker-specific stage. In the speaker-independent stage, we design three parallel networks to generate animation parameters of the mouth, upper face, and head from texts, separately. In the speaker-specific stage, we present a 3D face model guided attention network to synthesize videos tailored for different individuals. It takes the animation parameters as input and exploits an attention mask to manipulate facial expression changes for the input individuals. Furthermore, to better establish authentic correspondences between visual motions (i.e., facial expression changes and head movements) and audios, we leverage a high-accuracy motion capture dataset instead of relying on long videos of specific individuals. After attaining the visual and audio correspondences, we can effectively train our network in an end-to-end fashion. Extensive experiments on qualitative and quantitative results demonstrate that our algorithm achieves high-quality photo-realistic talking-head videos including various facial expressions and head motions according to speech rhythms and outperforms the state-of-the-art. Lincheng Li, Suzhen Wang 0001, Yu Ding 0001, Yixing Zheng, Xin Yu 0002, Changjie Fan |
AAAI | 4 |
| 2021 | Learning a Facial Expression Embedding Disentangled From IdentityabstractThe facial expression analysis requires a compact and identity-ignored expression representation. In this paper, we model the expression as the deviation from the identity by a subtraction operation, extracting a continuous and identity-invariant expression embedding. We propose a Deviation Learning Network (DLN) with a pseudo-siamese structure to extract the deviation feature vector. To reduce the optimization difficulty caused by additional fully connection layers, DLN directly provides high-order polynomial to nonlinearly project the high-dimensional feature to a low-dimensional manifold. Taking label noise into account, we add a crowd layer to DLN for robust embedding extraction. Also, to achieve a more compact representation, we use hierarchical annotation for data augmentation. We evaluate our facial expression embedding on the FEC validation set. The quantitative results prove that we achieve the state-of-the-art, both in terms of fine-grained and identity-invariant property. We further conduct extensive experiments to show that our expression embedding is of high quality for expression recognition, image retrieval, and face manipulation. Wei Zhang 0219, Xianpeng Ji, Yu Ding 0001, Changjie Fan |
CVPR | 4 |
| 2021 | Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual DatasetabstractOne-shot talking face generation should synthesize high visual quality facial videos with reasonable animations of expression and head pose, and just utilize arbitrary driving audio and arbitrary single face image as the source. Current works fail to generate over 256×256 resolution realistic-looking videos due to the lack of an appropriate high-resolution audio-visual dataset, and the limitation of the sparse facial landmarks in providing poor expression details. To synthesize high-definition videos, we build a large in-the-wild high-resolution audio-visual dataset and propose a novel flow-guided talking face generation framework. The new dataset is collected from youtube and consists of about 16 hours 720P or 1080P videos. We leverage the facial 3D morphable model (3DMM) to split the framework into two cascaded modules instead of learning a direct mapping from audio to video. In the first module, we propose a novel animation generator to produce the movements of mouth, eyebrow and head pose simultaneously. In the second module, we transform animation into dense flow to provide more expression details and carefully design a novel flow-guided video generator to synthesize videos. Our method is able to produce high-definition videos and outperforms state-of-the-art works in objective and subjective comparisons*. Lincheng Li, Yu Ding 0001, Changjie Fan |
CVPR | 3 |
| 2021 | Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head MotionabstractWe propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)} maintaining the appearance of a speaker in a large head motion while stabilizing the non-face regions. We first design a head pose predictor by modeling rigid 6D head movements with a motion-aware recurrent neural network (RNN). In this way, the predicted head poses act as the low-frequency holistic movements of a talking head, thus allowing our latter network to focus on detailed facial movement generation. To depict the entire image motions arising from audio, we exploit a keypoint based dense motion field representation. Then, we develop a motion field generator to produce the dense motion fields from input audio, head poses, and a reference image. As this keypoint based representation models the motions of facial regions, head, and backgrounds integrally, our method can better constrain the spatial and temporal consistency of the generated videos. Finally, an image generation network is employed to render photo-realistic talking-head videos from the estimated keypoint based motion fields and the input reference image. Extensive experiments demonstrate that our method produces videos with plausible head motions, synchronized facial expressions, and stable backgrounds and outperforms the state-of-the-art. Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Changjie Fan, Xin Yu 0002 |
IJCAI | 3 |
| 2021 | Build Your Own Bundle - A Neural Combinatorial Optimization MethodabstractIn the business domain,bundling is one of the most important marketing strategies to conduct product promotions, which is commonly used in online e-commerce and offline retailers. Existing recommender systems mostly focus on recommending individual items that users may be interested in, such as the considerable research work on collaborative filtering that directly models the interaction between users and items. In this paper, we target at a practical but less explored recommendation problem named personalized bundle composition, which aims to offer an optimal bundle (i.e., a combination of items) to the target user. To tackle this specific recommendation problem, we formalize it as a combinatorial optimization problem on a set of candidate items and solve it within a neural combinatorial optimization framework. Extensive experiments on public datasets are conducted to demonstrate the superiority of the proposed method. Kai Wang 0064, Minghao Zhao 0002, Runze Wu 0001, Yu Ding 0001, Zhene Zou, Jianrong Tao, Changjie Fan |
ACM Multimedia | 5 |
| 2021 | Learning a deep motion interpolation network for human skeleton animationsabstractAbstract Motion interpolation technology produces transition motion frames between two discrete movements. It is wildly used in video games, virtual reality and augmented reality. In the fields of computer graphics and animations, our data‐driven method generates transition motions of two arbitrary animations without additional control signals. In this work, we propose a novel carefully designed deep learning framework, named deep motion interpolation network (DMIN), to learn human movement habits from a real dataset and then to perform the interpolation function specific for human motions. It is a data‐driven approach to capture overall rhythm of two given discrete movements and generate natural in‐between motion frames. The sequence‐by‐sequence architecture allows completing all missing frames within single forward inference, which reduces computation time for interpolation. Experiments on human motion datasets show that our network achieves promising interpolation performance. The ablation study demonstrates the effectiveness of the carefully designed DMIN.1 Chi Zhou 0005, Zhangjiong Lai, Suzhen Wang 0001, Lincheng Li, Yu Ding 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2020 | FReeNet: Multi-Identity Face ReenactmentabstractThis paper presents a novel multi-identity face reenactment framework, named FReeNet, to transfer facial expressions from an arbitrary source face to a target face with a shared model. The proposed FReeNet consists of two parts: Unified Landmark Converter (ULC) and Geometry-aware Generator (GAG). The ULC adopts an encode-decoder architecture to efficiently convert expression in a latent landmark space, which significantly narrows the gap of the face contour between source and target identities. The GAG leverages the converted landmark to reenact the photorealistic image with a reference image of the target person. Moreover, a new triplet perceptual loss is proposed to force the GAG module to learn appearance and geometry information simultaneously, which also enriches facial details of the reenacted images. Further experiments demonstrate the superiority of our approach for generating photorealistic and expression-alike faces, as well as the flexibility for transferring facial expressions between identities. Jiangning Zhang, Xianfang Zeng, Mengmeng Wang 0005, Yusu Pan, Liang Liu 0007, Yong Liu 0007, Yu Ding 0001, Changjie Fan |
CVPR | 7 |
| 2020 | One-Shot Voice Conversion Using Star-GanabstractOur efforts are made on one-shot voice conversion where the target speaker is unseen in training dataset or both source and target speakers are unseen in the training dataset. In our work, StarGAN is employed to carry out voice conversion between speakers. An embedding vector is used to represent speaker ID. This work relies on two datasets in English and one dataset in Chinese, involving 38 speakers. A user study is conducted to validate our framework in terms of reconstruction quality and conversion quality. The results show that our framework is able to perform one-shot voice conversion and also outperforms state-of-the-art methods when the speaker in the test is seen in the training dataset. The exploration experiment demonstrates that our framework can be updated with incremental training when the data from new speakers is available. Ruobai Wang, Yu Ding 0001, Lincheng Li, Changjie Fan |
ICASSP | 2 |
| 2020 | Low-Level Characterization of Expressive Head Motion Through Frequency Domain AnalysisabstractFor the purpose of understanding how head motions contribute to the perception of emotion in an utterance, we aim to examine the perception of emotion based on Fourier transform-based static and dynamic features of head motion. Our work is to conduct intra-related objective analysis and perceptual experiments on the link between the perception of emotion and the static/dynamic features. The objective analysis outcome shows that the static and dynamic features are effective in characterizing and recognizing emotions. The perceptual experiments enable us to collect human perception of emotion through head motion. The collected perceptual data shows that humans are unable to reliably perceive emotion from head motion alone but reveals that humans are sensitive to the static feature (in reference to the averaged up-down rotation angle) and the dynamic features (which reflect the fluidity and speed of movement). It also indicates that humans perceive emotion carried in head motion and the naturalness of head motion in two different channels. Our work contributes to the understanding and the characterization of head motion in expressive speech through low-level descriptions of motion features, instead of commonly used high-level motion style (e.g., head nods, shakes, tilts, and raises). Yu Ding 0001, Lei Shi 0027, Zhigang Deng 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Text-driven Visual Prosody Generation for Embodied Conversational AgentsabstractIn face-to-face conversations, head motions play a crucial role in encoding information, and humans are very skilled at decoding multiple messages from interlocutors' head motions. It is of great importance to endow embodied conversational agents (ECAs) with the capability of conveying communicative intention through head movements. Our work is aimed at automatically synthesizing head motions for an ECA speaking Chinese. We propose to take only transcripts as input to compute head movements, based on a statistical framework. Subjective experiments are conducted to validate the proposed statistical framework. The results show that the generated head animation is able to improve human perception in terms of naturalness and demonstrate that the head animation is synchronized with the input of synthetic speech. Yong Liu 0007, Changjie Fan, Yu Ding 0001 |
IVA | 5 |
| 2017 | Perceptual enhancement of emotional mocap head motion: An experimental studyabstractMotion capture (mocap) systems have been widely used to collect various human behavior data. Despite existing numerous research efforts on mocap motion processing and understanding, to the best of our knowledge, to date few works have been dedicated to the investigation into whether the mo-cap human behavior data can be further enhanced to improve its perception. In this work, we investigate whether and how it is feasible to consistently manipulate mocap emotional head motion to enhance its perceived expressiveness. Our study relies on a mocap audiovisual dataset acquired in a laboratory setting. Participants are invited to view the animation clips of a virtual talking character displaying the original mocap head motion or manipulated head motion, and then to rate their perceived expressiveness. Statistical analysis of the rated perceptions shows that humans are sensitive to the mean of head pitch rotation (called up-down rotation) in an utterance and that the expressiveness of emotion could be improved by adjusting the mean of head pitch rotation in an utterance. Yu Ding 0001, Lei Shi 0027, Zhigang Deng 0001 |
ACII | 1 |
| 2017 | A Multifaceted Study on Eye Contact based Speaker Identification in Three-party ConversationsabstractTo precisely understand human gaze behaviors in three-party conversations, this work is dedicated to look into whether the speaker can be reliably identified from the interlocutors in a three-party conversation on the basis of the interactive behaviors of eye contact, where speech signals are not provided. Derived from a pre-recorded, multimodal, and three-party conversational behavior dataset, a statistical framework is pro- posed to determine who is the speaker from the interactive behaviors of eye contact. Additionally, with the aid of virtual human technologies, a user study is conducted to study whether subjects are capable of distinguishing the speaker from the listeners according to the gaze behaviors of the interlocutors alone. Our results show that eye contact provides a reliable cue for the identification of the speaker in three-party conversations. Yu Ding 0001, Meihua Xiao, Zhigang Deng 0001 |
CHI | 1 |
| 2017 | Audio-Driven Laughter Behavior ControllerabstractIt has been well documented that laughter is an important communicative and expressive signal in face-to-face conversations. Our work aims at building a laughter behavior controller for a virtual character which is able to generate upper body animations from laughter audio given as input. This controller relies on the tight correlations between laughter audio and body behaviors. A unified continuous-state statistical framework, inspired by Kalman filter, is proposed to learn the correlations between laughter audio and head/torso behavior from a recorded laughter human dataset. Due to the lack of shoulder behavior data in the recorded human dataset, a rule-based method is defined to model the correlation between laughter audio and shoulder behavior. In the synthesis step, these characterized correlations are rendered in the animation of a virtual character. To validate our controller, a subjective evaluation is conducted where participants viewed the videos of a laughing virtual character. It compares the animations of a virtual character using our controller and a state of the art method. The evaluation results show that the laughter animations computed with our controller are perceived as more natural, expressing amusement more freely and appearing more authentic than with the state of the art method. Yu Ding 0001, Jing Huang 0005, Catherine Pelachaud |
IEEE Trans. Affect. Comput. | 1 |
| 2017 | Implementing and Evaluating a Laughing Virtual CharacterabstractLaughter is a social signal capable of facilitating interaction in groups of people: it communicates interest, helps to improve creativity, and facilitates sociability. This article focuses on: endowing virtual characters with computational models of laughter synthesis, based on an expressivity-copying paradigm; evaluating how the physically co-presence of the laughing character impacts on the user’s perception of an audio stimulus and mood. We adopt music as a means to stimulate laughter. Results show that the character presence influences the user’s perception of music and mood. Expressivity-copying has an influence on the user’s perception of music, but does not have any significant impact on mood. Maurizio Mancini, Béatrice Biancardi, Florian Pecune, Giovanna Varni, Yu Ding 0001, Catherine Pelachaud, Gualtiero Volpe, Antonio Camurri |
ACM Trans. Internet Techn. | 5 |
| 2017 | Inverse kinematics using dynamic joint parameters: inverse kinematics animation synthesis learnt from sub-divided motion micro-segments
Jing Huang 0005, Marco Fratarcangeli, Yu Ding 0001, Catherine Pelachaud |
Vis. Comput. | 3 |
| 2015 | LOL - Laugh Out LoudabstractIn our demo, LoL, a user interacts with a virtual agentable to copy and to adapt its laughing and expressive behaviorson-the-fly. Our aim is to study copying capabilitiesparticipate in enhancing user’s experience in the interaction.User listens to funny audio stimuli in the presenceof a laughing agent: when funniness of audio increases, theagent laughs and the quality of its body movement (directionand amplitude of laughter movements) is modulated on-theflyby user’s body features. Florian Pecune, Béatrice Biancardi, Yu Ding 0001, Catherine Pelachaud, Maurizio Mancini, Giovanna Varni, Antonio Camurri, Gualtiero Volpe |
AAAI | 3 |
| 2015 | Perception of intensity incongruence in synthesized multimodal expressions of laughterabstractIn this paper, we study perception of intensity in-congruence between auditory and visual modalities of synthesized expressions of laughter. In particular, we investigate whether incongruent expressions are perceived as 1) regulated, and 2) unsuccessful in terms of animation synthesis. For this purpose, we conducted a perceptive study with the use of a virtual agent. Congruent and incongruent multimodal expressions of laughter were synthesized from natural audiovisual laughter episodes, using machine learning algorithms. Next, the intensity of facial expressions and body movements were systematically manipulated to check whether the resulting incongruent expressions are perceived differently compared to the corresponding congruent expressions. Results show that 1) intensity incongruence lowers the perception of believability and plausibility, and 2) the in-congruent laughter expressions displaying high intensity in the audio modality and low intensity in the body movement and facial expression are perceived as more fake than the corresponding congruent expressions. Such results have implications for both animation synthesis as well as expression regulation research. Radoslaw Niewiadomski, Yu Ding 0001, Maurizio Mancini, Catherine Pelachaud, Gualtiero Volpe, Antonio Camurri |
ACII | 2 |
| 2015 | Real-Time Visual Prosody for Interactive Virtual Agents
Herwin van Welbergen, Yu Ding 0001, Kai Sattler, Catherine Pelachaud, Stefan Kopp |
IVA | 2 |
| 2014 | Rhythmic Body Movements of LaughterabstractIn this paper we focus on three aspects of multimodal expressions of laughter. First, we propose a procedural method to synthesize rhythmic body movements of laughter based on spectral analysis of laughter episodes. For this purpose, we analyze laughter body motions from motion capture data and we reconstruct them with appropriate harmonics. Then we reduce the parameter space to two dimensions. These are the inputs of the actual model to generate a continuum of laughs rhythmic body movements. Radoslaw Niewiadomski, Maurizio Mancini, Yu Ding 0001, Catherine Pelachaud, Gualtiero Volpe |
ICMI | 3 |
| 2014 | Upper Body Animation Synthesis for a Laughing Character
Yu Ding 0001, Jing Huang 0005, Nesrine Fourati, Thierry Artières, Catherine Pelachaud |
IVA | 1 |
| 2013 | Speech-driven eyebrow motion synthesis with contextual Markovian modelsabstractNonverbal communicative behaviors during speech are important to model a virtual agent able to sustain a natural and lively conversation with humans. We investigate statistical frameworks for learning the correlation between speech prosody and eyebrow motion features. Such methods may be used to synthesize automatically accurate eyebrow movements from synchronized speech. Yu Ding 0001, Mathieu Radenen, Thierry Artières, Catherine Pelachaud |
ICASSP | 1 |
| 2013 | Modeling Multimodal Behaviors from Speech Prosody
Yu Ding 0001, Catherine Pelachaud, Thierry Artières |
IVA | 1 |