Zerrin Yumak

dblp:13/1974 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-0028-5806ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance Learning
abstract
Creating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the semantic context of the gestures. In this paper, we propose a novel approach for semantic grounding in co-speech gesture generation that integrates semantic information at both fine-grained and global levels. Our approach starts with learning the motion prior through a vector-quantized variational autoencoder. Built on this model, a second-stage module is applied to automatically generate gestures from speech, text-based semantics and speaker identity that ensures consistency between the semantic relevance of generated gestures and co-occurring speech semantics through semantic coherence and relevance modules. Experimental results demonstrate that our approach enhances the realism and coherence of semantic gestures. Extensive experiments and user studies show that our method outperforms state-of-the-art approaches across two benchmarks in co-speech gesture generation in both objective and subjective metrics. The qualitative results of our model, code, dataset and pre-trained models can be viewed at https://semgesture.github.io/.
Lanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin Yumak
ICCV4
2025 "I don't like my avatar": Investigating Human Digital Doubles
abstract
Creating human digital doubles is becoming easier and much more accessible to everyone using consumer grade devices. In this work, we investigate how avatar style (realistic vs cartoon) and avatar familiarity (self, acquaintance, unknown person) affect self/other-identification, perceived realism, affinity and social presence with a controlled offline experiment. We created two styles of avatars (realistic-looking MetaHumans and cartoon-looking ReadyPlayerMe avatars) and facial animations stimuli for them using performance capture. Questionnaire responses demonstrate that higher appearance realism leads to a higher level of identification, perceived realism and social presence. However, avatars with familiar faces, especially those with high appearance realism, lead to a lower levels of identification, perceived realism, and affinity. Although participants identified their digital doubles as their own, they consistently did not like their avatars, especially of realistic appearance. But they were less critical and more forgiving about their acquaintance’s or an unknown person’s digital double.
Siyi Liu 0003, Kazi Injamamul Haque, Zerrin Yumak
MIG3
2025 ProbTalk3D-X: Prosody enhanced non-deterministic emotion controllable speech-driven 3D facial animation synthesis
abstract
Audio-driven 3D facial animation synthesis has been an active field of research with attention from both academia and industry. While there are promising results in this area, recent approaches largely focus on lip-sync and identity control, neglecting the role of emotions and emotion control in the generative process. That is mainly due to the lack of emotionally rich facial animation data and algorithms that can synthesize speech animations with emotional expressions at the same time. In addition, the majority of the models are deterministic, meaning given the same audio input, they produce the same output motion. We argue that emotions and non-determinism are crucial to generate diverse and emotionally-rich facial animations. In this paper, we present ProbTalk3D-X by extending a prior work ProbTalk3D- a two staged VQ-VAE based non-deterministic model, by additionally incorporating prosody features for improved facial accuracy using an emotionally rich facial animation dataset, 3DMEAD. Further, we present a comprehensive comparison of non-deterministic emotion controllable models (including new extended experimental models) leveraging VQ-VAE, VAE and diffusion techniques. We provide an extensive comparative analysis of the experimental models against the recent 3D facial animation synthesis approaches, by evaluating the results objectively, qualitatively, and with a perceptual user study. We highlight several objective metrics that are more suitable for evaluating stochastic outputs and use both in-the-wild and ground truth data for subjective evaluation. Our evaluation demonstrates that ProbTalk3D-X and original ProbTalk3D achieve superior performance compared to state-of-the-art emotion-controlled, deterministic and non-deterministic models. We recommend watching the supplementary video for visual quality judgment. The entire codebase including the extended models is publicly available. 1
Kazi Injamamul Haque, Sichun Wu, Zerrin Yumak
Comput. Graph.3
2025 "Wild West" of Evaluating Speech-Driven 3D Facial Animation Synthesis: A Benchmark Study
abstract
Abstract Recent advancements in the field of audio‐driven 3D facial animation have accelerated rapidly, with numerous papers being published in a short span of time. This surge in research has garnered significant attention from both academia and industry with its potential applications on digital humans. Various approaches, both deterministic and non‐deterministic, have been explored based on foundational advancements in deep learning algorithms. However, there remains no consensus among researchers on standardized methods for evaluating these techniques. Additionally, rather than converging on a common set of datasets and objective metrics suited for specific methods, recent works exhibit considerable variation in experimental setups. This inconsistency complicates the research landscape, making it difficult to establish a streamlined evaluation process and rendering many cross‐paper comparisons challenging. Moreover, the common practice of A/B testing in perceptual studies focus only on two common metrics and not sufficient for non‐deterministic and emotion‐enabled approaches. The lack of correlations between subjective and objective metrics points out that there is a need for critical analysis in this space. In this study, we address these issues by benchmarking state‐of‐the‐art deterministic and non‐deterministic models, utilizing a consistent experimental setup across a carefully curated set of objective metrics and datasets. We also conduct a perceptual user study to assess whether subjective perceptual metrics align with the objective metrics. Our findings indicate that model rankings do not necessarily generalize across datasets, and subjective metric ratings are not always consistent with their corresponding objective metrics. The supplementary video, edited code scripts for training on different datasets and documentation related to this benchmark study are made publicly available‐ https://galib360.github.io/face-benchmark-project/ .
Kazi Injamamul Haque, Alkiviadis Pavlou, Zerrin Yumak
Comput. Graph. Forum3
2024 ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAE
abstract
Audio-driven 3D facial animation synthesis has been an active field of research with attention from both academia and industry. While there are promising results in this area, recent approaches largely focus on lip-sync and identity control, neglecting the role of emotions and emotion control in the generative process. That is mainly due to the lack of emotionally rich facial animation data and algorithms that can synthesize speech animations with emotional expressions at the same time. In addition, majority of the models are deterministic, meaning given the same audio input, they produce the same output motion. We argue that emotions and non-determinism are crucial to generate diverse and emotionally-rich facial animations. In this paper, we propose ProbTalk3D a non-deterministic neural network approach for emotion controllable speech-driven 3D facial animation synthesis using a two-stage VQ-VAE model and an emotionally rich facial animation dataset 3DMEAD. We provide an extensive comparative analysis of our model against the recent 3D facial animation synthesis approaches, by evaluating the results objectively, qualitatively, and with a perceptual user study. We highlight several objective metrics that are more suitable for evaluating stochastic outputs and use both in-the-wild and ground truth data for subjective evaluation. To our knowledge, that is the first non-deterministic 3D facial animation synthesis method incorporating a rich emotion dataset and emotion control with emotion labels and intensity levels. Our evaluation demonstrates that the proposed model achieves superior performance compared to state-of-the-art emotion-controlled, deterministic and non-deterministic models. We recommend watching the supplementary video for quality judgement. The entire codebase is publicly available1.
Sichun Wu, Kazi Injamamul Haque, Zerrin Yumak
MIG3
2023 Effect of Appearance and Animation Realism on the Perception of Emotionally Expressive Virtual Humans
abstract
3D Virtual Human technology is growing with several potential applications in health, education, business and telecommunications. Investigating the perception of these virtual humans can help guide to develop better and more effective applications. Recent developments show that the appearance of the virtual humans reached to a very realistic level. However, there is not yet adequate analysis on the perception of appearance and animation realism for emotionally expressive virtual humans. In this paper, we designed a user experiment and analyzed the effect of a realistic virtual human's appearance realism and animation realism in varying emotion conditions. We found that higher appearance realism and higher animation realism leads to higher social presence and higher attractiveness ratings. We also found significant effects of animation realism on perceived realism and emotion intensity levels. Our study sheds light into how appearance and animation realism effects the perception of highly realistic virtual humans in emotionally expressive scenarios and points out to future directions.
Nabila Amadou, Kazi Injamamul Haque, Zerrin Yumak
IVA3
2023 FaceDiffuser: Speech-Driven 3D Facial Animation Synthesis Using Diffusion
abstract
Speech-driven 3D facial animation synthesis has been a challenging task both in industry and research. Recent methods mostly focus on deterministic deep learning methods meaning that given a speech input, the output is always the same. However, in reality, the non-verbal facial cues that reside throughout the face are non-deterministic in nature. In addition, majority of the approaches focus on 3D vertex based datasets and methods that are compatible with existing facial animation pipelines with rigged characters is scarce. To eliminate these issues, we present FaceDiffuser, a non-deterministic deep learning model to generate speech-driven facial animations that is trained with both 3D vertex and blendshape based datasets. Our method is based on the diffusion technique and uses the pre-trained large speech representation model HuBERT to encode the audio input. To the best of our knowledge, we are the first to employ the diffusion method for the task of speech-driven 3D facial animation synthesis. We have run extensive objective and subjective analyses and show that our approach achieves better or comparable results in comparison to the state-of-the-art methods. We also introduce a new in-house dataset that is based on a blendshape based rigged character. The code and the dataset will be publicly available on the project page1.
Stefan Stan, Kazi Injamamul Haque, Zerrin Yumak
MIG3
2022 GENEA Workshop 2022: The 3rd Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents
abstract
Embodied agents benefit from using non-verbal behavior when communicating with humans. Despite several decades of non-verbal behavior-generation research, there is currently no well-developed benchmarking culture in the field. For example, most researchers do not compare their outcomes with previous work, and if they do, they often do so in their own way which frequently is incompatible with others. With the GENEA Workshop 2022, we aim to bring the community together to discuss key challenges and solutions, and find the most appropriate ways to move the field forward.
Pieter Wolfert, Taras Kucherenko, Carla Viegas, Zerrin Yumak, Youngwoo Yoon, Gustav Eje Henter
ICMI4
2021 GENEA Workshop 2021: The 2nd Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents
abstract
Embodied agents benefit from using non-verbal behavior when communicating with humans. Despite several decades of non-verbal behavior-generation research, there is currently no well-developed benchmarking culture in the field. For example, most researchers do not compare their outcomes with previous work, and if they do, they often do so in their own way which frequently is incompatible with others. With the GENEA Workshop 2021, we aim to bring the community together to discuss key challenges and solutions, and find the most appropriate ways to move the field forward.
Taras Kucherenko, Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, Zerrin Yumak, Gustav Eje Henter
ICMI5
2019 Data-driven Gaze Animation using Recurrent Neural Networks
abstract
We present a data-driven gaze animation method using recurrent neural networks. The neural network is trained with motion capture data including different poses such as standing, sitting, and lying down and is able to learn the constraints related with each particular pose. A simplified version of the neural network is also presented for Level of Detail (LOD) animation. We compare various neural network architectures and show that our method produces natural gaze motion in real-time. Results from a user study conducted among game industry professionals shows that our method has better perceived naturalness compared to the procedural gaze animation system of a well-known game company. Our approach is the first one to show the feasibility of gaze motions using deep neural networks.
Alex Klein, Zerrin Yumak, Arjen Beij, A. Frank van der Stappen
MIG2
2019 Audio-driven emotional speech animation for interactive virtual characters
abstract
Abstract We present a procedural audio‐driven speech animation method for interactive virtual characters. Given any audio with its respective speech transcript, we automatically generate lip‐synchronized speech animation that could drive any three‐dimensional virtual character. The realism of the animation is enhanced by studying the emotional features of the audio signal and its effect on mouth movements. We also propose a coarticulation model that takes into account various linguistic rules. The generated animation is configurable by the user by modifying the control parameters, such as viseme types, intensities, and coarticulation curves. We compare our approach against two lip‐synchronized speech animation generators. Our results show that our method surpasses them in terms of user preference.
Constantinos Charalambous, Zerrin Yumak, A. Frank van der Stappen
Comput. Animat. Virtual Worlds2
2018 Towards a generic framework for multi-party dialogue with virtual humans
abstract
Existing approaches and frameworks for modeling virtual dialogue tend to be designed with dyadic interactions in mind, and are often built to serve solely in task-oriented domains. However, modeling realistic action and turn-taking in more general scenarios remains a challenge. In this paper we propose a generic framework to aid in development of multi-modal, multi-party dialogue. It contains mechanisms inspired by social practice theory for both action selection and timing --- including handling of interruption. As a proof-of-concept, we employ these ideas in a virtual couples-therapy session, demonstrating their potential in modeling complex real-life situations.
Raoul Harel, Zerrin Yumak, Frank Dignum
CASA2
2017 Social Gaze Model for an Interactive Virtual Character
Bram van den Brink, Chris Christyowidiasmoro, Zerrin Yumak
IVA3
2017 Autonomous social gaze model for an interactive virtual character in real-life settings
abstract
Abstract This paper presents a gaze behavior model for an interactive virtual character situated in the real world. We are interested in estimating which user has an intention to interact, in other words which user is engaged with the virtual character. The model takes into account behavioral cues such as proximity, velocity, posture, and sound; estimates an engagement score; and drives the gaze behavior of the virtual character. Initially, we assign equal weights to these features. Using data collected in a real setting, we analyze which features have higher importance. We found that the model with weighted features correlates better with the ground‐truth data.
Zerrin Yumak, Bram van den Brink, Arjan Egges
Comput. Animat. Virtual Worlds1
2017 EHR: a Sensing Technology Readiness Model for Lifestyle Changes
Yu Chen 0008, Danni Le, Zerrin Yumak, Pearl Pu
Mob. Networks Appl.3
2014 Autonomous virtual humans and social robots in telepresence
abstract
Telepresence refers to the possibility of feeling present in a remote location through the use of technology. This can be achieved by immersing a user to a place reconstructed in 3D. The reconstructed place can be captured from the real world or can be completely virtual. Another way to realize telepresence is by using robots and virtual avatars that act as proxies for real people. In case a human-mediated interaction is not needed or not possible, the virtual human and the robot can rely on artificial intelligence to act and interact autonomously. In this paper, these forms of telepresence are discussed, how they are related and different from each other and how autonomy takes place in telepresence. The paper concludes with an overview of the ongoing research on autonomous virtual humans and social robots conducted in the BeingThere centre.
Nadia Magnenat-Thalmann, Zerrin Yumak, Aryel Beck
MMSP2
2013 Multi-party interaction with a virtual character and a human-like robot
abstract
Research on interactive virtual characters and social robots focuses mainly on one-to-one interactions and multi-party interactions concept are rather less explored. As we are developing these characters to be helpful to us in our daily lives as guides, companions, assistants or receptionists, they should be aware of the existence of multiple people and address their requirements in a natural way and act according to the social rules and norms. In contrast with previous work, we are interested in multi-party and multi-modal interactions between 3D virtual characters, real humans and social robots. This means that any of these participants can interact with each other. In this paper we present our on-going work, provide a discussion on multi-party interaction, describe the overall system architecture and mention our future work.
Zerrin Yumak, Nadia Magnenat-Thalmann
VRST1
2012 1st workshop on recommendation technologies for lifestyle change 2012
abstract
The workshop on Recommendation Technologies for Lifestyle Change will be an opportunity for discussing open issues, and propose technical solutions for the designing of intelligent information systems that can support and promote lifestyle change. The objective of these systems is to provide users with up-to-date information, and help them to make choices in every day life activities establishing a sustainable compromise between quality of life, individuality, and fun.
Bernd Ludwig, Francesco Ricci 0001, Zerrin Yumak
RecSys3