VLDB 2026 Research / reviewers in the wild / expert
Ylva Ferstl
dblp:183/0197
· DBLP profile ↗
15ranked-venue papers
12as first author
8since 2021 · last 2023
0000-0001-7259-0378ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Generating Emotionally Expressive Look-At AnimationabstractHumanoid characters in video games are generally animated using motion capture technology, enabling high-quality, realistic animation. This animation has to remain interactive and allow characters to react to their environment; one important component of this adaptation are look-at animations which direct the character’s torso towards an object or person of interest. Look-at animations are generated procedurally in order to handle any desired target direction, however, this procedural animation can have a robotic quality and negatively affect the overall perceived realism of the character. In this work, we present a neural network controller for generating look-at animations that are equally as appealing as motion capture while requiring minimal memory. Moreover, our controller can generate animations stylized by emotion, allowing characters to react to look-at targets depending on their context, and this style expressiveness is shown to be on par with motion captured samples. Ylva Ferstl |
MIG | 1 |
| 2023 | ZeroEGGS: Zero-shot Example-based Gesture Generation from SpeechabstractAbstract We present ZeroEGGS, a neural network framework for speech‐driven gesture generation with zero‐shot style control by example. This means style can be controlled via only a short example motion clip, even for motion styles unseen during training. Our model uses a Variational framework to learn a style embedding, making it easy to modify style through latent space manipulation or blending and scaling of style embeddings. The probabilistic nature of our framework further enables the generation of a variety of outputs given the input, addressing the stochastic nature of gesture motion. In a series of experiments, we first demonstrate the flexibility and generalizability of our model to new speakers and styles. In a user study, we then show that our model outperforms previous state‐of‐the‐art techniques in naturalness of motion, appropriateness for speech, and style portrayal. Finally, we release a high‐quality dataset of full‐body gesture motion including fingers, with speech, spanning across 19 different styles. Our code and data are publicly available at https://github.com/ubisoft/ubisoft‐laforge‐ZeroEGGS . Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje, Marc-André Carbonneau |
Comput. Graph. Forum | 2 |
| 2022 | Exemplar-based Stylized Gesture Generation from Speech: An Entry to the GENEA Challenge 2022abstractWe present our entry to the GENEA Challenge of 2022 on data-driven co-speech gesture generation. Our system is a neural network that generates gesture animation from an input audio file. The motion style generated by the model is extracted from an exemplar motion clip. Style is embedded in a latent space using a variational framework. This architecture allows for generating in styles unseen during training. Moreover, the probabilistic nature of our variational framework furthermore enables the generation of a variety of outputs given the same input, addressing the stochastic nature of gesture motion. The GENEA challenge evaluation showed that our model produces full-body motion with highly competitive levels of human-likeness. Saeed Ghorbani, Ylva Ferstl, Marc-André Carbonneau |
ICMI | 2 |
| 2022 | Investigating how speech and animation realism influence the perceived personality of virtual characters and agentsabstractThe portrayed personality of virtual characters and agents is understood to influence how we perceive and engage with digital applications. Understanding how the features of speech and animation drive portrayed personality allows us to intentionally design characters to be more personalized and engaging. In this study, we use performance capture data of unscripted conversations from a variety of actors to explore the perceptual outcomes associated with the modalities of speech and motion. Specifically, we contrast full performance-driven characters to those portrayed by generated gestures and synthesized speech, analysing how the features of each influence portrayed personality according to the Big Five personality traits. We find that processing speech and motion can have mixed effects on such traits, with our results highlighting motion as the dominant modality for portraying extraversion and speech as dominant for communicating agreeableness and emotional stability. Our results can support the Extended Reality (XR) community in development of virtual characters, social agents and 3D User Interface (3DUI) agents portraying a range of targeted personalities. Sean Thomas, Ylva Ferstl, Rachel McDonnell, Cathy Ennis |
VR | 2 |
| 2021 | Evaluating Study Design and Strategies for Mitigating the Impact of Hand Tracking LossabstractSocial virtual reality uses motion tracking to place people in virtual environments as animated avatars. Often this tracking only measures the position and orientation of the head and hands, and from this estimates the body pose. Optical hand tracking is an important technology to enable such avatars, but can frequently fail and cause motion errors when the hands are visually obscured. This paper presents three amelioration strategies to handle these errors and demonstrates experimentally that all three are effective in reducing their impact. This setting is also used to explore general issues around study design for motion perception. Different strategies for presenting stimuli and soliciting input are compared. The presence of a simultaneous recall task is shown to reduce but not eliminate sensitivity to motion errors. Finally, it is shown that motion errors are interpreted, at least in part, as a shift in interlocutor personality. Ylva Ferstl, Rachel McDonnell, Michael Neff |
SAP | 1 |
| 2021 | Human or Robot?: Investigating voice, appearance and gesture motion realism of conversational social agentsabstractResearch on creation of virtual humans enables increasing automatization of their behavior, including synthesis of verbal and nonverbal behavior. As the achievable realism of different aspects of agent design evolves asynchronously, it is important to understand if and how divergence in realism between behavioral channels can elicit negative user responses. Specifically, in this work, we investigate the question of whether autonomous virtual agents relying on synthetic text-to-speech voices should portray a corresponding level of realism in the non-verbal channels of motion and visual appearance, or if, alternatively, the best available realism of each channel should be used. In two perceptual studies, we assess how realism of voice, motion, and appearance influence the perceived match of speech and gesture motion, as well as the agent's likability and human-likeness. Our results suggest that maximizing realism of voice and motion is preferable even when this leads to realism mismatches, but for visual appearance, lower realism may be preferable. (A video abstract can be found at https://youtu.be/arfZZ-hxD1Y.) Ylva Ferstl, Sean Thomas, Cédric Guiard, Cathy Ennis, Rachel McDonnell |
IVA | 1 |
| 2021 | ExpressGesture: Expressive gesture generation from speech through database matchingabstractAbstract Co‐speech gestures are a vital ingredient in making virtual agents more human‐like and engaging. Automatically generated gestures based on speech‐input often lack realistic and defined gesture form. We present a database‐driven approach guaranteeing defined gesture form. We built a large corpus of over 23,000 motion‐captured co‐speech gestures and select individual gestures based on expressive gesture characteristics that can be estimated from speech audio. The expressive parameters are gesture velocity and acceleration, gesture size, arm swivel, and finger extension. Individual, parameter‐matched gestures are then combined into animated sequences. We evaluate our gesture generation system in two perceptual studies. The first study compares our method to the ground truth gestures as well as mismatched gestures. The second study compares our method to five current generative machine learning models. Our method outperformed mismatched gesture selection in the first study and showed competitive performance in the second. Ylva Ferstl, Michael Neff, Rachel McDonnell |
Comput. Animat. Virtual Worlds | 1 |
| 2021 | Facial Feature Manipulation for Trait Portrayal in Realistic and Cartoon-Rendered CharactersabstractPrevious perceptual studies on human faces have shown that specific facial features have consistent effects on perceived personality and appeal, but it remains unclear if and how findings relate to perception of virtual characters. For example, wider human faces have been found to appear more aggressive and dominant, whereas studies on virtual characters have shown opposite trends but have suffered from significant eeriness of exaggerated features. In this study, we use highly realistic virtual faces obtained from 3D scanning, as well as cartoon-rendered counterparts retaining facial proportions. We assess the effects of facial width and eye size on perceptions of appeal, trustworthiness, aggressiveness, dominance, and eeriness. Our manipulations did not affect eeriness, and we find the same perceptual trends previously reported for human faces. Ylva Ferstl, Michael McKay, Rachel McDonnell |
ACM Trans. Appl. Percept. | 1 |
| 2020 | Understanding the Predictability of Gesture Parameters from Speech and their Perceptual ImportanceabstractGesture behavior is a natural part of human conversation. Much work has focused on removing the need for tedious hand-animation to create embodied conversational agents by designing speech-driven gesture generators. However, these generators often work in a black-box manner, assuming a general relationship between input speech and output motion. As their success remains limited, we investigate in more detail how speech may relate to different aspects of gesture motion. We determine a number of parameters characterizing gesture, such as speed and gesture size, and explore their relationship to the speech signal in a two-fold manner. First, we train multiple recurrent networks to predict the gesture parameters from speech to understand how well gesture attributes can be modeled from speech alone. We find that gesture parameters can be partially predicted from speech, and some parameters, such as path length, being predicted more accurately than others, like velocity. Second, we design a perceptual study to assess the importance of each gesture parameter for producing motion that people perceive as appropriate for the speech. Results show that a degradation in any parameter was viewed negatively, but some changes, such as hand shape, are more impactful than others. A video summarization can be found at https://youtu.be/aw6-_5kmLjY. Ylva Ferstl, Michael Neff, Rachel McDonnell |
IVA | 1 |
| 2020 | Adversarial gesture generation with realistic gesture phasing
Ylva Ferstl, Michael Neff, Rachel McDonnell |
Comput. Graph. | 1 |
| 2019 | Multi-objective adversarial gesture generationabstractApplications for conversational virtual agents are on the rise, but producing realistic non-verbal behavior for spoken utterances remains an unsolved problem. We explore the use of a generative adversarial training paradigm to map speech to 3D gesture motion. We define the gesture generation problem as a series of smaller sub-problems, including plausible gesture dynamics, realistic joint configurations, and diverse and smooth motion. Each sub-problem is monitored by separate adversaries. For the problem of enforcing realistic gesture dynamics in our output, we train a classifier to automatically detect gesture phases. We find adversarial training to be superior to the use of a standard regression loss and discuss the benefit of each of our training objectives. We recorded a dataset of over 6 hours of natural, unrehearsed speech with high-quality motion capture, as well as audio and video recording. Ylva Ferstl, Michael Neff, Rachel McDonnell |
MIG | 1 |
| 2018 | Investigating the use of recurrent motion modelling for speech gesture generationabstractThe growing use of virtual humans demands generating increasingly realistic behavior for them while minimizing cost and time. Gestures are a key ingredient for realistic and engaging virtual agents and consequently automatized gesture generation has been a popular area of research. So far, good gesture generation has relied on explicit formulation of if-then rules and probabilistic modelling of annotated features. Machine learning approaches have yielded only marginal success, indicating a high complexity of the speech-to-motion learning task. In this work, we explore the use of transfer learning using previous motion modelling research to improve learning outcomes for gesture generation from speech. We use a recurrent network with an encoder-decoder structure that takes in prosodic speech features and generates a short sequence of gesture motion. We pre-train the network with a motion modelling task. We recorded a large multimodal database of conversational speech for the purpose of this work. Ylva Ferstl, Rachel McDonnell |
IVA | 1 |
| 2018 | A perceptual study on the manipulation of facial features for trait portrayal in virtual agentsabstractHuman perceptual studies have shown that facial characteristics affect judgments about the personality of a person. For example, larger facial width has been associated with judgments of aggressiveness, dominance, and untrustworthiness. Previous studies of virtual faces have not been able to reflect the same perceptual rules, but have used characters with unrealistic feature sizes or highly abstract characters. For this study, we created virtual characters with realistic feature dimensions and investigated the effects of facial width and eye size on personality perception. Our results indicate that virtual characters may indeed follow different perceptual rules for facial width, and care must be taken when manipulating eye size. These findings are useful for effective character design for video games, movies, and embodied virtual agents. Ylva Ferstl, Rachel McDonnell |
IVA | 1 |
| 2017 | Facial Features of Non-player Creatures Can Influence Moral Decisions in Video GamesabstractWith the development of increasingly sophisticated computer graphics, there is a continuous growth of the variety and originality of virtual characters used in movies and games. So far, however, their design has mostly been led by the artist’s preferences, not by perceptual studies. In this article, we explored how effective non-player character design can be used to influence gameplay. In particular, we focused on abstract virtual characters with few facial features. In experiment 1, we sought to find rules for how to use a character’s facial features to elicit the perception of certain personality traits, using prior findings for human face perception as a basis. In experiment 2, we then tested how perceived personality traits of a non-player character could influence a player’s moral decisions in a video game. We found that the appearance of the character interacting with the subject modulated aggressive behavior towards a non-present individual. Our results provide us with a better understanding of the perception of abstract virtual characters, their employment in video games, as well as giving us some insights about the factors underlying aggressive behavior in video games. Ylva Ferstl, Elena Kokkinara, Rachel McDonnell |
ACM Trans. Appl. Percept. | 1 |
| 2016 | Do I trust you, abstract creature?: a study on personality perception of abstract virtual facesabstractStudies in the field of social psychology have shown evidence that the dimensions of human facial features can directly impact the perception of personality of that human. Traits such as aggressiveness, trustworthiness and dominance have been directly correlated with facial features. If the same correlations were true for virtual faces, this could be a valuable design guideline to direct the creation of characters with intended personalities. In particular, this is relevant for extremely abstract characters that have minimal facial features (often seen in video games and movies), and rely heavily on these features for portraying personality. We conducted an exploratory study in order to retrieve insights about the way certain facial features affect the perceived personality, as well as affinity of very abstract virtual faces. We specifically tested the effect of different head shapes, eye shapes and eye sizes. Interestingly, our findings show that the same rules for real human faces do not apply to the perception of abstract faces, and in some cases are the complete reverse. These results provide us with a better understanding of the perception of abstract virtual faces, and a starting point for the creation of guidelines for how to portray personality using minimal facial cues. Ylva Ferstl, Elena Kokkinara, Rachel McDonnell |
SAP | 1 |