VLDB 2026 Research / reviewers in the wild / expert
Cathy Ennis
dblp:30/1113
· DBLP profile ↗
20ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-1274-5347ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Making Faces: Evaluating Facial Control Methods in VR for Live ConversationsabstractABSTRACT Facial expressions are a crucial part of affective communication in our daily lives, and the same applies to virtual humans. However, achieving a full range of believable and naturalistic emotional expression via VR devices still remains a challenge, particularly when constrained by the capabilities of midrange consumer VR devices without face tracking as opposed to the more expensive versions with facial tracking capabilities. In this study, we evaluated and compared three methods to control avatar facial expression in VR: continuous face tracking (FT); thumbstick label (TL), which is a VR controller's thumbstick‐based selection method with visually labeled expression; and thumbstick label with confirmation button (TB), which is a variation of TL with an additional step of pressing a button. Thirty participants took part in the experiment. Participants found FT to be significantly more usable than TL and TB. While TL was comparable to FT on co‐presence and several items of expressiveness, with significant gaps emerging mainly for naturalness and smoothness in expression change. These findings suggest that a well‐designed control method with low latency and low motor and cognitive load can serve as an alternative method of avatar control in VR devices without facial tracking technology. J. K. Sangeeth Chandran, Marisa Llorens-Salvador, Cathy Ennis |
Comput. Animat. Virtual Worlds | 3 |
| 2025 | Enhancing Synthetic Image Realism with Controlled Diffusion ModelsabstractIn this work, we present an innovative approach utilizing ControlNet-based diffusion models along with upscaling capabilities for domain adaptation and quality refinement of 3D modelled synthetic datasets, focusing on autonomous vehicle applications. A significant domain gap often exists between synthetic and real-world data, hindering the applicability of deep learning models trained on synthetic data for real-world scenarios. Our methodology leverages the strengths of Controlled Augmentation by simultaneously utilizing multiple ControlNet signals, including edge detection, depth information, segmentation maps, and tile resampling. To improve how synthetic data aligns with the desired domain specifications, these signals guide the generative process, and we also incorporate text-guided prompts extracted via Large Language Models (LLMs), to improve control over the synthesis of desired features and attributes. We test the approach on diverse environmental conditions from the VKITTI dataset, a well-known 3D modelled synthetic dataset generated in Unity for autonomous driving research. The refined data is validated using quantitative metrics including FID, SSIM, and LPIPS, and is also evaluated on downstream machine learning tasks of object detection and classification, using YOLO-v8 to ensure its utility and effectiveness. Experimental analysis demonstrates the effectiveness of this method in improving the realism and usability of synthetic data. Our approach contributes to fields that require high-quality data synthesis and domain adaptation. The experimental work, along with ControlNet models used in this project is available online.1 Iqra Nosheen, Peter Corcoran 0001, Cathy Ennis, Michael G. Madden |
IJCNN | 4 |
| 2025 | Synthetically Expressive: Evaluating gesture and voice for emotion and empathy in VR and 2D scenariosabstractThe creation of virtual humans increasingly leverages automated synthesis of speech and gestures, enabling expressive, adaptable agents that effectively engage users. However, the independent development of voice and gesture generation technologies, alongside the growing popularity of virtual reality (VR), presents significant questions about the integration of these signals and their ability to convey emotional detail in immersive environments. In this paper, we evaluate the influence of real and synthetic gestures and speech, alongside varying levels of immersion (VR vs. 2D displays) and emotional contexts (positive, neutral, negative) on user perceptions. We investigate how immersion affects the perceived match between gestures and speech and the impact on key aspects of user experience, including emotional and empathetic responses and the sense of co-presence. Our findings indicate that while VR enhances the perception of natural gesture–voice pairings, it does not similarly improve synthetic ones—amplifying the perceptual gap between them. These results highlight the need to reassess gesture appropriateness and refine AI-driven synthesis for immersive environments. Haoyang Du, Kiran Chhatre, Christopher Peters 0001, Brian Keegan, Rachel McDonnell, Cathy Ennis |
IVA | 6 |
| 2025 | Investigating How Text and Motion Style Shape Directness In Embodied Conversational AgentsabstractEmbodied Conversational Agents (ECAs) are becoming more widely used for various applications, notably in health. Some studies have demonstrated the effectiveness of delivering Motivational Interviews (MIs), a patient-centred behaviour change method, with ECAs. Despite showing promise, the effect of agent communication style on MI effectiveness remains underexplored. Directness has been shown to impact satisfaction and preference of dialogue systems, and effectiveness in some cases. However, outside of verbal style, how to control directness in ECAs to achieve these goals is not well understood. We designed an ECA to convey different levels of directness through language and Non-Verbal behaviours (NVBs), then evaluated the impact of language and NVB on directness with a perception study. The results showed that language influenced perceived directness, and that NVB contributed when aligned with indirect language. These findings suggest ways to shape conversational agents’ communication style to enhance their effectiveness. Michael O'Mahony, Cathy Ennis, Robert J. Ross |
MIG | 2 |
| 2022 | Investigating how speech and animation realism influence the perceived personality of virtual characters and agentsabstractThe portrayed personality of virtual characters and agents is understood to influence how we perceive and engage with digital applications. Understanding how the features of speech and animation drive portrayed personality allows us to intentionally design characters to be more personalized and engaging. In this study, we use performance capture data of unscripted conversations from a variety of actors to explore the perceptual outcomes associated with the modalities of speech and motion. Specifically, we contrast full performance-driven characters to those portrayed by generated gestures and synthesized speech, analysing how the features of each influence portrayed personality according to the Big Five personality traits. We find that processing speech and motion can have mixed effects on such traits, with our results highlighting motion as the dominant modality for portraying extraversion and speech as dominant for communicating agreeableness and emotional stability. Our results can support the Extended Reality (XR) community in development of virtual characters, social agents and 3D User Interface (3DUI) agents portraying a range of targeted personalities. Sean Thomas, Ylva Ferstl, Rachel McDonnell, Cathy Ennis |
VR | 4 |
| 2021 | Latent Dynamics for Artefact-Free Character Animation via Data-Driven Reinforcement Learning
Vihanga Gamage, Cathy Ennis, Robert J. Ross |
ICANN (4) | 2 |
| 2021 | Human or Robot?: Investigating voice, appearance and gesture motion realism of conversational social agentsabstractResearch on creation of virtual humans enables increasing automatization of their behavior, including synthesis of verbal and nonverbal behavior. As the achievable realism of different aspects of agent design evolves asynchronously, it is important to understand if and how divergence in realism between behavioral channels can elicit negative user responses. Specifically, in this work, we investigate the question of whether autonomous virtual agents relying on synthetic text-to-speech voices should portray a corresponding level of realism in the non-verbal channels of motion and visual appearance, or if, alternatively, the best available realism of each channel should be used. In two perceptual studies, we assess how realism of voice, motion, and appearance influence the perceived match of speech and gesture motion, as well as the agent's likability and human-likeness. Our results suggest that maximizing realism of voice and motion is preferable even when this leads to realism mismatches, but for visual appearance, lower realism may be preferable. (A video abstract can be found at https://youtu.be/arfZZ-hxD1Y.) Ylva Ferstl, Sean Thomas, Cédric Guiard, Cathy Ennis, Rachel McDonnell |
IVA | 4 |
| 2021 | Learned Dynamics Models and Online Planning for Model-Based Animation Agents
Vihanga Gamage, Cathy Ennis, Robert J. Ross |
KES-AMSTA | 2 |
| 2018 | Examining the effects of a virtual character on learning and engagement in serious gamesabstractVirtual characters have been employed for many purposes including interacting with players of serious games, with a purpose to increase engagement. These characters are often embodied conversational agents playing diverse roles, such as demonstrators, guides, teachers or interviewers. Recently, much research has been conducted into properties that affect the realism and plausibility of virtual characters, but it is less clear whether the inclusion of interactive agents in serious applications can enhance a user's engagement with the application, or indeed increase efficacy. In a first step towards answering these questions, we conducted a study where a Virtual Learning Environment was used to examine the effect of employing a virtual character to deliver a lesson. In order to investigate whether increased familiarity between the player and the character would help achieve learning outcomes, we allowed participants to customize the physical appearance of the character. We used direct and indirect measures to assess engagement and learning; we measured knowledge retention to ascertain learning via a test at the end of the lesson, and also measured participants' perceived engagement with the lesson. Our findings show that a virtual character can be an effective learning aid, causing heightened engagement and retention of knowledge. However, allowing participants to customize character appearance resulted in inhibited engagement, which was contrary to what we expected. Vihanga Gamage, Cathy Ennis |
MIG | 2 |
| 2016 | FrankenFolk: distinctiveness and attractiveness of voice and motionabstractNo abstract available. Jan Ondrej, Cathy Ennis, Niamh A. Merriman, Carol O'Sullivan |
SAP | 2 |
| 2016 | FrankenFolk: Distinctiveness and Attractiveness of Voice and MotionabstractIt is common practice in movies and games to use different actors for the voice and body/face motion of a virtual character. What effect does the combination of these different modalities have on the perception of the viewer? In this article, we conduct a series of experiments to evaluate the distinctiveness and attractiveness of human motions (face and body) and voices. We also create combination characters called FrankenFolks, where we mix and match the voice, body motion, face motion, and avatar of different actors and ask which modality is most dominant when determining distinctiveness and attractiveness or whether the effects are cumulative. Jan Ondrej, Cathy Ennis, Niamh A. Merriman, Carol O'Sullivan |
ACM Trans. Appl. Percept. | 2 |
| 2013 | The TARDIS Framework: Intelligent Virtual Agents for Social Coaching in Job Interviews
Keith Anderson, Elisabeth André, Tobias Baur 0001, Sara Bernardini, Mathieu Chollet, Evi Chryssafidou, Ionut Damian, Cathy Ennis, Arjan Egges, Patrick Gebhard, Hazaël Jones, Magalie Ochs, Catherine Pelachaud, Kaska Porayska-Pomsta, Paola Rizzo, Nicolas Sabouret |
Advances in Computer Entertainment | 8 |
| 2013 | Perception of complex emotional body language of a virtual character with limb modificationsabstractThis abstract discusses a perceptual study investigating the perception of emotion in conversational body language. Specifically, we wished to determine the parts of the body most important for the identification of emotional expressions. We conducted a perceptual study using motion captured clips of an actor conducting a number of emotional conversations, exhibiting a set of emotions with negative connotations. These motions were represented on a virtual character, showing either the full body, or in the absence of the motion of arms, legs or head. Participants were then asked to indicate the level of presence of both the correct emotion and a corresponding positive pair (e.g., relaxed/stressed). We found participants were able to identify the emotions correctly, but depended strongly on the motions of the arms when doing so. Jurgis Pamerneckas, Cathy Ennis, Arjan Egges |
SAP | 2 |
| 2013 | Perception of Approach and Reach in Combined Interaction TasksabstractOften in games, a virtual character is required to interact with objects in the surrounding environment. These interactions can occur in different locations, with different items, often in combination with environment navigation tasks. This results in switching and blending between different motions in order to fit to restrictions due to the position of the character and the interaction circumstances. In this paper, we conduct perceptual experiments to gain knowledge about such interactions and deduce important factors about their design for game animators. Our results identify at what point interaction information is obvious, and which body parts are most important to consider. We find that general information about target position is evident from early on in a combined navigation and manipulation task, and can be deduced from very few visual cues. We also learn that participants are highly sensitive to target positions during the interaction phase, relying mostly on indicators in the motion of the character's arm in the final steps. Cathy Ennis, Arjan Egges |
MIG | 1 |
| 2013 | Emotion Capture: Emotionally Expressive Characters for GamesabstractIt has been shown that humans are sensitive to the portrayal of emotions for virtual characters. However, previous work in this area has often examined this sensitivity using extreme examples of facial or body animation. Less is known about how attuned people are at recognizing emotions as they are expressed during conversational communication. In order to determine whether body or facial motion is a better indicator for emotional expression for game characters, we conduct a perceptual experiment using synchronized full-body and facial motion-capture data. We find that people can recognize emotions from either modality alone, but combining facial and body motion is preferable in order to create more expressive characters. Cathy Ennis, Ludovic Hoyet, Arjan Egges, Rachel McDonnell |
MIG | 1 |
| 2012 | Perception of Complex Emotional Body Language of a Virtual Character
Cathy Ennis, Arjan Egges |
MIG | 1 |
| 2012 | Perceptually plausible formations for virtual conversersabstractABSTRACT Recent progress in real‐time simulations has led to a higher demand for believability from virtual characters. Background characters are becoming a more integral part of games, with emphasis being placed in particular on interactions between them. Conversing groups can play a significant role in adding plausibility, or a sense of presence, to a real‐time simulation. However, it is not obvious how best to generate and vary these kinds of groups. In this paper, using anthropological standards for interacting distances and formations, we conduct a series of experiments to examine how these parameters inherent in human conversation are perceived for virtual characters. Our results show that, although participants were sensitive to both distance and orientation changes between talkers and listeners in a virtual conversation, they were not as sensitive to anomalous gesturing behaviours across different distances. Copyright © 2012 John Wiley & Sons, Ltd. Cathy Ennis, Carol O'Sullivan |
Comput. Animat. Virtual Worlds | 1 |
| 2011 | Perceptual effects of scene context and viewpoint for virtual pedestrian crowdsabstractIn this article, we evaluate the effects of position, orientation, and camera viewpoint on the plausibility of pedestrian formations. In a set of three perceptual studies, we investigated how humans perceive characteristics of virtual crowds in static scenes reconstructed from annotated still images, where the orientations and positions of the individuals have been modified. We found that by applying rules based on the contextual information of the scene, we improved the perceived realism of the crowd formations when compared to random formations. We also examined the effect of camera viewpoint on the plausibility of virtual pedestrian scenes, and we found that an eye-level viewpoint is more effective for disguising random behaviors, while a canonical viewpoint results in these behaviors being perceived as less realistic than an isometric or top-down viewpoint. Results from these studies can help in the creation of virtual crowds, such as computer graphics pedestrian models or architectural scenes, and identify situations when users' perception is less accurate. Cathy Ennis, Christopher Peters 0001, Carol O'Sullivan |
ACM Trans. Appl. Percept. | 1 |
| 2010 | Seeing is believing: body motion dominates in multisensory conversationsabstractIn many scenes with human characters, interacting groups are an important factor for maintaining a sense of realism. However, little is known about what makes these characters appear realistic. In this paper, we investigate human sensitivity to audio mismatches (i.e., when individuals' voices are not matched to their gestures) and visual desynchronization (i.e., when the body motions of the individuals in a group are mis-aligned in time) in virtual human conversers. Using motion capture data from a range of both polite conversations and arguments, we conduct a series of perceptual experiments and determine some factors that contribute to the plausibility of virtual conversing groups. We found that participants are more sensitive to visual desynchronization of body motions, than to mismatches between the characters' gestures and their voices. Furthermore, synthetic conversations can appear sufficiently realistic once there is an appropriate balance between talker and listener roles. This is regardless of body motion desynchronization or mismatched audio. Cathy Ennis, Rachel McDonnell, Carol O'Sullivan |
ACM Trans. Graph. | 1 |
| 2009 | Talking bodies: Sensitivity to desynchronization of conversationsabstractIn this article, we investigate human sensitivity to the coordination and timing of conversational body language for virtual characters. First, we captured the full body motions (excluding faces and hands) of three actors conversing about a range of topics, in either a polite (i.e., one person talking at a time) or debate/argument style. Stimuli were then created by applying the motion-captured conversations from the actors to virtual characters. In a 2AFC experiment, participants viewed paired sequences of synchronized and desynchronized conversations and were asked to guess which was the real one. Detection performance was above chance for both conversation styles but more so for the polite conversations, where desynchronization was more noticeable. Rachel McDonnell, Cathy Ennis, Simon Dobbyn, Carol O'Sullivan |
ACM Trans. Appl. Percept. | 2 |