VLDB 2026 Research / reviewers in the wild / expert
Jonathan Ehret
dblp:265/2818
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-6270-5119ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Objectifying Social Presence: Evaluating Multimodal Degraders in ECAs Using the Heard Text Recall ParadigmabstractEmbodied conversational agents (ECAs) are key social interaction partners in various virtual reality (VR) applications, with their perceived social presence significantly influencing the quality and effectiveness of user-ECA interactions. This article investigates the potential of a novel indirect objective proxy for evaluating social presence, which is traditionally assessed through subjective questionnaires. As proxy we use the Heard Text Recall (HTR) paradigm, originally designed to assess memory performance in listening tasks. In the task participants have to recall information from family stories being told. When combined with a secondary task in a dual-task paradigm, the HTR can also be used to evaluate cognitive spare capacity, which is then correlated with subjectively rated social presence. As a prerequisite for this investigation, we introduce various co-verbal gesture modification techniques and assess their impact on the perceived naturalness of the presenting ECA, a crucial aspect fostering social presence. The main study then explores the applicability of HTR as a proxy for social presence by examining its effectiveness under different multimodal degraders of ECA behavior, including degraded co-verbal gestures, omitted lip synchronization, and the use of synthetic voices. The findings suggest that while HTR shows potential as an objective measure of social presence, its effectiveness is primarily evident in response to substantial changes in ECA behavior. Additionally, the study also highlights the negative effects of synthetic voices and inadequate lip synchronization on perceived social presence, emphasizing the need for careful consideration of these elements in ECA design. Jonathan Ehret, Jonas Schüppen, Chinthusa Mohanathasan, Cosima A. Ermert, Janina Fels, Sabine J. Schlittmeier, Torsten W. Kuhlen, Andrea Bönsch |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | A Latency-Optimized LLM-based Multimodal Dialogue System for Embodied Conversational Agents in VRabstractFigure 1: Our LLM-driven guide assists users in exploring the museum and in learning about artworks and artists. Konstantin W. Kühlem, Jonathan Ehret, Torsten W. Kuhlen, Andrea Bönsch |
IVA | 2 |
| 2024 | German and Dutch Translations of the Artificial-Social-Agent Questionnaire Instrument for Evaluating Human-Agent InteractionsabstractEnabling the widespread utilization of the Artificial-Social-Agent (ASA) Questionnaire, a research instrument to comprehensively assess diverse ASA qualities while ensuring comparability, necessitates translations beyond the original English source language questionnaire. We thus present Dutch and German translations of the long and short versions of the ASA Questionnaire and describe the translation challenges we encountered. Summative assessments with 240 English-Dutch and 240 English-German bilingual participants show, on average, excellent correlations (Dutch ICC M = 0.82, SD = 0.07, range [0.58, 0.93]; German ICC M = 0.81, SD = 0.09, range [0.58, 0.94]) with the original long version on the construct and dimension level. Results for the short version show, on average, good correlations (Dutch ICC M = 0.65, SD = 0.12, range [0.39, 0.82]; German ICC M = 0.67, SD = 0.14, range [0.30, 0.91]). We hope these validated translations allow the Dutch and German-speaking populations to evaluate ASAs in their own language. Nele Albers, Andrea Bönsch, Jonathan Ehret, Boleslav A. Khodakov, Willem-Paul Brinkman |
IVA | 3 |
| 2023 | Where Do They Go?: Overhearing Conversing Pedestrian Groups during Scene ExplorationabstractOn entering an unknown immersive virtual environment, a user's first task is gaining knowledge about the respective scene, termed scene exploration. While many techniques for aided scene exploration exist, such as virtual guides, or maps, unaided wayfinding through pedestrians-as-cues is still in its infancy. We contribute to this research by indirectly guiding users through pedestrian groups conversing about their target location. A user who overhears the conversation without being a direct addressee can consciously decide whether to follow the group to reach an unseen point of interest. We outline our approach and give insights into the results of a first feasibility study in which we compared our new approach to non-talkative groups and groups conversing about random topics. Andrea Bönsch, Till Sittart, Jonathan Ehret, Torsten W. Kuhlen |
IVA | 3 |
| 2023 | Whom Do You Follow?: Pedestrian Flows Constraining the User's Navigation during Scene ExplorationabstractIn this work-in-progress, we strive to combine two wayfinding techniques supporting users in gaining scene knowledge, namely (i) the River Analogy, in which users are considered as boats automatically floating down predefined rivers, e.g., streets in an urban scene, and (ii) virtual pedestrian flows as social cues indirectly guiding users through the scene. In our combined approach, the pedestrian flows function as rivers. To navigate through the scene, users leash themselves to a pedestrian of choice, considered as boat, and are dragged along the flow towards an area of interest. Upon arrival, users can detach themselves to freely explore the site without navigational constraints. We briefly outline our approach, and discuss the results of an initial study focusing on various leashing visualizations. Andrea Bönsch, Lukas B. Zimmermann, Jonathan Ehret, Torsten W. Kuhlen |
IVA | 3 |
| 2023 | Who's next?: Integrating Non-Verbal Turn-Taking Cues for Embodied Conversational AgentsabstractTaking turns in a conversation is a delicate interplay of various signals, which we as humans can easily decipher. Embodied conversational agents (ECAs) communicating with humans should leverage this ability for smooth and enjoyable conversations. Extensive research has analyzed human turn-taking cues, and attempts have been made to predict turn-taking based on observed cues. These cues vary from prosodic, semantic, and syntactic modulation over adapted gesture and gaze behavior to actively used respiration. However, when generating such behavior for social robots or ECAs, often only single modalities were considered, e.g., gazing. We strive to design a comprehensive system that produces cues for all non-verbal modalities: gestures, gaze, and breathing. The system provides valuable cues without requiring speech content adaptation. We evaluated our system in a VR-based user study with N = 32 participants executing two subsequent tasks. First, we asked them to listen to two ECAs taking turns in several conversations. Second, participants engaged in taking turns with one of the ECAs directly. We examined the system's usability and the perceived social presence of the ECAs' turn-taking behavior, both with respect to each individual non-verbal modality and their interplay. While we found effects of gesture manipulation in interactions with the ECAs, no effects on social presence were found. Jonathan Ehret, Andrea Bönsch, Patrick Nossol, Cosima A. Ermert, Chinthusa Mohanathasan, Sabine J. Schlittmeier, Janina Fels, Torsten W. Kuhlen |
IVA | 1 |
| 2021 | Being Guided or Having Exploratory Freedom: User Preferences of a Virtual Agent's Behavior in a MuseumabstractA virtual guide in an immersive virtual environment allows users a structured experience without missing critical information. However, although being in an interactive medium, the user is only a passive listener, while the embodied conversational agent (ECA) fulfills the active roles of wayfinding and conveying knowledge. Thus, we investigated for the use case of a virtual museum, whether users prefer a virtual guide or a free exploration accompanied by an ECA who imparts the same information compared to the guide. Results of a small within-subjects study with ahead-mounted display are given and discussed, resulting in the idea of combining benefits of both conditions for a higher user acceptance. Furthermore, the study indicated the feasibility of the carefully designed scene and ECA's appearance. Andrea Bönsch, David Hashem, Jonathan Ehret, Torsten W. Kuhlen |
IVA | 3 |
| 2021 | Do Prosody and Embodiment Influence the Perceived Naturalness of Conversational Agents' Speech?abstractFor conversational agents’ speech, either all possible sentences have to be prerecorded by voice actors or the required utterances can be synthesized. While synthesizing speech is more flexible and economic in production, it also potentially reduces the perceived naturalness of the agents among others due to mistakes at various linguistic levels. In our article, we are interested in the impact of adequate and inadequate prosody, here particularly in terms of accent placement, on the perceived naturalness and aliveness of the agents. We compare (1) inadequate prosody, as generated by off-the-shelf text-to-speech (TTS) engines with synthetic output; (2) the same inadequate prosody imitated by trained human speakers; and (3) adequate prosody produced by those speakers. The speech was presented either as audio-only or by embodied, anthropomorphic agents, to investigate the potential masking effect by a simultaneous visual representation of those virtual agents. To this end, we conducted an online study with 40 participants listening to four different dialogues each presented in the three Speech levels and the two Embodiment levels. Results confirmed that adequate prosody in human speech is perceived as more natural (and the agents are perceived as more alive) than inadequate prosody in both human (2) and synthetic speech (1). Thus, it is not sufficient to just use a human voice for an agents’ speech to be perceived as natural—it is decisive whether the prosodic realisation is adequate or not. Furthermore, and surprisingly, we found no masking effect by speaker embodiment, since neither a human voice with inadequate prosody nor a synthetic voice was judged as more natural, when a virtual agent was visible compared to the audio-only condition. On the contrary, the human voice was even judged as less “alive” when accompanied by a virtual agent. In sum, our results emphasize, on the one hand, the importance of adequate prosody for perceived naturalness, especially in terms of accents being placed on important words in the phrase, while showing, on the other hand, that the embodiment of virtual agents plays a minor role in the naturalness ratings of voices. Jonathan Ehret, Andrea Bönsch, Lukas Aspöck, Christine T. Röhr, Stefan Baumann, Martine Grice, Janina Fels, Torsten W. Kuhlen |
ACM Trans. Appl. Percept. | 1 |
| 2020 | Inferring a User's Intent on Joining or Passing by Social GroupsabstractModeling the interactions between users and social groups of virtual agents (VAs) is vital in many virtual-reality-based applications. However, only little research on group encounters has been conducted yet. We intend to close this gap by focusing on the distinction between joining and passing-by a group. To enhance the interactive capacity of VAs in these situations, knowing the user's objective is required to show reasonable reactions. To this end, we propose a classification scheme which infers the user's intent based on social cues such as proxemics, gazing and orientation, followed by triggering believable, non-verbal actions on the VAs. We tested our approach in a pilot study with overall promising results and discuss possible improvements for further studies. Andrea Bönsch, Alexander R. Bluhm, Jonathan Ehret, Torsten W. Kuhlen |
IVA | 3 |
| 2020 | Immersive Sketching to Author Crowd Movements in Real-timeabstractSketch-based interfaces in 2D screen space allow to efficiently author the flow of virtual crowds in a direct and interactive manner. Here, options to redirect a flow by sketching barriers, or guiding entities based on a sketched network of connected sections are provided. Andrea Bönsch, Sebastian J. Barton, Jonathan Ehret, Torsten W. Kuhlen |
IVA | 3 |
| 2020 | The Impact of a Virtual Agent's Non-Verbal Emotional Expression on a User's Personal Space PreferencesabstractVirtual-reality-based interactions with virtual agents (VAs) are likely subject to similar influences as human-human interactions. In either real or virtual social interactions, interactants try to maintain their personal space (PS), an ubiquitous, situative, flexible safety zone. Building upon larger PS preferences to humans and VAs with angry facial expressions, we extend the investigations to whole-body emotional expressions. In two immersive settings-HMD and CAVE-66 males were approached by an either happy, angry, or neutral male VA. Subjects preferred a larger PS to the angry VA when being able to stop him at their convenience (Sample task), replicating previous findings, and when being able to actively avoid him (Pass By task). In the latter task, we also observed larger distances in the CAVE than in the HMD. Andrea Bönsch, Sina Radke, Jonathan Ehret, Ute Habel, Torsten W. Kuhlen |
IVA | 3 |
| 2020 | Evaluating the Influence of Phoneme-Dependent Dynamic Speaker Directivity of Embodied Conversational Agents' SpeechabstractGenerating natural embodied conversational agents within virtual spaces crucially depends on speech sounds and their directionality. In this work, we simulated directional filters to not only add directionality, but also directionally adapt each phoneme. We therefore mimic reality where changing mouth shapes have an influence on the directional propagation of sound. We conducted a study (n = 32) evaluating naturalism ratings, preference and distinguishability of omnidirectional speech auralization compared to static and dynamic, phoneme-dependent directivities. The results indicated that participants cannot distinguish dynamic from static directivity. Furthermore, participants' preference ratings aligned with their naturalism ratings. There was no unanimity, however, with regards to which auralization is the most natural. Jonathan Ehret, Jonas Stienen, Chris Brozdowski, Andrea Bönsch, Irene Mittelberg, Michael Vorländer, Torsten W. Kuhlen |
IVA | 1 |