VLDB 2026 Research / reviewers in the wild / expert
Peter Andrews
dblp:86/4082
· DBLP profile ↗
5ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0007-2853-2569ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "Same Voice, Different Language": An Exploration of Voice-Cloned Translation to Support Non-Native Speakers in Online MeetingsabstractCross-lingual meetings have become essential for global collaboration, yet current translation technologies often strip away vocal identity — the unique speaker characteristics that convey nuance and social presence. While generic text-to-speech (TTS) provides basic intelligibility, it creates a disconnect between speakers and their translated voices, potentially undermining engagement and comprehension. This paper investigates whether voice cloning technology can bridge this gap by preserving speaker identity in real-time translation. We present a controlled study comparing four voice conditions in meeting interpretation: original speech, gender-neutral TTS, gender-matched TTS, and voice cloning. Through a within-subjects experiment with 45 participants, we demonstrate that voice cloning significantly reduces mental workload (p <.001) and enhances user experience across pragmatic quality (p <.001), hedonic quality (p <.001), and overall satisfaction (p <.001) compared to traditional TTS. While original speech maintained advantages in naturalness, voice cloning achieved superior intelligibility, social impression, and user preference. Qualitative analysis revealed that participants valued voice cloning for preserving speaker identity and improving conversation tracking in multi-speaker scenarios. Our findings suggest that identity-preserving translation represents a significant advancement for cross-lingual communication systems, offering both cognitive and social benefits. We conclude with design implications for integrating voice cloning into meeting platforms while addressing ethical considerations around consent and transparency. Yong Ma 0003, Yuchong Zhang 0001, Peter Andrews, Zhikun Wu, Stephanie Zubicueta Portales, Morten Fjeld |
IUI | 3 |
| 2025 | AiModerator: A Co-Pilot for Hyper-Contextualization in Political Debate Video
Peter Andrews, Njål Borch, Morten Fjeld |
IUI | 1 |
| 2024 | AiCommentator: A Multimodal Conversational Agent for Embedded Visualization in Football ViewingabstractTraditionally, sports commentators provide viewers with diverse information, encompassing in-game developments and player performances. Yet young adult football viewers increasingly use mobile devices for deeper insights during football matches. Such insights into players on the pitch and performance statistics support viewers’ understanding of game stakes, creating a more engaging viewing experience. Inspired by commentators’ traditional roles and to incorporate information into a single platform, we developed AiCommentator, a Multimodal Conversational Agent (MCA) for embedded visualization and conversational interactions in football broadcast video. AiCommentator integrates embedded visualization, either with an automated non-interactive or with a responsive interactive commentary mode. Our system builds upon multimodal techniques, integrating computer vision and large language models, to demonstrate ways for designing tailored, interactive sports-viewing content. AiCommentator’s event system infers game states based on a multi-object tracking algorithm and computer vision backend, facilitating automated responsive commentary. We address three key topics: evaluating young adults’ satisfaction and immersion across the two viewing modes, enhancing viewer understanding of in-game events and players on the pitch, and devising methods to present this information in a usable manner. In a mixed-method evaluation (n=16) of AiCommentator, we found that the participants appreciated aspects of both system modes but preferred the interactive mode, expressing a higher degree of engagement and satisfaction. Our paper reports on our development of AiCommentator and presents the results from our user study, demonstrating the promise of interactive MCA for a more engaging sports viewing experience. Systems like AiCommentator could be pivotal in transforming the interactivity and accessibility of sports content, revolutionizing how sports viewers engage with video content. Peter Andrews, Oda Elise Nordberg, Stephanie Zubicueta Portales, Njål Borch, Frode Guribye, Kazuyuki Fujita, Morten Fjeld |
IUI | 1 |
| 2024 | Designing for Automated Sports Commentary SystemsabstractAdvancements in Natural Language Processing (NLP) and Computer Vision (CV) are revolutionizing how we experience sports broadcasting. Traditionally, sports commentary has played a crucial role in enhancing viewer understanding and engagement with live games. Yet, the prospects of automated commentary, especially in light of these technological advancements and their impact on viewers’ experience, remain largely unexplored. This paper elaborates upon an innovative automated commentary system that integrates NLP and CV to provide a multimodal experience, combining auditory feedback through text-to-speech and visual cues, known as italicizing, for real-time in-game commentary. The system supports color commentary, which aims to inform the viewer of information surrounding the game by pulling additional content from a database. Moreover, it also supports play-by-play commentary covering in-game developments derived from an event system based on CV. As the system reinvents the role of commentary in sports video, we must consider the design and implications of multimodal artificial commentators. A focused user study with eight participants aimed at understanding the design implications of such multimodal artificial commentators reveals critical insights. Key findings emphasize the importance of language precision, content relevance, and delivery style in automated commentary, underscoring the necessity for personalization to meet diverse viewer preferences. Our results validate the potential value and effectiveness of multimodal feedback and derive design considerations, particularly in personalizing content to revolutionize the role of commentary in sports broadcasts. Peter Andrews, Oda Elise Nordberg, Njål Borch, Frode Guribye, Morten Fjeld |
IMX | 1 |
| 2006 | Multimedia signal processing for behavioral quantification in neuroscienceabstractWhile there have been great advances in quantification of the genotype of organisms, including full genomes for many species, the quantification of phenotype is at a comparatively primitive stage. Part of the reason is technical difficulty: the phenotype covers a wide range of characteristics, ranging from static morphological features, to dynamic behavior. The latter poses challenges that are in the area of multimedia signal processing. Automated analysis of video and audio recordings of animal and human behavior is a growing area of research, ranging from the behavioral phenotyping of genetically modified mice or drosophila to the study of song learning in birds and speech acquisition in human infants. This paper reviews recent advances and identifies key problems for a range of behavior experiments that use audio and video recording. This research area offers both research challenges and an application domain for advanced multimedia signal processing. There are a number of MMSP tools that now exist which are directly relevant for behavioral quantification, such as speech recognition, video analysis and more recently, wired and wireless sensor networks for surveillance. The research challenge is to adapt these tools and to develop new ones required for studying human and animal behavior in a high throughput manner while minimizing human intervention. In contrast with consumer applications, in the research arena there is less of a penalty for computational complexity, so that algorithmic quality can be maximized through the utilization of larger computational resources that are available to the biomedical researcher. Peter Andrews, Dan Valente, Jihène Serkhane, Partha P. Mitra, Sigal Saar, Ofer Tchernichovski, Ilan Golani |
ACM Multimedia | 1 |