VLDB 2026 Research / reviewers in the wild / expert
Riccardo Bovo
dblp:229/6488
· DBLP profile ↗
14ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-0634-0260ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating AI Assistance in Human-in-the-Loop Scenographic Prop AuthoringabstractIn scenographic productions, the selection of stage props is a structured process that follows a careful reading of the script and unfolds through three different phases: a visual exploration aimed at stimulating and shaping the creative vision, the selection of the props to be used on stage and their customization. While recent AI-driven approaches enable the automatic generation of 3D content, relatively few systems support sustained user involvement and iterative interpretive control throughout the authoring workflow. We present a WebXR-based Human-in-the-Loop system designed to support scenographic prop authoring by aligning AI assistance with this canonical workflow. The system provides AI-generated alternatives—images, retrieved 3D models, and textures—that users can iteratively explore, select, and refine while retaining control over creative decisions. We report an exploratory within-subject study with 20 experts in theater and cinema, comparing an AI-assisted workflow with a non-AI baseline. Results indicate that AI assistance was associated with higher perceived technology acceptance and creative support, alongside increased cognitive workload. Differences in the perceived coherence of the final props were not statistically significant, although descriptive trends suggest potential support for representing object type and symbolic meaning, indicating that AI assistance may be more effective in supporting early-stage exploration than precise semantic refinement. Giacomo Vallasciani, Pasquale Cascarano, Jacopo Meglioraldi, Daniele Giunchi, Riccardo Bovo, Gustavo Marfia |
AVI | 5 |
| 2026 | Tuning Immersion and Performance with Adaptive Generative Music in VRabstractMusic in virtual environments has long been treated as a temporal evolving element, enhancing atmosphere and game pace but rarely considered as a performance adaptive element. Recent advances in artificial intelligence (AI) and procedural audio make it possible to generate music that adapts in real time to player actions and system state. Yet, despite its potential, the behavioural impact of such adaptive generative soundtracks in head-mounted display-based virtual reality (VR) remains largely unexplored. To address this gap, we introduce a VR archery system that integrates Google MusicFXDJ with Ubiq-Genie to deliver continuous AI-generated adaptive music driven by gameplay events. In a within-subjects experiment (N = 22), participants completed trials with either a stylistically-matched fixed soundtrack or an adaptive soundtrack that escalated tension across four phases as arrows depleted. Measures combined self-reported ratings of presence, focus, stress, and emotional impact, with performance metrics of accuracy and aiming time. Results reveal that adaptive generative music not only heightens immersion and emotional salience but also modulates motor precision in an arousal-dependent inverted-U pattern: moderate musical tension improved accuracy and speed, whereas excessive tension impaired them. These findings establish AI-generated music as a powerful behavioural feedback modality in VR, opening pathways for training, rehabilitation, and next-generation immersive entertainment. Jiayuan Wen, Daniele Giunchi, Pasquale Cascarano, Riccardo Bovo, Eyal Ofek, Anthony Steed |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | A Mixed-Methods Investigation of XR Security Warnings - Lessons LearnedabstractAs immersive XR environments become more prevalent, timely and effective security warnings are essential to protect users from cyberattacks that compromise performance and well-being. This paper investigates how users perceive and respond to in-headset alerts triggered during Denial-of-Service (DoS) attacks. We developed a real-time warning system and evaluated its effectiveness across three pilot studies ($n$= 46) in healthcare and industrial training scenarios. Using self-report measures (IDSQ, SAM) and behavioural categorization, we assessed alert comprehension, urgency perception, and user action. We distil three design lessons emphasizing the importance of visual salience, modality coordination, and urgency calibration. These findings offer practical guidance for designing effective XR security notifications that support user awareness and action during immersive threats. Junyi Zou, Riccardo Bovo, George Loukas |
CBMI | 2 |
| 2025 | Embodiment in Smartphone Augmented Reality: Effects on User PerformanceabstractEmbodiment in mixed reality describes the sensation of experiencing a virtual representation as an extension of one’s own body. While research has extensively examined embodiment in virtual reality (VR) and head-mounted augmented reality (AR), its impact on smartphones remains underexplored. This study examines how smartphone-based AR embodiment affects user engagement and cognitive performance in a comprehension task. A study involving 24 participants explored whether using a smartphone AR face-filter to embody a virtual audience member influenced the recall of a historical speech. Findings show that participants in the AR condition scored higher on a factual quiz than those in the control group. At the same time, stronger perceived embodiment, especially self-location, was negatively associated with quiz performance, consistent with Cognitive Load Theory. These results should be interpreted cautiously: our comparison contrasted a static image (no AR) with AR that included facial embodiment, so we did not include an “AR without embodiment” condition to fully separate AR novelty from embodiment. Stimuli were also restricted to a single speech and a single historical scene presented as a static image, limiting generalizability to other content and to dynamic or interactive AR. Finally, the sample was modest (N=24), so estimates are preliminary and warrant replication. We discuss implications for designing smartphone AR that balances engagement with cognitive efficiency. Han Loong Low, Daniele Giunchi, Riccardo Bovo, Pasquale Cascarano, Nick Ritchie, Enrico Costanza, Anthony Steed |
MUM | 3 |
| 2025 | EmBARDiment: an Embodied AI Agent for Productivity in XRabstractXR devices running chat-bots powered by Large Language Models (LLMs) have the to become always-on agents that enable much better productivity scenarios. Current screen based chat-bots do not take advantage of the the full-suite of natural inputs available in XR, including inward facing sensor data, instead they over-rely on explicit voice or text prompts, sometimes paired with multi-modal data dropped as part of the query. We propose a solution that leverages an attention framework that derives context implicitly from user actions, eye-gaze, and contextual memory within the XR environment. Our work minimizes the need for engineered explicit prompts, fostering grounded and intuitive interactions that glean user insights for the chat-bot. Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez, Li-Te Cheng, Mar González-Franco |
VR | 1 |
| 2025 | See It and Hear It: Multimodal Guidance in MR-Based Neurosurgical Simulation for Skill RetentionabstractExternal Ventricular Drain (EVD) placement is a complex neurosurgical task that requires identifying a target point within the brain and accurately positioning a catheter at the appropriate angle. While Mixed Reality (MR) technologies have seen limited adoption in the operating room, they offer significant potential for developing training systems that enhance skill acquisition and retention in unaided conditions. A current gap in research concerns the effectiveness of multimodal guidance systems that incorporate both visual and audio-based MR cues. In this paper, we present an MR-based simulator for EVD placement training and evaluate the impact of three MR-guided training modalities: (1) a baseline condition using only 2D CT scans and a 2D catheter projection; (2) a visual guidance modality incorporating a 3D trajectory overlay; and (3) an embodied-audio guidance modality featuring a virtual agent delivering spoken instructions and feedback. Participants underwent a digital training phase using one of the three modalities, followed by an unaided EVD placement on a physical phantom with a real catheter to evaluate skill transfer and retention. Results indicate that both advanced MR modalities significantly improve procedural accuracy, execution speed and receive higher scores in usability and technology acceptance compared to the baseline. Notably, training with 3D visual trajectory guidance led to significantly higher unaided placement accuracy, indicating stronger skill retention. However, multimodal guidance demonstrated equivalent execution speed, while showing a trend toward lower overall cognitive load. Pasquale Cascarano, Andrea Loretti, Luca Zanuttini, Daniele Giunchi, Riccardo Bovo, Shirin Hajahmadi, Giacomo Vallasciani, Matteo Martinoni, Gustavo Marfia |
VRST | 5 |
| 2024 | Visualisations with semantic icons: Assessing engagement with distracting elementsabstractAs visualisations reach a broad range of audiences, designing visualisations that attract and engage becomes more critical. Prior work suggests that semantic icons entice and immerse the reader; however, little is known about their impact with informational tasks and when the viewer’s attention is divided because of a distracting element. To address this gap, we first explored a variety of semantic icons with various visualisation attributes. The findings of this exploration shaped the design of our primary comparative online user studies, where participants saw a target visualisation with a distracting visualisation on a web page and were asked to extract insights. Their engagement was measured through three dependent variables: (1) visual attention, (2) effort to write insights, and (3) self-reported engagement. In Study 1, we discovered that visualisations with semantic icons were consistently perceived to be more engaging than the plain version. However, we found no differences in visual attention and effort between the two versions. Thus, we ran Study 2 using visualisations with more salient semantic icons to achieve maximum contrast. The results were consistent with our first Study. Furthermore, we found that semantic icons elevated engagement with visualisations depicting less interesting and engaging topics from the participant’s perspective. We extended prior work by demonstrating the semantic value after performing an informational task (extracting insights) and reflecting on the visualisation, besides its value to the first impression. Our findings may be helpful to visualisation designers and storytellers keen on designing engaging visualisations with limited resources. We also contribute reflections on engagement measurements with visualisations and provide future directions. Muna Alebri, Enrico Costanza, Georgia Panagiotidou 0001, Duncan P. Brumby, Fatima Althani, Riccardo Bovo |
Int. J. Hum. Comput. Stud. | 6 |
| 2023 | Speech-Augmented Cone-of-Vision for Exploratory Data AnalysisabstractMutual awareness of visual attention is crucial for successful collaboration. Previous research has explored various ways to represent visual attention, such as field-of-view visualizations and cursor visualizations based on eye-tracking, but these methods have limitations. Verbal communication is often utilized as a complementary strategy to overcome such disadvantages. This paper proposes a novel method that combines verbal communication with the Cone of Vision to improve gaze inference and mutual awareness in VR. We conducted a within-group study with pairs of participants who performed a collaborative analysis of data visualizations in VR. We found that our proposed method provides a better approximation of eye gaze than the approximation provided by head direction. Furthermore, we release the first collaborative head, eyes, and verbal behaviour dataset. The results of this study provide a foundation for investigating the potential of verbal communication as a tool for enhancing visual cues for joint attention. Riccardo Bovo, Daniele Giunchi, Ludwig Sidenmark, Joshua Newn, Hans-Werner Gellersen, Enrico Costanza, Thomas Heinis |
CHI | 1 |
| 2022 | Shall I describe it or shall I move closer? Verbal references and locomotion in VR collaborative search tasks
Riccardo Bovo, Daniele Giunchi, Enrico Costanza, Anthony Steed, Thomas Heinis |
ECSCW | 1 |
| 2022 | Real-time head-based deep-learning model for gaze probability regions in collaborative VRabstractEye behavior has gained much interest in the VR research community as an interactive input and support for collaboration. Researchers used head behavior and saliency to implement gaze inference models when eye-tracking is missing. However, these solutions are resource-demanding and thus unfit for untethered devices, and their angle accuracy is around 7°, which can be a problem in high-density informative areas. To address this issue, we propose a lightweight deep learning model that generates the probability density function of the gaze as a percentile contour. This solution allows us to introduce a visual attention representation based on a region rather than a point. In this way, we manage the trade-off between the ambiguity of a region and the error of a point. We tested our model in untethered devices with real-time performances; we evaluated its accuracy, outperforming our identified baselines (average fixation map and head direction). Riccardo Bovo, Daniele Giunchi, Ludwig Sidenmark, Hans-Werner Gellersen, Enrico Costanza, Thomas Heinis |
ETRA | 1 |
| 2022 | Cone of Vision as a Behavioural Cue for VR CollaborationabstractMutual awareness of visual attention is essential for collaborative work. In the field of collaborative virtual environments (CVE), it has been proposed to use Field-of-View (FoV) frustum visualisations as a cue to support mutual awareness during collaboration. Recent studies on FoV frustum visualisations focus on asymmetric collaboration with AR/VR hardware setups and 3D reconstructed environments. In contrast, we focus on the general-purpose CVEs (i.e., VR shared offices), whose popularity is increasing due to the availability of low-cost headsets, and the restrictions imposed by the pandemic. In these CVEs collaboration roles are symmetrical, and the same 2D content available on desktop computers is displayed on 2D surfaces in a 3D space (VR screens). We prototyped one such CVE to evaluate FoV frustrum visualisation within this collaboration scenario. We also implement a FoV visualisation generated from an average fixation map (AFM), therefore directly generated by users' gaze behaviour which we call Cone of Vision (CoV). Our approach to displaying the frustum visualisations is tailored for 2D surfaces in 3D space and allows for self-awareness of this visual cue. We evaluate CoV in the context of a general exploratory data analysis (EDA) with 10 pairs of participants. Our findings indicate that CoV is beneficial during shifts between independent and collaborative work and supports collaborative progression across the visualisation. Self-perception of the CoV improves visual attention coupling, reduces the number of times users watch the collaborator's avatars and offers a consistent representation of the shared reality. Riccardo Bovo, Daniele Giunchi, Muna Alebri, Anthony Steed, Enrico Costanza, Thomas Heinis |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | Perceived Realism of Pedestrian Crowds Trajectories in VRabstractCrowd simulation algorithms play an essential role in populating Virtual Reality (VR) environments with multiple autonomous humanoid agents. The generation of plausible trajectories can be a significant computational cost for real-time graphics engines, especially in untethered and mobile devices such as portable VR devices. Previous research explores the plausibility and realism of crowd simulations on desktop computers but fails to account the impact it has on immersion. This study explores how the realism of crowd trajectories affects the perceived immersion in VR. We do so by running a psychophysical experiment in which participants rate the realism of real/synthetic trajectories data, showing similar level of perceived realism. Daniele Giunchi, Riccardo Bovo, Panayiotis Charalambous, Fotis Liarokapis, Alastair Shipman, Stuart James, Anthony Steed, Thomas Heinis |
VRST | 2 |
| 2020 | Detecting errors in pick and place procedures: detecting errors in multi-stage and sequence-constrained manual retrieve-assembly proceduresabstractMany human activities, such as manufacturing and assembly, are sequence-constrained procedural tasks (SPTs): they consist of a series of steps that must be executed in a specific spatial/temporal order. However, these tasks can be error prone - steps can be missed out, executed out-of-order, and repeated. The ability to automatically predict if a person is about to commit an error could greatly help in these cases. The prediction could be used, for example, to provide feedback to prevent mistakes or mitigate their effects. In this paper, we present a novel approach for real-time error prediction for multi-step sequence tasks which uses a minimum viable set of behavioural signals. We have three main contributions. The first we present an architecture for real-time error prediction based on task tracking and intent prediction. The second is to explore the effectiveness of using hand position and eye-gaze tracking for task tracking. We confirm that eye-gaze is more effective for intent prediction, hand tracking is more accurate for task tracking and that combining the two provides the best overall response. We show that using Hands and Gaze tracking data we can predict selection/placement errors with an F1 score of 97%, approximately 300ms before the error would occur. Finally, we discuss the application of this hand-gaze error detection architecture used in conjunction with head-mounted AR displays, to support industrial manual assembly. Riccardo Bovo, Nicola Binetti, Duncan P. Brumby, Simon J. Julier |
IUI | 1 |
| 2018 | A Taxonomy for Combining Activity Recognition and Process Discovery in Industrial Environments
Felix Mannhardt, Riccardo Bovo, Manuel Fradinho, Simon J. Julier |
IDEAL (2) | 2 |