Daniele Giunchi

dblp:122/6463 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-1674-8876ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating AI Assistance in Human-in-the-Loop Scenographic Prop Authoring
abstract
In scenographic productions, the selection of stage props is a structured process that follows a careful reading of the script and unfolds through three different phases: a visual exploration aimed at stimulating and shaping the creative vision, the selection of the props to be used on stage and their customization. While recent AI-driven approaches enable the automatic generation of 3D content, relatively few systems support sustained user involvement and iterative interpretive control throughout the authoring workflow. We present a WebXR-based Human-in-the-Loop system designed to support scenographic prop authoring by aligning AI assistance with this canonical workflow. The system provides AI-generated alternatives—images, retrieved 3D models, and textures—that users can iteratively explore, select, and refine while retaining control over creative decisions. We report an exploratory within-subject study with 20 experts in theater and cinema, comparing an AI-assisted workflow with a non-AI baseline. Results indicate that AI assistance was associated with higher perceived technology acceptance and creative support, alongside increased cognitive workload. Differences in the perceived coherence of the final props were not statistically significant, although descriptive trends suggest potential support for representing object type and symbolic meaning, indicating that AI assistance may be more effective in supporting early-stage exploration than precise semantic refinement.
Giacomo Vallasciani, Pasquale Cascarano, Jacopo Meglioraldi, Daniele Giunchi, Riccardo Bovo, Gustavo Marfia
AVI4
2026 Belt and whistles - adding lower body collision awareness for MR experiences
abstract
Users of Virtual Reality (VR) primarily sense their environment through audiovisual cues. The lack of haptic feedback on their body can make them unaware of virtual obstacles outside their field of view. This lack of sensing can cause the user to unknowingly penetrate virtual objects, breaking the scene’s plausibility and disrupting the experience of other users in the same virtual space. We propose a haptic belt that increases the user’s scene awareness by rendering signals of collisions and proximity to virtual objects around the user.
Diar Abdlkarim, Devika Mukherjee, Daniele Giunchi, Massimiliano Di Luca, Eyal Ofek
CHI3
2026 Tuning Immersion and Performance with Adaptive Generative Music in VR
abstract
Music in virtual environments has long been treated as a temporal evolving element, enhancing atmosphere and game pace but rarely considered as a performance adaptive element. Recent advances in artificial intelligence (AI) and procedural audio make it possible to generate music that adapts in real time to player actions and system state. Yet, despite its potential, the behavioural impact of such adaptive generative soundtracks in head-mounted display-based virtual reality (VR) remains largely unexplored. To address this gap, we introduce a VR archery system that integrates Google MusicFXDJ with Ubiq-Genie to deliver continuous AI-generated adaptive music driven by gameplay events. In a within-subjects experiment (N = 22), participants completed trials with either a stylistically-matched fixed soundtrack or an adaptive soundtrack that escalated tension across four phases as arrows depleted. Measures combined self-reported ratings of presence, focus, stress, and emotional impact, with performance metrics of accuracy and aiming time. Results reveal that adaptive generative music not only heightens immersion and emotional salience but also modulates motor precision in an arousal-dependent inverted-U pattern: moderate musical tension improved accuracy and speed, whereas excessive tension impaired them. These findings establish AI-generated music as a powerful behavioural feedback modality in VR, opening pathways for training, rehabilitation, and next-generation immersive entertainment.
Jiayuan Wen, Daniele Giunchi, Pasquale Cascarano, Riccardo Bovo, Eyal Ofek, Anthony Steed
IEEE Trans. Vis. Comput. Graph.2
2025 Embodiment in Smartphone Augmented Reality: Effects on User Performance
abstract
Embodiment in mixed reality describes the sensation of experiencing a virtual representation as an extension of one’s own body. While research has extensively examined embodiment in virtual reality (VR) and head-mounted augmented reality (AR), its impact on smartphones remains underexplored. This study examines how smartphone-based AR embodiment affects user engagement and cognitive performance in a comprehension task. A study involving 24 participants explored whether using a smartphone AR face-filter to embody a virtual audience member influenced the recall of a historical speech. Findings show that participants in the AR condition scored higher on a factual quiz than those in the control group. At the same time, stronger perceived embodiment, especially self-location, was negatively associated with quiz performance, consistent with Cognitive Load Theory. These results should be interpreted cautiously: our comparison contrasted a static image (no AR) with AR that included facial embodiment, so we did not include an “AR without embodiment” condition to fully separate AR novelty from embodiment. Stimuli were also restricted to a single speech and a single historical scene presented as a static image, limiting generalizability to other content and to dynamic or interactive AR. Finally, the sample was modest (N=24), so estimates are preliminary and warrant replication. We discuss implications for designing smartphone AR that balances engagement with cognitive efficiency.
Han Loong Low, Daniele Giunchi, Riccardo Bovo, Pasquale Cascarano, Nick Ritchie, Enrico Costanza, Anthony Steed
MUM2
2025 See It and Hear It: Multimodal Guidance in MR-Based Neurosurgical Simulation for Skill Retention
abstract
External Ventricular Drain (EVD) placement is a complex neurosurgical task that requires identifying a target point within the brain and accurately positioning a catheter at the appropriate angle. While Mixed Reality (MR) technologies have seen limited adoption in the operating room, they offer significant potential for developing training systems that enhance skill acquisition and retention in unaided conditions. A current gap in research concerns the effectiveness of multimodal guidance systems that incorporate both visual and audio-based MR cues. In this paper, we present an MR-based simulator for EVD placement training and evaluate the impact of three MR-guided training modalities: (1) a baseline condition using only 2D CT scans and a 2D catheter projection; (2) a visual guidance modality incorporating a 3D trajectory overlay; and (3) an embodied-audio guidance modality featuring a virtual agent delivering spoken instructions and feedback. Participants underwent a digital training phase using one of the three modalities, followed by an unaided EVD placement on a physical phantom with a real catheter to evaluate skill transfer and retention. Results indicate that both advanced MR modalities significantly improve procedural accuracy, execution speed and receive higher scores in usability and technology acceptance compared to the baseline. Notably, training with 3D visual trajectory guidance led to significantly higher unaided placement accuracy, indicating stronger skill retention. However, multimodal guidance demonstrated equivalent execution speed, while showing a trend toward lower overall cognitive load.
Pasquale Cascarano, Andrea Loretti, Luca Zanuttini, Daniele Giunchi, Riccardo Bovo, Shirin Hajahmadi, Giacomo Vallasciani, Matteo Martinoni, Gustavo Marfia
VRST4
2024 DreamCodeVR: Towards Democratizing Behavior Design in Virtual Reality with Speech-Driven Programming
abstract
Virtual Reality (VR) has revolutionized how we interact with digital worlds. However, programming for VR remains a complex and challenging task, requiring specialized skills and knowledge. Powered by large language models (LLMs), DreamCodeVR is designed to assist users, irrespective of their coding skills, in crafting basic object behavior in VR environments by translating spoken language into code within an active application. This approach seeks to simplify the process of defining behaviors visual changes through speech. Our preliminary user study indicated that the system’s speech interface supports elementary programming tasks, highlighting its potential to improve accessibility for users with varying technical skills. However, it also uncovered a wide range of challenges and opportunities. In an extensive discussion, we detail the system’s strengths, weaknesses, and areas for future research.
Daniele Giunchi, Nels Numan, Elia Gatti, Anthony Steed
VR1
2023 Speech-Augmented Cone-of-Vision for Exploratory Data Analysis
abstract
Mutual awareness of visual attention is crucial for successful collaboration. Previous research has explored various ways to represent visual attention, such as field-of-view visualizations and cursor visualizations based on eye-tracking, but these methods have limitations. Verbal communication is often utilized as a complementary strategy to overcome such disadvantages. This paper proposes a novel method that combines verbal communication with the Cone of Vision to improve gaze inference and mutual awareness in VR. We conducted a within-group study with pairs of participants who performed a collaborative analysis of data visualizations in VR. We found that our proposed method provides a better approximation of eye gaze than the approximation provided by head direction. Furthermore, we release the first collaborative head, eyes, and verbal behaviour dataset. The results of this study provide a foundation for investigating the potential of verbal communication as a tool for enhancing visual cues for joint attention.
Riccardo Bovo, Daniele Giunchi, Ludwig Sidenmark, Joshua Newn, Hans-Werner Gellersen, Enrico Costanza, Thomas Heinis
CHI2
2022 Shall I describe it or shall I move closer? Verbal references and locomotion in VR collaborative search tasks
Riccardo Bovo, Daniele Giunchi, Enrico Costanza, Anthony Steed, Thomas Heinis
ECSCW2
2022 Real-time head-based deep-learning model for gaze probability regions in collaborative VR
abstract
Eye behavior has gained much interest in the VR research community as an interactive input and support for collaboration. Researchers used head behavior and saliency to implement gaze inference models when eye-tracking is missing. However, these solutions are resource-demanding and thus unfit for untethered devices, and their angle accuracy is around 7°, which can be a problem in high-density informative areas. To address this issue, we propose a lightweight deep learning model that generates the probability density function of the gaze as a percentile contour. This solution allows us to introduce a visual attention representation based on a region rather than a point. In this way, we manage the trade-off between the ambiguity of a region and the error of a point. We tested our model in untethered devices with real-time performances; we evaluated its accuracy, outperforming our identified baselines (average fixation map and head direction).
Riccardo Bovo, Daniele Giunchi, Ludwig Sidenmark, Hans-Werner Gellersen, Enrico Costanza, Thomas Heinis
ETRA2
2022 Fast Blue-Noise Generation via Unsupervised Learning
abstract
Blue noise is known for its uniformity in the spatial domain, avoiding the appearance of structures such as voids and clusters. Because of this characteristic, it has been adopted in a wide range of visual computing applications, such as image dithering, rendering and visualisation. This has motivated the development of a variety of generative methods for blue noise, with different trade-offs in terms of accuracy and computational performance. We propose a novel unsupervised learning approach that leverages a neural network architecture to generate blue noise masks with high accuracy and real-time performance, starting from a white noise input. We train our model by combining three unsupervised losses that work by conditioning the Fourier spectrum and intensity histogram of noise masks predicted by the network. We evaluate our method by leveraging the generated noise for two applications: grayscale blue noise masks for image dithering, and blue noise samples for Monte Carlo integration.
Daniele Giunchi, Alejandro Sztrajman, Anthony Steed
IJCNN1
2022 Cone of Vision as a Behavioural Cue for VR Collaboration
abstract
Mutual awareness of visual attention is essential for collaborative work. In the field of collaborative virtual environments (CVE), it has been proposed to use Field-of-View (FoV) frustum visualisations as a cue to support mutual awareness during collaboration. Recent studies on FoV frustum visualisations focus on asymmetric collaboration with AR/VR hardware setups and 3D reconstructed environments. In contrast, we focus on the general-purpose CVEs (i.e., VR shared offices), whose popularity is increasing due to the availability of low-cost headsets, and the restrictions imposed by the pandemic. In these CVEs collaboration roles are symmetrical, and the same 2D content available on desktop computers is displayed on 2D surfaces in a 3D space (VR screens). We prototyped one such CVE to evaluate FoV frustrum visualisation within this collaboration scenario. We also implement a FoV visualisation generated from an average fixation map (AFM), therefore directly generated by users' gaze behaviour which we call Cone of Vision (CoV). Our approach to displaying the frustum visualisations is tailored for 2D surfaces in 3D space and allows for self-awareness of this visual cue. We evaluate CoV in the context of a general exploratory data analysis (EDA) with 10 pairs of participants. Our findings indicate that CoV is beneficial during shifts between independent and collaborative work and supports collaborative progression across the visualisation. Self-perception of the CoV improves visual attention coupling, reduces the number of times users watch the collaborator's avatars and offers a consistent representation of the shared reality.
Riccardo Bovo, Daniele Giunchi, Muna Alebri, Anthony Steed, Enrico Costanza, Thomas Heinis
Proc. ACM Hum. Comput. Interact.2
2021 Mixing Modalities of 3D Sketching and Speech for Interactive Model Retrieval in Virtual Reality
abstract
Sketch and speech are intuitive interaction methods that convey complementary information and have been independently used for 3D model retrieval in virtual environments. While sketch has been shown to be an effective retrieval method, not all collections are easily navigable using this modality alone. We design a new challenging database for sketch comprised of 3D chairs where each of the components (arms, legs, seat, back) are independently colored. To overcome this, we implement a multimodal interface for querying 3D model databases within a virtual environment. We base the sketch on the state-of-the-art for 3D Sketch Retrieval, and use a Wizard-of-Oz style experiment to process the voice input. In this way, we avoid the complexities of natural language processing which frequently requires fine-tuning to be robust. We conduct two user studies and show that hybrid search strategies emerge from the combination of interactions, fostering the advantages provided by both modalities.
Daniele Giunchi, Alejandro Sztrajman, Stuart James, Anthony Steed
IMX1
2021 Perceived Realism of Pedestrian Crowds Trajectories in VR
abstract
Crowd simulation algorithms play an essential role in populating Virtual Reality (VR) environments with multiple autonomous humanoid agents. The generation of plausible trajectories can be a significant computational cost for real-time graphics engines, especially in untethered and mobile devices such as portable VR devices. Previous research explores the plausibility and realism of crowd simulations on desktop computers but fails to account the impact it has on immersion. This study explores how the realism of crowd trajectories affects the perceived immersion in VR. We do so by running a psychophysical experiment in which participants rate the realism of real/synthetic trajectories data, showing similar level of perceived realism.
Daniele Giunchi, Riccardo Bovo, Panayiotis Charalambous, Fotis Liarokapis, Alastair Shipman, Stuart James, Anthony Steed, Thomas Heinis
VRST1
2019 Selecting texture resolution using a task-specific visibility metric
abstract
Abstract In real‐time rendering, the appearance of scenes is greatly affected by the quality and resolution of the textures used for image synthesis. At the same time, the size of textures determines the performance and the memory requirements of rendering. As a result, finding the optimal texture resolution is critical, but also a non‐trivial task since the visibility of texture imperfections depends on underlying geometry, illumination, interactions between several texture maps, and viewing positions. Ideally, we would like to automate the task with a visibility metric, which could predict the optimal texture resolution. To maximize the performance of such a metric, it should be trained on a given task. This, however, requires sufficient user data which is often difficult to obtain. To address this problem, we develop a procedure for training an image visibility metric for a specific task while reducing the effort required to collect new data. The procedure involves generating a large dataset using an existing visibility metric followed by refining that dataset with the help of an efficient perceptual experiment. Then, such a refined dataset is used to retune the metric. This way, we augment sparse perceptual data to a large number of per‐pixel annotated visibility maps which serve as the training data for application‐specific visibility metrics. While our approach is general and can be potentially applied for different image distortions, we demonstrate an application in a game‐engine where we optimize the resolution of various textures, such as albedo and normal maps.
Krzysztof Wolski, Daniele Giunchi, Shinichi Kinuwaki, Piotr Didyk, Karol Myszkowski, Anthony Steed, Rafal Mantiuk
Comput. Graph. Forum2
2018 Model Retrieval by 3D Sketching in Immersive Virtual Reality
abstract
We describe a novel method for searching 3D model collections using free-form sketches within a virtual environment as queries. As opposed to traditional Sketch Retrieval, our queries are drawn directly onto an example model. Using immersive virtual reality the user can express their query through a sketch that demonstrates the desired structure, color and texture. Unlike previous sketch-based retrieval methods, users remain immersed within the environment without relying on textual queries or 2D projections which can disconnect the user from the environment. We show how a convolutional neural network (CNN) can create multi-view representations of colored 3D sketches. Using such a descriptor representation, our system is able to rapidly retrieve models and in this way, we provide the user with an interactive method of navigating large object datasets. Through a preliminary user study we demonstrate that by using our VR 3D model retrieval system, users can perform quick and intuitive search. Using our system users can rapidly populate a virtual environment with specific models from a very large database, and thus the technique has the potential to be broadly applicable in immersive editing systems.
Daniele Giunchi, Stuart James, Anthony Steed
VR1
2018 Dataset and Metrics for Predicting Local Visible Differences
abstract
A large number of imaging and computer graphics applications require localized information on the visibility of image distortions. Existing image quality metrics are not suitable for this task as they provide a single quality value per image. Existing visibility metrics produce visual difference maps, and are specifically designed for detecting just noticeable distortions but their predictions are often inaccurate. In this work, we argue that the key reason for this problem is the lack of large image collections with a good coverage of possible distortions that occur in different applications. To address the problem, we collect an extensive dataset of reference and distorted image pairs together with user markings indicating whether distortions are visible or not. We propose a statistical model that is designed for the meaningful interpretation of such data, which is affected by visual search and imprecision of manual marking. We use our dataset for training existing metrics and we demonstrate that their performance significantly improves. We show that our dataset with the proposed statistical model can be used to train a new CNN-based metric, which outperforms the existing solutions. We demonstrate the utility of such a metric in visually lossless JPEG compression, super-resolution and watermarking.
Krzysztof Wolski, Daniele Giunchi, Nanyang Ye 0001, Piotr Didyk, Karol Myszkowski, Radoslaw Mantiuk, Hans-Peter Seidel, Anthony Steed, Rafal Mantiuk
ACM Trans. Graph.2