EDBT 2026 Demo / reviewers in the wild / expert
Irene Viola 0001
dblp:199/0539-1
· DBLP profile ↗
43ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0001-8990-2665ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 11 first-author · 33 since 2021Human-computer interaction and ubiquitous computing · 13 · 4 first-author · 8 since 2021Computer networks · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WebRTC-Based Volumetric Video Conferencing: SFU Architecture Evaluation and BenchmarkingabstractImmersive technologies promise to revolutionize communication through enhanced sense of presence and interactivity. To enable interaction, reliable low-latency transport mechanisms are needed to handle the large volumes of data created by complex 3D objects. In this paper, we propose an open-source, codec-independent, selective forwarding unit (SFU) for real-time volumetric video streaming using WebRTC. For evaluation purposes, we provide a reference client implementation by extending VR2Gather, a TCP-based system for immersive communication. We conduct extensive evaluations using both new and existing datasets to compare the performance of WebRTC against TCP-based protocols in an emulated testbed environment. The evaluations demonstrate that WebRTC outperforms other protocols in high-latency scenarios and adapts video quality to user movement 13% and 36% faster than its TCP-based counterparts in networks with 5 ms and 10 ms of network latency, respectively. Matthias De Fré, Jeroen van der Hooft, Jack Jansen 0001, Silvia Rossi 0001, Thomas Röggla, Tim Wauters, Filip De Turck, Irene Viola 0001, Pablo César |
NOSSDAV | 8 |
| 2026 | Social XR for Pre-Production Meetings: An In-the-Wild Study of Early Stage Communication Between XR Producers and Clients
Sueyoon Lee, Irene Viola 0001, Jack Jansen 0001, Ashutosh Singla, Karolina Wylezek, Thomas Röggla, Pablo César |
IMX | 2 |
| 2026 | The Influence of Context on Learning in a Social VR Historical Fashion Exhibition
Karolina Wylezek, Irene Viola 0001, Silvia Rossi 0001, Jack Jansen 0001, Thomas Röggla, Pablo César |
IMX | 2 |
| 2025 | QoE Evaluation of Remote Physiotherapy in Volumetric Video and Video-Based Real-Time CommunicationabstractIn recent years, video conferencing platforms have become powerful tools for remote communication. There has also been an increase in the use of VR systems for communication. However, very few of these systems utilize photorealistic human representation. This paper investigates the strengths, challenges, and limitations of a novel 3D communication prototype (VR2Gather) and a well-established video conferencing system (Zoom). Specifically, we explore whether the 3D communication prototype can achieve comparable performance levels in a remote physiotherapy use case. By assessing various aspects, such as audio-visual quality, presence, and interaction, we aim to determine if the current prototype is comparable with commercial systems in some dimensions while exceeding expectations in others. Our results indicated that VR2Gather has the potential for a better sense of connection and higher concentration. However, challenges like improving 3D rendering quality and communication ease still need to be overcome to make it suitable for physiotherapy. Ashutosh Singla, Irene Viola 0001, Jack Jansen 0001, Pablo César |
ICME | 2 |
| 2025 | IXR '25: 3rd International Workshop on Interactive eXtended RealityabstractDespite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, telesurgery, interconnection in new buildings project chain, cultural heritage, and museum contents communication, are well on their way to integrating immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human-centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks, and end-user infrastructures to enable the next generation of interactive XR applications and services. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746269 Irene Viola 0001, Silvia Rossi 0001, Marta Orduna, Maria Torres Vega |
ACM Multimedia | 1 |
| 2025 | UVG-CWI-DQPC: Dual-Quality Point Cloud Dataset for Volumetric Video ApplicationsabstractVolumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP). Guillaume Gautier, Xuemei Zhou, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat, Pablo César, Jarno Vanne |
ACM Multimedia | 9 |
| 2025 | RCQoEA-360VR: Real-time Continuous QoE Scores for HMD-based 360° VR DatasetabstractAs immersive 360° video experiences through head-mounted displays (HMDs) gain widespread adoption, the need for real-time, fine-grained assessment of Quality of Experience (QoE) becomes increasingly critical for optimising user engagement and system performance. This paper introduces RCQoEA-360VR, a novel multi-modal dataset designed for continuous QoE evaluation in virtual reality (VR) environments. In a controlled study (N=32), participants watched five selected 360° video sequences across eight different video quality configurations (from the VQEG database) using a Vive Pro Eye while providing continuous QoE annotations via a touchpad-based input method, enhanced by the DotMorph peripheral visualisation technique. The dataset also includes synchronised physiological signals (electrocardiogram and galvanic skin response), behavioural data (eye and head movements) and post-viewing QoE ratings gathered through a within-VR interface. RCQoEA-360VR addresses a critical gap in existing public datasets by providing a fine-grained, synchronised multimodal data for immersive QoE analysis. It offers a unique and valuable resource for the research community, supporting a wide range of research applications, including QoE prediction, behavioural modelling, adaptive streaming, and implicit perceptual analysis. Sowmya Vijayakumar, Tong Xue, Abdallah El Ali, Irene Viola 0001, Ronan Flynn, Peter Corcoran 0001, Pablo César, Niall Murray |
ACM Multimedia | 4 |
| 2025 | Enhancing the Audience Experience for VR and AR Theatre with AI-generated SubtitlesabstractRecent technological developments on AI and immersive media are transforming the artistic landscape, providing novel mechanisms for artists and audiences. Following a human-centric approach, together with a theatre company in Greece, this paper investigates how subtitle placement affects user experience and cognitive load in a live theatre performance enhanced by AR glasses. To do so, we design and develop a system for displaying subtitles in VR and AR. We evaluated the system in two conditions (N = 19;N = 12), both in a controlled environment (VR) and an actual theatre (AR). In the latter, we integrate AI solutions to provide automatic captioning and translation in real time, and VFX to further augment the experience. Our quantitative and qualitative results showed no difference between subtitle placements in terms of cognitive load and user experience, with users equally liking the two proposed approaches. Results also highlighted the perceived usefulness of AR to enhance theatre performances, indicating new paths for wider accessibility and further immersion. Irene Viola 0001, Moonisa Ahsan, Olga Chatzifoti, Atanas Yonkov, Eleni Oikonomou, Ioannis Radin, Pawel Maka, Abderrahmane Issam, Pablo César |
VRST | 1 |
| 2025 | VRD: A multi-lingual translation and Ai-Assisted Navigation experience for VR Conference ApplicationabstractThis paper presents an AI-assisted VR conference application with multilingual translation and navigation agent capabilities. A pilot study with 18 participants (11 females, 7 males) was conducted to assess the system’s usability. AI-assisted navigation worked smoothly, but the AI translation had issues that prevented the users from having a good experience, nonetheless, participants expressed positive attitudes toward the system, and future work will focus on achieving better user experience. Moonisa Ahsan, Irene Viola 0001, Manuel Toledo, Dimitris Kontopoulos, Pablo César |
VRST | 2 |
| 2025 | PointPCA+: A full-reference Point Cloud Quality Assessment metric with PCA-based features
Xuemei Zhou, Evangelos Alexiou, Irene Viola 0001, Pablo César |
Signal Process. Image Commun. | 3 |
| 2025 | A Clustering Approach to Unveil User Similarities in 6 df Extended Reality ApplicationsabstractThe advent in our daily life of Extended Reality (XR) technologies, such as Virtual and Augmented Reality, has led to the rise of user-centric systems, offering higher level of interaction and presence in virtual environments. In this context, understanding the actual interactivity of users is still an open challenge and a key step to enabling user-centric system. In this work, our goal is to construct an efficient clustering tool for 6 df navigation trajectories by extending the applicability of existing behavioural tool. Specifically, we first compare the navigation in 6 df with its 3 df counterpart, highlighting the main differences and novelties. Then, we investigate new metrics aimed at better modelling behavioural similarities between users in a 6 df system. More concretely, we define and compare 11 similarity metrics which are based on different distance features (i.e., user positions in the 3D space, user viewing directions) and distance measurements (i.e., Euclidean, Geodesic, angular distance). Our solutions are validated and tested on real navigation paths of users interacting with dynamic volumetric media in both 6 df Virtual Reality and Augmented Reality conditions. Results show that metrics based on both user position and viewing direction better perform in detecting user similarity while navigating in a 6 df system. Such easy-to-use but robust metrics allow us to answer a fundamental question for user-centric systems: ‘How do we detect if users look at the same content in 6 df?’, opening the gate to new solutions based on users interactivity, such as viewport prediction, live streaming services optimised based on users behaviour but also for user-based quality assessment methods. Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Subjective and Objective Quality Assessment for Dynamic Point Cloud with Visual Attention in 6 DoFabstractPerceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user. Understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data are scarcely available. Besides, testing methods and procedures for collecting visual attention data are still to be agreed on. This article presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. Both the quality score and eye-tracking data were collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 Degrees of Freedom (DoF) inspection, using a head-mounted display. Qualitative interview analysis was also conducted after the experiment. The dataset consists of 50 DPCs, including 5 reference DPCs, with each reference encoded at 3 distortion levels using 3 different codecs (namely G-PCC, V-PCC, CWI-PCL), amounting to a total of 9 degraded version per reference. Additionally, it incorporates 1,000 gaze trials from 40 participants, yielding a total of 15,000 visual attention maps across all the DPCs. We additionally benchmark objective quality metrics originally designed for static point clouds, evaluating their performance in our dataset using two temporal pooling strategies. Furthermore, we employ the visual attention data that are retrieved during our experiment to evaluate whether the performance of widely used objective quality metrics is improved by considering subjective measurements of visual attention. This dataset establishes a link between quality assessment and visual attention within the context of DPC. Moreover, thematic analysis of the interviews helps uncover user behavior and factors impacting perceptual quality for DPC in 6 DoF. This work deepens our understanding of DPC quality assessment and visual attention, driving progress in the realm of VR experiences and perception. Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependentabstractThis paper presents a Task-Free eye-tracking dataset for Dynamic Point Clouds (TF-DPC) aimed at investigating visual attention. The dataset is composed of eye gaze and head movements collected from 24 participants observing 19 scanned dynamic point clouds in a Virtual Reality (VR) environment with 6 degrees of freedom. We compare the visual saliency maps generated from this dataset with those from a prior task-dependent experiment (focused on quality assessment) to explore how high-level tasks influence human visual attention. To measure the similarity between these visual saliency maps, we apply the well-known Pearson correlation coefficient and an adapted version of the Earth Mover's Distance metric, which takes into account both spatial information and the degrees of saliency. Our experimental results provide both qualitative and quantitative insights, revealing significant differences in visual attention due to task influence. This work enhances our understanding of the visual attention for dynamic point cloud (specifically human figures) in VR from gaze and human movement trajectories, and highlights the impact of task-dependent factors, offering valuable guidance for advancing visual saliency models and improving VR perception. Xuemei Zhou, Irene Viola 0001, Silvia Rossi 0001, Pablo César |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Open-Sourcing VR2Gather: A Collaborative Social VR System for Adaptive Multi-Party Real Time CommunicationabstractSocial Virtual Reality is envisioned to transform how individu- als communicate remotely, offering a sense of immersion and co- presence within a virtual space. Current platforms enabling remote social interactions rely on synthetic user representations. We ad- dress this limitation by enabling realistic human representation through volumetric content capture, encoding and transmission. Specifically, we present an extended version of VR2Gather, now a fully open source Unity package, available at https://github.com/ cwi-dis/VR2Gather-acmmm-oss. Our platform is a customisable system to transmit volumetric content in a multi-party real-time environment, easy to integrate into existing applications. Jack Jansen 0001, Thomas Röggla, Silvia Rossi 0001, Irene Viola 0001, Pablo César |
ACM Multimedia | 4 |
| 2024 | Deciphering Perceptual Quality in Colored Point Cloud: Prioritizing Geometry or Texture Distortion?abstractPoint clouds represent one of the prevalent formats for 3D content. Distortions introduced at various stages in the point cloud processing pipeline affect the visual quality, altering their geometric composition, texture information, or both. Understanding and quantifying the impact of the distortion domain on visual quality is vital to driving rate optimization and guiding post-processing steps to improve the quality of experience. In this paper, we propose a multi-task guided multi-modality no reference metric (M3-Unity), which utilizes 4 types of modalities across attributes and dimensionalities to represent point clouds. An attention mechanism establishes inter/intra associations among 3D/2D patches, which can complement each other, yielding local and global features, to fit the highly nonlinear property of the human vision system. A multi-task decoder involving distortion type classification selects the best association among 4 modalities, aiding the regression task and enabling the in-depth analysis of the interplay between geometrical and textural distortions. Furthermore, our framework design and attention strategy enable us to measure the impact of individual attributes and their combinations, providing insights into how these associations contribute particularly in relation to distortion type. Extensive experimental results on 4 datasets consistently outperform the state-of-the-art metrics by a large margin. The code is available at https://github.com/cwi-dis/ACMMM2024-Oral. Xuemei Zhou, Irene Viola 0001, Yunlu Chen, Jiahuan Pei, Pablo César |
ACM Multimedia | 2 |
| 2024 | ComPEQ-MR: Compressed Point Cloud Dataset with Eye Tracking and Quality Assessment in Mixed RealityabstractPoint clouds (PCs) have attracted researchers and developers due to their ability to provide immersive experiences with six degrees of freedom (6DoF). However, there are still several open issues in understanding the Quality of Experience (QoE) and visual attention of end users while experiencing 6DoF volumetric videos. First, encoding and decoding point clouds require a significant amount of both time and computational resources. Second, QoE prediction models for dynamic point clouds in 6DoF have not yet been developed due to the lack of visual quality databases. Third, visual attention in 6DoF is hardly explored, which impedes research into more sophisticated approaches for adaptive streaming of dynamic point clouds. In this work, we provide an open-source Compressed Point cloud dataset with Eye-tracking and Quality assessment in Mixed Reality (ComPEQ--MR). The dataset comprises four compressed dynamic point clouds processed by Moving Picture Experts Group (MPEG) reference tools (i.e., VPCC and GPCC), each with 12 distortion levels. We also conducted subjective tests to assess the quality of the compressed point clouds with different levels of distortion. The rating scores are attached to ComPEQ--MR so that they can be used to develop QoE prediction models in the context of MR environments. Additionally, eye-tracking data for visual saliency is included in this dataset, which is necessary to predict where people look when watching 3D videos in MR experiences. We collected opinion scores and eye-tracking data from 41 participants, resulting in 2132 responses and 164 visual attention maps in total. The dataset is available at https://ftp.itec.aau.at/datasets/ComPEQ-MR/. Minh Nguyen 0006, Shivi Vats, Xuemei Zhou, Irene Viola 0001, Pablo César, Christian Timmerer, Hermann Hellwagner |
MMSys | 4 |
| 2024 | Enhancing Immersive Experiences through 3D Point Cloud Analysis: A Novel Framework for Applying 2D Visual Saliency Models to 3D Point CloudsabstractIn the new area of immersive multimedia environments, understanding and manipulating visual attention are crucial for enhancing user experience. This study introduces an innovative framework that extends traditional 2D saliency maps to the analysis of 3D point clouds, a step forward in adapting saliency prediction to more complex and immersive environments. Our framework centers on the orthographic projection of 3D point clouds onto 2D planes, enabling the application of established 2D saliency models to this novel context. We further delve into the evaluation of these models on a 3D point cloud eye-tracking dataset, exploring various projection settings and thresholding techniques to maintain the integrity of saliency information in the transition from 2D to 3D. This research not only bridges a gap in applying visual attention models to 3D data but also offers insights into the optimization of quality of experience in immersive multimedia systems. Marouane Tliba, Xuemei Zhou, Irene Viola 0001, Pablo César, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
QoMEX | 3 |
| 2024 | Communication Challenges between Clients and Producers of Immersive Media Applications: can Social XR help?abstractExtended Reality (XR) has emerged as a transformative and immersive technology with versatile applications in content creation and consumption. As XR gains popularity, companies eager to adopt it often possess a surface-level understanding, investing significant resources without effectively addressing the genuine needs of end-users. This study explores the current workflows of XR production companies, and the potential of social XR in mitigating challenges throughout the XR production workflow. We present the outcomes of three respective focus group workshops conducted with three XR production companies and their experts (N=17). The results indicate that at every stage of the production, namely pre-production, production, post-production, and post-release, there are communication challenges between producers and clients, as well as different production and post-production specialists. We discuss various aspects of XR concerning the problem and propose novel opportunities offered by social XR to ameliorate those challenges, improving communication and making development more agile. Sueyoon Lee, Irene Viola 0001, Ashutosh Singla, Pablo César |
IMX | 2 |
| 2024 | Delay Threshold for Social Interaction in Volumetric eXtended Reality CommunicationabstractImmersive technologies like eXtended Reality (XR) are the next step in videoconferencing. In this context, understanding the effect of delay on communication is crucial. This article presents the first study on the impact of delay on collaborative tasks using a realistic Social XR system. Specifically, we design an experiment and evaluate the impact of end-to-end delays of 300, 600, 900, 1,200, and 1,500 ms on the execution of a standardized task involving the collaboration of two remote users that meet in a virtual space and construct block-based shapes. To measure the impact of the delay in this communication scenario, objective and subjective data were collected. As objective data, we measured the time required to execute the tasks and computed conversational characteristics by analyzing the recorded audio signals. As subjective data, a questionnaire was prepared and completed by every user to evaluate different factors such as overall quality, perception of delay, annoyance using the system, level of presence, cybersickness, and other subjective factors associated with social interaction. The results show a clear influence of the delay on the perceived quality and a significant negative effect as the delay increases. Specifically, the results indicate that the acceptable threshold for end-to-end delay should not exceed 900 ms. This article additionally provides guidelines for developing standardized XR tasks for assessing interaction in Social XR environments. Carlos Cortés 0001, Irene Viola 0001, Jesús Gutiérrez 0001, Jack Jansen 0001, Shishir Subramanyam, Evangelos Alexiou, Pablo Pérez 0001, Narciso García, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Designing and Evaluating a VR Lobby for a Socially Enriching Remote Opera Watching ExperienceabstractThe latest social VR technologies have enabled users to attend traditional media and arts performances together while being geographically removed, making such experiences accessible despite budget, distance, and other restrictions. In this work, we aim at improving the way remote performances are shared by designing and evaluating a VR theatre lobby which serves as a space for users to gather, interact, and relive the common experience of watching a virtual opera. We conducted an initial test with experts ($\mathrm{N}=10$, i.e., designers and opera enthusiasts) in pairs using our VR lobby prototype, developed based on the theoretical lobby design concept. A unique aspect of our experience is its highly realistic representation of users in the virtual space. The test results guided refinements to the VR lobby structure and implementation, aiming to improve the user experience and align it more closely with the social VR lobby's intended purpose. With the enhanced prototype, we ran a between-subject controlled study ($\mathrm{N}=40$) to compare the user experience in the social VR lobby between individuals and paired participants. To do so, we designed and validated a questionnaire to measure the user experience in the VR lobby. Results of our mixed-methods analysis, including interviews, questionnaire results, and user behavior, reveal the strength of our social VR lobby in connecting with other users, consuming the opera in a deeper manner, and exploring new possibilities beyond what is common in real life. All supplemental materials are available at https://github.com/cwi-dis/IEEEVR2024-VRLobby. Sueyoon Lee, Irene Viola 0001, Silvia Rossi 0001, Zhirui Guo, Ignacio Reimat, Kinga Lawicka, Alina Striner, Pablo César |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | QAVA-DPC: Eye-Tracking Based Quality Assessment and Visual Attention Dataset for Dynamic Point Cloud in 6 DoFabstractPerceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user, understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data is scarcely available. This paper presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. The data was collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 degrees-of-freedom inspection, using a head-mounted display. The dataset comprises 5 reference DPC contents, with each reference encoded at 3 distortion levels using 3 different codecs, amounting to a total of 9 degraded DPC contents. Moreover, it includes 1,000 gaze trials from 40 participants, resulting in 15,000 visual attention maps in total. The curated dataset can serve as authentic benchmark data for assessing the performance of objective DPC quality metrics. Additionally, it establishes a link between quality assessment and visual attention within the context of DPC. This work deepens our understanding of DPC quality and visual attention, driving progress in the realm of VR experiences and perception. Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César |
ISMAR | 2 |
| 2023 | IXR '23: 2nd International Workshop on Interactive eXtended RealityabstractDespite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, telesurgery, interconnection in new buildings project chain, Cultural Heritage, and Museum contents communication, are well on their way to integrating immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human-centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks, and end-user infrastructures to enable the next generation of interactive XR applications and services. Irene Viola 0001, Hadi Amirpour, Stephanie Arevalo, Maria Torres Vega |
ACM Multimedia | 1 |
| 2023 | On the Impact of Interactive eXtended Reality: Challenges and Opportunities for Multimedia ResearchabstractExtended Reality (XR) has been hailed as the new frontier of media, ushering new possibilities for societal areas such as communications, training, entertainment, gaming, and cultural heritage. However, despite the remarkable technical advances, current XR applications are in their majority local and individual experiences. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely content realism, by means of Artificial Intelligence (AI) techniques, motion-to-photon latency, and accurate human-centric driven experiences able to map real and virtual worlds seamlessly. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. In this panel, together with the leading experts of the SIGMM community, we will explore the challenges and opportunities to unlock the next generation of interactive XR applications and services. Irene Viola 0001, Maria Torres Vega |
ACM Multimedia | 1 |
| 2023 | Extending 3-DoF Metrics to Model User Behaviour Similarity in 6-DoF Immersive ApplicationsabstractImmersive reality technologies, such as Virtual and Augmented Reality, have ushered a new era of user-centric systems, in which every aspect of the coding-delivery-rendering chain is tailored to the interaction of the users. Understanding the actual interactivity and behaviour of the users is still an open challenge and a key step to enabling such a user-centric system. Our main goal is to extend the applicability of existing behavioural methodologies for studying user navigation in the case of 6 Degree-of-Freedom (DoF). Specifically, we first compare the navigation in 6-DoF with its 3-DoF counterpart highlighting the main differences and novelties. Then, we define new metrics aimed at better modelling behavioural similarities between users in a 6-DoF system. We validate and test our solutions on real navigation paths of users interacting with dynamic volumetric media in 6-DoF Virtual Reality conditions. Our results show that metrics that consider both user position and viewing direction better perform in detecting user similarity while navigating in a 6-DoF system. Having easy-to-use but robust metrics that underpin multiple tools and answer the question "how do we detect if two users look at the same content?" open the gate to new solutions for a user-centric system. Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César |
MMSys | 2 |
| 2023 | Evaluation of point cloud features for no-reference visual quality assessmentabstractThe development and widespread adoption of immersive XR applications has led to a renewed interest in representations that are capable of reproducing real-world objects and scenes with high fidelity. Among such representations, point clouds have attracted the interest of industry and academia alike, and new compression solutions have been developed to facilitate their adoption in mainstream applications. To ensure the best quality of experience for the end-user in limited bandwidth scenarios, new full-reference objective quality metrics have been proposed, promoting features designed specifically for point cloud contents. However, the performance of such features to predict the quality of point cloud contents when the reference is not available is largely unexplored. In this paper, we evaluate the performance of features commonly used to model point cloud distortions in a no-reference framework. The obtained features are integrated into a quality value through a support vector regression model. Results demonstrate the potential of full-reference features for no-reference assessment. Gwennan Smitskamp, Irene Viola 0001, Pablo César |
QoMEX | 2 |
| 2022 | Impact of Self-View Latency on Quality of Experience: Analysis of Natural Interaction in XR EnvironmentsabstractThe rise of eXtended Reality (XR) has led to multiple ways of including the user’s body in interactive experiences. However, the delay limits of self-view rendering in interactive XR remain unexplored. This article presents a minimum self-view latency system and an interactive task-based experiment to study the influence of different levels of self-view delay on Quality of Experience (QoE) and task performance. During the experiment, 23 users tested 8 delay conditions (from 190 to 597 ms) while building block-based models. The results show a hard threshold in terms of involvement and overall quality around 450ms. However, the impact on adaptation and execution time was less pronounced. This indicates that although users adapted to the task in a certain way, their immersion was severely affected above a certain self-view delay value. Carlos Cortés 0001, Jesús Gutiérrez 0001, Pablo Pérez 0001, Irene Viola 0001, Pablo César, Narciso García |
ICIP | 4 |
| 2022 | Mediascape XR: A Cultural Heritage Experience in Social VRabstractSocial virtual reality (VR) allows multiple remote users to interact in a shared space, unveiling new possibilities for communication in immersive environments. Mediascape XR presents a social VR experience that teleports 3D representations of remote users, using volumetric video, to a virtual museum. It enables visitors to interact with cultural heritage artifacts while allowing social interactions in real time between them. The application is designed following a human-centered approach, enabling an interactive, educating, and entertaining experience. Ignacio Reimat, Yanni Mei, Evangelos Alexiou, Jack Jansen 0001, Jie Li 0064, Shishir Subramanyam, Irene Viola 0001, Johan Oomen, Pablo César |
ACM Multimedia | 7 |
| 2022 | Evaluating the Impact of Tiled User-Adaptive Real-Time Point Cloud Streaming on VR Remote CommunicationabstractRemote communication has rapidly become a part of everyday life in both professional and personal contexts. However, popular video conferencing applications present limitations in terms of quality of communication, immersion and social meaning. VR remote communication applications offer a greater sense of co-presence and mutual sensing of emotions between remote users. Previous research on these applications has shown that realistic point cloud user reconstructions offer better immersion and communication as compared to synthetic user avatars. However, photorealistic point clouds require a large volume of data per frame and are challenging to transmit over bandwidth-limited networks. Recent research has demonstrated significant improvements to perceived quality by optimizing the usage of bandwidth based on the position and orientation of the user's viewport with user-adaptive streaming. In this work, we developed a real-time VR communication application with an adaptation engine that features tiled user-adaptive streaming based on user behaviour. The application also supports traditional network adaptive streaming. The contribution of this work is to evaluate the impact of tiled user-adaptive streaming on quality of communication, visual quality, system performance and task completion in a functional live VR remote communication system. We performed a subjective evaluation with 33 users to compare the different streaming conditions with a neck exercise training task. As a baseline, we use uncompressed streaming requiring approximately 300 megabits per second and our solution achieves similar visual quality with tiled adaptive streaming at 14 megabits per second. We also demonstrate statistically significant gains in the quality of interaction and improvements to system performance and CPU consumption with tiled adaptive streaming as compared to the more traditional network adaptive streaming. Shishir Subramanyam, Irene Viola 0001, Jack Jansen 0001, Evangelos Alexiou, Alan Hanjalic, Pablo César |
ACM Multimedia | 2 |
| 2022 | IXR '22: 1st Workshop on Interactive eXtended RealityabstractDespite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, tele-surgery, interconnection in new buildings project chain, Cultural Heritage and Museum contents communication, are well on their way to integrate immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely the (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks and end-user infrastructures to enable the next generation of interactive XR applications and services. The complete IXR'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552483 Irene Viola 0001, Hadi Amirpour, Maria Torres Vega |
ACM Multimedia | 1 |
| 2022 | Subjective QoE Evaluation of User-Centered Adaptive Streaming of Dynamic Point CloudsabstractTechnological advances in head-mounted displays and novel real-time 3D acquisition and reconstruction solutions have fostered the development of 6 Degrees of Freedom (6DoF) teleimmersive systems for social VR applications. Point clouds have emerged as a popular format for such applications, owing to their simplicity and versatility; yet, dense point cloud contents are too large to deliver directly over bandwidth-limited networks. In this context, user-adaptive delivery mechanisms are a promising solution to exploit the increased range of motion offered by 6DoF VR applications to yield gains in perceived quality of 3D point cloud user representations, while reducing their bandwidth requirements. In this paper, we perform a user study in VR to quantify the gains adaptive tile selection strategies can bring with respect to non-adaptive solutions. In particular, we define an auxiliary utility function, we employ established methods from the literature and newly-proposed schemes for distributing the bit budget across the tiles, and we evaluate them together with non-adaptive streaming baselines through subjective QoE assessment. Results confirm that considerable gains can be obtained with user-adaptive streaming, achieving bit rate gains of up to 65% with respect to a non-adaptive approach to deliver comparable quality. Our analysis provides useful insights for the design and development of social VR applications. Shishir Subramanyam, Irene Viola 0001, Jack Jansen 0001, Evangelos Alexiou, Alan Hanjalic, Pablo César |
QoMEX | 2 |
| 2022 | Designing Real-time, Continuous QoE Score Acquisition Techniques for HMD-based 360°VR Video WatchingabstractWatching HMD-based 360° video has become in-creasing popular as a medium for immersive viewing of photo-realistic content. To evaluate subjective video quality, researchers typically prompt users to provide an overall Quality of Experience (QoE) score after viewing a stimulus. However, since users can adjust their viewport throughout a 360° video, a higher level of spatiotemporal granularity is needed for adaptive 360° video streaming. To address this, we design several real-time, continuous QoE annotation input and peripheral visualization techniques, with the goal of minimizing mental workload and distraction during score acquisition. Drawing on two parallel co-design sessions with seven experts, we find that touchpad and joystick are most suitable for continuous input, with DotMorph (circle with tick label that varies in filling) for peripheral state feedback. We contribute design findings for testing QoE score acquisition techniques during HMD-based 360° video watching, which enable more precise optimization of adaptive video streaming quality. Tong Xue, Abdallah El Ali, Irene Viola 0001, Pablo César |
QoMEX | 3 |
| 2022 | Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919abstractRecently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community. Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García |
IEEE Trans. Multim. | 7 |
| 2021 | A New Challenge: Behavioural Analysis Of 6-DOF User When Consuming Immersive MediaabstractThanks to recent advances in computer graphics, wearable technology, and connectivity, Virtual Reality (VR) has landed in our daily life. A key novelty in VR is the role of the user, which has turned from merely passive to entirely active. Thus, improving any aspect of the coding-delivery-rendering chain starts with the need for understanding user behaviour. To do so, we investigate the navigation trajectories of users within a 6-Degrees-of-Freedom (DoF) VR environment. Specifically, we investigate the main differences and similarities between 3 and 6-DoF navigation through existing methodologies adopted to study user behaviour in 3-DoF settings. Our simulation results, based on real navigation paths of users while displaying dynamic volumetric media in 6-DoF conditions, show the limitations of clustering algorithms for 3-DoF in assessing user similarity in 6-DoF. Given these observations, we state the need for developing new solutions for the analysis of 6-DoF trajectories. Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César |
ICIP | 2 |
| 2021 | CWIPC-SXR: Point Cloud dynamic human dataset for Social XRabstractReal-time, immersive telecommunication systems are quickly becoming a reality, thanks to the advances in acquisition, transmission, and rendering technologies. Point clouds in particular serve as a promising representation in these type of systems, offering photorealistic rendering capabilities with low complexity. Further development of transmission, coding, and quality evaluation algorithms, though, is currently hindered by the lack of publicly available datasets that represent realistic scenarios of remote communication between people in real-time. In this paper, we release a dynamic point cloud dataset that depicts humans interacting in social XR settings. Using commodity hardware, we capture a total of 45 unique sequences, according to several use cases for social XR. As part of our release, we provide annotated raw material, resulting point cloud sequences, and an auxiliary software toolbox to acquire, process, encode, and visualize data, suitable for real-time applications. The dataset can be accessed via the following link: https://www.dis.cwi.nl/cwipc-sxr-dataset/. Ignacio Reimat, Evangelos Alexiou, Jack Jansen 0001, Irene Viola 0001, Shishir Subramanyam, Pablo César |
MMSys | 4 |
| 2020 | User Centered Adaptive Streaming of Dynamic Point Clouds with Low Complexity TilingabstractIn recent years, the development of devices for acquisition and rendering of 3D contents have facilitated the diffusion of immersive virtual reality experiences. In particular, the point cloud representation has emerged as a popular format for volumetric photorealistic reconstructions of dynamic real world objects, due to its simplicity and versatility. To optimize the delivery of the large amount of data needed to provide these experiences, adaptive streaming over HTTP is a promising solution. In order to ensure the best quality of experience within the bandwidth constraints, adaptive streaming is combined with tiling to optimize the quality of what is being visualized by the user at a given moment; as such, it has been successfully used in the past for omnidirectional contents. However, its adoption to the point cloud streaming scenario has only been studied to optimize multi-object delivery. In this paper, we present a low-complexity tiling approach to perform adaptive streaming of point cloud content. Tiles are defined by segmenting each point cloud object in several parts, which are then independently encoded. In order to evaluate the approach, we first collect real navigation paths, obtained through a user study in 6 degrees of freedom with 26 participants. The variation in movements and interaction behaviour among users indicate that a user-centered adaptive delivery could lead to sensible gains in terms of perceived quality. Evaluation of the performance of the proposed tiling approach against state of the art solutions for point cloud compression, performed on the collected navigation paths, confirms that considerable gains can be obtained by exploiting user-adaptive streaming, achieving bitrate gains up to 57% with respect to a non-adaptive approach with the same codec. Moreover, we demonstrate that the selection of navigation data has an impact on the relative objective scores. Shishir Subramanyam, Irene Viola 0001, Alan Hanjalic, Pablo César |
ACM Multimedia | 2 |
| 2020 | A Color-Based Objective Quality Metric for Point Cloud ContentsabstractIn recent years, point clouds have gained popularity as a promising representation for volumetric contents in immersive scenarios. Standardization bodies such as MPEG have been developing new compression standards for point cloud contents to reduce the volume of data, while maintaining an acceptable level of visual quality. To do so, reliable metrics are needed in order to automatically estimate the perceptual quality of degraded point cloud contents. Whereas several objective metrics have been developed to assess the geometrical impairment of degraded point cloud contents, fewer publications have been devoted to evaluating color artifacts. In this paper, we propose new color-based objective metrics for quality evaluation of point cloud contents. Our work extracts color statistics from both reference and degraded point cloud contents, in order to assess the level of impairment. Using publicly available ground-truth data, we compare the performance of our proposed work with state-of-the-art metrics, and we demonstrate how the color metrics are able to achieve comparable results with respect to widely adopted solutions. Moreover, we combine color- and geometry-based metrics in order to provide a global quality score. The novelty of our works resides in simultaneously taking both degradation types into account, while being independent of the rendering process. Results show that our solution is able to overcome the limitations of focusing on only one type of degradation, achieving better performance with respect to current metrics. Irene Viola 0001, Shishir Subramanyam, Pablo César |
QoMEX | 1 |
| 2020 | Comparing the Quality of Highly Realistic Digital Humans in 3DoF and 6DoF: A Volumetric Video Case StudyabstractVirtual Reality (VR) and Augmented Reality (AR) applications have seen a drastic increase in commercial popularity. Different representations have been used to create 3D reconstructions for AR and VR. Point clouds are one such representation characterized by their simplicity and versatility, making them suitable for real time applications, such as reconstructing humans for social virtual reality. In this study, we evaluate how the visual quality of digital humans, represented using point clouds, is affected by compression distortions. We compare the performance of the upcoming point cloud compression standard against an octree-based anchor codec. Two different VR viewing conditions enabling 3- and 6 degrees of freedom are tested, to understand how interacting in the virtual space affects the perception of quality. To the best of our knowledge, this is the first work performing user quality evaluation of dynamic point clouds in VR; in addition, contributions of the paper include quantitative data and empirical findings. Results highlight how perceived visual quality is affected by the tested content, and how current data sets might not be sufficient to comprehensively evaluate compression solutions. Moreover, shortcomings in how point cloud encoding solutions handle visually-lossless compression are discussed. Shishir Subramanyam, Jie Li 0064, Irene Viola 0001, Pablo César |
VR | 3 |
| 2020 | A Reduced Reference Metric for Visual Quality Evaluation of Point Cloud ContentsabstractPoint cloud representation has seen a surge of popularity in recent years, thanks to its capability to reproduce volumetric scenes in immersive scenarios. New compression solutions for streaming of point cloud contents have been proposed, which require objective quality metrics to reliably assess the level of degradation introduced by coding and transmission distortions. In this context, reduced reference metrics aim to predict the visual quality of the transmitted contents, while requiring only a small set of features to be sent in addition to the streamed media. In this paper, we propose a reduced reference metric to predict the quality of point cloud contents under compression distortions. To do so, we extract a small set of statistical features from the reference point cloud in the geometry, color and normal vector domain, which can be used at the receiver side to assess the visual degradation of the content. Using publicly available ground-truth datasets, we compare the performance of our metric to widely-used full reference metrics. Results demonstrate that our metric is able to effectively predict the level of distortion in the degraded point cloud contents, achieving high correlation values with respect to subjective scores. Irene Viola 0001, Pablo César |
IEEE Signal Process. Lett. | 1 |
| 2019 | Light field compression using translation-assisted view estimationabstractLight field technology has recently been gaining traction in the research community. Several acquisition technologies have been demonstrated to properly capture light field information, and portable devices have been commercialized to the general public. However, new and efficient compression algorithms are needed to sensibly reduce the amount of data that needs to be stored and transmitted, while maintaining an adequate level of perceptual quality. In this paper, we propose a novel light field compression scheme that uses view estimation to recover the entire light field from a small subset of encoded views. Experimental results on a widely used light field dataset show that our method achieves good coding efficiency with average rate savings of 54.83% with respect to HEVC. Baptiste Hériard-Dubreuil, Irene Viola 0001, Touradj Ebrahimi |
PCS | 2 |
| 2019 | An in-depth analysis of single-image subjective quality assessment of light field contentsabstractQuality assessment of light field images poses new questions and challenges, due to the enriched nature of the content and the possibilities it offers at the rendering step. Image-based rendering is conventionally used to showcase the increased capabilities of light field contents on traditional 2D screens. However, the range of possibilities for rendering parameters is virtually endless, which poses the problem of what rendered images should be used when performing visual quality assessments, as well as how to properly present them to subjects during quality evaluations. Single-image assessment has been used in the past to conduct subjective quality evaluations. Since this type of assessment generates a large number of stimuli to be evaluated, which increases the complexity, length, and cost of the test, it is fundamental to analyze whether the added strain on the evaluation procedure is compensated by statistically relevant results. In this paper, we analyze the results of a subjective evaluation campaign that used single-image assessment by means of statistical tools, to understand whether the advantages of evaluating light field contents through separately rendered images counterbalance the increase in complexity. In particular, we test whether different types of rendering lead to statistically different ratings, and if testing a variety of rendering parameters through single-image assessment is advisable. Results provide useful guidelines to designs more efficient subjective quality assessment for light field contents. Irene Viola 0001, Touradj Ebrahimi |
QoMEX | 1 |
| 2018 | VALID: Visual quality Assessment for Light field Images DatasetabstractIn the last years, light field imaging has experienced a surge of popularity among the scientific community for its capability of rendering the 3D world in a more immersive way. In particular, several compression algorithms have been proposed to efficiently reduce the amount of data generated in the acquisition process, and different methodologies have been designed to reliably evaluate the visual quality of compressed contents. In this paper we propose a dataset for visual quality assessment of light field images (VALID). The dataset contains five contents compressed at various bitrates, using both off-the-shelf solutions and state-of-the-art algorithms. Results of objective quality evaluation using popular image metrics are included, as well as annotated subjective scores using three different methodologies and two types of visualization setups. The proposed dataset will help develop new objective metrics to predict visual quality, design new subjective assessment methodologies and compare them to existing ones, as well as produce novel analysis approaches to interpret the results. Irene Viola 0001, Touradj Ebrahimi |
QoMEX | 1 |
| 2017 | Impact of interactivity on the assessment of quality of experience for light field contentabstractThe recent advances in light field imaging are changing the way in which visual content is captured, processed and consumed. Storage and delivery systems for light field images rely on efficient compression algorithms. Such algorithms must additionally take into account the feature-rich rendering for light field content. Therefore, a proper evaluation of visual quality is essential to design and improve coding solutions for light field content. Consequently, the design of subjective tests should also reflect the light field rendering process. This paper aims at presenting and comparing two methodologies to assess the quality of experience in light field imaging. The first methodology uses an interactive approach, allowing subjects to engage with the light field content when assessing it. The second, on the other hand, is completely passive to ensure all the subjects will have the same experience. Advantages and drawbacks of each approach are compared by relying on statistical analysis of results and conclusions are drawn. The obtained results provide useful insights for future design of evaluation techniques for light field content. Irene Viola 0001, Martin Rerábek, Touradj Ebrahimi |
QoMEX | 1 |
| 2016 | Objective and subjective evaluation of light field image compression algorithmsabstractThis paper reports results of subjective and objective quality assessments of responses to a grand challenge on light field image compression. The goal of the challenge was to collect and evaluate new compression algorithms for light field images. In total seven proposals were received, out of which five were accepted for further evaluations. For objective evaluations, conventional metrics were used, whereas the double stimulus continuous quality scale method was selected to perform subjective assessments. Results show competitive performance among submitted proposals. However, in low bitrates, one proposal outperforms the others. Irene Viola 0001, Martin Rerábek, Tim Bruylants, Peter Schelkens, Fernando Pereira 0001, Touradj Ebrahimi |
PCS | 1 |