Silvia Rossi 0001

dblp:19/1326-1 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-2779-2314ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 12 since 2021Computer networks · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 WebRTC-Based Volumetric Video Conferencing: SFU Architecture Evaluation and Benchmarking
abstract
Immersive technologies promise to revolutionize communication through enhanced sense of presence and interactivity. To enable interaction, reliable low-latency transport mechanisms are needed to handle the large volumes of data created by complex 3D objects. In this paper, we propose an open-source, codec-independent, selective forwarding unit (SFU) for real-time volumetric video streaming using WebRTC. For evaluation purposes, we provide a reference client implementation by extending VR2Gather, a TCP-based system for immersive communication. We conduct extensive evaluations using both new and existing datasets to compare the performance of WebRTC against TCP-based protocols in an emulated testbed environment. The evaluations demonstrate that WebRTC outperforms other protocols in high-latency scenarios and adapts video quality to user movement 13% and 36% faster than its TCP-based counterparts in networks with 5 ms and 10 ms of network latency, respectively.
Matthias De Fré, Jeroen van der Hooft, Jack Jansen 0001, Silvia Rossi 0001, Thomas Röggla, Tim Wauters, Filip De Turck, Irene Viola 0001, Pablo César
NOSSDAV4
2026 The Influence of Context on Learning in a Social VR Historical Fashion Exhibition
Karolina Wylezek, Irene Viola 0001, Silvia Rossi 0001, Jack Jansen 0001, Thomas Röggla, Pablo César
IMX3
2025 IXR '25: 3rd International Workshop on Interactive eXtended Reality
abstract
Despite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, telesurgery, interconnection in new buildings project chain, cultural heritage, and museum contents communication, are well on their way to integrating immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human-centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks, and end-user infrastructures to enable the next generation of interactive XR applications and services. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746269
Irene Viola 0001, Silvia Rossi 0001, Marta Orduna, Maria Torres Vega
ACM Multimedia2
2025 From Individual QoE to Shared Mental Models: A Novel Evaluation Paradigm for Collaborative XR
abstract
Extended Reality (XR) systems are rapidly shifting from isolated, single-user applications towards collaborative and social multi-user experiences. To evaluate the quality and effectiveness of such interactions, it is therefore required to move beyond traditional individual metrics such as Quality-of-Experience (QoE) or Sense of Presence (SoP). Instead, group-level dynamics such as effective communication, coordination etc. need to be encompassed to assess the shared understanding of goals and procedures. In psychology, this is referred to as a Shared Mental Model (SMM). The strength and congruence of such an SMM are known to be key for effective team collaboration and performance. In an immersive XR setting, though, novel Influence Factors (IFs) emerge that are not considered in a setting of physical co-location. Evaluations on the impact of these novel factors on SMM formation in XR, however, are close to non-existent. Therefore, this work proposes SMMs as a novel evaluation tool for collaborative and social XR experiences. To better understand how to explore this construct, we ran a prototypical experiment based on ITU recommendations in which the influence of asymmetric end-to-end latency is evaluated through a collaborative, two-user block building task. The results show how also in an XR context strong SMM formation can take place even when collaborators have fundamentally different responsibilities and behavior. Moreover, the study confirms previous findings by showing in an XR context that a teams’ SMM strength is positively associated with its performance.
Sam Van Damme, Jack Jansen 0001, Silvia Rossi 0001, Pablo César
QoMEX3
2025 A Clustering Approach to Unveil User Similarities in 6 df Extended Reality Applications
abstract
The advent in our daily life of Extended Reality (XR) technologies, such as Virtual and Augmented Reality, has led to the rise of user-centric systems, offering higher level of interaction and presence in virtual environments. In this context, understanding the actual interactivity of users is still an open challenge and a key step to enabling user-centric system. In this work, our goal is to construct an efficient clustering tool for 6 df navigation trajectories by extending the applicability of existing behavioural tool. Specifically, we first compare the navigation in 6 df with its 3 df counterpart, highlighting the main differences and novelties. Then, we investigate new metrics aimed at better modelling behavioural similarities between users in a 6 df system. More concretely, we define and compare 11 similarity metrics which are based on different distance features (i.e., user positions in the 3D space, user viewing directions) and distance measurements (i.e., Euclidean, Geodesic, angular distance). Our solutions are validated and tested on real navigation paths of users interacting with dynamic volumetric media in both 6 df Virtual Reality and Augmented Reality conditions. Results show that metrics based on both user position and viewing direction better perform in detecting user similarity while navigating in a 6 df system. Such easy-to-use but robust metrics allow us to answer a fundamental question for user-centric systems: ‘How do we detect if users look at the same content in 6 df?’, opening the gate to new solutions based on users interactivity, such as viewport prediction, live streaming services optimised based on users behaviour but also for user-based quality assessment methods.
Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependent
abstract
This paper presents a Task-Free eye-tracking dataset for Dynamic Point Clouds (TF-DPC) aimed at investigating visual attention. The dataset is composed of eye gaze and head movements collected from 24 participants observing 19 scanned dynamic point clouds in a Virtual Reality (VR) environment with 6 degrees of freedom. We compare the visual saliency maps generated from this dataset with those from a prior task-dependent experiment (focused on quality assessment) to explore how high-level tasks influence human visual attention. To measure the similarity between these visual saliency maps, we apply the well-known Pearson correlation coefficient and an adapted version of the Earth Mover's Distance metric, which takes into account both spatial information and the degrees of saliency. Our experimental results provide both qualitative and quantitative insights, revealing significant differences in visual attention due to task influence. This work enhances our understanding of the visual attention for dynamic point cloud (specifically human figures) in VR from gaze and human movement trajectories, and highlights the impact of task-dependent factors, offering valuable guidance for advancing visual saliency models and improving VR perception.
Xuemei Zhou, Irene Viola 0001, Silvia Rossi 0001, Pablo César
IEEE Trans. Vis. Comput. Graph.3
2024 Open-Sourcing VR2Gather: A Collaborative Social VR System for Adaptive Multi-Party Real Time Communication
abstract
Social Virtual Reality is envisioned to transform how individu- als communicate remotely, offering a sense of immersion and co- presence within a virtual space. Current platforms enabling remote social interactions rely on synthetic user representations. We ad- dress this limitation by enabling realistic human representation through volumetric content capture, encoding and transmission. Specifically, we present an extended version of VR2Gather, now a fully open source Unity package, available at https://github.com/ cwi-dis/VR2Gather-acmmm-oss. Our platform is a customisable system to transmit volumetric content in a multi-party real-time environment, easy to integrate into existing applications.
Jack Jansen 0001, Thomas Röggla, Silvia Rossi 0001, Irene Viola 0001, Pablo César
ACM Multimedia3
2024 AGAR - Attention Graph-RNN for Adaptative Motion Prediction of Point Clouds of Deformable Objects
abstract
This article focuses on motion prediction for point cloud sequences in the challenging case of deformable 3D objects, such as human body motion. First, we investigate the challenges caused by deformable shapes and complex motions present in this type of representation, with the ultimate goal of understanding the technical limitations of state-of-the-art models. From this understanding, we propose an improved architecture for point cloud prediction of deformable 3D objects. Specifically, to handle deformable shapes, we propose a graph-based approach that learns and exploits the spatial structure of point clouds to extract more representative features. Then, we propose a module able to combine the learned features in aadaptativemanner according to the point cloud movements. The proposed adaptative module controls the composition of local and global motions for each point, enabling the network to model complex motions in deformable 3D objects more effectively. We tested the proposed method on the following datasets: MNIST moving digits, theMixamohuman bodies motions [ 15 ], JPEG [ 5 ] and CWIPC-SXR [ 32 ] real-world dynamic bodies. Simulation results demonstrate that our method outperforms the current baseline methods given its improved ability to model complex movements as well as preserve point cloud shape. Furthermore, we demonstrate the generalizability of the proposed framework for dynamic feature learning by testing the framework for action recognition on the MSRAction3D dataset [ 19 ] and achieving results on par with state-of-the-art methods.
Silvia Rossi 0001, Laura Toni
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Designing and Evaluating a VR Lobby for a Socially Enriching Remote Opera Watching Experience
abstract
The latest social VR technologies have enabled users to attend traditional media and arts performances together while being geographically removed, making such experiences accessible despite budget, distance, and other restrictions. In this work, we aim at improving the way remote performances are shared by designing and evaluating a VR theatre lobby which serves as a space for users to gather, interact, and relive the common experience of watching a virtual opera. We conducted an initial test with experts ($\mathrm{N}=10$, i.e., designers and opera enthusiasts) in pairs using our VR lobby prototype, developed based on the theoretical lobby design concept. A unique aspect of our experience is its highly realistic representation of users in the virtual space. The test results guided refinements to the VR lobby structure and implementation, aiming to improve the user experience and align it more closely with the social VR lobby's intended purpose. With the enhanced prototype, we ran a between-subject controlled study ($\mathrm{N}=40$) to compare the user experience in the social VR lobby between individuals and paired participants. To do so, we designed and validated a questionnaire to measure the user experience in the VR lobby. Results of our mixed-methods analysis, including interviews, questionnaire results, and user behavior, reveal the strength of our social VR lobby in connecting with other users, consuming the opera in a deeper manner, and exploring new possibilities beyond what is common in real life. All supplemental materials are available at https://github.com/cwi-dis/IEEEVR2024-VRLobby.
Sueyoon Lee, Irene Viola 0001, Silvia Rossi 0001, Zhirui Guo, Ignacio Reimat, Kinga Lawicka, Alina Striner, Pablo César
IEEE Trans. Vis. Comput. Graph.3
2023 Extending 3-DoF Metrics to Model User Behaviour Similarity in 6-DoF Immersive Applications
abstract
Immersive reality technologies, such as Virtual and Augmented Reality, have ushered a new era of user-centric systems, in which every aspect of the coding-delivery-rendering chain is tailored to the interaction of the users. Understanding the actual interactivity and behaviour of the users is still an open challenge and a key step to enabling such a user-centric system. Our main goal is to extend the applicability of existing behavioural methodologies for studying user navigation in the case of 6 Degree-of-Freedom (DoF). Specifically, we first compare the navigation in 6-DoF with its 3-DoF counterpart highlighting the main differences and novelties. Then, we define new metrics aimed at better modelling behavioural similarities between users in a 6-DoF system. We validate and test our solutions on real navigation paths of users interacting with dynamic volumetric media in 6-DoF Virtual Reality conditions. Our results show that metrics that consider both user position and viewing direction better perform in detecting user similarity while navigating in a 6-DoF system. Having easy-to-use but robust metrics that underpin multiple tools and answer the question "how do we detect if two users look at the same content?" open the gate to new solutions for a user-centric system.
Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César
MMSys1
2022 Explaining Hierarchical Features in Dynamic Point Cloud Processing
abstract
This paper aims at bringing some light and understanding to the field of deep learning for dynamic point cloud processing. Specifically, we focus on the hierarchical features learning aspect, with the ultimate goal of understanding which features are learned at the different stages of the process and what their meaning is. Last, we bring clarity on how hierarchical components of the network affect the learned features and their importance for a successful learning model. This study is conducted for point cloud prediction tasks, useful for predicting coding applications.
Silvia Rossi 0001, Laura Toni
PCS2
2021 Spatio-Temporal Graph-RNN for Point Cloud Prediction
abstract
In this paper, we propose an end-to-end learning network to predict future frames in a point cloud sequence. As main novelty, an initial layer learns topological information of point clouds as geometric features, to form representative spatiotemporal neighborhoods. This module is followed by multiple Graph-RNN cells. Each cell learns point dynamics (i.e., RNN states) by processing each point jointly with its spatiotemporal neighbours. We tested the network performance with a MNIST dataset of moving digits, a synthetic human bodies motions, and JPEG dynamic bodies datasets. Simulation results demonstrate that our method outperforms baseline ones, which neglect geometry features information.
Silvia Rossi 0001, Laura Toni
ICIP2
2021 A New Challenge: Behavioural Analysis Of 6-DOF User When Consuming Immersive Media
abstract
Thanks to recent advances in computer graphics, wearable technology, and connectivity, Virtual Reality (VR) has landed in our daily life. A key novelty in VR is the role of the user, which has turned from merely passive to entirely active. Thus, improving any aspect of the coding-delivery-rendering chain starts with the need for understanding user behaviour. To do so, we investigate the navigation trajectories of users within a 6-Degrees-of-Freedom (DoF) VR environment. Specifically, we investigate the main differences and similarities between 3 and 6-DoF navigation through existing methodologies adopted to study user behaviour in 3-DoF settings. Our simulation results, based on real navigation paths of users while displaying dynamic volumetric media in 6-DoF conditions, show the limitations of clustering algorithms for 3-DoF in assessing user similarity in 6-DoF. Given these observations, we state the need for developing new solutions for the analysis of 6-DoF trajectories.
Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César
ICIP1
2020 Do Users Behave Similarly in VR? Investigation of the User Influence on the System Design
abstract
With the overarching goal of developing user-centric Virtual Reality (VR) systems, a new wave of studies focused on understanding how users interact in VR environments has recently emerged. Despite the intense efforts, however, current literature still does not provide the right framework to fully interpret and predict users’ trajectories while navigating in VR scenes. This work advances the state-of-the-art on both the study of users’ behaviour in VR and the user-centric system design. In more detail, we complement current datasets by presenting a publicly available dataset that provides navigation trajectories acquired for heterogeneous omnidirectional videos and different viewing platforms—namely, head-mounted display, tablet, and laptop. We then present an exhaustive analysis on the collected data to better understand navigation in VR across users, content, and, for the first time, across viewing platforms. The novelty lies in the user-affinity metric, proposed in this work to investigate users’ similarities when navigating within the content. The analysis reveals useful insights on the effect of device and content on the navigation, which could be precious considerations from the system design perspective. As a case study of the importance of studying users’ behaviour when designing VR systems, we finally propose a user-centric server optimisation. We formulate an integer linear program that seeks the best stored set of omnidirectional content that minimises encoding and storage cost while maximising the user’s experience. This is posed while taking into account network dynamics, type of video content, and also user population interactivity. Experimental results prove that our solution outperforms common company recommendations in terms of experienced quality but also in terms of encoding and storage, achieving a savings up to 70%. More importantly, we highlight a strong correlation between the storage cost and the user-affinity metric, showing the impact of the latter in the system architecture design.
Silvia Rossi 0001, Cagri Ozcinar, Aljoscha Smolic, Laura Toni
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Spherical Clustering of Users Navigating 360° Content
abstract
In Virtual Reality (VR) applications, understanding how users explore the omnidirectional content is important to optimize content creation, to develop user-centric services, or even to detect disorders in medical applications. Clustering users based on their common navigation patterns is a first direction to understand users behavior. However, classical clustering techniques fail in identifying this common paths, since they are usually focused on minimizing a simple distance metric. In this paper, we argue that minimizing the distance metric does not necessarily guarantee to identify users that experience similar navigation path in the VR domain. Therefore, we propose a graph-based method to identify clusters of users who are attending the same portion of the spherical content over time. The proposed solution takes into account the spherical geometry of the content and aims at clustering users based on the actual overlap of displayed content among users. Our method is tested on real VR user navigation patterns. Results show that our solution leads to clusters in which at least 85% of the content displayed by one user is shared among the other users belonging to the same cluster.
Silvia Rossi 0001, Francesca De Simone, Pascal Frossard, Laura Toni
ICASSP1
2017 Navigation-aware adaptive streaming strategies for omnidirectional video
abstract
Virtual reality (VR) applications target high-quality and zero-latency scene navigation to provide users with a full-immersion sensation within a scene. From a network perspective, this requires transmission of the omnidirectional content in its entirety, at a high resolution, which is not always feasible in bandwidth-limited networks. In this work, we propose an optimal transmission strategy for virtual reality applications able to fulfill the bandwidth requirements, while optimizing the end-user quality experienced in the navigation. In further detail, we consider a tile-based coded content for adaptive streaming systems, and we propose a navigation-aware transmission strategy at the clientside (i.e., adaptation logic), which is able to optimize the rate at which each tile is downloaded. First, we introduce the viewport- quality as metric that reflects the quality of any portion of the sphere displayed by the end-user. Then, we cast the tile-rate optimization as an integer linear programming problem and show that the proposed solution achieves substantial quality gains when compared to state-of-the-art adaptation logic methods.
Silvia Rossi 0001, Laura Toni
MMSP1