Maria Torres Vega

dblp:147/1075 · DBLP profile ↗
← Back
47ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0002-5656-6607ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 4 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 16 · 13 since 2021Computer networks · 6 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 MultiSenseVR: An open multimodal dataset for human pose estimation and perception in interactive VR
abstract
Current Virtual Reality (VR) systems rely on inside-out visual-inertial tracking, which enables accurate localization but provides only a partial representation of the user's body. This limitation restricts embodiment and interaction fidelity in interactive VR scenarios requiring full-body awareness and expressive gestures. To capture both global body motion and fine-grained interaction cues within a single sensing framework, we introduce MultiSenseVR, the first open multimodal dataset that jointly captures synchronized millimeter-wave (mmWave) Wi-Fi, Surface Electromyography (sEMG), inertial signals, and high-precision 3D motion capture for ground truth in an immersive VR setting. The dataset includes recordings from 24 participants interacting with a custom fast-food simulation designed to elicit natural full-body movement. In addition to objective sensing data, MultiSenseVR provides subjective measures of presence and cybersickness. Baseline evaluations show that mmWave Wi-Fi sensing supports 3D pose estimation with accuracy comparable to camera-based approaches, while sEMG enables accurate subject-specific grasp classification. The dataset and supporting code are publicly available at https://osf.io/f6r7d.
Javad Sameri, Nabeel Nisar Bhat, Filip De Turck, Rafael Berkvens, Jeroen Famaey, Maria Torres Vega
MMSys6
2026 Virtual Chemistry: A Pilot Study on Physiological Synchrony in Collaborative and Competitive VR
abstract
As Extended Reality (XR) transitions from individual experiences to multi-user collaborative environments, understanding the dynamics of team connection and cooperation becomes critical. This pilot study investigates the potential of Physiological Synchrony (PS) as an objective measure of team chemistry in Collaborative Virtual Reality (CVR). We conduct a within-subjects study where participant pairs engage in a pizza-making task both in a collaborative and in a competitive scenario. Physiological data, i.e. Galvanic Skin Response (GSR), Photoplethysmogram (PPG), and Interbeat Interval (IBI), are collected and analyzed to quantify synchrony levels and compared to subjective questionnaires. Results confirm that participants perceive significantly higher team chemistry in the collaborative scenario. Objectively, PPG shows a preliminary tendency towards synchrony compared to subjective scores.
Sam Van Damme, Jannes Bryon, Javad Sameri, Filip De Turck, Maria Torres Vega
QoMEX5
2026 Impact of Interaction-Induced Body Movement on Cybersickness in Interactive Virtual Reality Environments
abstract
This paper investigates the impact of interactioninduced body movement on cybersickness in Virtual Reality (VR). Therefore we designed a controlled VR environment with three interaction conditions of increasing movement complexity. Data from 24 participants, combining subjective measures and motionderived features, show a consistent increase in cybersickness with higher movement demands, particularly in conditions involving vertical displacement and full-body interaction. Moreover, correlation analysis reveals a significant negative relationship between the coefficient of variation of head velocity and cybersickness. This suggests that more variable and adaptive motion may mitigate the discomfort. These findings highlight the importance of movement characteristics, beyond movement intensity alone, for designing full-body interactive VR experiences.
Javad Sameri, Nabeel Nisar Bhat, Filip De Turck, Rafael Berkvens, Jeroen Famaey, Maria Torres Vega
QoMEX6
2026 Multimodal Haptics for Realistic Texture Rendering in Extended Reality
abstract
Realistic texture perception in Extended Reality (XR) requires tactile feedback that simultaneously addresses the fundamental perceptual dimensions of hardness, roughness, and warmth. However, current haptic solutions often rely solely on vibrotactile feedback, which alone cannot convey the full spectrum of tactile sensations. This paper presents a multimodal texture rendering method that combines three distinct haptic modalities—force, vibration, and temperature—to provide all required perceptual dimensions. We developed a multimodal fingertip haptic display to deliver these stimuli, and evaluated its performance through perceptual user studies within a Virtual Reality (VR) experimental framework. This framework assesses the modalities, both in isolation and in combination, to examine the robustness of the mappings between stimuli and their corresponding dimensions. Furthermore, we demonstrate our multimodal method by rendering real-world materials such as wood, metal, and water. A subjective evaluation with 48 participants indicates that users can reliably identify most perceptual dimensions, and that the integration of more modalities consistently improved texture realism compared to representations with fewer. These findings validate the effectiveness of our multimodal texture rendering method and advance the design of immersive, texture-rich XR experiences.
Vivian Tsang, Jonathan Valgaeren, Marlon Rodríguez, Carlos Rodriguez Guerrero, Maria Torres Vega
IMX5
2025 Reinforcement Learning-based Orchestration of XR applications in Distributed 6G Cloud Infrastructures
abstract
eXtended Reality (XR) and holographic telepresence place stringent Quality of Service (QoS) demands on network infrastructure, requiring ultra-low latency, high throughput, and reliable connectivity. Meeting such QoS demands is critical in dynamic, distributed cloud environments, but does not always guarantee a satisfactory user experience. Quality of Experience (QoE) captures the user’s perception of service performance, which may be influenced by factors not fully reflected in systemlevel metrics. Thus, novel orchestration strategies must consider both QoS and QoE. This paper proposes a Reinforcement Learning (RL)-driven approach to edge-cloud orchestration capable of adapting to dynamic network conditions, leveraging a multiobjective reward function, including both QoS and QoE aspects, to guide service placement decisions. Evaluation shows that our RL approach reaches a 21.3% QoE gain over heuristics and 14.7% over balanced strategies, with 100% request acceptance. The results highlight the robustness and scalability of RL-driven orchestration, particularly for latency-sensitive 6G applications. Our findings also reveal the limitations of traditional heuristics under complex objectives and highlight the potential of RL as a transformative tool for intelligent network and service management in next-generation communication systems.
Javad Sameri, José Santos 0001, Sam Van Damme, Susanna Schwarzmann, Qing Wei 0001, Riccardo Trivisonno, Filip De Turck, Maria Torres Vega
CNSM8
2025 Towards a Hybrid Hierarchical Digital Twin Architecture for the 6G Compute Continuum
abstract
The emergence of the 6 G era demands seamless orchestration across an increasingly heterogeneous and distributed compute continuum-spanning edge, fog, and cloud resources. Moreover, the next generation of the mobile network is poised to redefine the digital landscape by enabling pervasive intelligence, ultra-low latency communication, and extreme heterogeneity across the entire network infrastructure. This transformation introduces unprecedented orchestration challenges due to the dynamic, multi-domain, and resource-constrained nature of emerging workloads such as Generative Artificial Intelligence (GenAI) inference, immersive eXtended Reality (XR), and autonomous systems. To tackle this complexity, we advocate for a Hybrid Hierarchical Digital Twin (DT) architecture that serves as a foundation for intelligent, adaptive, and real-time orchestration in 6 G environments. We present a comprehensive vision for integrating DTs as enablers of intelligent, context-aware, and adaptive orchestration mechanisms that span across multiple domains. The proposed architecture introduces a multi-layered DT hierarchy combining local and global views, enabling scalable coordination and real-time decision-making. We highlight key architectural enhancements required to realize this vision, including inter-twin interoperability and behavioral modeling for QoE estimation. This work aims to guide researchers and practitioners in shaping the foundations of resilient and efficient orchestration frameworks for 6 G systems.
José Santos 0001, Javad Sameri, Sam Van Damme, Susanna Schwarzmann, Qing Wei 0001, Riccardo Trivisonno, Maria Torres Vega, Filip De Turck
CNSM7
2025 IXR '25: 3rd International Workshop on Interactive eXtended Reality
abstract
Despite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, telesurgery, interconnection in new buildings project chain, cultural heritage, and museum contents communication, are well on their way to integrating immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human-centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks, and end-user infrastructures to enable the next generation of interactive XR applications and services. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3746269
Irene Viola 0001, Silvia Rossi 0001, Marta Orduna, Maria Torres Vega
ACM Multimedia4
2025 Redirected Walking for Multi-User eXtended Reality Experiences with Confined Physical Spaces
abstract
eXtended Reality (XR) applications allow the user to explore nearly infinite virtual worlds in a truly immersive way. However, wandering around through these Virtual Environments (VE)s while physically walking in reality is heavily constrained by the size of the Physical Environment (PE). Therefore, in the last years different techniques have been devised to improve Locomotion in XR. One of these is Redirected Walking (RDW), which aims to find a balance between immersion and PE requirements by steering users away from the boundaries of the PE while allowing for arbitrary motion in the VE. However, current RDW methods still require large PEs, as to avoid obstacles and other users. Moreover, they introduce unnatural alterations in the natural path of the user, which can trigger perception anomalies, such as cybersickness or break of presence. These circumstances limit their usage in real life scenarios. This paper introduces a novel RDW algorithm, with the focus on allowing multiple users to explore an infinite VE in a confined space (6x6 m2). To evaluate it, we designed a multi-user Virtual Reality (VR) maze game, and benchmarked it against the state-of-the-art. A subjective study (20 participants) was conducted, where objective metrics, e.g., the path and the speed of the user, were combined wit subjective perception analysis in terms of their cybersickness levels. Our results show that our method reduces the appearance of cybersickness appearance in 80% of participants compared to the state-of-the-art. These findings show the applicability of RDW to multi-user VR with constrained environments.
Gijs Fiten, Jit Chatterjee, Kobe Vanhaeren, Mattis Martens, Maria Torres Vega
QoMEX5
2025 Virtual Gravity: Enhancing Weight Sensation in Virtual Reality with Haptic Interfaces
abstract
Virtual Reality (VR) training offers the possibility of simulating diverse environments and scenarios in a safe setting. One interesting use case is for the training of astronauts, who can use VR to train for demanding space missions. However, experiences where the perception goes beyond the visual, such as for instance the impact of gravity in astronaut training, multi-modality becomes crucial. This work investigates the impact of including haptics, in the form of an exoskeleton, to simulate extremity weight sensations for astronaut training in VR. Focusing on the forearm, we introduce Virtual Gravity, a novel multi-modal (haptic-visual) approach to emulates the gravitational forces in planets. By means of a subjective study, we evaluated the comfort, portability, as well as the overall perception of our solution. Our results reveal that the system effectively simulates weight and different gravitational conditions, enhancing VR immersion through forearm interaction. Moreover, they highlight the impact of multi-modal interactions on enhancing VR experiences. Our codes are available at https://github.com/Jit-INP/virtualGravity.
Marta Rózycka, Sindja Lika, Jit Chatterjee, Carlos Rodriguez, Maria Torres Vega
QoMEX5
2025 Extending QoE Research: Addressing Harassment in Social VR Through Feminist Digital Ethnography
abstract
Social Virtual Reality (SVR) enables immersive, avatar-based interaction, introducing new dimensions of presence, co-presence, and social engagement. While existing Quality of Experience (QoE) frameworks capture technical and perceptual aspects of user experience, they often overlook the social, relational, and power-sensitive dimensions that shape users’ realities in SVR. This paper proposes Feminist Digital Ethnography (FDE) as a complementary methodology for QoE research, foregrounding reflexivity, intersectionality, and ethical care. Applied to a case study in VRChat, FDE helps surface lived experiences often missed by traditional approaches—revealing how platform affordances, digital embodiment, and social structures interact. By demonstrating FDE’s potential to expand QoE frameworks, this work contributes to more inclusive and context-aware understandings of user experience in SVR.
Aleksandra Zheleva, Katleen Gabriels, Maria Torres Vega
QoMEX3
2025 PanoLoRA: 360-Degree Content Generation via Low-Rank Adaptation of Diffusion Models
abstract
Transforming a narrow Field of View (nFoV) image into a seamless 360-degree panoramic scene represents one of the most challenging tasks in computer vision. It requires models to hallucinate realistic content while maintaining perfect spatial continuity across wraparound boundaries. Powerful generative AI models have transformed this field by enabling the creation of realistic content that fills missing areas while maintaining visual and spatial consistency. However, current diffusion models struggle with panoramic imagery due to boundary discontinuities, the computational cost of full model fine-tuning, and their tendency to hallucinate unrealistic artifacts in regions with limited visual context. This work presents PanoLoRA, an efficient method for generating panoramic images from nFoV inputs using Low-Rank Adaptation (LoRA) of Stable Diffusion models. Our approach introduces two key innovations: (i) spherical-2D modules replacing standard convolutions with circular padding to ensure seamless wrap-around at panoramic boundaries; (ii) LoRA fine-tuning that adapts only 0.8% of model parameters while preserving pre-trained semantic knowledge. Results show that our PanoLoRA outperforms the leading state-of-the-art method in generating panoramic images from nFoV inputs by 20 points in terms of Fréchet Inception Distance (FID) score.
Jit Chatterjee, Maria Torres Vega
VCIP2
2025 3D-Scene-Former: 3D scene generation from a single RGB image using Transformers
Jit Chatterjee, Maria Torres Vega
Vis. Comput.2
2024 Collaborative Cooking in VR: Effects of Network Distortion in Multi-User Virtual Environments
abstract
The future of human interaction is virtual. Thus it will require effective collaboration on tasks among users in remote settings. eXtended Reality (XR) is playing a leading role in this transition, offering a realm where virtual collaboration becomes not just possible but essential in situations where physical presence is limited by risk, cost, or complexity. However, while networks are continuously evolving, they can still introduce unexpected impairments that potentially degrade the user perception, i.e., the Quality-of-Experience (QoE), of such Collaborative Virtual Reality (CVR) scenarios. In response to this challenge, this paper presents a demonstrator designed to explicitly showcase the effects of network conditions on CVR. Our platform, centered around a pizza-making game, allows for exploration of the real-time impact of different network parameters, such as packet delay, loss, and throttling on the user engagement and perception in CVR. The framework employs a combination of subjective, objective, and physiological assessments, including the capture of heart rate and skin conductivity, to gain comprehensive insights into user experiences. Our platform not only allows users to directly experience the impact of network impairments on CVR interactions but also provides initial evidence of how such distortions affect both subjective perceptions and objective performance metrics.
Javad Sameri, Sam Van Damme, Susanna Schwarzmann, Qing Wei 0001, Riccardo Trivisonno, Filip De Turck, Maria Torres Vega
MMSys7
2024 Enhancing Virtual Reality Stress Relief with Haptics: The Virtual Rage Room Use Case
abstract
EXtended Reality (XR) is already demonstrating its potential beyond the entertainment and gaming industry. One sector clearly benefiting from the advantages of XR is the treatment of stress related mental illnesses by means of Virtual Reality (VR). This is a form of therapy using VR which seeks to help decrease the intensity of the stress responses and anxiety levels due to the various modern-day pressures (e.g., situations, thoughts, or memories which provoke anxiety or fear). While showing promising results, the audiovisual essence of Virtual Reality (VR) can limit the effectiveness of this type of virtual therapy, as the patient’s interaction with the environment is constrained to their visual or at most also their audio senses. As such, including in the immersive therapy tactile therapy could enhance the experience and thus the effectiveness of the therapy. However, this has been largely unexplored. The purpose of this paper is to explore the impact of haptic feedback in reducing anxiety for stress relief treatment. Therefore, we present a haptic-enabled subjective methodology. As use case, we selected the booming case of the Virtual Rage Room (VRR), where participants can vent their rage by (virtually) destroying objects. The results of our study highlight the significantly positive impact of incorporating haptic feedback in mitigating anxiety within this context. Moreover, the analysis reassures the intrinsic value of this treatment as a potent tool for anxiety alleviation.
Javad Sameri, Flor Neufkens, Sam Van Damme, Filip De Turck, Maria Torres Vega
QoMEX5
2023 IXR '23: 2nd International Workshop on Interactive eXtended Reality
abstract
Despite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, telesurgery, interconnection in new buildings project chain, Cultural Heritage, and Museum contents communication, are well on their way to integrating immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human-centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks, and end-user infrastructures to enable the next generation of interactive XR applications and services.
Irene Viola 0001, Hadi Amirpour, Stephanie Arevalo, Maria Torres Vega
ACM Multimedia4
2023 On the Impact of Interactive eXtended Reality: Challenges and Opportunities for Multimedia Research
abstract
Extended Reality (XR) has been hailed as the new frontier of media, ushering new possibilities for societal areas such as communications, training, entertainment, gaming, and cultural heritage. However, despite the remarkable technical advances, current XR applications are in their majority local and individual experiences. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely content realism, by means of Artificial Intelligence (AI) techniques, motion-to-photon latency, and accurate human-centric driven experiences able to map real and virtual worlds seamlessly. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. In this panel, together with the leading experts of the SIGMM community, we will explore the challenges and opportunities to unlock the next generation of interactive XR applications and services.
Irene Viola 0001, Maria Torres Vega
ACM Multimedia2
2023 Immersive and Interactive Subjective Quality Assessment of Dynamic Volumetric Meshes
abstract
Dynamic point cloud delivery can provide the required interactivity and realism to six degrees of freedom (6DoF) interactive applications. However, dynamic point cloud rendering imposes stringent requirements (e.g., frames per second (FPS) and quality) that current hardware cannot handle. A possible solution is to convert point cloud into meshes before rendering on the head-mounted display (HMD). However, this conversion can induce degradation in quality perception such as a change in depth, level of detail, or presence of artifacts. This paper, as one of the first, presents an extensive subjective study of the effects of converting point cloud to meshes with different quality representations. In addition, we provide a novel in-session content rating methodology, providing a more accurate assessment as well as avoiding post-study bias. Our study shows that both compression level and observation distance have their influence on subjective perception. However, the degree of influence is heavily entangled with the content and geometry at hand. Furthermore, we also noticed that while end users are clearly aware of quality switches, the influence on their quality perception is limited. As a result, this has the potential to open up possibilities in bringing the adaptive video streaming paradigm to the 6DoF environment.
Sam Van Damme, Imen Mahdi, Hemanth Kumar Ravuri, Jeroen van der Hooft, Filip De Turck, Maria Torres Vega
QoMEX6
2023 Are we ready for Haptic Interactivity in VR? An Experimental Comparison of Different Interaction Methods in Virtual Reality Training
abstract
In recent years, Virtual Reality (VR) has gained attention as a tool for a plethora of applications such as first-aid, firefighting and in the automotive industry. End-user immersion is a key factor in these applications to make the experience representative for its real-life counterpart. By enhancing the traditional audiovisual cues with additional sensory inputs in terms of haptic vibro-tactile and kinesthetic feedback, this immersion can be improved. But are current haptic implementations sufficient to provide the required added value? And how do they compare to other types of VR interaction? In this paper, we present a multi-modal VR training framework able to provide subjective and objective comparisons among three different interaction options: (i) haptic gloves, (ii) traditional VR controllers, and (iii) non-haptic handtracking. We performed a user test where the different interactivity flavours were compared in terms of their influence on both subjective perception and objective performance of the end-user by means of three VR training scenarios. The subjective results show an aversion towards non-haptic handtracking for constrained, cognitively light tasks while a preference towards controllers exist for more cognitively heavy multi-tasking. This is however not reflected in objective results, where differences between interaction methods are far less pronounced.
Sam Van Damme, Jordy Tack, Glenn Van Wallendael, Filip De Turck, Maria Torres Vega
QoMEX5
2023 Impact of Quality and Distance on the Perception of Point Clouds in Mixed Reality
abstract
Point Cloud (PC) streaming has recently attracted research attention as it has the potential to provide six degrees of freedom (6DoF), which is essential for truly immersive media. PCs require high-bandwidth connections, and adaptive streaming is a promising solution to cope with fluctuating bandwidth conditions. Thus, understanding the impact of different factors in adaptive streaming on the Quality of Experience (QoE) becomes fundamental. Mixed Reality (MR) is a novel technology and has recently become popular. However, quality evaluations of PCs in MR environments are still limited to static images. In this paper, we perform a subjective study on four impact factors on the QoE of PC video sequences in MR conditions, including quality switches, viewing distance, and content characteristics. The experimental results show that these factors significantly impact QoE. The QoE decreases if the sequence switches to lower quality and/or is viewed at a shorter distance, and vice versa. Additionally, the end user might not distinguish the quality differences between two quality levels at a specific viewing distance. Regarding content characteristics, objects with lower contrast seem to provide better quality scores.
Minh Nguyen 0006, Shivi Vats, Sam Van Damme, Jeroen van der Hooft, Maria Torres Vega, Tim Wauters, Christian Timmerer, Hermann Hellwagner
QoMEX5
2023 A Platform for Subjective Quality Assessment in Mixed Reality Environments
abstract
3D objects are important components in Mixed Reality (MR) environments as they allow users to inspect and interact with them in a six degrees of freedom (6DoF) system. Point clouds (PCs) and meshes are two common 3D object representations that can be compressed to reduce the delivered data at the cost of quality degradation. In addition, as the end users can move around in 6DoF applications, the viewing distance can vary. Quality assessment is necessary to evaluate the impact of the compressed representation and viewing distance on the Quality of Experience (QoE) of end users. This paper presents a demonstrator for subjective quality assessment of dynamic PC and mesh objects under different conditions in MR environments. Our platform allows conducting subjective tests to evaluate various QoE influence factors, including encoding parameters, quality switching, viewing distance, and content characteristics, with configurable settings for these factors.
Shivi Vats, Minh Nguyen 0006, Sam Van Damme, Jeroen van der Hooft, Maria Torres Vega, Tim Wauters, Christian Timmerer, Hermann Hellwagner
QoMEX5
2023 Human-Centered and AI-driven Generation of 6-DoF Extended Reality
abstract
In order to unlock the full potential of Extended Reality (XR) and its application to societal sectors such as health (e.g., training) or Industry 5.0 (e.g., remote control of infrastructure) there is a need for very realistic environments to enhance the presence of the user. However, current photo-realistic content generation methods (such as Light Fields) require a massive amount of data transmission (i.e., ultra-high bandwidths) and extreme computational power for displaying. Thus, they are not suited for interactive immersive and realistic applications. In this research, we hypothesize that is possible to generate realistic dynamic 3D environments by means of Deep Generative Networks. The work will consist of two parts: (1) a computer vision system that generates the 3D environment based on 2D images, and (2) a Human-Computer Interaction system (HCI) that predicts Region of Interest (RoI) for efficient 3D rendering, subjective and objective assessment of user perception (by means of presence) to enhance the 3D scene quality. This work aims to gain insights into how well deep generative methods can create realistic and immersive environments. This will significantly help future developments in realistic and immersive XR content creation.
Jit Chatterjee, Maria Torres Vega
IMX2
2022 Clustering-Based Psychometric No-Reference Quality Model for Point Cloud Video
abstract
Point cloud video streaming is a fundamental application of immersive multimedia. In it, objects represented as sets of points are streamed and displayed to remote users. Given the high bandwidth requirements of this content, small changes in the network and/or encoding can affect the users' perceived quality in unexpected manners. To tackle the degradation of the service as fast as possible, real-time Quality of Experience (QoE) assessment is needed. As subjective evaluations are not feasible in real time due to their inherent costs and duration, low-complexity objective quality assessment is a must. Traditional No-Reference (NR) objective metrics at client side are best suited to fulfill the task. However, they lack on accuracy to human perception. In this paper, we present a cluster-based objective NR QoE assessment model for point cloud video. By means of Machine Learning (ML)-based clustering and prediction techniques combined with NR pixel-based features (e.g., blur and noise), the model shows high correlations (up to a 0.977 Pearson Linear Correlation Coefficient (PLCC)) and low Root Mean Squared Error (RMSE) (down to 0.077 on a zero-to-one scale) towards objective benchmarks after evaluation on an adaptive streaming point cloud dataset consisting of sixteen source videos and 453 sequences in total.
Sam Van Damme, Maria Torres Vega, Jeroen van der Hooft, Filip De Turck
ICIP2
2022 IXR '22: 1st Workshop on Interactive eXtended Reality
abstract
Despite remarkable advances, current Extended Reality (XR) applications are in their majority local and individual experiences. A plethora of interactive applications, such as teleconferencing, tele-surgery, interconnection in new buildings project chain, Cultural Heritage and Museum contents communication, are well on their way to integrate immersive technologies. However, interconnected, and interactive XR, where participants can virtually interact across vast distances, remains a distant dream. In fact, three great barriers stand between current technology and remote immersive interactive life-like experiences, namely the (i) content realism, (ii) motion-to-photon latency, and accurate (iii) human centric quality assessment and control. Overcoming these barriers will require novel solutions at all elements of the end-to-end transmission chain. This workshop focuses on the challenges, applications, and major advancements in multimedia, networks and end-user infrastructures to enable the next generation of interactive XR applications and services. The complete IXR'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552483
Irene Viola 0001, Hadi Amirpour, Maria Torres Vega
ACM Multimedia3
2022 Machine Learning Based Content-Agnostic Viewport Prediction for 360-Degree Video
abstract
Accurate and fast estimations or predictions of the (near) future location of the users of head-mounted devices within the virtual omnidirectional environment open a plethora of opportunities in application domains such as interactive immersive gaming and tele-surgery. Therefore, the past years have seen growing attention to models for viewport prediction in 360֯ environments. Among the approaches, content-agnostic, trajectory-based methods have the potential to provide very fast solutions, as they do not require complex analysis of the videos to provide a prediction. However, accurate trajectory-based viewport prediction is rather difficult due to the intrinsic variability in user behaviour. Furthermore, even when making use of machine learning, current approaches tend to be brute-force and heavily tailored to specific datasets with little comparison to existing benchmarks or publicly available studies. This article presents a generic, content-agnostic viewport prediction method consisting of a window-based approach combined with a preprocessing system to classify behavioural patterns in terms of user clustering and trajectory correlation. Moreover, as the state of the art does not provide a comparative analysis of different approaches, this work contributes to this. Based on the obtained results, a combined prediction model is proposed and evaluated. Our method shows a 36.8% to 53.9% improvement when compared to the static prediction baseline for a prediction horizon of 8 seconds. In addition, a 11.5% to 24.0% improvement to a brute-force machine learning prediction approach is obtained. As such, this work contributes towards the creation of more generic and structured solutions for content-agnostic viewport prediction in terms of data representation, preprocessing and modelling.
Sam Van Damme, Maria Torres Vega, Filip De Turck
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Efficient Orchestration of Service Chains in Fog Computing for Immersive Media
abstract
Immersive media services, such as Augmented and Virtual Reality (AR/VR) are getting significant attention in recent years with the promise of bringing immersive experiences to end users. However, despite the remarkable advances in the field, AR/VR applications are mostly local and individual experiences. The main obstacle between current technology and future remote, multi-user AR/VR applications is the stringent end-to-end (E2E) latency requirement, which cannot exceed 20 ms to avoid motion sickness. Emerging AR/VR services put even more pressure on current network infrastructures, calling for considerable advancements toward fully cloud-native architectures. Cloud-based VR services, where participants can virtually interact across vast distances, remain a distant dream. Several challenges still arise concerning the deployment and management of VR services. This paper presents a Mixed-Integer Linear Programming (MILP) formulation for the efficient orchestration of VR services in fog-cloud infrastructures. The model considers Fog Computing (FC), an extension of cloud computing, and Segment Routing (SR), which leverages the source routing paradigm. The evaluation of realistic VR container-based service chains shows that deploying VR components hosted in a fog-cloud infrastructure can satisfy the 20 ms latency boundary.
José Santos 0001, Jeroen van der Hooft, Maria Torres Vega, Tim Wauters, Bruno Volckaert, Filip De Turck
CNSM3
2021 SRFog: A flexible architecture for Virtual Reality content delivery through Fog Computing and Segment Routing
José Santos 0001, Jeroen van der Hooft, Maria Torres Vega, Tim Wauters, Bruno Volckaert, Filip De Turck
IM3
2021 A Full- and No-Reference Metrics Accuracy Analysis for Volumetric Media Streaming
abstract
Volumetric media streaming will be one of the fundamental technologies to enable near future immersive multimedia experiences. In it, objects represented as sets of points (i.e. point-clouds), are presented to remote users wearing Head-Mounted Displays (HMDs). Due to the stringent bandwidth and latency requirements of such applications, small changes in the network can affect the user in unexpected manners (physical discomfort, lack of concentration, etc.). Therefore, there is a need for assessing the perceived quality of this type of applications in real-time, i.e, the Quality of Experience (QoE). Given that subjective evaluations are not feasible for (near) real-time applications, objectively measuring this quality will be a must. While traditional objective metrics could potentially be used to fulfill the task, it is still unclear how accurate they are to assess volumetric media. To this end, this paper presents a thorough correlation analysis of both Full Reference (FR) and No Reference (NR) objective metrics to subjective Mean Opinion Scores (MOS) for different volumetric streaming scenarios. To enhance the accuracy, multiple Region-Of-Interest (ROI) selection and weighting procedures have been applied and their influence on the results have been investigated. Our results show that the classical video quality metric Video Multimethod Assessment Fusion (VMAF) is well-suited as an objective benchmark for volumetric media streaming in terms of correlation to subjective scores, while a combination of NR features could provide a suitable real-time assessment. Finally, ROI selection proves to widen the range of objective metrics, which is an important issue to apply traditional objective metrics to volumetric media.
Sam Van Damme, Maria Torres Vega, Filip De Turck
QoMEX2
2020 Human-centric Quality Management of Immersive Multimedia Applications
abstract
Augmented Reality (AR) and Virtual Reality (VR) multimodal systems are the latest trend within the field of multimedia. As they emulate the senses by means of omnidirectional visuals, 360° sound, motion tracking and touch simulation, they are able to create a strong feeling of presence and interaction with the virtual environment. These experiences can be applied for virtual training (Industry 4.0), tele-surgery (healthcare) or remote learning (education). However, given the strong time and task sensitiveness of these applications, it is of great importance to sustain the end-user quality, i.e. the Quality-of-Experience (QoE), at all times. Lack of synchronization and quality degradation need to be reduced to a minimum to avoid feelings of cybersickness or loss of immersiveness and concentration. This means that there is a need to shift the quality management from system-centered performance metrics towards a more human, QoE-centered approach. However, this requires for novel techniques in the three areas of the QoE-management loop (monitoring, modelling and control). This position paper identifies open areas of research to fully enable human-centric driven management of immersive multimedia. To this extent, four main dimensions are put forward: (1) Task and well-being driven subjective assessment; (2) Real-time QoE modelling; (3) Accurate viewport prediction; (4) Machine Learning (ML)-based quality optimization and content recreation. This paper discusses the state-of-the-art, and provides with possible solutions to tackle the open challenges.
Sam Van Damme, Maria Torres Vega, Filip De Turck
NetSoft2
2020 Objective and Subjective QoE Evaluation for Adaptive Point Cloud Streaming
abstract
Volumetric media has the potential to provide the six degrees of freedom (6DoF) required by truly immersive media. However, achieving 6DoF requires ultra-high bandwidth transmissions, which real-world wide area networks cannot provide today. Therefore, recent efforts have started to target efficient delivery of volumetric media, using a combination of compression and adaptive streaming techniques. It remains, however, unclear how the effects of such techniques on the user perceived quality can be accurately evaluated. In this paper, we present the results of an extensive objective and subjective quality of experience (QoE) evaluation of volumetric 6DoF streaming. We use PCC-DASH, a standards-compliant means for HTTP adaptive streaming of scenes comprising multiple dynamic point cloud objects. By means of a thorough analysis, we investigate the perceived quality impact of the available bandwidth, rate adaptation algorithm, viewport prediction strategy and user's motion within the scene. We determine which of these aspects has more impact on the user's QoE, and to what extent subjective and objective assessments are aligned.
Jeroen van der Hooft, Maria Torres Vega, Christian Timmerer, Ali C. Begen, Filip De Turck, Raimund Schatz
QoMEX2
2020 A low-complexity psychometric curve-fitting approach for the objective quality assessment of streamed game videos
Sam Van Damme, Maria Torres Vega, Joris Heyse, Femke De Backere, Filip De Turck
Signal Process. Image Commun.2
2020 Dissecting the Performance of VR Video Streaming through the VR-EXP Experimentation Platform
abstract
To cope with the massive bandwidth demands of Virtual Reality (VR) video streaming, both the scientific community and the industry have been proposing optimization techniques such as viewport-aware streaming and tile-based adaptive bitrate heuristics. As most of the VR video traffic is expected to be delivered through mobile networks, a major problem arises: both the network performance and VR video optimization techniques have the potential to influence the video playout performance and the Quality of Experience (QoE). However, the interplay between them is neither trivial nor has it been properly investigated. To bridge this gap, in this article, we introduce VR-EXP, an open-source platform for carrying out VR video streaming performance evaluation. Furthermore, we consolidate a set of relevant VR video streaming techniques and evaluate them under variable network conditions, contributing to an in-depth understanding of what to expect when different combinations are employed. To the best of our knowledge, this is the first work to propose a systematic approach, accompanied by a software toolkit, which allows one to compare different optimization techniques under the same circumstances. Extensive evaluations carried out using realistic datasets demonstrate that VR-EXP is instrumental in providing valuable insights regarding the interplay between network performance and VR video streaming optimization techniques.
Roberto Irajá Tavares da Costa Filho, Marcelo Caggiani Luizelli, Stefano Petrangeli, Maria Torres Vega, Jeroen van der Hooft, Tim Wauters, Filip De Turck, Luciano Paschoal Gaspary
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Tile-based Adaptive Streaming for Virtual Reality Video
abstract
The increasing popularity of head-mounted devices and 360° video cameras allows content providers to provide virtual reality (VR) video streaming over the Internet, using a two-dimensional representation of the immersive content combined with traditional HTTP adaptive streaming (HAS) techniques. However, since only a limited part of the video (i.e., the viewport) is watched by the user, the available bandwidth is not optimally used. Recent studies have shown the benefits of adaptive tile-based video streaming; rather than sending the whole 360° video at once, the video is cut into temporal segments and spatial tiles, each of which can be requested at a different quality level. This allows prioritization of viewable video content and thus results in an increased bandwidth utilization. Given the early stages of research, there are still a number of open challenges to unlock the full potential of adaptive tile-based VR streaming. The aim of this work is to provide an answer to several of these open research questions. Among others, we propose two tile-based rate adaptation heuristics for equirectangular VR video, which use the great-circle distance between the viewport center and the center of each of the tiles to decide upon the most appropriate quality representation. We also introduce a feedback loop in the quality decision process, which allows the client to revise prior decisions based on more recent information on the viewport location. Furthermore, we investigate the benefits of parallel TCP connections and the use of HTTP/2 as an application layer optimization. Through an extensive evaluation, we show that the proposed optimizations result in a significant improvement in terms of video quality (more than twice the time spent on the highest quality layer), compared to non-tiled HAS solutions.
Jeroen van der Hooft, Maria Torres Vega, Stefano Petrangeli, Tim Wauters, Filip De Turck
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Optimizing Adaptive Tile-Based Virtual Reality Video Streaming
Jeroen van der Hooft, Maria Torres Vega, Stefano Petrangeli, Tim Wauters, Filip De Turck
IM2
2019 Exploring New York in 8K: an adaptive tile-based virtual reality video streaming experience
abstract
Adapting and tiling the streaming of virtual reality (VR) video content has the potential to reduce the ultra-high bandwidth requirements of this type of multimedia services. Towards that goal, the optimization of a number of aspects is currently actively being researched. Novel rate adaptation heuristics, sophisticated viewport prediction algorithms and streaming protocol optimizations have proven their value to improve certain aspect of the VR streaming chain. However, the interplay between all these different optimizations as well as their tradeoff has not yet been explored in an experimental playground. The purpose of this demonstrator is to provide a full end-to-end adaptive tile-based VR video streaming system where each of the optimization aspects can be tuned with and their effect illustrated on-site.
Maria Torres Vega, Jeroen van der Hooft, Joris Heyse, Femke De Backere, Tim Wauters, Filip De Turck, Stefano Petrangeli
MMSys1
2019 An Experimental Evaluation of Flow Setup Latency in Distributed Software Defined Networks
abstract
Next generation application domains such as Virtual Reality (VR), Augmented Reality (AR) together with the Tactile Internet paradigm impose ultra-low latency requirements on the networks (1 to 5 ms end-to-end latency). Towards this objective, networks are undergoing a tremendous transformation from the current packet switching models to Software Defined Networking (SDN) architectures, which provide programmability to configure the network. In its simplest variant, one single centralized controller orchestrates the whole SDN infrastructure. However, the fully centralized architecture (one single controller) can become a performance bottleneck, especially in terms of response throughput and flow setup latency. Furthermore, it suffers from massive scalability issues. In this direction, a number of more sophisticated SDN architectures are currently under research. While their theoretical advantages have been thoroughly discussed in the state-of-the-art, a comparative experimental analysis of these architectures is still missing. This study aims at providing such experimental performance comparison. Herein, we put to test a set of SDN architectures ranging from a fully centralized to a completely distributed control plane, comparing them in terms of flow setup latency. Overall results show that completely distributed architectures provide significantly better performance with almost 31% gain in terms of average flow setup latency over the centralized case.
Hemanth Kumar Ravuri, Maria Torres Vega, Tim Wauters, Bin Da, Alexander Clemm, Filip De Turck
NetSoft2
2019 A personalized Virtual Reality Experience for Relaxation Therapy
abstract
Virtual Reality (VR) has the potential to change not only to the way we consume and perceive entertainment but also to improve other important areas of society. One sector that is starting to benefit from the advantages of VR is the treatment of stress related mental illnesses. VR is able to bring relaxation therapy to the next level in which solutions can be scalable (without the need for real-time dedicated professionals) and personalized. This paper presents VRelax, a personalized VR relaxation therapy approach. By means of semantic methodologies and online learning techniques, VRelax provides a personalized, relaxing virtual environment to the user.
Joris Heyse, Thomas De Jonge, Maria Torres Vega, Femke De Backere, Filip De Turck
QoMEX3
2019 Contextual Bandit Learning-Based Viewport Prediction for 360 Video
abstract
Accurately predicting where the user of a Virtual Reality (VR) application will be looking at in the near future improves the perceive quality of services, such as adaptive tile-based streaming or personalized online training. However, because of the unpredictability and dissimilarity of user behavior it is still a big challenge. In this work, we propose to use reinforcement learning, in particular contextual bandits, to solve this problem. The proposed solution tackles the prediction in two stages: (1) detection of movement; (2) prediction of direction. In order to prove its potential for VR services, the method was deployed on an adaptive tile-based VR streaming testbed, for benchmarking against a 3D trajectory extrapolation approach. Our results showed a significant improvement in terms of prediction error compared to the benchmark. This reduced prediction error also resulted in an enhancement on the perceived video quality.
Joris Heyse, Maria Torres Vega, Femke De Backere, Filip De Turck
VR2
2018 Enabling Virtual Reality for the Tactile Internet: Hurdles and Opportunities
Maria Torres Vega, Taha Mehmli, Jeroen van der Hooft, Tim Wauters, Filip De Turck
CNSM1
2018 Predicting the performance of virtual reality video streaming in mobile networks
abstract
The demand of Virtual Reality (VR) video streaming to mobile devices is booming, as VR becomes accessible to the general public. However, the variability of conditions of mobile networks affects the perception of this type of high-bandwidth-demanding services in unexpected ways. In this situation, there is a need for novel performance assessment models fit to the new VR applications. In this paper, we present PERCEIVE, a two-stage method for predicting the perceived quality of adaptive VR videos when streamed through mobile networks. By means of machine learning techniques, our approach is able to first predict adaptive VR video playout performance, using network Quality of Service (QoS) indicators as predictors. In a second stage, it employs the predicted VR video playout performance metrics to model and estimate end-user perceived quality. The evaluation of PERCEIVE has been performed considering a real-world environment, in which VR videos are streamed while subjected to LTE/4G network condition. The accuracy of PERCEIVE has been assessed by means of the residual error between predicted and measured values. Our approach predicts the different performance metrics of the VR playout with an average prediction error lower than 3.7% and estimates the perceived quality with a prediction error lower than 4% for over 90% of all the tested cases. Moreover, it allows us to pinpoint the QoS conditions that affect adaptive VR streaming services the most.
Roberto Irajá Tavares da Costa Filho, Marcelo Caggiani Luizelli, Maria Torres Vega, Jeroen van der Hooft, Stefano Petrangeli, Tim Wauters, Filip De Turck, Luciano Paschoal Gaspary
MMSys3
2017 Unsupervised deep learning for real-time assessment of video streaming services
abstract
Evaluating quality of experience in video streaming services requires a quality metric that works in real time and for a broad range of video types and network conditions. This means that, subjective video quality assessment studies, or complex objective video quality assessment metrics, which would be best suited from the accuracy perspective, cannot be used for this tasks (due to their high requirements in terms of time and complexity, in addition to their lack of scalability). In this paper we propose a light-weight No Reference (NR) method that, by means of unsupervised machine learning techniques and measurements on the client side is able to assess quality in real-time, accurately and in an adaptable and scalable manner. Our method makes use of the excellent density estimation capabilities of the unsupervised deep learning techniques, the restricted Boltzmann machines, and light-weight video features computed just on the impaired video to provide a delta of quality degradation. We have tested our approach in two network impaired video sets, the LIMP and the ReTRiEVED video quality databases, benchmarking the results of our method against the well-known full reference metric VQM. We have obtained levels of accuracy of at least 85% in both datasets using all possible cases.
Maria Torres Vega, Decebal Constantin Mocanu, Antonio Liotta
Multim. Tools Appl.1
2017 Predictive no-reference assessment of video quality
Maria Torres Vega, Decebal Constantin Mocanu, Stavros Stavrou, Antonio Liotta
Signal Process. Image Commun.1
2017 Deep Learning for Quality Assessment in Live Video Streaming
abstract
Video content providers put stringent requirements on the quality assessment methods realized on their services. They need to be accurate, real-time, adaptable to new content, and scalable as the video set grows. In this letter, we introduce a novel automated and computationally efficient video assessment method. It enables accurate real-time (online) analysis of delivered quality in an adaptable and scalable manner. Offline deep unsupervised learning processes are employed at the server side and inexpensive no-reference measurements at the client side. This provides both real-time assessment and performance comparable to the full reference counterpart, while maintaining its no-reference characteristics. We tested our approach on the LIMP Video Quality Database (an extensive packet loss impaired video set) obtaining a correlation between 78% and 91% to the FR benchmark (the video quality metric). Due to its unsupervised learning essence, our method is flexible and dynamically adaptable to new content and scalable with the number of videos.
Maria Torres Vega, Decebal Constantin Mocanu, Jeroen Famaey, Stavros Stavrou, Antonio Liotta
IEEE Signal Process. Lett.1
2016 A Regression Method for real-time video quality evaluation
Maria Torres Vega, Decebal Constantin Mocanu, Antonio Liotta
MoMM1
2016 Resource allocation in optical beam-steered indoor networks
abstract
Optical Wireless (OW) technologies deploying narrow multiwavelength light beams offer a promising alternative to traditional wireless indoor communications as they provide higher bandwidths and overcome the radio spectrum congestion typical of the 2.4 and 5GHz frequency bands. However, unlocking their full potential requires exploring novel control and management techniques. Specifically, there is a need for efficient and intelligent resource management and localization techniques that allot wavelengths and capacity to devices. In this paper we present a resource allocation model for one such indoor optical wireless approach, a Beam-steered Reconfigurable Optical-Wireless System for Energy-efficient communication (BROWSE). BROWSE aims to supply each user within a room with its own downstream infrared light beam with at least 10Gbps throughput, while providing a 60GHz radio channel upstream. Using Integer Linear Programming (ILP) techniques, we have designed and implemented a resource allocation model for the BROWSE OW downstream connection. The designed model optimises the trade-off between energy-consumption and throughput, while providing TDM capabilities to effectively serve densely deployed devices with a limited number of simultaneous available wavelengths. Through several test-scenarios we have assessed the model's performance, as well as its applicability to future ultra-high bandwidth video streaming applications.
Maria Torres Vega, Jeroen Famaey, Antonius M. J. Koonen, Antonio Liotta
NOMS1
2015 Cognitive streaming on android devices
abstract
As the number of mobile devices increases, so do the complexity of wireless networks and the user's requirements. This tendency makes necessary for Multimedia Services to take the needed actions to adapt to the upcoming technology. A prominent example of this type of services is HTTP Adaptive Video Streaming Applications. In this research, we have studied how the latest HTTP Adaptive Streaming techniques, mainly developed for standard computers, could be adapted and used in mobile wireless devices. Furthermore, inspired by these solutions, which usually make use of Reinforcement Learning (RL) algorithms to find the suitable streaming rate, we have conceived a novel smart video player client in Java for Android platform using the Dynamic Adaptive Streaming over HTTP (DASH) protocol. We have assessed the performance of our proposed solution in a self-developed wireless test-bed under different network conditions. Thus, we have seen that by including in the reward function contributions regarding the download speed of the video segments, especially needed due to the fluctuating nature of the wireless networks, and the segments already buffered, improves drastically the overall performance of the video client. Besides that, we have discovered that, in a cognitive adaptive approach, bandwidth constraints affect the user's experience more substantially, while impairments such as packet loss can be prevented.
Maria Torres Vega, Decebal Constantin Mocanu, Rosario Barresi, Giancarlo Fortino, Antonio Liotta
IM1
2015 Accuracy of No-Reference Quality Metrics in Network-impaired Video Streams
abstract
The Video Quality Metric (VQM) is nowadays one of the most used objective methods to assess video quality, thanks to its high correlation with both the human visual system (HVS) and subjective methods. VQM is, however, not viable in real-time deployments such as mobile streaming, not only due to its high computational demands but, specifically, because it is a Full-Reference (FR) metric, which requires as input both the original video and its impaired counterpart. On the other hand, No-Reference (NR) objective algorithms operate directly on the impaired video and are considerably faster, but loose out when it comes to accuracy. In this research, we assess a range of NR metrics, alongside a lightweight FR metric, using VQM as benchmark. Our study covers a range of methods, a diverse set of video types and encoding conditions, and a range of network impairment test-cases. We show the extent by which packet loss affects different video types, correlating the accuracy of NR metrics to the FR benchmark. Our study helps identifying the conditions under which simple metrics may be used effectively and indicates an avenue to control the quality of streaming systems in line with human perception.
Maria Torres Vega, Vittorio Sguazzo, Decebal Constantin Mocanu, Antonio Liotta
MoMM1
2014 When does lower bitrate give higher quality in modern video services?
abstract
Due to the difficulties on approximating the human perception with algorithms, increasing the users Quality of Experience (QoE) in modern video services is a challenging task. But more than that, prior to estimating QoE, it is important to know how different types of network impairments actually affect the video quality. This paper takes a closer look at the relation between the network quality of service (QoS) and the video QoE degradation. Using a sophisticated network emulation environment, we benchmark a range of video types and video quality levels under controlled network conditions. Our analysis shows that, along with a number of expected situations come also some counterintuitive QoS-to-QoE conditions. We discuss ways in which a better understanding of the mutual influence between networks and video streams could lead to more efficient utilization of the Internet.
Decebal Constantin Mocanu, Antonio Liotta, Arianna Ricci, Maria Torres Vega, Georgios Exarchakos
NOMS4