Sara Baldoni

dblp:230/0806 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-5642-3430ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 6 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Beyond Pixels: Assessing Image Quality through Semantic Information Loss
Annalisa Gallina, Sara Baldoni, Federica Battisti
QoMEX2
2026 When milliseconds matter: the impact of network latency on VR gaming experience
abstract
In recent years, Virtual Reality systems moved towards network-based applications. This implies that the network transmission performance directly impacts on users’ experience, thus potentially impairing users’ immersion, engagement, and effectiveness. In this work, we present a work-in-progress study involving an experimental campaign aimed at testing the impact of network latency on users’ Quality of Experience during Virtual Reality gaming. We present the designed experimental protocol and the preliminary results.
Sara Baldoni, Federica Battisti
IMX1
2026 Quality assessment of 3D reconstructed meshes: Bridging objective metrics, subjective perception, and behavioral cues
abstract
Assessing the quality of 3D reconstructed models remains a key challenge in multimedia applications, especially in the context of cultural heritage, where visual fidelity and perceptual realism are equally crucial. This study investigates how reconstruction parameters, as well as existing objective quality metrics, align with human perception. In addition, we analyze how perceived quality and user interaction are related. A dataset of 3D models was generated by varying the number of input images, mesh complexity, and texture resolution. Results from a subjective study show that texture resolution significantly affects perceived quality, whereas variations in number of images and mesh complexity have a limited impact. Furthermore, interaction behavior was found to vary with perceived quality, with participants spending more time and exploring larger viewing angles for models receiving higher scores. These findings highlight the need for perceptually grounded, interaction-aware evaluation methodologies and provide guidelines for future perceptual optimization of 3D reconstruction pipelines.
Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, David Barbero García, Daniel Berjón, Francisco Morán, Narciso García, Federica Battisti, Jesús Gutiérrez 0001, Marco Carli, Julián Cabrera
Signal Process. Image Commun.4
2026 Markerless emotion recognition from full-body movements for Social XR
abstract
In this work, an emotion recognition system for enhancing social XR applications is presented. Although several techniques for emotion recognition have been proposed in the literature, they either require invasive and advanced equipment or exploit facial expressions, speech excerpts, physiological data, and text. In this contribution, on the contrary, an approach for markerless emotion classification through body language is designed. More specifically, human movements are analyzed over time by extracting the skeleton joints in videos acquired by consumer cameras. A normalization procedure has been introduced to provide a depth-independent skeleton representation without distorting the skeleton shape. The performance of the proposed method have been assessed using a dataset of videos recorded from multiple points of view. An ad-hoc learning-based emotion classifier has been trained to recognize four emotions (happiness, boredom, interest, and disgust) achieving an average accuracy of 72.5%. The pre-processed dataset, code, and demo with pre-trained models are available at https://github.com/michaelneri/emotion-recognition-human-movements . • We define a human skeleton representation that does not change with the distance between the user and the camera. • We introduce a new approach for classifying emotions based exclusively on body movements. • We extend an existing dataset employing multiple points of view for classifying four emotional states: happiness, interest, boredom, and disgust.
Michael Neri, Sara Baldoni, Marco Carli, Federica Battisti
Signal Process. Image Commun.2
2025 Sphere-GAN: a GAN-based approach for saliency estimation in 360° videos
abstract
The recent success of immersive applications is pushing the research community to define new approaches to process 360° images and videos and optimize their transmission. Among these, saliency estimation provides a powerful tool that can be used to identify visually relevant areas and, consequently, adapt processing algorithms. Although saliency estimation has been widely investigated for 2D content, very few algorithms have been proposed for 360° saliency estimation. Towards this goal, we introduce Sphere-GAN, a saliency detection model for 360° videos that leverages a Generative Adversarial Network with spherical convolutions. Extensive experiments were conducted using a public 360° video saliency dataset, and the results demonstrate that Sphere-GAN outperforms state-of-the-art models in accurately predicting saliency maps.
Mahmoud Z. A. Wahba, Sara Baldoni, Federica Battisti
MMSP2
2025 Improved RAHT-based Compression of 3D Gaussian Splats
Annalisa Gallina, Giuseppe Valenzise, Sara Baldoni, Federica Battisti
PCS3
2025 Analysis of Objective 3D Mesh Quality Metrics for Cultural Heritage
abstract
Extended reality technologies are increasingly used in cultural heritage for preserving and accessing sites and artworks, where 3D model acquisition and rendering are key. Despite progress in reconstruction methodologies, a standardized approach to quality assessment is still missing. This study aims to evaluate objective quality metrics —both image-based and model-based, Full Reference and No Reference— applied to 3D models generated using the Structure from Motion algorithm. By varying parameters such as the number of images, number of triangles, and texture resolution, we examine the impact of these factors on metric outcomes, aiming to assess their reliability in cultural heritage applications.
Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, Jesús Gutiérrez 0001, Daniel Berjón, Francisco Morán, Federica Battisti, Narciso García, Marco Carli, Julián Cabrera
QoMEX4
2025 Movement- and Traffic-based User Identification in Commercial Virtual Reality Applications: Threats and Opportunities
abstract
With the unprecedented diffusion of virtual reality, the number of application scenarios is continuously growing. As commercial and gaming applications become pervasive, the need for the secure and convenient identification of users, often overlooked by the research in immersive media, is becoming more and more pressing. Networked scenarios such as Cloud gaming or cooperative virtual training and teleoperation require both a user-friendly and streamlined experience and user privacy and security. In this work, we investigate the possibility of identifying users from their movement patterns and data traffic traces while playing four commercial games, using a publicly available dataset. If, on the one hand, this paves the way for easy identification and automatic customization of the virtual reality content, it also represents a serious threat to users’ privacy due to network analysis-based fingerprinting. Based on this, we analyze the threats and opportunities for virtual reality users’ security and privacy.
Sara Baldoni, Salim Benhamadi, Federico Chiariotti, Michele Zorzi, Federica Battisti
VR1
2025 Histogram-based network traffic representation for anomaly detection through PCA
abstract
The constant increase of the number of connected devices, as well as of their heterogeneity, has greatly expanded the security threat landscape. For this reason, the prompt and effective detection of network traffic anomalies has become critical. In this work, we propose a new network traffic representation that aims at providing a compact and constantly updated summary of the current network condition. In addition, we propose an anomaly detection method based on the Principal Component Analysis of the aforementioned network representation. The proposed method exploits one-second time windows of network traffic, thus allowing an immediate reaction to anomalies. It is completely unsupervised, thus enabling the detection of zero-day attacks, and it has a low computational complexity, thus reducing the required capabilities of the monitoring nodes. The performance analysis showed that the proposed approach achieves comparable results with respect to state-of-the-art methods. • A new traffic representation providing updated and compact summaries of the network. • An unsupervised, low-complexity, and prompt anomaly detector based on PCA. • An in-depth comparison between the proposed approach and state-of-the-art methods.
Sara Baldoni, Federica Battisti
Comput. Networks1
2024 Quality of Experience for immersive media: from content creation to rendering
abstract
This paper presents the key phases that need to be addressed for enhancing Quality of Experience for immersive media. The overall immersive media pipeline is analyzed, from creation to rendering and fruition, in order to understand how each phase impacts on the user’s Quality of Experience. The relation and inter-dependencies between the different steps are highlighted and the current challenges are discussed.
Sara Baldoni
ISCC1
2024 Questset: A VR Dataset for Network and Quality of Experience Studies
abstract
The rapid development of Virtual Reality (VR) technology has led the industry and research community to look at its major challenges with increased interest. The main challenge in ensuring a high Quality of Experience (QoE) for users is represented by cybersickness, a phenomenon similar to motion sickness experienced by many VR users, while at the same time, the high data rates needed by VR require the definition of traffic models for network optimization. These two problems are intertwined, but have never been studied jointly before due to the lack of suitable datasets. In this paper, we present Questset, the first dataset designed for this purpose. Questset contains over 40 hours of VR traces from 70 users playing commercially available video games, and includes both traffic data for network optimization, and movement and user experience data for cybersickness analysis. Therefore, Questset represents an enabler to jointly address the main VR challenges in the near future.
Sara Baldoni, Federica Battisti, Federico Chiariotti, Fabio Mistrorigo, Alfi Baqiatus Shofi, Paolo Testolina, Alessandro Traspadini, Andrea Zanella, Michele Zorzi
MMSys1
2024 On the identification of the leading sensory cue in mulsemedia VR applications
abstract
This work aims to investigate the existence of a leading sensory cue in a mulsemedia Virtual Reality application involving three senses: vision, hearing, and touch. On the one hand this study can help in gaining insights into how the different senses contribute to mulsemedia applications in Virtual Reality. On the other, the identification of the leading cue could drive the optimization of Virtual Reality applications in terms of Quality of Experience, stimuli definition, and transmission requirements. In this paper, we present a mulsemedia experimental protocol for a subjective test aimed at identifying the material of a virtual object. We present the encountered challenges and describe and discuss the obtained results.
Anna Ferrarotti, Sara Baldoni, Marco Carli, Federica Battisti
QoMEX2
2024 Interaction goes virtual: towards collaborative XR
abstract
This demo presents an interactive communication system based on immersive media for analyzing the users’ Quality of Experience in collaborative tasks. The interaction between users is studied in an asymmetric scenario where a peer-to-peer communication has been set up between a PC and a Virtual Reality headset. Two application scenarios have been considered: a Block Building task and a Treasure Hunt game. The two users will cooperate to perform the two tasks. The goal is to study the relation between the type of transmitted information (i.e., audio and video or audio only) and the quality and quantity of interaction. During the demo, participants will have the opportunity to try one of the designed applications.
Federica Battisti, Anna Ferrarotti, Marco Carli, Sara Baldoni
IMX4
2024 Characterizing the Geometric Complexity of G-PCC Compressed Point Clouds
abstract
Measuring the complexity of visual content is crucial in various applications, such as selecting sources to test processing algorithms, designing subjective studies, and efficiently determining the appropriate encoding parameters and bandwidth allocation for streaming. While spatial and temporal complexity measures exist for 2D videos, a geometric complexity measure for 3D content is still lacking. In this paper, we present the first study to characterize the geometric complexity of 3D point clouds. Inspired by existing complexity measures, we propose several compression-based definitions of geometric complexity derived from the rate-distortion curves obtained by compressing a dataset of point clouds using G-PCC. Additionally, we introduce density-based and geometry-based descriptors to predict complexity. Our initial results show that even simple density measures can accurately predict the geometric complexity of point clouds.
Annalisa Gallina, Hadi Amirpour, Sara Baldoni, Giuseppe Valenzise, Federica Battisti
VCIP3
2024 Definition of guidelines for virtual reality application design based on visual attention
abstract
Abstract In virtual reality applications, head-mounted displays allow users to explore virtual surroundings, thus creating a high sense of immersion. However, due to the novelty of the technology and the possibility of freely enjoying a $$360^\circ $$ 360 ∘ virtual world, users can get distracted and divert their attention from the content of the application. In this work, we define a set of guidelines for the design of virtual reality applications for enhancing the users’ attention. To the best of our knowledge, this is one of the first attempts to provide general guidelines for virtual application design based on visual attention. More specifically, we analyze the different categories of factors that contribute to the user’s responsiveness and define a set of experiments for measuring the user’s promptness with respect to visual stimuli with different features and in the presence of audio/visual distractions. Experimental tests have been carried out with 36 volunteers. The users’ reaction time has been recorded and the performed analysis allowed the definition of a set of guidelines based on individual, operational, and technological factors for the design of virtual reality applications optimized in terms of user attention. In particular, statistical tests demonstrated that the presence of distractions leads to significantly different reaction times with respect to the case of no distractions, and that users belonging to different age intervals have significantly different behaviors. Moreover, the optimal placement of objects has been identified and the impact of cybersickness has been analyzed.
Sara Baldoni, M. Saifeddine Hadj Sassi, Marco Carli, Federica Battisti
Multim. Tools Appl.1
2024 Stress Assessment for Augmented Reality Applications Based on Head Movement Features
abstract
Augmented reality is one of the enabling technologies of the upcoming future. Its usage in working and learning scenarios may lead to a better quality of work and training by helping the operators during the most crucial stages of processes. Therefore, the automatic detection of stress during augmented reality experiences can be a valuable support to prevent consequences on people's health and foster the spreading of this technology. In this work, we present the design of a non-invasive stress assessment approach. The proposed system is based on the analysis of the head movements of people wearing a Head Mounted Display while performing stress-inducing tasks. First, we designed a subjective experiment consisting of two stress-related tests for data acquisition. Then, a statistical analysis of head movements has been performed to determine which features are representative of the presence of stress. Finally, a stress classifier based on a combination of Support Vector Machines has been designed and trained. The proposed approach achieved promising performances thus paving the way for further studies in this research direction.
Anna Ferrarotti, Sara Baldoni, Marco Carli, Federica Battisti
IEEE Trans. Vis. Comput. Graph.2
2018 A feature-based approach for saliency estimation of omni-directional images
Federica Battisti, Sara Baldoni, Michele Brizzi, Marco Carli
Signal Process. Image Commun.2