VLDB 2026 Research / reviewers in the wild / expert
Marco Carli
dblp:56/4366
· DBLP profile ↗
59ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-7489-3767ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 3 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 12 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Designing Virtual Reality-Mediated Instructional Activities: An Operational Framework for Educational Practitioners
Angela D'Angelo, Marco Carli |
CSEDU (2) | 2 |
| 2026 | Beyond Digital Replication: Designing an Interactive and Multisensory VR Museum Experience
Anna Ferrarotti, Manuela Viscontini, Alessia Lorenzi, Ivan Legnaioli, Vincenzo Mauro, Valentina Cancilla, Andrea di Meo, Federico Capriuoli, Matias Guerra, Marcos Valdes, Marco Carli |
IMX | 11 |
| 2026 | The Impact of Haptics on User Presence in Virtual RealityabstractThis work explores the integration of advanced haptic feedback within an immersive Virtual Reality educational framework, specifically applied to planetary science. The purpose of this demo is to showcase how haptic technologies can enhance learning experiences in VR-based learning environments. The proposed experience is structured around an effective storyline and utilizes state-of-the-art haptic gloves to present a four-stage interactive lesson on Mercury. The system leverages thermal modulation, vibrotactile texturing, and force feedback to allow learners to physically perceive the planet’s extreme temperature gradients and cratered topography. Matteo Pestelli, Erica Rocchi, Marco Carli |
IMX | 3 |
| 2026 | Quality assessment of 3D reconstructed meshes: Bridging objective metrics, subjective perception, and behavioral cuesabstractAssessing the quality of 3D reconstructed models remains a key challenge in multimedia applications, especially in the context of cultural heritage, where visual fidelity and perceptual realism are equally crucial. This study investigates how reconstruction parameters, as well as existing objective quality metrics, align with human perception. In addition, we analyze how perceived quality and user interaction are related. A dataset of 3D models was generated by varying the number of input images, mesh complexity, and texture resolution. Results from a subjective study show that texture resolution significantly affects perceived quality, whereas variations in number of images and mesh complexity have a limited impact. Furthermore, interaction behavior was found to vary with perceived quality, with participants spending more time and exploring larger viewing angles for models receiving higher scores. These findings highlight the need for perceptually grounded, interaction-aware evaluation methodologies and provide guidelines for future perceptual optimization of 3D reconstruction pipelines. Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, David Barbero García, Daniel Berjón, Francisco Morán, Narciso García, Federica Battisti, Jesús Gutiérrez 0001, Marco Carli, Julián Cabrera |
Signal Process. Image Commun. | 11 |
| 2026 | Markerless emotion recognition from full-body movements for Social XRabstractIn this work, an emotion recognition system for enhancing social XR applications is presented. Although several techniques for emotion recognition have been proposed in the literature, they either require invasive and advanced equipment or exploit facial expressions, speech excerpts, physiological data, and text. In this contribution, on the contrary, an approach for markerless emotion classification through body language is designed. More specifically, human movements are analyzed over time by extracting the skeleton joints in videos acquired by consumer cameras. A normalization procedure has been introduced to provide a depth-independent skeleton representation without distorting the skeleton shape. The performance of the proposed method have been assessed using a dataset of videos recorded from multiple points of view. An ad-hoc learning-based emotion classifier has been trained to recognize four emotions (happiness, boredom, interest, and disgust) achieving an average accuracy of 72.5%. The pre-processed dataset, code, and demo with pre-trained models are available at https://github.com/michaelneri/emotion-recognition-human-movements . • We define a human skeleton representation that does not change with the distance between the user and the camera. • We introduce a new approach for classifying emotions based exclusively on body movements. • We extend an existing dataset employing multiple points of view for classifying four emotional states: happiness, interest, boredom, and disgust. Michael Neri, Sara Baldoni, Marco Carli, Federica Battisti |
Signal Process. Image Commun. | 3 |
| 2025 | 5GVIREH: a 5G-enabled Virtual Reality based solution for telerehabilitationabstractIntegrating 5G technology with Virtual Reality creates an advanced human-machine interface applicable to various fields. With its high-speed, stable connection, 5G enables seamless real-time user communication. Virtual Reality enhances personalized interaction and allows continuous monitoring without complex additional equipment. This study presents a use case where a clinician, connected through a PC, interacts with a patient using a Virtual Reality headset during physiotherapy sessions. The goal is to showcase the potential of 5G and Virtual Reality in enhancing telerehabilitation. While initial tests focused on patients recovering from rotator cuff surgery, the system can be easily adapted to support a wide range of treatment and rehabilitation protocols. Anna Ferrarotti, Michele Brizzi, Erica Rocchi, Alessia Fabrizio, Arianna Carnevale, Umile Giuseppe Longo, Marco Carli |
QoMEX | 7 |
| 2025 | Analysis of Objective 3D Mesh Quality Metrics for Cultural HeritageabstractExtended reality technologies are increasingly used in cultural heritage for preserving and accessing sites and artworks, where 3D model acquisition and rendering are key. Despite progress in reconstruction methodologies, a standardized approach to quality assessment is still missing. This study aims to evaluate objective quality metrics —both image-based and model-based, Full Reference and No Reference— applied to 3D models generated using the Structure from Motion algorithm. By varying parameters such as the number of images, number of triangles, and texture resolution, we examine the impact of these factors on metric outcomes, aiming to assess their reliability in cultural heritage applications. Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, Jesús Gutiérrez 0001, Daniel Berjón, Francisco Morán, Federica Battisti, Narciso García, Marco Carli, Julián Cabrera |
QoMEX | 10 |
| 2024 | On the identification of the leading sensory cue in mulsemedia VR applicationsabstractThis work aims to investigate the existence of a leading sensory cue in a mulsemedia Virtual Reality application involving three senses: vision, hearing, and touch. On the one hand this study can help in gaining insights into how the different senses contribute to mulsemedia applications in Virtual Reality. On the other, the identification of the leading cue could drive the optimization of Virtual Reality applications in terms of Quality of Experience, stimuli definition, and transmission requirements. In this paper, we present a mulsemedia experimental protocol for a subjective test aimed at identifying the material of a virtual object. We present the encountered challenges and describe and discuss the obtained results. Anna Ferrarotti, Sara Baldoni, Marco Carli, Federica Battisti |
QoMEX | 3 |
| 2024 | Interaction goes virtual: towards collaborative XRabstractThis demo presents an interactive communication system based on immersive media for analyzing the users’ Quality of Experience in collaborative tasks. The interaction between users is studied in an asymmetric scenario where a peer-to-peer communication has been set up between a PC and a Virtual Reality headset. Two application scenarios have been considered: a Block Building task and a Treasure Hunt game. The two users will cooperate to perform the two tasks. The goal is to study the relation between the type of transmitted information (i.e., audio and video or audio only) and the quality and quantity of interaction. During the demo, participants will have the opportunity to try one of the designed applications. Federica Battisti, Anna Ferrarotti, Marco Carli, Sara Baldoni |
IMX | 3 |
| 2024 | Definition of guidelines for virtual reality application design based on visual attentionabstractAbstract In virtual reality applications, head-mounted displays allow users to explore virtual surroundings, thus creating a high sense of immersion. However, due to the novelty of the technology and the possibility of freely enjoying a $$360^\circ $$ 360 ∘ virtual world, users can get distracted and divert their attention from the content of the application. In this work, we define a set of guidelines for the design of virtual reality applications for enhancing the users’ attention. To the best of our knowledge, this is one of the first attempts to provide general guidelines for virtual application design based on visual attention. More specifically, we analyze the different categories of factors that contribute to the user’s responsiveness and define a set of experiments for measuring the user’s promptness with respect to visual stimuli with different features and in the presence of audio/visual distractions. Experimental tests have been carried out with 36 volunteers. The users’ reaction time has been recorded and the performed analysis allowed the definition of a set of guidelines based on individual, operational, and technological factors for the design of virtual reality applications optimized in terms of user attention. In particular, statistical tests demonstrated that the presence of distractions leads to significantly different reaction times with respect to the case of no distractions, and that users belonging to different age intervals have significantly different behaviors. Moreover, the optimal placement of objects has been identified and the impact of cybersickness has been analyzed. Sara Baldoni, M. Saifeddine Hadj Sassi, Marco Carli, Federica Battisti |
Multim. Tools Appl. | 3 |
| 2024 | Speaker Distance Estimation in Enclosures From Single-Channel AudioabstractDistance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where distances are discretized into distinct categories, enabling smoother model training and achieving higher accuracy but imposing restrictions on the precision of the obtained sound source position. Towards this direction, in this paper we propose a novel approach for continuous distance estimation from audio signals using a convolutional recurrent neural network with an attention module. The attention mechanism enables the model to focus on relevant temporal and spectral features, enhancing its ability to capture fine-grained distance-related information. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using audio recordings in controlled environments with three levels of realism (synthetic room impulse response, measured response with convolved speech, and real recordings) on four datasets (our synthetic dataset, QMULTIMIT, VoiceHome-2, and STARSS23). Experimental results show that the model achieves an absolute error of 0.11 meters in a noiseless synthetic scenario. Moreover, the results showed an absolute error of about 1.30 meters in the hybrid scenario. The algorithm's performance in the real scenario, where unpredictable environmental factors and noise are prevalent, yields an absolute error of approximately 0.50 meters. For reproducible research purposes we make model, code, and synthetic datasets available at https://github.com/michaelneri/audio-distance-estimation Michael Neri, Archontis Politis, Daniel Krause 0001, Marco Carli, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Stress Assessment for Augmented Reality Applications Based on Head Movement FeaturesabstractAugmented reality is one of the enabling technologies of the upcoming future. Its usage in working and learning scenarios may lead to a better quality of work and training by helping the operators during the most crucial stages of processes. Therefore, the automatic detection of stress during augmented reality experiences can be a valuable support to prevent consequences on people's health and foster the spreading of this technology. In this work, we present the design of a non-invasive stress assessment approach. The proposed system is based on the analysis of the head movements of people wearing a Head Mounted Display while performing stress-inducing tasks. First, we designed a subjective experiment consisting of two stress-related tests for data acquisition. Then, a statistical analysis of head movements has been performed to determine which features are representative of the presence of stress. Finally, a stress classifier based on a combination of Support Vector Machines has been designed and trained. The proposed approach achieved promising performances thus paving the way for further studies in this research direction. Anna Ferrarotti, Sara Baldoni, Marco Carli, Federica Battisti |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Objective Metrics Definition for QoE Assessment for Extended Reality ApplicationsabstractThe increasing advancement of Virtual and Augmented Reality technologies opens new research perspectives in various fields. While traditional multimedia content quality assessment has been extensively investigated, different issues need to be addressed for these novel technologies. In particular, evaluating the quality of experience must consider several aspects, involving, on the one hand, the quality of the displayed multimedia content and, on the other hand, human factors. Due to its inherent subjectivity, defining objective metrics for Quality of Experience is complex. This paper aims to frame the problem of objective Quality of Experience assessment and propose research directions to be pursued in this field. Anna Ferrarotti, Marco Carli |
IMX | 2 |
| 2023 | Artificial Intelligence Techniques for Quality Assessments of Immersive MultimediaabstractArtificial Intelligence techniques are being applied in the quality assessment of immersive multimedia content, such as virtual and augmented reality scenarios. The immersive nature of these applications poses a unique challenge to traditional quality assessment methods. In fact, estimating user acceptance of immersive technologies is complex due to multiple aspects, such as usability, enjoyment, and cyber sickness. Artificial Intelligence-based approaches offer a promising solution to this problem, enabling objective evaluations of immersive multimedia such as spatial audios, point clouds, and light field images. This work presents an overview of different artificial intelligence techniques that have been used for quality assessments of immersive multimedia content, including machine learning algorithms, deep learning, and computer vision. The advantages of these techniques and some examples of practical application are provided. Future works are presented, underlining the possible outcomes of a Ph.D. study in this field. Michael Neri, Marco Carli |
IMX | 2 |
| 2023 | Selective video enhancement in the Laguerre-Gauss domainabstractTraditional image and video enhancement techniques do not consider the subjective preferences of the user, who might be interested in modifying or highlighting specific elements of a scene. In fact, most image or video post-processing techniques are applied to the entire image or frame. In this work, a framework for selective video enhancement is presented. It is based on an adaptive multi-resolution edge enhancement technique performed in the Laguerre–Gauss complex wavelet domain. In the proposed scheme, edges are selectively enhanced or attenuated according to the inputs of the user, taking into account the characteristics of the content, the response of the human visual system, and the masking effects induced by textured background. Two scenarios have been implemented and tested: (i) an interactive video editing system, where the user selects the objects to enhance through a graphical interface and the system automatically propagates the selection through the following video frames belonging to the same shot, (ii) a perceptually-driven technique to improve the quality of RGB plus Depth videos, which uses depth maps as additional features to guide the enhancement. Michele Brizzi, Federica Battisti, Marco Carli, Alessandro Neri 0001 |
Signal Process. Image Commun. | 3 |
| 2023 | A CNN-based no reference image quality metric exploiting content saliencyabstractAssessing the quality of images is a challenging task. To achieve this goal, images must be evaluated by a pool of subjects following a well-defined protocol or an objective quality metric must be defined. In this work, an objective quality metric based on deep neural network is proposed. The metric takes into account the human vision system by computing the saliency map and natural scene statistics features of the image under test. The neural network is composed by two modules: the convolutional layers and the regression units. The first one is trained by using preprocessed distorted images. The feature weights of the first module are smoothed by exploiting the estimated saliency map. The latter module is fit with the ground truth quality scores of the input image and the scaled feature weights obtained from first module by using visual sensitivity factor of image obtained using natural scene statistics features. The performances of the proposed metric have been evaluated by using four datasets: LIVEIQA, TID2013, CSIQ, and KADID10K. The achieved results show the effectiveness of the proposed system in closely matching the predicted quality scores with the ground truth ones. Kamal Lamichhane, Marco Carli, Federica Battisti |
Signal Process. Image Commun. | 2 |
| 2022 | Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919abstractRecently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community. Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García |
IEEE Trans. Multim. | 26 |
| 2021 | Analysis of the influence of human faces for the estimation of salience in omnidirectional imagesabstractIn this contribution a study dedicated to understanding the influence of the presence of human faces in a 360° image on human perception is presented. Extensive research on saliency estimation in 2D images has shown that the presence of faces attracts human attention. Following these studies, recent 2D image quality assessment methods exploit face detection systems in their models. The application of these concepts to 360° image is not straightforward. Furthermore, existing literature lacks a comparative study between the performance of face detection algorithms on various types of images (2D, fisheye, and omnidirectional) and how detected faces affect the procedure of saliency estimation. In this direction, we analyze the importance of person faces in a scene, by performing a set of subjective tests. From the performed analysis, it results that, even in 360° images, human faces represent an important factor for image saliency. However, giving equal importance to all the detected faces does not lead to a better saliency estimation. Therefore, in this work, a study on the possible use of the face detector in estimating the salience of the 360° image is performed. Pramit Mazumdar, Giuliano Arru, Marco Carli, Federica Battisti |
MMSP | 3 |
| 2021 | Exploiting saliency in quality assessment for light field imagesabstractThe evaluation of the quality of light field images is a demanding task given the peculiarities of this media. In the literature, attempts to assess quality have been done by considering specific coding approaches or visualization techniques. In this paper we intend to i) investigate whether the distortions of the light field are reflected in the distortion of the saliency map and ii) propose a metric for image quality assessment of light fields based on a convolutional neural network that exploits the measure of the distortion of the saliency map. In our tests, the annotated SMART dataset has been used. The achieved results confirm the importance of saliency for improving the performance of quality metrics. Kamal Lamichhane, Federica Battisti, Pradip Paudyal, Marco Carli |
PCS | 4 |
| 2020 | A Physiology-based Driver Readiness Estimation Model for Tuning ISO 26262 ControllabilityabstractWhen a hazardous situation approaches, the semi-autonomous vehicle opts for the driver as a fallback solution, unaware of the driver's readiness. During such a situation, autonomy misuse can occur when a driver becomes over-reliant on autonomous driving. For handling the hazardous event, controllability is paramount. We postulate that semi-autonomous vehicles decline their consideration in understanding the drivers' focus on the vehicle and the road. To examine the drivers' focus on the vehicle and the road we uphold that the vehicle must initiate exploring the drivers' situation awareness for the readiness, which could feasibly tune the ISO 26262 controllability. In this paper, we propose a physiology-based driver situation awareness for the readiness model through the driver's stress and drowsiness estimation. In addition, we boost the situation awareness for the readiness of the driver by enabling frequent interaction between the driver and the vehicle managing system. Moses Mariajoseph, Barbara Gallina, Marco Carli, Daniele Bibbo |
VTC Spring | 3 |
| 2019 | A New Security Approach in Telecom Infrastructures: The RESISTO ConceptabstractCommunications play a fundamental role in the economic and social well-being of the citizens and on operations of most of the critical infrastructures (CIs). Extreme weather events, natural disasters and criminal attacks represent a challenge due to their increase in frequency and intensity requiring smarter resilience of the Communication CIs, which are extremely vulnerable due to the ever-increasing complexity of the architecture also in light of the evolution towards 5G, the extensive use of programmable platforms and exponential growth of connected devices. In this paper, we present the aim of RESISTO H2020 EU-funded project, which constitutes an innovative solution for Communication CIs holistic situation awareness and enhanced resilience. Maria Belesioti, Rodoula Makri, Mirjam Fehling-Kaschek, Marco Carli, Alexandros Kostopoulos, Ioannis P. Chochliouros, Alberto Neri, Federico Frosali |
DCOSS | 4 |
| 2019 | Contactless approach for heart rate estimation for QoE assessment
Mattia Bonomi, Federica Battisti, Giulia Boato, Miguel Barreda-Ángeles, Marco Carli, Patrick Le Callet |
Signal Process. Image Commun. | 5 |
| 2019 | A Maximum Likelihood Approach for Depth Field Estimation Based on Epipolar Plane ImagesabstractIn this paper, a multi-resolution method for depth estimation from dense image arrays is presented. Recent progress in consumer electronics has enabled the development of low cost hand-held plenoptic cameras. In these systems, multiple views of a scene are captured in a single shot by means of a micro-lens array placed on the focal point of the first camera lens, in front of the imaging sensor. These views can be processed jointly to obtain accurate depth maps. In this contribution, to reduce the computational complexity associated to global optimization schemes based on match cost functions, we make a local estimate based on the maximization of the total log-likelihood spatial density aggregated along the epipolar lines corresponding to each view pair. This method includes the local maximum likelihood estimation of the depth field based on epipolar plane images. To face the potential accuracy losses associated to the ambiguity problem that arises in flat surface regions while preserving bandwidth in correspondence of the edges, we adopt a multi-resolution scheme. In practice, the depth map resolution is reduced in regions where maximizing the higher resolution functional is ill-conditioned. The main benefits of the proposed system are in a reduced computational complexity and a high accuracy of the estimated depth. Experimental results show that the proposed scheme represents a good tradeoff among accuracy, robustness, and discontinuities handling. Alessandro Neri 0001, Marco Carli, Federica Battisti |
IEEE Trans. Image Process. | 2 |
| 2018 | A non-intrusive system for seated posture identificationabstractIn this contribution a system for seated posture identification is presented. The assessment tools is based on an office-chair equipped with sensors. In more details, a set of textile pressure sensors has been placed on a chair both on the chair backrest and on the seat. The position of the sensors has been selected for maximizing the possibility of sensing minimum variations of the subject's posture. To validate the system, an extensive subjective experiment has been performed in which the subject undergoes an increasing stress-level test. The collected results show that this instrument is effective in assessing the attention/fatigue of a subject in seating condition by the analysis of body posture. Daniele Bibbo, Federica Battisti, Silvia Conforto, Marco Carli |
HealthCom | 4 |
| 2018 | A feature-based approach for saliency estimation of omni-directional images
Federica Battisti, Sara Baldoni, Michele Brizzi, Marco Carli |
Signal Process. Image Commun. | 4 |
| 2017 | Enhancing audio surveillance with hierarchical recurrent neural networksabstractThe need for effective and reliable surveillance techniques is getting nowadays more and more of primary importance, especially in the actual scenario in which safety and security have become a priority. While classical techniques rely on video-based surveillance systems, such as Close-Circuit television, many studies show that also the audio signal can be effectively used for these purposes. There are many characteristics that make the audio signal particularly suited for this task and, above all, the fact that the analysis of the audio signal can greatly improve thanks to the introduction of automatic classification. Recently, a large focus has been on the use of Deep Neural Networks for classifying audio data and, in this work, we aim to test their performance in the audio surveillance field. In this contribution we propose an algorithm for audio events detection in noisy environments based on the use of deep recurrent neural network. The achieved results show satisfactory and improved performances with respect to state-of-the-art techniques. Federico Colangelo, Federica Battisti, Marco Carli, Alessandro Neri 0001, Francesco Calabrò |
AVSS | 3 |
| 2017 | Effect of visualization techniques on subjective quality of light field imagesabstractLight Field imaging derives from the fundamentals of light field sampling, where the spatial information about a scene can be captured with angular information. That is, the Light Field imaging is based on a camera recording information about the intensity of light from the scene and about the direction of the light rays. The acquired data can be shown to the user in different ways, such as image with digitally extended depth of field, 3D, parallax, and 360 degree display. In this contribution, the impact of different rendering techniques on the Quality of Experience is addressed. The achieved results show that the visualization techniques may have different impact on the perceived quality even when the same content is considered. Pradip Paudyal, Federica Battisti, Marco Carli |
ICIP | 3 |
| 2017 | Unsupervised video orchestration based on aesthetic featuresabstractIn this work, the problem of dynamic video scene creation obtained by combining information extracted from multiple video sequences is considered. The main novelty of the proposed approach relies on the use of aesthetic features for automatically aggregating the inputs from different cameras in a unique video. While prior methodologies have separately addressed the issues of aesthetic feature extraction from videos and video orchestration, in this work we exploit selected features of a scene for automatically selecting the shots being characterized by the best aesthetic score. In order to evaluate the effectiveness of the proposed method, a subjective experiment has been carried out with experts from the audiovisual field. The achieved results are encouraging and show that there is space for improving the performances. Alessandro Neri 0001, Federica Battisti, Federico Colangelo, Marco Carli |
ISCAS | 4 |
| 2017 | Characterization and selection of light field content for perceptual assessmentabstractLight field technology may have a positive impact on several multimedia applications thanks to novel ways to explore the captured scenes, such as changing the parallax (horizontally and vertically) and refocusing the content. These innovative use cases require new considerations that affect the whole processing chain, from content acquisition to visualization, as well as the methodologies for quality evaluation. In particular, capturing and selecting the appropriate content is crucial for a successful development and evaluation of audiovisual technologies. Thus, this paper presents a framework addressing the reconsideration of the space of attributes for an adequate characterization of light field data. Firstly, an exhaustive characterization of light field content is described, based on various particular features, including depth and refocusing properties to traditional spatial, temporal, and color indicators. Then, based on this characterization, specific techniques are proposed for an effective selection of light field content for perceptual quality assessment. Pradip Paudyal, Jesús Gutiérrez 0001, Patrick Le Callet, Marco Carli, Federica Battisti |
QoMEX | 4 |
| 2017 | Securing cyber physical systems from injection attacks by exploiting random sequencesabstractThe security of Cyber Physical Systems is a key factor since these architectures are being applied in many critical scenarios, such as power or water plants, hospitals, etc. In the state of the art, several approaches have been proposed to deal with the possible attacks that could be inferred to the Cyber Physical Systems. This contribution proposes a method for counter fighting the injection of tampered data in the communication channel. If this attack is not timely detected, it may result in severe disruption of the system or even in its complete damage. The proposed approach is based on coding the physical output of the system through permutation matrices whose scheme varies based on a randomly generated sequence. The strength of this method is in reducing the possibility of a successful injection attack while limiting the computational complexity. The experimental tests prove the effectiveness of the proposed scheme in detecting the performed attack while granting the real time constraints of the Cyber Physical Systems. Federica Battisti, Marco Carli, Federica Pascucci |
WiMob | 2 |
| 2016 | SMART: a light field image quality datasetabstractIn this contribution, the design of a Light Field image dataset is presented. It can be useful for design, testing, and benchmarking Light Field image processing algorithms. As first step, image content selection criteria have been defined based on selected image quality key-attributes, i.e. spatial information, colorfulness, texture key features, depth of field, etc. Next, image scenes have been selected and captured by using the Lytro Illum Light Field camera. Performed analysis shows that the proposed set of images is sufficient for addressing a wide range of attributes relevant for assessing Light Field image quality. Pradip Paudyal, Roger Olsson, Mårten Sjöström, Federica Battisti, Marco Carli |
MMSys | 5 |
| 2016 | Free viewpoint video quality assessment based on morphological multiscale metricsabstractIn this paper, two image quality metrics based on morphological multiscale decompositions have been applied for the evaluation of free viewpoint video sequences. These sequences are synthesized using decompressed depth maps in the Depth-Image-Based Rendering synthesis process. Since the synthesis introduces edge distortion in the synthesized image/video, morphological filters are used for their ability to maintain important geometric information (i.e., edges) across different resolution levels. The edges displacement in different resolution scales is evaluated by means of the Mean Square Error. The adopted metrics show higher correlation with human judgment than state-of-the-art image quality measures used in this context. Dragana Sandic-Stankovic, Federica Battisti, Dragan Kukolj, Patrick Le Callet, Marco Carli |
QoMEX | 5 |
| 2016 | Impact of video content and transmission impairments on quality of experience
Pradip Paudyal, Federica Battisti, Marco Carli |
Multim. Tools Appl. | 3 |
| 2015 | A multi-resolution approach to depth field estimation in dense image arraysabstractIn this paper a multi-resolution depth field estimation algorithm for plenoptic cameras is presented. To face the potential accuracy losses originated from ambiguity problems arising in flat surface regions, still preserving bandwidth in correspondence of edges, a multi-resolution scheme is proposed. The achieved results show that the proposed local optimization method outperforms state of the art more complex global optimization based methods. Alessandro Neri 0001, Marco Carli, Federica Battisti |
ICIP | 2 |
| 2015 | Objective image quality assessment of 3D synthesized views
Federica Battisti, Emilie Bosc, Marco Carli, Patrick Le Callet, Simone Perugia |
Signal Process. Image Commun. | 3 |
| 2015 | Image database TID2013: Peculiarities, results and perspectivesabstractThis paper describes a recently created image database, TID2013, intended for evaluation of full-reference visual quality assessment metrics. With respect to TID2008, the new database contains a larger number (3000) of test images obtained from 25 reference images, 24 types of distortions for each reference image, and 5 levels for each type of distortion. Motivations for introducing 7 new types of distortions and one additional level of distortions are given; examples of distorted images are presented. Mean opinion scores (MOS) for the new database have been collected by performing 985 subjective experiments with volunteers (observers) from five countries (Finland, France, Italy, Ukraine, and USA). The availability of MOS allows the use of the designed database as a fundamental tool for assessing the effectiveness of visual quality. Furthermore, existing visual quality metrics have been tested with the proposed database and the collected results have been analyzed using rank order correlation coefficients between MOS and considered metrics. These correlation indices have been obtained both considering the full set of distorted images and specific image subsets, for highlighting advantages and drawbacks of existing, state of the art, quality metrics. Approaches to thorough performance analysis for a given metric are presented to detect practical situations or distortion types for which this metric is not adequate enough to human perception. The created image database and the collected MOS values are freely available for downloading and utilization for scientific purposes. Nikolay N. Ponomarenko, Lina Jin, Oleg Ieremeiev, Vladimir Lukin 0001, Karen Egiazarian, Jaakko Astola, Benoît Vozel, Kacem Chehdi, Marco Carli, Federica Battisti, C.-C. Jay Kuo |
Signal Process. Image Commun. | 9 |
| 2014 | Design of a Non-intrusive Augmented Trumpet
Claudia Rinaldi, Federica Battisti, Marco Carli, Luigi Pomante |
ArtsIT | 3 |
| 2014 | Exploiting perceptual quality issues in countering SIFT-based Forensic methodsabstractScale Invariant Feature Transform (SIFT) has been widely employed in several image application domains, including Image Forensics (e.g. detection of copy-move forgery or near duplicates). Recently, a number of methods allowing to remove SIFT keypoints from an original image have been devised studying the problem of SIFT security against malicious procedures. Such techniques are quite effective in producing an attacked image with very few (or no) keypoints, but at the expense of an image distortion. Final perceptual quality has been taken in account very roughly so far. In this paper, effectiveness of the attacking methods is evaluated also from the side of perceptual image quality; a new version of a SIFT keypoint removal method, based on a perceptual metric, is presented and an extended series of perceptive experiments is reported. Irene Amerini, Federica Battisti, Roberto Caldelli, Marco Carli, Andrea Costanzo |
ICASSP | 4 |
| 2014 | A joint routing and localization algorithm for emergency scenario
Marco Carli, Stefano Panzieri, Federica Pascucci |
Ad Hoc Networks | 1 |
| 2013 | A New Color Image Database TID2013: Innovations and Results
Nikolay N. Ponomarenko, Oleg Ieremeiev, Vladimir Lukin 0001, Lina Jin, Karen Egiazarian, Jaakko Astola, Benoît Vozel, Kacem Chehdi, Marco Carli, Federica Battisti, C.-C. Jay Kuo |
ACIVS | 9 |
| 2011 | A commutative digital image watermarking and encryption method in the tree structured Haar transform domain
Michela Cancellaro, Federica Battisti, Marco Carli, Giulia Boato, Francesco G. B. De Natale, Alessandro Neri 0001 |
Signal Process. Image Commun. | 3 |
| 2010 | Near lossless reversible data hiding based on adaptive predictionabstractIn this paper we present a new near lossless reversible watermarking algorithm using adaptive prediction for embedding. The prediction is based on directional first-order differences of pixel intensities within a suitably selected neighborhood. The proposed scheme results to be computationally efficient and allows achieving high embedding capacity while preserving a high image quality. Extensive experimental results demonstrate the effectiveness of the proposed approach. Valentina Conotter, Giulia Boato, Marco Carli, Karen Egiazarian |
ICIP | 3 |
| 2010 | Impact of contrast modification on human feeling: an objective and subjective assessmentabstractImages are powerful means of communication. Adding images to documents, websites, magazines, helps attracting the attention and providing an immediate feeling about the content of the document itself. At the same time, images can be used to influence the attitude of the reader. In the digital era it is easier than ever to create and share multimedia documents making extensive use of visual data. Furthermore, it is extremely easy to modify and adapt the images in order to make them more suitable to convey a concept. These modifications may simply consist in focusing the attention on a specific part of the image, to heavier manipulations such as changing the visual appearance or the contents to bias the opinion of the observer. This creates a rising interest for the availability of automatic blind tools able to detect possible modifications of images, and to correlate these modifications with the relevant impact on the viewer. This paper proposes an example of application of these concepts that exploits a tool for automatic detection of contrast modifications and analyzes their impact on human feeling. The results of instrumental and subjective studies are presented and discussed. Pamela Zontone, Marco Carli, Giulia Boato, Francesco G. B. De Natale |
ICIP | 2 |
| 2010 | Localization services in hybrid self-organizing networksabstractIn this paper we present a localization technique for tracking and positioning services in self-organizing hybrid networks, where both indoor and outdoor scenarios coexist. In the proposed framework, outdoor anchor nodes act as reference nodes for estimating the position of mobile nodes, which are moving in a neighboring indoor region. The mobile nodes share some location information for estimating their own position by using flooding communication scheme. Communications between mobile nodes are allowed by IEEE 802.16e technology. The outdoor reference nodes are equipped with Global Positioning Service network interface card. The cooperation among mobile and reference nodes, and the self-organizing network behavior guarantee fast positioning estimate. The accuracy of localization measurement - in terms of position and speed uncertainty - is assessed through the Dilution of Precision factor. Moreover, to minimize position and speed estimation errors of each node, an Extended Kalman Filtering technique is adopted. Simulation results demonstrate the effectiveness of the proposed approach. Anna Maria Vegni, Marco Carli, Alessandro Neri 0001 |
IPIN | 2 |
| 2009 | Quality evaluation of motion estimation algorithms based on structural distortionsabstractIn this paper a methodology for understanding the effectiveness of motion estimation techniques is presented. Unlike other performances evaluation systems, that are based on measuring the errors between the actual and the predicted displacements, the proposed technique is inspired to the Human Visual System. More in detail, the perceptual impact of geometric distortions induced by non accurate motion estimation is considered by means of an objective measure of the perceived distortion impact. Some of the most common block-based motion estimation algorithms have been tested. For each of them the performances have been evaluated by comparing the proposed metric with the state of the art metrics. A subjective experiment has been performed to assess the effectiveness of the estimation algorithms from a perceptual point of view. The obtained results show that the scores obtained with the tested metrics generally do not match with the perceived quality, while the proposed methodology does. Therefore, the presented tool can be used in the design and in the verification of a generic motion estimation algorithm. Angela D'Angelo, Marco Carli, Mauro Barni |
MMSP | 2 |
| 2009 | Markerless Human Motion Analysis in Gauss-Laguerre Transform Domain: An Application to Sit-To-Stand in Young and Elderly PeopleabstractA markerless computer vision technique specifically designed to track natural elements on the human body surface is presented. The method implements the estimate of translation, rotation, and scaling by means of a maximum likelihood approach carried out in the Gauss-Laguerre transform domain. The approach is particularly suitable for human movement analysis in clinical contexts, where kinematics is at present performed by means of marker-based systems. Specific drawbacks of these latter systems, such as the burden of time for marker placement and the intrinsic intrusive nature, would be removed by the proposed method. Experimental results in terms of tracking performance are obtained by analyzing video sequences capturing the execution of the sit-to-stand task in two groups of young and elderly volunteers. The results are compared with clinical studies that used marker-based systems, and are particularly encouraging for a future extension of the approach to other motor tasks and to predict scores obtained from the physical performance batteries that are widely and regularly used by clinicians and physical therapists. Michela Goffredo, Maurizio Schmid, Silvia Conforto, Marco Carli, Alessandro Neri 0001, Tommaso D'Alessio |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2008 | Color image database for evaluation of image quality metricsabstractIn this contribution, a new image database for testing full-reference image quality assessment metrics is presented. It is based on 1700 test images (25 reference images, 17 types of distortions for each reference image, 4 levels for each type of distortion). Using this image database, 654 observers from three different countries (Finland, Italy, and Ukraine) have carried out about 400000 individual human quality judgments (more than 200 judgments for each distorted image). The obtained mean opinion scores for the considered images can be used for evaluating the performances of visual quality metrics as well as for comparison and for the design of new metrics. The database, with testing results, is freely available. Nikolay N. Ponomarenko, Vladimir Lukin 0001, Karen Egiazarian, Jaakko Astola, Marco Carli, Federica Battisti |
MMSP | 5 |
| 2007 | Robust Data Hiding Technique for Video Error Concealment over DVB-H channelabstractIn this paper a method for providing additional information for error concealment during the transmission of a video file over an error-prone channel is presented. The transmission of the additional information is performed without increasing the bandwidth occupation, by hiding this information into the video itself at the encoder side. The considered transmission system is the DVB-H system. The performance of the method has been investigated by simulations and experimental tests in order to evaluate the perceived quality of the processed video and the robustness of the data hiding method against the H.264/AVC coding. Francesca De Simone, Marco Carli, Alessandro Neri 0001, Adrian Hornsby, Irek Defée |
MMSP | 2 |
| 2005 | Quality assessment using data hiding on perceptually important areasabstractIn this paper, we present a no-reference video quality metric that blindly estimates the quality of a video. The proposed approach makes use of a data hiding technique to embed a fragile mark into perceptually important areas of the video frame. To estimate the importance of an area, we take into account three perceptual features that are known to attract visual attention: motion, contrast, and color. At the receiver, the mark is extracted from the perceptually important areas of the decoded video. Then, a quality measure of the video is obtained by computing the degradation of the extracted mark. Simulation results indicate that the proposed video quality metric outperforms standard peak signal to noise ratio (PSNR) in estimating the perceived quality of a video. Additionally, results from a subjective experiment show that the metric output values increase monotonically with the mean annoyance scores gathered from the human observers. Marco Carli, Mylène C. Q. Farias, Elisa Drelie Gelasca, Roberto Tedesco, Alessandro Neri 0001 |
ICIP (3) | 1 |
| 2005 | Data Hiding Driven by Perceptual FeaturesabstractIn this paper, a new methodology for embedding data in a video sequence is presented. To guarantee the imperceptibility of the embedded data, we propose a method to select frame regions that are considered perceptually non relevant. For each frame a salience analysis is performed based on features that are relevant to the human vision system. In particular, the local contrast, the color and the areas of motion have been considered. By weighting all these feature at once, an importance map is built to drive the sequent embedding procedure. Subjective experiment results show that the artifacts caused by this localized embedding procedure are considered less annoying than if the embedding is performed on the whole frame Marco Carli, Patrizio Campisi, Alessandro Neri 0001 |
MMSP | 1 |
| 2005 | A robust error concealment technique using data hiding for image and video transmission over lossy channelsabstractA robust error concealment scheme using data hiding which aims at achieving high perceptual quality of images and video at the end-user despite channel losses is proposed. The scheme involves embedding a low-resolution version of each image or video frame into itself using spread-spectrum watermarking, extracting the embedded watermark from the received video frame, and using it as a reference for reconstruction of the parent image or frame, thus detecting and concealing the transmission errors. Dithering techniques have been used to obtain a binary watermark from the low-resolution version of the image/video frame. Multiple copies of the dithered watermark are embedded in frequencies in a specific range to make it more robust to channel errors. It is shown experimentally that, based on the frequency selection and scaling factor variation, a high-quality watermark can be extracted from a low-quality lossy received image/video frame. Furthermore, the proposed technique is compared to its two-part variant where the low-resolution version is encoded and transmitted as side information instead of embedding it. Simulation results show that the proposed concealment technique using data hiding outperforms existing approaches in improving the perceptual quality, especially in the case of higher loss probabilities. Chowdary Adsumilli, Mylène C. Q. Farias, Sanjit K. Mitra, Marco Carli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | Annoyance of spatio-temporal artifacts in segmentation quality assessmentabstractThis paper describes the results of a series of subjective experiments that investigated the annoyance caused by the most common artifacts present in segmented video sequences. Various types of artifacts were inserted into a reference segmented video, considered as ideal, and shown to our test subjects. The artifacts varied in their location, size, appearance and duration. Annoyance of segmentation artifacts are found to be tied up with their intrinsic characteristics (e.g., size, position) but only weakly related to the video content. The results identify the characteristics that should be taken into account in the design of a perceptually driven objective metric. Elisa Drelie Gelasca, Touradj Ebrahimi, Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra |
ICIP | 4 |
| 2003 | A hybrid constrained unequal error protection and data hiding scheme for packet video transmissionabstractA novel hybrid scheme with constrained unequal error protection (UEP) and data hiding is proposed; it maximizes the perceptual quality of the video at the end user to compensate for the effects of channel losses. The technique involves: (1) implementing a forcing function which weighs the objective perceptual quality of the video frame based on a hidden mark signal to give optimum protection level in the packet; (2) utilizing a data hiding mechanism to embed second level wavelet approximation coefficients of the frame in itself. An optimum UEP in the transmitted packets is sought using a constrained optimization approach. Simulation results show that the proposed technique outperforms existing non-adaptive error concealment approaches in improving the perceptual quality, especially for higher loss probabilities. Chowdary Adsumilli, Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra |
ICASSP (5) | 3 |
| 2003 | An accurate billing mechanism for multimedia communicationsabstractA novel billing mechanism is presented in this paper to provide telecommunication service users and advertisers with an effective billing mechanism based on the actual amount of advertisement data being transmitted/displayed. Experimental results show the effectiveness of the proposed system in terms of computational cost and performance. José Gabriel R. C. Gomes, Mylène C. Q. Farias, Sanjit K. Mitra, Marco Carli |
ICME | 4 |
| 2002 | Tracing watermarking for multimedia communication quality assessmentabstractMultimedia data hiding by digital watermarking is usually employed for copyright protection purposes. In this contribution, a new application of watermarking is presented. Specifically, watermarking is here employed as a technique for testing the quality of service in multimedia mobile communications. A fragile known watermark is embedded in a MPEG-like host data video transport stream using a spread-spectrum technique to avoid visual interference. Like a tracing signal, a (known) tracing watermark tracks the (unknown) information stream that follows the same communication link. The detection of the tracing watermark allows dynamically evaluating the effective quality of the provided video services, depending on the whole physical layer (including the employed image co/decoder). The performed method is based on the mean-square-error between estimated and actual watermarks. The devised technique has been usefully applied to typical scenarios of mobile wireless multimedia communication systems, in presence of multipath channel and interfering users. Patrizio Campisi, Marco Carli, Gaetano Giunta, Alessandro Neri 0001 |
ICC | 2 |
| 2002 | A comparison between an objective quality measure and the mean annoyance values of watermarked videosabstractA comparison between an objective quality measure and the perceived mean annoyance values of watermarked videos is presented. A psychophysical experiment has been performed to measure the detection threshold and mean annoyance values of several watermarked videos, using two different marks. The results of this experiment were then compared with an objective quality measure, obtained through a tracing watermarking system. An estimation of the detection threshold of the watermarked videos was found. Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra, Alessandro Neri 0001 |
ICIP (3) | 2 |
| 2002 | Algorithm for viewpoint estimation and registration
Fabiola Calcopietro, Marco Carli, Luca Lucchese, Alessandro Neri 0001 |
VCIP | 2 |
| 2001 | Mobile IP and cellular IP integration for inter access networks handoffabstractIn this contribution two solutions for the management of cellular intranet, based on the mobile IP and cellular IP protocols integration are investigated. The first solution adopts a centralized architecture built over the gateway and the home agent. It is most suited for security needs and client/server traffic. The second solution utilizes the mobile IP with routing optimization for macro mobility management. It offers optimized routing, speeds-up the handoff procedures, supports real time traffic and is therefore oriented toward the quality of service. Marco Carli, Alessandro Neri 0001, Andrea Rem Picci |
ICC | 1 |
| 2001 | Markovian motion field regularization based on the Gauss-Laguerre transform
Marco Carli, Giovanni Jacovitti, Alessandro Neri 0001 |
VCIP | 1 |