VLDB 2026 Research / reviewers in the wild / expert
Antonio Servetti
dblp:98/5416
· DBLP profile ↗
20ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-0159-5718ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 6 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Modeling Subject Scoring Behaviors in Subjective Experiments Based on a Discrete Quality ScaleabstractSeveral approaches have been proposed to estimate quality in subjective experiments while highlighting peculiar subject behaviors. However, there is some room for improvement in existing approaches, both in terms of robustness to noise and the ability to accurately indicate several peculiar subject behaviors in subjective experiments. This work advances the state-of-the-art in three main directions: i) A new approach to estimate the subjective quality from noisy ratings is proposed and is shown to be more robust to noise than are four state-of-the-art approaches; ii) a novel subject scoring model is proposed that makes it possible to highlight several peculiar behaviors typically observed in subjective experiments; and iii) our proposed probabilistic subject scoring model results from the proof of a theorem, whereas in previous approaches a probabilistic scoring model is assumed apriori. This represents an important first step toward models supported by a stronger theoretical foundation. Numerical experiments conducted on several datasets highlight the effectiveness of our proposal. Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Enrico Masala |
IEEE Trans. Multim. | 2 |
| 2024 | Multiple Image Distortion DNN Modeling Individual Subject Quality AssessmentabstractA recent research direction is focused on training Deep Neural Networks (DNNs) to replicate individual subject assessments of media quality. These DNNs are referred to as Artificial Intelligence-based Observers (AIOs). An AIO is designed to simulate, in real-time, the quality ratings of a specific individual, enabling an automatic quality assessment that accounts for subjects characteristics and preferences. Training AIOs is a promising but challenging research area due to the greater noise in individual raw opinion scores compared to the Mean Opinion Score. Effective learning from noisy labels necessitates the training of complex models on large-scale datasets. Unfortunately, this is challenging for AIOs as the media quality assessment community lacks extensive datasets that include individual opinion scores. To address the complexity of the task, we first created a dataset comprising two million samples, with synthetic labels derived from human annotation. We then trained a customized network for image quality assessment, named Multi-Distortion ResNet50 (MDResNet50), on this dataset. The weights of the MDResNet50 were subsequently utilized to initialize the learning process of each AIO, thereby avoiding the need to train a complex model from scratch on a small-scale dataset with raw individual opinion scores. Computational experiments show that our approach significantly advances the state-of-the-art in the AIO research. In particular: (i) we demonstrate through a simulation the ability of AIOs to mimic two well-known behavioral characteristics of a subject, i.e., bias and inconsistency, when scoring the media quality; (ii) we train and release DNN-based AIOs that, compared to the state-of-the-art, exhibit a higher performance with a statistical significance in assessing multiple image distortions; (iii) we train AIOs that more accurately mimic the sensitivity of real subjects to noise and color saturation and also better predict the opinion score distribution compared to the state-of-the-art AIOs. Lohic Fotio Tiotsop, Antonio Servetti, Peter Pocta, Glenn Van Wallendael, Marcus Barkowsky, Enrico Masala |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | A Scoring Model Considering the Variability of Subjects' Characteristics in Subjective ExperimentsabstractMany authors argued that the scoring behavior of a subject in a subjective quality evaluation experiment can be modeled by two main characteristics, i.e., the subject's bias and the subject's inconsistency. However, for simplicity's sake, they disregarded the fact that subjects are usually less inconsistent when evaluating stimuli with very low or very high quality. This work addresses this shortcoming by providing an analytical formulation about how to link subjects' bias and inconsistency to the ground truth subjective quality of the stimulus under evaluation. By integrating this formulation into a state-of-the-art subject scoring model we obtain a more realistic model to recover the ground truth subjective quality of each stimulus. An iterative algorithm able to estimate the model parameters is also provided. Computational experiments show that our proposed model yields more realistic confidence intervals for the recovered ground truth subjective quality values and exhibits more robustness to synthetically added noise in several testing conditions. Lohic Fotio Tiotsop, Antonio Servetti, Enrico Masala |
QoMEX | 2 |
| 2023 | Predicting individual quality ratings of compressed images through deep CNNs-based artificial observers
Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Peter Pocta, Tomas Mizdos, Glenn Van Wallendael, Enrico Masala |
Signal Process. Image Commun. | 2 |
| 2022 | Regularized Maximum Likelihood Estimation of the Subjective Quality from Noisy Individual RatingsabstractDespite several approaches to recover the ground truth subjective quality score from noisy individual ratings in subjective experiments have been explored in the literature, there is still room for improvement, in particular in terms of robustness to noise. This paper proposes a new approach that combines the traditional maximum likelihood estimation framework with a newly proposed regularization term, based on information theory concepts, that is meant to underweight surprising ratings of the quality of a given stimulus, looked at as a noise manifestation, in the final analytical expression of the recovered subjective quality. Computational experiments show the higher robustness to noise of our proposal when compared to three state-of-the-art methods. Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Enrico Masala |
QoMEX | 2 |
| 2022 | Mimicking Individual Media Quality Perception with Neural Network based Artificial ObserversabstractThe media quality assessment research community has traditionally been focusing on developing objective algorithms to predict the result of a typical subjective experiment in terms of Mean Opinion Score (MOS) value. However, the MOS, being a single value, is insufficient to model the complexity and diversity of human opinions encountered in an actual subjective experiment. In this work we propose a complementary approach for objective media quality assessment that attempts to more closely model what happens in a subjective experiment in terms of single observers and, at the same time, we perform a qualitative analysis of the proposed approach while highlighting its suitability. More precisely, we propose to model, using neural networks (NNs) , the way single observers perceive media quality. Once trained, these NNs, one for each observer, are expected to mimic the corresponding observer in terms of quality perception. Then, similarly to a subjective experiment, such NNs can be used to simulate the users’ single opinions, which can be later aggregated by means of different statistical indicators such as average, standard deviation, quantiles, etc. Unlike previous approaches that consider subjective experiments as a black box providing reliable ground truth data for training, the proposed approach is able to consider human factors by analyzing and weighting individual observers. Such a model may therefore implicitly account for users’ expectations and tendencies, that have been shown in many studies to significantly correlate with visual quality perception. Furthermore, our proposal also introduces and investigates an index measuring how much inconsistency there would be if an observer was asked to rate many times the same stimulus. Simulation experiments conducted on several datasets demonstrate that the proposed approach can be effectively implemented in practice and thus yielding a more complete objective assessment of end users’ quality of experience. Lohic Fotio Tiotsop, Tomas Mizdos, Marcus Barkowsky, Peter Pocta, Antonio Servetti, Enrico Masala |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Full Reference Video Quality Measures Improvement Using Neural NetworksabstractThe accuracy of video quality metrics (VQMs) is an important issue for several applications. In this work, first we observe that the accuracy of several video quality metrics (VQMs) is strongly related to the spatial complexity index (SI) of the source. In particular, our investigation suggests that the VQMs are more likely to inaccurately predict the subjective quality of the processed video sequences derived from sources characterized by low SI. To address such a situation, we propose a machine learning based improvement for each of the VQMs considered in this work and a video quality metric fusion index (VQMFI) that jointly exploits all the VQMs considered in the study as well as spatiotemporal features to produce a better estimation of the subjective quality. Computational results demonstrate the superiority of our proposals on several datasets. Lohic Fotio Tiotsop, Antonio Servetti, Enrico Masala |
ICASSP | 2 |
| 2016 | Discovering users with similar internet access performance through cluster analysis
Tania Cerquitelli, Antonio Servetti, Enrico Masala |
Expert Syst. Appl. | 2 |
| 2014 | Measuring DASH streaming performance from the end users perspective using neubotabstractThe popularity of DASH streaming is rapidly increasing and a number of commercial streaming services are adopting this new standard. While the benefits of building streaming services on top of the HTTP protocol are clear, further work is still necessary to evaluate and enhance the system performance from the perspective of the end user. Here we present a novel framework to evaluate the performance of rate-adaptation algorithms for DASH streaming using network measurements collected from more than a thousand Internet clients. Data, which have been made publicly available, are collected by a DASH module built on top of Neubot, an open source tool for the collection of network measurements. Some examples about the possible usage of the collected data are given, ranging from simple analysis and performance comparisons of download speeds to the performance simulation of alternative adaptation strategies using, e.g., the instantaneous available bandwidth values. Simone Basso, Antonio Servetti, Enrico Masala, Juan Carlos De Martin |
MMSys | 2 |
| 2013 | Challenges and Issues on Collecting and Analyzing Large Volumes of Network Data Measurements
Enrico Masala, Antonio Servetti, Simone Basso, Juan Carlos De Martin |
ADBIS (2) | 2 |
| 2011 | A Semantic Web Annotation Tool for a Web-Based Audio Sequencer
Luca Restagno, Vincent Akkermans, Giuseppe Rizzo 0002, Antonio Servetti |
ICWE | 4 |
| 2011 | The network neutrality bot architecture: A preliminary approach for self-monitoring of Internet access QoSabstractThe “network neutrality bot” (Neubot) is an evolving software architecture for distributed Internet access quality and network neutrality measurements. The core of this architecture is an open-source agent that ordinary users may install on their computers to gain a deeper understanding of their Internet connections. The agent periodically monitors the quality of service provided to the user, running background active transmission tests that emulate different application-level protocols. The results are then collected on a central server and made publicly available to allow constant monitoring of the state of the Internet by interested parties. In this article we describe how we enhanced Neubot architecture both to deploy a distributed broadband speed test and to allow the development of plug-in transmission tests. In addition, we start a preliminary discussion on the results we have collected in the first three months after the first public release of the software. Simone Basso, Antonio Servetti, Juan Carlos De Martin |
ISCC | 2 |
| 2009 | Text-independent compressed domain speaker verification for digital communication networks call monitoringabstractIn this paper we present a text-independent automatic speaker verification system that works in the compressed domain using GSM AMR coded speech. While traditional approaches process LPC-based cepstral coefficients extracted from LPC related bitstream coefficients, our objective is to study the feasibility of a system that directly processes the raw bitstream transmitted over a digital communication network. In a text-independent closet-set task, with a database of 100 speakers, the proposed system achieves an EER equal to 5.93%, 4.41% and 3.80% for 10, 20 and 30 second long test speech segments respectively. Matteo Petracca, Antonio Servetti, Juan Carlos De Martin |
ICME | 2 |
| 2006 | Performance Analysis of Compressed-Domain Automatic Speaker Recognition as a Function of Speech Coding Technique and Bit RateabstractCompressed-domain automatic speaker recognition is based on the analysis of the compressed parameters of speech coders. The objective is to perform low-complexity on-line speaker recognition for VoIP in the compressed domain, without the need to decode or resynthesize the speech bitstream. In this paper, we present initial results in determining the recognition accuracy that can be achieved with five widely used speech coding standards. Experiments with a database of 14 speakers obtain a recognition ratio close to 100% after the analysis of 30 seconds of active speech for most of the considered speech coders and rates. In particular, the results show that performance does not strictly depend on coding rate or codec speech quality Matteo Petracca, Antonio Servetti, Juan Carlos De Martin |
ICME | 2 |
| 2005 | Standard Compatible Error Correction for Multimedia Transmissions Over 802.11 WLANabstractIn this paper, we analyze a standard compatible error correction technique for multimedia transmissions over 802.11 WLANs that exploits, when available, the information of previous erroneous transmissions. The basic idea is to store erroneous frames for error correction purposes. More specifically, at the receiver each bit is estimated with a majority criterion. The performance of different standard compliant error recovery techniques have been evaluated using actual transmission experiments in various channel conditions. Optimal tradeoffs between complexity, memory and perceived quality have been determined studying the quality gains that can be achieved for the specific case of multimedia applications. Perceived quality has been evaluated using objective measures, e.g. ITU-T PESQ for voice and PSNR for video. Results show that the majority combining approach is particularly effective for multimedia communications, even in very noisy scenarios. Gains up to about one unit on the MOS scale for speech and up to 5-6 dB PSNR in case of video have been measured with respect to the standard ARQ technique Enrico Masala, Antonio Servetti, Juan Carlos De Martin |
ICME | 2 |
| 2005 | Low-complexity automatic speaker recognition in the compressed GSM AMR domainabstractThis paper presents an experimental implementation of a low-complexity speaker recognition algorithm working in the compressed speech domain. The goal is to perform speaker modeling and identification without decoding the speech bitstream to extract speaker dependent features, thus saving important system resources, for instance, in mobile devices. The compressed bitstream values of the widely used GSM AMR speech coding standard are studied to identify statistics enabling fair recognition after a few seconds of speech. Using Euclidean distance measures on elementary statistical values such as coefficient of variation and skewness of nine standard GSM AMR parameters delivers recognition accuracies close to 100% after about 20 seconds of active speech for a database of 14 speakers recorded in a normal room environment. Matteo Petracca, Antonio Servetti, Juan Carlos De Martin |
ICME | 2 |
| 2004 | Fast implementation of the MPEG-4 AAC main and low complexity decoderabstractWe present a fast software implementation of the MPEG-4 AAC (advanced audio coding) main and low complexity (LC) decoder. The reference implementation is analyzed and selected algorithms are presented to improve the performance of most of its building blocks, i.e., bitstream de-formatter, noiseless decoding, prediction, and filterbank. The code is further optimized by means of assembler procedures for the filterbank and prediction tools to exploit the Intel Pentium streaming SIMD extension (SSE) instruction set. SSE performs operations over four single precision floating-point numbers in a single instruction using 128-bit long registers. Results show that the presented decoder proves to be almost five times faster than the reference implementation, while preserving its full compatibility with the MPEG-4 standard. Antonio Servetti, Alessandro Rinotti, Juan Carlos De Martin |
ICASSP (5) | 1 |
| 2003 | Frequency-selective partial encryption of compressed audioabstractThe widespread adoption of compressed digital audio increasingly demands effective ways to conjugate ease of distribution with content protection. In this paper a low-complexity partial-encryption scheme for MPEG audio is presented. The aim is to provide listeners with sample-quality audio material that can be upgraded to full-quality by simply acquiring a key and decrypting a few selected bits (1-10% of the total bitstream) of the already available data. Sample-quality is obtained by limiting the frequency content of the signal, and it is achieved without altering the compatibility of the compressed audio bitstream. Audio material partially encrypted with the proposed scheme can thus be freely distributed for evaluation and then easily unlocked to achieve full quality without any further transmission of audio data. Antonio Servetti, Cristiano Testa, Juan Carlos De Martin |
ICASSP (5) | 1 |
| 2002 | Perception-based selective encryption of G.729 speechabstractMobile multimedia applications, the focus of many forth-coming wireless services, increasingly demand low-power techniques implementing content protection and customer privacy. In this paper a low-complexity, perception-based partial encryption scheme for telephone-bandwidth speech is presented. Speech compressed by a widely-used speech coding algorithm, ITU-T G.729 CS-ACELP at 8 kb/s, is partitioned in two classes, one, the most perceptually relevant, to be encrypted, the other, to be left unprotected. Encryption of about 45% of the bitstream achieves content protection equivalent to full encryption of the bitstream, as verified by both objective measures and formal listening tests. Low-power, portable devices can, therefore, implement very high levels of speech-content protection at a fraction of the computational load of current techniques, freeing resources for other tasks and enabling longer battery life. Antonio Servetti, Juan Carlos De Martin |
ICASSP | 1 |
| 2002 | Perception-based partial encryption of compressed speechabstractMobile multimedia applications, the focus of many forthcoming wireless services, increasingly demand low-power techniques implementing content protection and customer privacy. In this paper low complexity perception-based partial encryption schemes for speech are presented. Speech compressed by a widely-used speech coding algorithm, the ITU-T G.729 standard at 8 kb/s, is partitioned in two classes, one, the most perceptually relevant, to be encrypted, the other, to be left unprotected. Two partial-encryption techniques are developed, a low-protection scheme, aimed at preventing most kinds of eavesdropping and a high-protection scheme, based on the encryption of a larger share of perceptually important bits and meant to perform as well as full encryption of the compressed bitstream. The high-protection scheme, based on the encryption of about 45% of the bitstream, achieves content protection comparable to that obtained by full encryption, as verified by both objective measures and formal listening tests. For the low-protection scheme, encryption of as little as 30% of the bitstream virtually eliminates intelligibility as well as most of the remaining perceptual information. Low-power, portable devices could therefore achieve very high levels of speech-content protection at only 30-45% of the computational load of current techniques, freeing resources for other tasks and enabling longer battery life. Antonio Servetti, Juan Carlos De Martin |
IEEE Trans. Speech Audio Process. | 1 |