EDBT 2026 Demo / reviewers in the wild / expert
Enrico Masala
dblp:21/2038
· DBLP profile ↗
54ranked-venue papers
14as first author
8since 2021 · last 2024
0000-0001-8906-354XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 12 first-author · 8 since 2021Computer networks · 13 · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Modeling Subject Scoring Behaviors in Subjective Experiments Based on a Discrete Quality ScaleabstractSeveral approaches have been proposed to estimate quality in subjective experiments while highlighting peculiar subject behaviors. However, there is some room for improvement in existing approaches, both in terms of robustness to noise and the ability to accurately indicate several peculiar subject behaviors in subjective experiments. This work advances the state-of-the-art in three main directions: i) A new approach to estimate the subjective quality from noisy ratings is proposed and is shown to be more robust to noise than are four state-of-the-art approaches; ii) a novel subject scoring model is proposed that makes it possible to highlight several peculiar behaviors typically observed in subjective experiments; and iii) our proposed probabilistic subject scoring model results from the proof of a theorem, whereas in previous approaches a probabilistic scoring model is assumed apriori. This represents an important first step toward models supported by a stronger theoretical foundation. Numerical experiments conducted on several datasets highlight the effectiveness of our proposal. Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Enrico Masala |
IEEE Trans. Multim. | 4 |
| 2024 | Multiple Image Distortion DNN Modeling Individual Subject Quality AssessmentabstractA recent research direction is focused on training Deep Neural Networks (DNNs) to replicate individual subject assessments of media quality. These DNNs are referred to as Artificial Intelligence-based Observers (AIOs). An AIO is designed to simulate, in real-time, the quality ratings of a specific individual, enabling an automatic quality assessment that accounts for subjects characteristics and preferences. Training AIOs is a promising but challenging research area due to the greater noise in individual raw opinion scores compared to the Mean Opinion Score. Effective learning from noisy labels necessitates the training of complex models on large-scale datasets. Unfortunately, this is challenging for AIOs as the media quality assessment community lacks extensive datasets that include individual opinion scores. To address the complexity of the task, we first created a dataset comprising two million samples, with synthetic labels derived from human annotation. We then trained a customized network for image quality assessment, named Multi-Distortion ResNet50 (MDResNet50), on this dataset. The weights of the MDResNet50 were subsequently utilized to initialize the learning process of each AIO, thereby avoiding the need to train a complex model from scratch on a small-scale dataset with raw individual opinion scores. Computational experiments show that our approach significantly advances the state-of-the-art in the AIO research. In particular: (i) we demonstrate through a simulation the ability of AIOs to mimic two well-known behavioral characteristics of a subject, i.e., bias and inconsistency, when scoring the media quality; (ii) we train and release DNN-based AIOs that, compared to the state-of-the-art, exhibit a higher performance with a statistical significance in assessing multiple image distortions; (iii) we train AIOs that more accurately mimic the sensitivity of real subjects to noise and color saturation and also better predict the opinion score distribution compared to the state-of-the-art AIOs. Lohic Fotio Tiotsop, Antonio Servetti, Peter Pocta, Glenn Van Wallendael, Marcus Barkowsky, Enrico Masala |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | A Scoring Model Considering the Variability of Subjects' Characteristics in Subjective ExperimentsabstractMany authors argued that the scoring behavior of a subject in a subjective quality evaluation experiment can be modeled by two main characteristics, i.e., the subject's bias and the subject's inconsistency. However, for simplicity's sake, they disregarded the fact that subjects are usually less inconsistent when evaluating stimuli with very low or very high quality. This work addresses this shortcoming by providing an analytical formulation about how to link subjects' bias and inconsistency to the ground truth subjective quality of the stimulus under evaluation. By integrating this formulation into a state-of-the-art subject scoring model we obtain a more realistic model to recover the ground truth subjective quality of each stimulus. An iterative algorithm able to estimate the model parameters is also provided. Computational experiments show that our proposed model yields more realistic confidence intervals for the recovered ground truth subjective quality values and exhibits more robustness to synthetically added noise in several testing conditions. Lohic Fotio Tiotsop, Antonio Servetti, Enrico Masala |
QoMEX | 3 |
| 2023 | Predicting individual quality ratings of compressed images through deep CNNs-based artificial observers
Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Peter Pocta, Tomas Mizdos, Glenn Van Wallendael, Enrico Masala |
Signal Process. Image Commun. | 7 |
| 2022 | Regularized Maximum Likelihood Estimation of the Subjective Quality from Noisy Individual RatingsabstractDespite several approaches to recover the ground truth subjective quality score from noisy individual ratings in subjective experiments have been explored in the literature, there is still room for improvement, in particular in terms of robustness to noise. This paper proposes a new approach that combines the traditional maximum likelihood estimation framework with a newly proposed regularization term, based on information theory concepts, that is meant to underweight surprising ratings of the quality of a given stimulus, looked at as a noise manifestation, in the final analytical expression of the recovered subjective quality. Computational experiments show the higher robustness to noise of our proposal when compared to three state-of-the-art methods. Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Enrico Masala |
QoMEX | 4 |
| 2022 | Mimicking Individual Media Quality Perception with Neural Network based Artificial ObserversabstractThe media quality assessment research community has traditionally been focusing on developing objective algorithms to predict the result of a typical subjective experiment in terms of Mean Opinion Score (MOS) value. However, the MOS, being a single value, is insufficient to model the complexity and diversity of human opinions encountered in an actual subjective experiment. In this work we propose a complementary approach for objective media quality assessment that attempts to more closely model what happens in a subjective experiment in terms of single observers and, at the same time, we perform a qualitative analysis of the proposed approach while highlighting its suitability. More precisely, we propose to model, using neural networks (NNs) , the way single observers perceive media quality. Once trained, these NNs, one for each observer, are expected to mimic the corresponding observer in terms of quality perception. Then, similarly to a subjective experiment, such NNs can be used to simulate the users’ single opinions, which can be later aggregated by means of different statistical indicators such as average, standard deviation, quantiles, etc. Unlike previous approaches that consider subjective experiments as a black box providing reliable ground truth data for training, the proposed approach is able to consider human factors by analyzing and weighting individual observers. Such a model may therefore implicitly account for users’ expectations and tendencies, that have been shown in many studies to significantly correlate with visual quality perception. Furthermore, our proposal also introduces and investigates an index measuring how much inconsistency there would be if an observer was asked to rate many times the same stimulus. Simulation experiments conducted on several datasets demonstrate that the proposed approach can be effectively implemented in practice and thus yielding a more complete objective assessment of end users’ quality of experience. Lohic Fotio Tiotsop, Tomas Mizdos, Marcus Barkowsky, Peter Pocta, Antonio Servetti, Enrico Masala |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | How to Train No Reference Video Quality Measures for New Coding Standards using Existing Annotated Datasets?abstractSubjective experiments are important for developing objective Video Quality Measures (VQMs). However, they are time-consuming and resource-demanding. In this context, being able to reuse existing subjective data on previous video coding standards to train models capable of predicting the perceptual quality of video content processed with newer codecs acquires significant importance. This paper investigates the possibility of generating an HEVC encoded Processed Video Sequence (PVS) in such a way that its perceptual quality is as similar as possible to that of an AVC encoded PVS whose quality has already been assessed by human subjects. In this way, the perceptual quality of the newly generated HEVC encoded PVS may be annotated approximately with the Mean Opinion Score (MOS) of the related AVC encoded PVS. To show the effectiveness of our approach, we compared the performance of a simple and low complexity but yet effective no reference hybrid model trained on the data generated with our approach with the same model trained on data collected in the context of a pristine subjective experiment. In addition, we merged seven subjective experiments such that they can be used as one aligned dataset containing either original HEVC bitstreams or the newly generated data explained in our proposed approach. The merging process accounts for the differences in terms of quality scale, chosen assessment method and context influence factors. This yields a large annotated dataset of HEVC sequences that is made publicly available for the design and training of no reference hybrid VQMs for HEVC encoded content. Lohic Fotio Tiotsop, Tomas Mizdos, Enrico Masala, Marcus Barkowsky, Peter Pocta |
MMSP | 3 |
| 2021 | Modeling and estimating the subjects' diversity of opinions in video quality assessment: a neural network based approachabstractAbstract Subjective experiments are considered the most reliable way to assess the perceived visual quality. However, observers’ opinions are characterized by large diversity: in fact, even the same observer is often not able to exactly repeat his first opinion when rating again a given stimulus. This makes the Mean Opinion Score (MOS) alone, in many cases, not sufficient to get accurate information about the perceived visual quality. To this aim, it is important to have a measure characterizing to what extent the observed or predicted MOS value is reliable and stable. For instance, the Standard deviation of the Opinions of the Subjects (SOS) could be considered as a measure of reliability when evaluating the quality subjectively. However, we are not aware of the existence of models or algorithms that allow to objectively predict how much diversity would be observed in subjects’ opinions in terms of SOS. In this work we observe, on the basis of a statistical analysis made on several subjective experiments, that the disagreement between the quality as measured by means of different objective video quality metrics (VQMs) can provide information on the diversity of the observers’ ratings on a given processed video sequence (PVS). In light of this observation we: i) propose and validate a model for the SOS observed in a subjective experiment; ii) design and train Neural Networks (NNs) that predict the average diversity that would be observed among the subjects’ ratings for a PVS starting from a set of VQMs values computed on such a PVS; iii) give insights into how the same NN based approach can be used to identify potential anomalies in the data collected in subjective experiments. Lohic Fotio Tiotsop, Tomas Mizdos, Miroslav Uhrina, Marcus Barkowsky, Peter Pocta, Enrico Masala |
Multim. Tools Appl. | 6 |
| 2020 | Full Reference Video Quality Measures Improvement Using Neural NetworksabstractThe accuracy of video quality metrics (VQMs) is an important issue for several applications. In this work, first we observe that the accuracy of several video quality metrics (VQMs) is strongly related to the spatial complexity index (SI) of the source. In particular, our investigation suggests that the VQMs are more likely to inaccurately predict the subjective quality of the processed video sequences derived from sources characterized by low SI. To address such a situation, we propose a machine learning based improvement for each of the VQMs considered in this work and a video quality metric fusion index (VQMFI) that jointly exploits all the VQMs considered in the study as well as spatiotemporal features to produce a better estimation of the subjective quality. Computational results demonstrate the superiority of our proposals on several datasets. Lohic Fotio Tiotsop, Antonio Servetti, Enrico Masala |
ICASSP | 3 |
| 2020 | Cloud Gaming with Foveated Video EncodingabstractCloud gaming enables playing high-end games, originally designed for PC or game console setups, on low-end devices such as netbooks and smartphones, by offloading graphics rendering to GPU-powered cloud servers. However, transmitting the high-resolution video requires a large amount of network bandwidth, even though it is a compressed video stream. Foveated video encoding (FVE) reduces the bandwidth requirement by taking advantage of the non-uniform acuity of human visual system and by knowing where the user is looking. Based on a consumer-grade real-time eye tracker and an open source cloud gaming platform, we provide a cloud gaming FVE prototype that is game-agnostic and requires no modifications to the underlying game engine. In this article, we describe the prototype and its evaluation through measurements with representative games from different genres to understand the effect of parametrization of the FVE scheme on bandwidth requirements and to understand its feasibility from the latency perspective. We also present results from a user study on first-person shooter games. The results suggest that it is possible to find a “sweet spot” for the encoding parameters so the users hardly notice the presence of foveated encoding but at the same time the scheme yields most of the achievable bandwidth savings. Gazi Karam Illahi, Thomas Van Gemert, Matti Siekkinen, Enrico Masala, Antti Oulasvirta, Antti Ylä-Jääski |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Computing Quality-of-Experience Ranges for Video Quality EstimationabstractTypically, the measurement of the Quality of Experience for video sequences aims at a single value, in most cases the Mean Opinion Score (MOS). Predicting this value using various algorithms has been widely studied. However, deviation from the MOS is often handled as an unpredictable error. The approach in this contribution estimates intervals of video quality instead of the single valued MOS. Well-known video quality estimators are fused together to output a lower and upper border for the expected video quality, on the basis of a model derived from a well-known subjectively annotated dataset. Results on different datasets provide insight on the suitability of the well-known estimators for this particular approach. Lohic Fotio Tiotsop, Enrico Masala, Ahmed Aldahdooh, Glenn Van Wallendael, Marcus Barkowsky |
QoMEX | 2 |
| 2019 | Improving relevant subjective testing for validation: Comparing machine learning algorithms for finding similarities in VQA datasets using objective measures
Ahmed Aldahdooh, Enrico Masala, Glenn Van Wallendael, Peter Lambert, Marcus Barkowsky |
Signal Process. Image Commun. | 2 |
| 2019 | Improved Performance Measures for Video Quality Assessment Algorithms Using Training and Validation SetsabstractThe training and performance analysis of objective video quality assessment algorithms is complex due to the huge variety of possible content classes and transmission distortions. Several secondary issues such as free parameters in machine learning algorithms and alignment of subjective datasets put an additional burden on the developer. In this paper, three subsequent steps are presented to address such issues. First, the content and coding parameter space of a large-scale database is used to select dedicated subsets for training objective algorithms. This aims at providing a method for selecting the most significant contents and coding parameters from all imaginable combinations. In the practical case where only a limited set is available, it also helps us to avoid redundancy in the training subset selection. The second step is a discussion on performance measures for algorithms that employ machine-learning methods. The particularity of the performance measures is that the quality of the training and verification datasets is taken into consideration. Common issues that often use existing measures are presented, and improved or complementary methods are proposed. The measures are applied to two examples of no-reference objective assessment algorithms using the aforementioned subsets of the large-scale database. While limited in terms of practical applications, this sandbox approach of objectively predicting an objectively evaluated video sequences allows for eliminating additional influence factors from subjective studies. In the third step, the proposed performance measures are applied to the practical case of training and analyzing assessment algorithms on readily available subjectively annotated image datasets. The presentation method in this part of the paper can also be used as an exemplified recommendation for reporting in-depth information on the performance. Using this presentation method, future publications presenting newly developed quality assessment algorithms may be significantly improved. Ahmed Aldahdooh, Enrico Masala, Olivier Janssens, Glenn Van Wallendael, Marcus Barkowsky, Patrick Le Callet |
IEEE Trans. Multim. | 2 |
| 2018 | Can You See What I See? Quality-of-Experience Measurements of Mobile Live Video BroadcastingabstractBroadcasting live video directly from mobile devices is rapidly gaining popularity with applications like Periscope and Facebook Live. The quality of experience (QoE) provided by these services comprises many factors, such as quality of transmitted video, video playback stalling, end-to-end latency, and impact on battery life, and they are not yet well understood. In this article, we examine mainly the Periscope service through a comprehensive measurement study and compare it in some aspects to Facebook Live. We shed light on the usage of Periscope through analysis of crawled data and then investigate the aforementioned QoE factors through statistical analyses as well as controlled small-scale measurements using a couple of different smartphones and both versions, Android and iOS, of the two applications. We report a number of findings including the discrepancy in latency between the two most commonly used protocols, RTMP and HLS, surprising surges in bandwidth demand caused by the Periscope app’s chat feature, substantial variations in video quality, poor adaptation of video bitrate to available upstream bandwidth at the video broadcaster side, and significant power consumption caused by the applications. Matti Siekkinen, Teemu Kämäräinen, Leonardo Favario, Enrico Masala |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Foveated video streaming for cloud gamingabstractGood user experience with interactive cloud-based multimedia applications, such as cloud gaming and cloud-based VR, requires low end-to-end latency and large amounts of downstream network bandwidth at the same time. In this paper, we present a foveated video streaming system for cloud gaming. The system adapts video stream quality by adjusting the encoding parameters on the fly to match the player's gaze position. We conduct measurements with a prototype that we developed for a cloud gaming system in conjunction with eye tracker hardware. Evaluation results suggest that such foveated streaming can reduce bandwidth requirements by even more than 50% depending on parametrization of the foveated video coding and that it is feasible from the latency perspective. Gazi Karam Illahi, Matti Siekkinen, Enrico Masala |
MMSP | 3 |
| 2017 | Analysis of HEVC transform throughput requirements for hardware implementations
Maurizio Masera, Lorenzo Re Fiorentin, Enrico Masala, Guido Masera, Maurizio Martina |
Signal Process. Image Commun. | 3 |
| 2017 | Optimized Upload Strategies for Live Scalable Video Transmission from Mobile DevicesabstractSharing live multimedia content is becoming increasingly popular among mobile users. In this article, we study the problem of optimizing video quality in such a scenario using scalable video coding (SVC) and chunked video content. We consider using only standard stateless HTTP servers that do not need to perform additional processing of the video content. Our key contribution is to provide close to optimal algorithms for scheduling video chunk upload for multiple clients having different viewing delays. Given such a set of clients, the problem is to decide which chunks to upload and in which order to upload them so that the quality-delay tradeoff can be optimally balanced. We show by means of simulations that the proposed algorithms can achieve notably better performance than naive solutions in practical cases. Especially the heuristic-based greedy algorithm is a good candidate for deployment on mobile devices because it is not computationally intensive but it still delivers in most cases on-par video quality compared to the more complex local optimization algorithm. We also show that using shorter video segments and being able to predict bandwidth and video chunk properties improve the delivered video quality in certain cases. Matti Siekkinen, Enrico Masala, Jukka K. Nurminen |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | Work-in-progress: Integrating a remote laboratory system in an online learning environmentabstractThe possibility to include practical training sessions in online learning systems is important to increase the value of engineering courses. This paper presents a work in progress focusing on integrating a remote laboratory system in an existing online learning environment so that, through the use of a simple web browser, students can develop both theoretical and practical skills at the same time. The proposed architecture builds on our Free Architecture for Remote Education (FARE), which is a learning management system designed to foster collaboration and sharing of resources between participants. We seamlessly integrate in FARE all the necessary software for practical remote laboratory experiences. We believe that this synergy can provide significant advantages on traditional remote laboratory approaches: integration and automatic resource reuse can help students to have a simplified interaction with the laboratory and focus on the most important parts of the practical activities. Leonardo Favario, Enrico Masala |
EDUCON | 2 |
| 2016 | A First Look at Quality of Mobile Live Streaming Experience: the Case of Periscope
Matti Siekkinen, Enrico Masala, Teemu Kämäräinen |
Internet Measurement Conference | 2 |
| 2016 | Mobile live streaming: Insights from the periscope serviceabstractLive video streaming from mobile devices is quickly becoming popular through services such as Periscope, Meerkat, and Facebook Live. Little is known, however, about how such services tackle the challenges of the live mobile streaming scenario. This work addresses such gap by investigating in details the characteristics of the Periscope service. A large number of publicly available streams have been captured and analyzed in depth, in particular studying the characteristics of the encoded streams and the communication evolution over time. Such an investigation allows to get an insight into key performance parameters such as bandwidth, latency, buffer levels and freezes, as well as the limits and strategies adopted by Periscope to deal with this challenging application scenario. Leonardo Favario, Matti Siekkinen, Enrico Masala |
MMSP | 3 |
| 2016 | Comparing temporal behavior of fast objective video quality measures on a large-scale databaseabstractIn many application scenarios, video quality assessment is required to be fast and reasonably accurate. The characterization of objective algorithms by subjective assessment is well established but limited due to the small number of test samples. Verification using large-scale objectively annotated databases provides a complementary solution. In this contribution, three simple but fast measures are compared regarding their agreement on a large-scale database. In contrast to subjective experiments, not only sequence-wise but also framewise agreement can be analyzed. Insight is gained into the behavior of the measures with respect to 5952 different coding configurations of High Efficiency Video Coding (HEVC). Consistency within a video sequence is analyzed as well as across video sequences. The results show that the occurrence of discrepancies depends mostly on the configured coding structure and the source content. The detailed observations stimulate questions on the combined usage of several video quality measures for encoder optimization. Ahmed Aldahdooh, Enrico Masala, Glenn Van Wallendael, Marcus Barkowsky |
PCS | 2 |
| 2016 | Comparing simple video quality measures for loss-impaired video sequences on a large-scale databaseabstractThe performance of objective video quality measures is usually identified by comparing their predictions to subjective assessment results which are regarded as the ground truth. In this work we propose a complementary approach for this performance evaluation by means of a large-scale database of test sequences evaluated with several objective measurement algorithms. Such an approach is expected to detect performance anomalies that could highlight shortcomings in current objective measurement algorithms. Using realistic coding and network transmission conditions, we investigate the consistency of the prediction of different measures as well as how much their behavior can be predicted by content, coding and transmission features, discussing unexpected and peculiar behaviors, and highlighting how a large-scale database can help in identifying anomalies not easily found by means of subjective testing. We expect that this analysis will shed light on directions to pursue in order to overcome some of the limitations of existing reliability assessment methods for objective video quality measures. Ahmed Aldahdooh, Enrico Masala, Olivier Janssens, Glenn Van Wallendael, Marcus Barkowsky |
QoMEX | 2 |
| 2016 | Discovering users with similar internet access performance through cluster analysis
Tania Cerquitelli, Antonio Servetti, Enrico Masala |
Expert Syst. Appl. | 3 |
| 2015 | A New Platform for Cross-Repository Creation and Sharing of Educational Resources: Architecture and a Case StudyabstractCurrently there is a large amount of educational resources available, developed by many educational institutions at all level. One of the main challenges is to make use of such wealth of material in a simple and effective way, especially when pre-university education is involved, ranging from primary schools to high schools. The main difficulties experienced by the teachers willing to tailor the available material to the specific needs of their class are typically the lack of available time and the difficulties in learning the peculiar procedures that each repository system (including CMS, LMS etc.) requires. This work presents an architecture to integrate different repository systems using the Content Management Interoperability Services (CMIS) API, as well as an integration layer that provides a much more simplified interface suitable for the needs of the content creators (i.e., Teachers), and the users of the contents (i.e., The learners). Both the technical aspects of integration and the usability issues from the point of view of the teachers are described and considered in the design. The first experiments, evaluated by means of performance indicators and some user feedbacks, show that the platform has the potential for widespread adoption in the Italian educational environment. Leonardo Favario, Angelo Raffaele Meo, Enrico Masala |
COMPSAC | 3 |
| 2015 | Seamless cross-platform integration of educational resources for improved learning experiencesabstractThe main contribution of this work in progress is an open source architecture for seamless integration of open access educational resources located in heterogeneous learning management systems (LMS) distributed across the network. One of the main advantages of the proposal is the simplicity in creating, managing and making available new content that builds on the existing material in different repositories. Some previous works addressed similar issues but either in simpler environments, i.e., when the distributed systems were homogeneous, or they require extensive knowledge to take advantages of the integration, from the point of view of both the content creator and the manager of each learning system. Our solution exploits the possibilities offered by the Content Management Interoperability Services (CMIS) interface typically present in many content repository platforms allowing users to browse and interact with contents stored in several heterogeneous places. Some preliminary results show that it is possible to efficiently combine engineering learning material taken from our technical university with other online resources made available from different educational institutions. Moreover, two extensions, currently in development, aim at providing seamless videoconferencing integrations between the teacher and the learners as well as at supporting remote laboratory sessions by means of controlling programmable objects in remote laboratories with real-time video feedbacks for the students. Leonardo Favario, Angelo Raffaele Meo, Enrico Masala |
FIE | 3 |
| 2015 | TDuCSMA: Efficient support for triple-play services in wireless home networksabstractThe recently proposed Time-Division Unbalanced Carrier Sense Multiple Access (TDuCSMA) coordination function has been shown to cope very efficiently with many different data flows with widely different bitrates and packet lengths. This case may happen in a typical wireless home network setting where HD video, voice, videosurveillance and videoconference applications are used concurrently. This paper illustrates the advantages of TDuCSMA in such a scenario compared to the Enhanced Distributed Channel Access (EDCA), currently provided by the IEEE 802.11 standard, in terms of both performance from the end user's point of view and network resource utilization. The results show that TDuCSMA provides much better performance than EDCA, especially regarding video quality, delay and packet loss rate for voice streams, and stability over time. Moreover, the TDuCSMA performance is consistent as the network load increases while the EDCA performance significantly decreases. Andrea Vesco, Riccardo Scopigno, Enrico Masala |
ICC | 3 |
| 2015 | A new quality optimization framework for dash streaming over wireless channelsabstractMobile devices are increasingly used as terminals for playback of multimedia content. However, maximizing the user's quality of experience is challenging due to the highly variable conditions of the wireless channels. A possibility to cope with such a variability is to dynamically adapt the source coding rate during the transmission, which is the underlying idea of the DASH standard. This work proposes a new framework to improve the quality of the DASH-based streaming experience by allowing to adjust the tradeoff between the quality of received content and the risk of playback freeze due to an empty buffer, which is a strong quality-disruptive event. The problem is analytically formulated and an efficient method to compute the playback freeze probability as a function of the representation choices over time is presented. Numerous simulation results using real download rate traces of 3G channels show the performance improvement compared to other bandwidth-adaptive algorithms as well as the robustness of the framework to variations of its most important parameters. Leonardo Favario, Enrico Masala |
ICME | 2 |
| 2015 | On the effects of sender-receiver concealment mismatch on multimedia communication optimization
Enrico Masala, Fabio De Vito, Juan Carlos De Martin |
Multim. Tools Appl. | 1 |
| 2014 | Measuring DASH streaming performance from the end users perspective using neubotabstractThe popularity of DASH streaming is rapidly increasing and a number of commercial streaming services are adopting this new standard. While the benefits of building streaming services on top of the HTTP protocol are clear, further work is still necessary to evaluate and enhance the system performance from the perspective of the end user. Here we present a novel framework to evaluate the performance of rate-adaptation algorithms for DASH streaming using network measurements collected from more than a thousand Internet clients. Data, which have been made publicly available, are collected by a DASH module built on top of Neubot, an open source tool for the collection of network measurements. Some examples about the possible usage of the collected data are given, ranging from simple analysis and performance comparisons of download speeds to the performance simulation of alternative adaptation strategies using, e.g., the instantaneous available bandwidth values. Simone Basso, Antonio Servetti, Enrico Masala, Juan Carlos De Martin |
MMSys | 3 |
| 2014 | Supporting triple-play communications with TDuCSMA and first experimentsabstractThis work addresses the implications of using the Time-Division Unbalanced Carrier Sense Multiple Access (TDuCSMA) coordination function to support triple-play services. Firstly, the theoretical background of TDuCSMA is reported, presenting its advantages and discussing its full compliance with the IEEE 802.11 standard. Secondly, a prototype of TDuCSMA is discussed in details. Then, a set of experiments with the prototype implementation of TDuCSMA is presented, showing for the first time the advantages of TDuCSMA in a realistic setting with audio, video and elastic data applications. Experimental results show the superiority of TDuCSMA over the legacy 802.11 Medium Access Control (MAC) in terms of both channel utilization and Quality of Experience (QoE) as measured at the application level. Andrea Vesco, Riccardo Scopigno, Enrico Masala |
WCNC | 3 |
| 2014 | Content-based group-of-picture size control in distributed video coding
Enrico Masala, Yan-Mei Yu, Xiaohai He |
Signal Process. Image Commun. | 1 |
| 2013 | Challenges and Issues on Collecting and Analyzing Large Volumes of Network Data Measurements
Enrico Masala, Antonio Servetti, Simone Basso, Juan Carlos De Martin |
ADBIS (2) | 1 |
| 2013 | Sensor-based real-time adaptation of 3D video encoding quality for remote control applicationsabstractThe availability of stereoscopic mobile devices, such as mobile phones, on the consumer market allows to attempt the development of low-cost remote control systems that can provide a real-time 3D video feedback. In this work we show how implement such a communication system by considering the stringent latency constraints of the remote control scenario. To reduce the impact of this issue, we observe that part of the latency is due to the limited processing power of the mobile device that cannot sustain video transmission at high quality with low latency. Thus, we propose to dynamically change the latency-quality trade-off at the transmitter to optimize the quality of experience as perceived by the operator of the remote control system, by taking into account, in real-time, the dynamics of the control operations. In more details, low-cost accelerometer and gyroscopic sensors are employed to decide in real-time how much latency has to be privileged over quality and vice versa, by selectively reducing the quality of one of the views in favor of a reduced overall latency. Comparisons with a non-adaptive higher-quality but also higher-latency system show that the operators prefer the adaptive system despite the video quality is slightly reduced in dynamic control conditions. Enrico Masala |
MMSP | 1 |
| 2013 | Low-complexity driving event detection from side information of a 3D video encoderabstractMobile phones are often found in cars, for instance when they are used as navigation assistants. This work propose to use their camera, which is often already pointed to the road, to perform some low-complexity analysis of the driving context, with the final aim to detect potentially unsafe conditions. Since content understanding algorithms are typically too complex to run in real time on a mobile device, a driving event detection algorithm is presented based on the side information available from video encoders, which are a highly optimized application in mobile phones. A set of interesting and easy-to-extract features has been identified in the side information and then further reduced and adapted to the specific events of interest. A detection algorithm based on support vector machines has been designed and trained on several hours of video annotated by a human operator to extract the events of interest. The detection algorithm is shown to achieve a good identification rate for the considered events and feature sets. Moreover, results also show that the use of a stereoscopic camera significantly improves the performance of the detection algorithm in most cases. Ruiliang Wang, Enrico Masala |
MMSP | 3 |
| 2012 | Moving Multimedia Simulations into the Cloud: A Cost-effective SolutionabstractResearchers often demand bursts of computing power to quickly obtain the results of certain simulation activities. Multimedia communication simulations usually belong to such category. They may require several days on a generic PC to test a comprehensive set of conditions depending on the complexity of the scenario. This paper proposes to use a cloud computing framework to accelerate these simulations and, consequently, research activities, while at the same time reducing the overall costs. A practical simulation example is shown, representative of a typical simulation of H.264/AVC video communications over a wireless channel. This work shows that, by means of a commercial cloud computing provider, the gains of the proposed technique compared to more traditional solutions using dedicated computers can be significant in terms of speed and cost reduction. Enrico Masala |
ISPA | 1 |
| 2012 | A cost-effective cloud computing framework for accelerating multimedia communication simulations
Daniele Angeli, Enrico Masala |
J. Parallel Distributed Comput. | 2 |
| 2010 | Rate-distortion optimized low-delay 3D video communicationsabstractThis paper focuses on the rate-distortion optimization of low-delay 3D video communications based on the latest H.264/MVC video coding standard. The first part of the work proposes a new low-complexity model for distortion estimation suitable for low-delay stereoscopic video communication scenarios such as 3D videoconferencing. The distortion introduced by the loss of a given frame is investigated and a model is designed in order to accurately estimate the impact that the loss of each frame would have on future frames. The model is then employed in a rate-distortion optimized framework for video communications over a generic QoS-enabled network. Simulations results show consistent performance gains, up to 1.7 dB PSNR, with respect to a traditional a priori technique based on frame dependency information only. Moreover, the performance is shown to be consistently close to the one of the prescient technique that has perfect knowledge of the distortion characteristics of future frames. Enrico Masala |
MMSP | 1 |
| 2010 | Content-adaptive traffic prioritization of spatio-temporal scalable video for robust communications over QoS-provisioned 802.11e networks
Attilio Fiandrotti, Dario Gallucci, Enrico Masala, Juan Carlos De Martin |
Signal Process. Image Commun. | 3 |
| 2009 | Content-Adaptive Robust H.264/SVC Video Communications over 802.11e NetworksabstractIn this paper we present a low-complexity traffic prioritization strategy for video transmission using the H.264 scalable video coding (SVC) standard over 802.11e wireless networks.The first part of this work focuses on assessing the perceptual impact of data loss in the various enhancement layers using a wide set of H.264/SVC encoded videos.The analysis shows that perceptual impairments are highly correlated with the motion activity in the video sequence.Thus, we propose an adaptive unequal error protection strategy which identifies the most perceptually important parts of the enhancement layers in the video sequence by means of a low complexity macroblock motion analysis process.The algorithm is tested by simulating a realistic 802.11e based home network scenario.Results obtained on a large set of video sequences show that the proposed content-aware traffic prioritization strategy enables PSNR gains up to 2.5 dB as well as noticeable visual quality improvements with respect to a traditional prioritization strategy aiming at minimizing error propagation. Dario Gallucci, Attilio Fiandrotti, Enrico Masala, Juan Carlos De Martin |
AINA | 3 |
| 2009 | Distortion Prediction for Video Quality Optimization over Packet Switched NetworksabstractScheduling techniques are often deployed at the network edge to maximize the quality of the video communication while satisfying a given constraint on the maximum high priority usable bandwidth. For the case of video communications, the importance of each packet, in terms of the distortion that would be caused by its loss, can be used to decide which packets should be prioritized in order to maximize the expected video quality. However, the resulting performance strictly depends on the length of the video stream segment considered, at each time instant, by the distortion-aware scheduling algorithms. This work focuses on improving the performance of those algorithms by predicting the characteristics of the near-future part of video streams, exploiting the short and long term dependencies in packet distortion values. This approach could be particularly valuable in live video scenarios where accessing the future part of the video would imply the insertion of an additional delay. Several distortion prediction models are derived and their performance evaluated with actual scheduling algorithms. Results show that the best predictions are provided by neural network models, yielding significant improvements of the video communication quality. Andrea Vesco, Enrico Masala, Carlo Novara |
GLOBECOM | 2 |
| 2009 | Optimized H.264 Video Encoding and Packetization for Video Transmission Over Pipeline Forwarding NetworksabstractPrevious works showed that the quality-of-service (QoS) requirements of multimedia applications can be optimally satisfied by pipeline forwarding (PF) by providing end-to-end delay guarantees as well as high network resource utilization. However, the unavoidable mismatch between reserved resources and the unpredictable traffic profile of a video stream has an impact on the resulting application layer quality. Therefore, a new low-complexity H.264 video encoding and packetization scheme based on a distortion-optimized macroblock grouping technique is designed here to maximize the performance of video transmission on PF networks. The scheme considers the perceptual importance of the different parts of the video data to group the most important information in few packets that are the natural candidates to receive the deterministic service provided by PF. Results show peak signal-to-noise ratio (PSNR) gains up to 2.5 dB over traditional video encoding and packetization schemes, as well as more graceful degradation in case of high network load. Enrico Masala, Andrea Vesco, Mario Baldi, Juan Carlos De Martin |
IEEE Trans. Multim. | 1 |
| 2008 | Traffic Prioritization of H.264/SVC Video over 802.11e Ad Hoc Wireless NetworksabstractThe H.264/SVC video codec extends the H.264/AVC standard with scalability features. In this paper we introduce a traffic prioritization algorithm suitable for the transmission of both H.264/SVC and H.264/AVC video over 802.11e ad hoc wireless networks. The proposed algorithm exploits the traffic prioritization capabilities offered by 802.11e to provide better protection to the most perceptually important parts of a video while achieving efficient network resource usage. We evaluate the algorithm by simulating video transmissions in an ad hoc network scenario. Results show that the H.264/SVC codec particularly benefits from the proposed algorithm, which enables a graceful video quality degradation in congested network conditions, as well as PSNR gains up to 2 dB with respect to the H.264/AVC codec using the same amount of network resources. Attilio Fiandrotti, Dario Gallucci, Enrico Masala, Enrico Magli |
ICCCN | 3 |
| 2008 | Application-aware optimization of packet scheduling for video communications over intervehicle ad hoc networksabstractThis paper focuses on optimizing real-time intervehicle video communications from an image analysis perspective. Differently from traditional multimedia communication schemes which are typically optimized to maximize perceptual quality as perceived by human users, in this work we propose a new packet importance estimation method designed to capture the contribution of each packet to the performance of the image analysis process. Packet importance values are used to optimize video packet scheduling and retransmission in the context of ad hoc 802.11 intervehicle communications. Simulations using actual intervehicle transmission traces show PSNR gains up to 2.3 dB and 0.8 dB compared with the standard MAC-layer ARQ and a distortion based optimized ARQ technique respectively, as well as up to 10% accuracy increase with the tested image analysis algorithm. Enrico Masala |
ICME | 1 |
| 2008 | Variable Time Scale Multimedia Streaming Over IP NetworksabstractThis paper presents a comprehensive analysis of a variable time-scale streaming technique, VTSS, according to which rate changes are obtained by varying the inter-packet transmission interval, rather than altering, as in most cases, the source coding rate. Instead of constraining the transmitter to operate in real-time, the time scale of the packet scheduler can vary between zero, when the network is congested, to as faster than real-time as the channel bandwidth allows, when the network is lightly loaded. Although this approach is reportedly used in commercial streaming products, so far the technique has not yet been analyzed in a rigorous fashion, nor it has been compared to other state-of-the-art streaming techniques. This work first presents a theoretical analysis of the performance achievable by the VTSS approach, and it shows that, for the same channel conditions, VTSS yields a total distortion which is lower or, in the worst case, equal than the distortion of the standard real-time source-rate adaptive approach. A lower bound on receiver buffer size is also derived. Network simulations then analyze the performance of a TCP-friendly test implementation of VTSS compared with an ideal real-time source rate-adaptive technique, whose performance, being ideal, represents the upper bound of any transmission scheme based on source rate adaptation. The simulation results, also based on actual network traces, show that the VTSS approach delivers higher perceptual quality (up to 1.2 dB PSNR in the considered scenarios) and reduced video quality fluctuations (1.6 dB standard deviation PSNR, instead of 4.9 dB) for a wide range of standard video sequences. Perceptual quality evaluation by means of PVQM confirms such results. The gains, as expected, are even more pronounced (7.6 dB PSNR on average) if compared to real-time constant bit-rate video transmission. Enrico Masala, Davide Quaglia, Juan Carlos De Martin |
IEEE Trans. Multim. | 1 |
| 2006 | Distortion-aware video communication with pipeline forwardingabstractThis paper tackles the issue of optimizing the transport of video over packet networks with respect to both resource utilization and user perceived quality. Previous work showed that the quality of service requirements of multimedia applications can be satisfied by pipeline forwarding of packets. However, the current Internet is not based on such technology and its incremental introduction raises questions on how to handle video packets generated by pipeline forwarding unaware sources at the interface between a subnetwork deploying conventional packet scheduling techniques and one implementing pipeline forwarding. This work proposes to use the perceptual importance of the carried video samples to determine which packets shall be transferred with pipeline forwarding - thus receiving deterministic service - and which with a traditional, e.g., best effort or differentiated, service. Simulation results with the first implemented variants of this solution are presented. Mario Baldi, Juan Carlos De Martin, Enrico Masala, Andrea Vesco |
ACM Multimedia | 3 |
| 2005 | Perceptually optimized MPEG compression of synthetic video sequencesabstractThis paper addresses the problem of improving the quality performance of synthetic video sequences by means of standard frame-based coders. The proposed technique can exploit both the knowledge of the 3D model and the intermediate information computed during the rendering process. Firstly, objects are classified, either semantically or automatically, according to their importance. Then the object classification is translated into a macroblock classification, with particular attention to object boundaries. The classification influences the encoder parameters selection, for instance, the quantization parameter. In order to maximize the performance, we propose a rate-distortion formulation of the problem. Experimental results compared with model-unaware encoding show that the proposed techniques can deliver consistent visual quality improvements for different synthetic scenarios using the same bitrate or even less. Demo sequences are available at http://media.polito.it/perceptual3d. Enrico Masala, Davide Quaglia |
ICIP (1) | 1 |
| 2005 | Standard Compatible Error Correction for Multimedia Transmissions Over 802.11 WLANabstractIn this paper, we analyze a standard compatible error correction technique for multimedia transmissions over 802.11 WLANs that exploits, when available, the information of previous erroneous transmissions. The basic idea is to store erroneous frames for error correction purposes. More specifically, at the receiver each bit is estimated with a majority criterion. The performance of different standard compliant error recovery techniques have been evaluated using actual transmission experiments in various channel conditions. Optimal tradeoffs between complexity, memory and perceived quality have been determined studying the quality gains that can be achieved for the specific case of multimedia applications. Perceived quality has been evaluated using objective measures, e.g. ITU-T PESQ for voice and PSNR for video. Results show that the majority combining approach is particularly effective for multimedia communications, even in very noisy scenarios. Gains up to about one unit on the MOS scale for speech and up to 5-6 dB PSNR in case of video have been measured with respect to the standard ARQ technique Enrico Masala, Antonio Servetti, Juan Carlos De Martin |
ICME | 1 |
| 2005 | Performance Evaluation of H.264 Video Streaming over Inter-Vehicular 802.11 Ad Hoc NetworksabstractThis paper evaluates the performance of video streaming in inter-vehicular environments using the 802.11 ad hoc network protocol. We performed transmission experiments while driving two cars equipped with 802.11b standard devices in urban and highway scenarios. Different sequences, bitrates and packctization policies have been tested. The experiments show that each scenario presents peculiar characteristics in terms of average link availability and SNR, which can be exploited to develop more efficient applications. In this paper we also determine the best packetization policies for the two scenarios, showing that large packets lead to better performance in the highway scenario and vice versa. Perceptual quality results indicate that the best packetization policy achieves consistent gains in terms of PSNR values (up to 5 dB), and reduced quality variations, with respect to a fixed-policy transmission technique Paolo Bucciol, Enrico Masala, Nobuo Kawaguchi, Kazuya Takeda, Juan Carlos De Martin |
PIMRC | 2 |
| 2004 | Rate-Distortion Optimized Slicing, Packetization and Coding for Error Resilient Video TransmissionabstractThis paper presents an algorithm to optimize the tradeoff between rate and expected end-to-end distortion of a video sequence transmitted over a packet network. The approach optimizes the source coding parameters, slicing, network QoS class selection and/or error control coding parameters, and accounts for the effects of compression, packetization, error propagation, and concealment at the decoder. It builds on, and substantially extends the applicability of, the recursive optimal per-pixel estimate (ROPE) technique for end-to-end distortion estimation. A trellis-based algorithm is introduced in order to overcome macroblock interdependencies in the estimation procedure, and allow adaptive slicing. Moreover, we propose a complementary packetization scheme to efficiently arrange the slices into packets for FEC protection while minimizing rate loss due to padding. Simulations demonstrate consistent gains over currently used techniques. Enrico Masala, Kenneth Rose, Juan Carlos De Martin |
Data Compression Conference | 1 |
| 2004 | Cross-layer perceptual ARQ for H.264 video streaming over 802.11 wireless networksabstractWe present a new cross-layer ARQ algorithm for video streaming over 802.11 wireless networks. The algorithm combines application-level information about the perceptual and temporal importance of each packet into a single priority value, which drives packet selection at each retransmission opportunity. Hence, only the most perceptually important packets are retransmitted, delivering higher perceptual quality and less bandwidth usage compared to the standard 802.11 MAC-layer ARQ scheme. H.264 video streaming based on the proposed technique has been simulated using ns in a realistic home network scenario, using the standard ARQ technique for all interfering traffic. Results show that the proposed method consistently outperforms the standard MAC-layer 802.11 retransmission scheme, delivering more than 1.5 dB PSNR gains using approximately half of the retransmission bandwidth. Paolo Bucciol, Gabriele Davini, Enrico Masala, Enrica Filippi, Juan Carlos De Martin |
GLOBECOM | 3 |
| 2004 | Perceptual ARQ for H.264 video streaming over 3G wireless networksabstractThis paper presents a new ARQ algorithm for video streaming over wireless channels. The algorithm takes into account the perceptual and temporal importance of each packet to determine the packet scheduling which maximizes the perceived quality. A simple and flexible function to combine the perceptual importance with the real-lime constraints of each packet to determine which is the best packet to transmit at each transmission opportunity is proposed. The perceptual importance is evaluated using the analysis-by-synthesis technique. The performance of the proposed algorithm has been analyzed by simulating the transmission of H.264 encoded sequences over a 144 kbit/s UMTS channel and compares the proposed method with time-driven ARQ techniques, using PSNR as distortion measure. The results show that for the considered channel conditions the proposed method delivers gains up to 2 dB with respect to the time-driven ARQ technique. Paolo Bucciol, Enrico Masala, Juan Carlos De Martin |
ICC | 2 |
| 2003 | A simulative study of analysis-by-synthesis perceptual video classification and transmission over DiffServ IP networksabstractThis paper presents the results of transmission of video data on 2-class DiffServ IP networks using perceptual packet classification and slicing. An analysis-by-synthesis technique to identify perceptually important video regions, to create optimal video slices and to assign the resulting packets to the appropriate DiffServ classes is described. The proposed technique was implemented using the ISO/IEC MPEG-2 video coding standard. Several transmission scenarios, including homogeneous video traffic and interfering FTP traffic, were simulated using network simulator (NS). The proposed perception-based video transmission approach outperformed classical data partitioning in all tested network usage and potential to match time-varying channels. Substantially higher PSNR values than the regular best-effort case were also obtained assigning to the high-QoS class as little as 10% of the traffic. Fabio D'Agostino, Enrico Masala, Laura Farinetti, Juan Carlos De Martin |
ICC | 2 |
| 2003 | Analysis-by-synthesis distortion computation for rate-distortion optimized multimedia streamingabstractThis paper presents an analysis-by-synthesis technique to evaluate the perceptual importance of multimedia packets for rate-distortion optimized streaming. The proposed technique, instead of relying on a priori information, computes the distortion that would be caused by the loss of each single packet, including the effects of error propagation and receiver-side error concealment. A rate-distortion optimized streaming algorithm is presented to compare the perceptual performance obtained using content-adaptive analysis-by-synthesis distortion values versus distortion values obtained using a priori knowledge of the statistical importance of the elements of the compressed multimedia bitstream. Simulations with video test sequences compressed with the MPEG-2 coding standard show that the proposed technique delivers substantial and consistent PSNR gains (1.2-2.8 dB) with respect to ideal frame type-driven a priori distortion evaluation for a wide range of channel conditions. Compared to distortion-agnostic streaming techniques such as SoftARQ, the gain is even more pronounced. Enrico Masala, Juan Carlos De Martin |
ICME | 1 |
| 2001 | Adaptive picture slicing for distortion-based classification of video packetsabstractWe present a new algorithm for the dynamic identification of perceptually important regions in pictures of a video sequence. For each macroblock, the distortion that its loss would cause at the decoder is computed. High-distortion macroblocks are then grouped together into slices to be protected with forward error correction or to be sent as "premium" packets. We developed the approach for the case of video sequences encoded with the ISO MPEG-2 video coding standard. We then applied the algorithm to classify video packets within a 1-bit differentiated services architecture: slices were grouped either into premium packets, to be sent on a "virtual wire", or into regular packets, to be sent as best-effort traffic. In packet losses, the proposed distortion-based classification scheme outperforms source-transparent packet-marking techniques and provides substantially higher PSNR values than the regular best-effort case sending as little as 10% of the packets as premium traffic. Video samples are available at http://multimedia.polito.it/mmsp2001/. Enrico Masala, Davide Quaglia, Juan Carlos De Martin |
MMSP | 1 |