Peter Pocta

dblp:11/9174 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0001-6791-1325ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Computer networks · 5 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A robust deep learning based model for denoising phonocardiogram signals in clinical environments
Maros Jakubec, Eva Lieskovska, Peter Pocta
Eng. Appl. Artif. Intell.3
2024 Multiple Image Distortion DNN Modeling Individual Subject Quality Assessment
abstract
A recent research direction is focused on training Deep Neural Networks (DNNs) to replicate individual subject assessments of media quality. These DNNs are referred to as Artificial Intelligence-based Observers (AIOs). An AIO is designed to simulate, in real-time, the quality ratings of a specific individual, enabling an automatic quality assessment that accounts for subjects characteristics and preferences. Training AIOs is a promising but challenging research area due to the greater noise in individual raw opinion scores compared to the Mean Opinion Score. Effective learning from noisy labels necessitates the training of complex models on large-scale datasets. Unfortunately, this is challenging for AIOs as the media quality assessment community lacks extensive datasets that include individual opinion scores. To address the complexity of the task, we first created a dataset comprising two million samples, with synthetic labels derived from human annotation. We then trained a customized network for image quality assessment, named Multi-Distortion ResNet50 (MDResNet50), on this dataset. The weights of the MDResNet50 were subsequently utilized to initialize the learning process of each AIO, thereby avoiding the need to train a complex model from scratch on a small-scale dataset with raw individual opinion scores. Computational experiments show that our approach significantly advances the state-of-the-art in the AIO research. In particular: (i) we demonstrate through a simulation the ability of AIOs to mimic two well-known behavioral characteristics of a subject, i.e., bias and inconsistency, when scoring the media quality; (ii) we train and release DNN-based AIOs that, compared to the state-of-the-art, exhibit a higher performance with a statistical significance in assessing multiple image distortions; (iii) we train AIOs that more accurately mimic the sensitivity of real subjects to noise and color saturation and also better predict the opinion score distribution compared to the state-of-the-art AIOs.
Lohic Fotio Tiotsop, Antonio Servetti, Peter Pocta, Glenn Van Wallendael, Marcus Barkowsky, Enrico Masala
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Predicting individual quality ratings of compressed images through deep CNNs-based artificial observers
Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Peter Pocta, Tomas Mizdos, Glenn Van Wallendael, Enrico Masala
Signal Process. Image Commun.4
2022 The Effects of Vehicle-to-Infrastructure Communication Reliability on Performance of Signalized Intersection Traffic Control
abstract
Vehicle-to-infrastructure communications can inform an intersection controller about the location and speed of connected vehicles. Recently, the design of adaptive intersection control algorithms that utilize this information received substantial research attention. These studies typically assume perfect communications. This study explores the possible effects of a temporal decrease in the reliability of the communication channel, on the intersection throughput. Road traffic and DSRC-VANET communications are modelled by integrating traffic and communication simulation tools (Vissim and OMNeT++, respectively). Simulations of scenarios with challenging, but realistic communication distortions conditions show significantly larger average delays to vehicles compared to scenarios with perfect communication conditions. These additional delays are largely independent of whether all or only some of the intersection approaches are affected by the communication distortions. Furthermore, delays do not increase uniformly on all signal groups. They may even decrease for some, which causes unfair allocation of green times. The control may be corrected for the lost communications using the information received in previous time intervals and simple assumptions about the vehicle movements. This correction decreases delays in all scenarios both for isolated and connected intersections. It performs similarly to the case with perfect communications when the communication distortions are distributed uniformly among all intersection approaches. Overall, the results demonstrate that the impact of the communication distortions should be considered in the design of the adaptive intersection control algorithms.
Ilya Finkelberg, Tibor Petrov, Ayelet Gal-Tzur, Nina Zarkhin, Peter Pocta, Tatiana Kovacikova, Lubos Buzna, Milan Dado, Tomer Toledo
IEEE Trans. Intell. Transp. Syst.5
2022 Mimicking Individual Media Quality Perception with Neural Network based Artificial Observers
abstract
The media quality assessment research community has traditionally been focusing on developing objective algorithms to predict the result of a typical subjective experiment in terms of Mean Opinion Score (MOS) value. However, the MOS, being a single value, is insufficient to model the complexity and diversity of human opinions encountered in an actual subjective experiment. In this work we propose a complementary approach for objective media quality assessment that attempts to more closely model what happens in a subjective experiment in terms of single observers and, at the same time, we perform a qualitative analysis of the proposed approach while highlighting its suitability. More precisely, we propose to model, using neural networks (NNs) , the way single observers perceive media quality. Once trained, these NNs, one for each observer, are expected to mimic the corresponding observer in terms of quality perception. Then, similarly to a subjective experiment, such NNs can be used to simulate the users’ single opinions, which can be later aggregated by means of different statistical indicators such as average, standard deviation, quantiles, etc. Unlike previous approaches that consider subjective experiments as a black box providing reliable ground truth data for training, the proposed approach is able to consider human factors by analyzing and weighting individual observers. Such a model may therefore implicitly account for users’ expectations and tendencies, that have been shown in many studies to significantly correlate with visual quality perception. Furthermore, our proposal also introduces and investigates an index measuring how much inconsistency there would be if an observer was asked to rate many times the same stimulus. Simulation experiments conducted on several datasets demonstrate that the proposed approach can be effectively implemented in practice and thus yielding a more complete objective assessment of end users’ quality of experience.
Lohic Fotio Tiotsop, Tomas Mizdos, Marcus Barkowsky, Peter Pocta, Antonio Servetti, Enrico Masala
ACM Trans. Multim. Comput. Commun. Appl.4
2021 How to Train No Reference Video Quality Measures for New Coding Standards using Existing Annotated Datasets?
abstract
Subjective experiments are important for developing objective Video Quality Measures (VQMs). However, they are time-consuming and resource-demanding. In this context, being able to reuse existing subjective data on previous video coding standards to train models capable of predicting the perceptual quality of video content processed with newer codecs acquires significant importance. This paper investigates the possibility of generating an HEVC encoded Processed Video Sequence (PVS) in such a way that its perceptual quality is as similar as possible to that of an AVC encoded PVS whose quality has already been assessed by human subjects. In this way, the perceptual quality of the newly generated HEVC encoded PVS may be annotated approximately with the Mean Opinion Score (MOS) of the related AVC encoded PVS. To show the effectiveness of our approach, we compared the performance of a simple and low complexity but yet effective no reference hybrid model trained on the data generated with our approach with the same model trained on data collected in the context of a pristine subjective experiment. In addition, we merged seven subjective experiments such that they can be used as one aligned dataset containing either original HEVC bitstreams or the newly generated data explained in our proposed approach. The merging process accounts for the differences in terms of quality scale, chosen assessment method and context influence factors. This yields a large annotated dataset of HEVC sequences that is made publicly available for the design and training of no reference hybrid VQMs for HEVC encoded content.
Lohic Fotio Tiotsop, Tomas Mizdos, Enrico Masala, Marcus Barkowsky, Peter Pocta
MMSP5
2021 How to reuse existing annotated image quality datasets to enlarge available training data with new distortion types
Tomas Mizdos, Marcus Barkowsky, Miroslav Uhrina, Peter Pocta
Multim. Tools Appl.4
2021 Modeling and estimating the subjects' diversity of opinions in video quality assessment: a neural network based approach
abstract
Abstract Subjective experiments are considered the most reliable way to assess the perceived visual quality. However, observers’ opinions are characterized by large diversity: in fact, even the same observer is often not able to exactly repeat his first opinion when rating again a given stimulus. This makes the Mean Opinion Score (MOS) alone, in many cases, not sufficient to get accurate information about the perceived visual quality. To this aim, it is important to have a measure characterizing to what extent the observed or predicted MOS value is reliable and stable. For instance, the Standard deviation of the Opinions of the Subjects (SOS) could be considered as a measure of reliability when evaluating the quality subjectively. However, we are not aware of the existence of models or algorithms that allow to objectively predict how much diversity would be observed in subjects’ opinions in terms of SOS. In this work we observe, on the basis of a statistical analysis made on several subjective experiments, that the disagreement between the quality as measured by means of different objective video quality metrics (VQMs) can provide information on the diversity of the observers’ ratings on a given processed video sequence (PVS). In light of this observation we: i) propose and validate a model for the SOS observed in a subjective experiment; ii) design and train Neural Networks (NNs) that predict the average diversity that would be observed among the subjects’ ratings for a PVS starting from a set of VQMs values computed on such a PVS; iii) give insights into how the same NN based approach can be used to identify potential anomalies in the data collected in subjective experiments.
Lohic Fotio Tiotsop, Tomas Mizdos, Miroslav Uhrina, Marcus Barkowsky, Peter Pocta, Enrico Masala
Multim. Tools Appl.5
2021 An Adaptive Bitrate Switching Algorithm for Speech Applications in Context of WebRTC
abstract
Web Real-Time Communication (WebRTC) combines a set of standards and technologies to enable high-quality audio, video, and auxiliary data exchange in web browsers and mobile applications. It enables peer-to-peer multimedia sessions over IP networks without the need for additional plugins. The Opus codec, which is deployed as the default audio codec for speech and music streaming in WebRTC, supports a wide range of bitrates. This range of bitrates covers narrowband, wideband, and super-wideband up to fullband bandwidths. Users of IP-based telephony always demand high-quality audio. In addition to users’ expectation, their emotional state, content type, and many other psychological factors; network quality of service; and distortions introduced at the end terminals could determine their quality of experience. To measure the quality experienced by the end user for voice transmission service, the E-model standardized in the ITU-T Rec. G.107 (a narrowband version), ITU-T Rec. G.107.1 (a wideband version), and the most recent ITU-T Rec. G.107.2 extension for the super-wideband E-model can be used. In this work, we present a quality of experience model built on the E-model to measure the impact of coding and packet loss to assess the quality perceived by the end user in WebRTC speech applications. Based on the computed Mean Opinion Score, a real-time adaptive codec parameter switching mechanism is used to switch to the most optimum codec bitrate under the present network conditions. We present the evaluation results to show the effectiveness of the proposed approach when compared with the default codec configuration in WebRTC.
Mohannad Alahmadi, Peter Pocta, Hugh Melvin
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Improved Jitter Buffer Management for WebRTC
abstract
This work studies the jitter buffer management algorithm for Voice over IP in WebRTC. In particular, it details the core concepts of WebRTC’s jitter buffer management. Furthermore, it investigates how jitter buffer management algorithm behaves under network conditions with packet bursts. It also proposes an approach, different from the default WebRTC algorithm, to avoid distortions that occur under such network conditions. Under packet bursts, when the packet buffer becomes full, the WebRTC jitter buffer algorithm may discard all the packets in the buffer to make room for incoming packets. The proposed approach offers a novel strategy to minimize the number of packets discarded in the presence of packet bursts. Therefore, voice quality as perceived by the user is improved. ITU-T Rec. P.863, which also confirms the improvement, is employed to objectively evaluate the listening quality.
Yusuf Cinar, Peter Pocta, Desmond Chambers, Hugh Melvin
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Audiovisual quality of live music streaming over mobile networks using MPEG-DASH
Rafael Rodrigues, Peter Pocta, Hugh Melvin, Marco V. Bernardo, Manuela Pereira, António M. G. Pinheiro
Multim. Tools Appl.2
2016 MPEG DASH - some QoE-based insights into the tradeoff between audio and video for live music concert streaming under congested network conditions
abstract
The rapid adoption of MPEG-DASH is testament to its core design principles that enable the client to make the informed decision relating to media encoding representations, based on network conditions, device type and preferences. Typically, the focus has mostly been on the different video quality representations rather than audio. However, for device types with small screens, the relative bandwidth budget difference allocated to the two streams may not be that large. This is especially the case if high quality audio is used, and in this scenario, we argue that increased focus should be given to the bit rate representations for audio. Arising from this, we have designed and implemented a subjective experiment to evaluate and analyses the possible effect of using different audio quality levels. In particular, we investigate the possibility of providing reduced audio quality so as to free up bandwidth for video under certain conditions. Thus, the experiment was implemented for live music concert scenarios transmitted over mobile networks, and we suggest that the results will be of significant interest to DASH content creators when considering bandwidth tradeoff between audio and video.
Rafael Rodrigues, Peter Pocta, Hugh Melvin, Manuela Pereira, António M. G. Pinheiro
QoMEX2
2015 Can context monitoring improve QoE? A case study of video flash crowds in the internet of services
abstract
Over the last decade or so, significant research has focused on defining Quality of Experience (QoE) of Multimedia Systems and identifying the key factors that collectively determine it. Some consensus thus exists as to the role of System Factors, Human Factors and Context Factors. In this paper, the notion of context is broadened to include information gleaned from simultaneous out-of-band channels, such as social network trend analytics, that can be used if interpreted in a timely manner, to help further optimise QoE. A case study involving simulation of HTTP adaptive streaming (HAS) and load balancing in a content distribution network (CDN) in a flash crowd scenario is presented with encouraging results.
Tobias Hoßfeld, Lea Skorin-Kapov, Yoram Haddad 0001, Peter Pocta, Vasilios A. Siris, Andrej Zgank, Hugh Melvin
IM4
2015 Subjective and objective measurement of synthesized speech intelligibility in modern telephone conditions
Peter Pocta, John G. Beerends
Speech Commun.1