Marcus Barkowsky

dblp:69/217 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-2739-3708ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 5 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Computer networks · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Computational Attention-based Modeling of Individually Perceived Quality of Compressed Images
abstract
Achieving personalized Quality of Experience (QoE) prediction and optimization is a long-term goal in media quality assessment, and recent research has shifted from predicting Mean Opinion Scores (MOS) toward modeling and predicting the quality judgments of individual subjects. In this work, we propose a model of individual image quality assessment that explicitly accounts for two key components of human visual perception: the spatial allocation of attention and the extraction of perceptual information from different regions of the visual scene. The model describes individual quality perception as a multi-stage process in which perceptual information and computational attention are iteratively combined and updated, followed by a probabilistic decision stage. From a theoretical perspective, the proposed formulation is defined at an abstract level and does not rely on any specific neural network architecture, and it provides a framework for explaining how computational attention and perceptual information interact to produce individual quality judgments. To illustrate the practical applicability of the proposed model, we instantiate it using a Vision Transformer (ViT) architecture as one possible implementation, training one AI-based Observer (AIO) per subject with observer-specific learned parameters. Experimental results show that the resulting ViT-based AIOs (i) compare favorably with state-of-the-art deep CNN-based individual models; (ii) outperform state-of-the-art no-reference image quality assessment metrics when used to predict individual opinion scores; (iii) tend to allocate greater attention to regions with severe quality degradation.
Lohic Fotio Tiotsop, Max Geissler, Marcus Barkowsky
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Modeling Subject Scoring Behaviors in Subjective Experiments Based on a Discrete Quality Scale
abstract
Several approaches have been proposed to estimate quality in subjective experiments while highlighting peculiar subject behaviors. However, there is some room for improvement in existing approaches, both in terms of robustness to noise and the ability to accurately indicate several peculiar subject behaviors in subjective experiments. This work advances the state-of-the-art in three main directions: i) A new approach to estimate the subjective quality from noisy ratings is proposed and is shown to be more robust to noise than are four state-of-the-art approaches; ii) a novel subject scoring model is proposed that makes it possible to highlight several peculiar behaviors typically observed in subjective experiments; and iii) our proposed probabilistic subject scoring model results from the proof of a theorem, whereas in previous approaches a probabilistic scoring model is assumed apriori. This represents an important first step toward models supported by a stronger theoretical foundation. Numerical experiments conducted on several datasets highlight the effectiveness of our proposal.
Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Enrico Masala
IEEE Trans. Multim.3
2024 Multiple Image Distortion DNN Modeling Individual Subject Quality Assessment
abstract
A recent research direction is focused on training Deep Neural Networks (DNNs) to replicate individual subject assessments of media quality. These DNNs are referred to as Artificial Intelligence-based Observers (AIOs). An AIO is designed to simulate, in real-time, the quality ratings of a specific individual, enabling an automatic quality assessment that accounts for subjects characteristics and preferences. Training AIOs is a promising but challenging research area due to the greater noise in individual raw opinion scores compared to the Mean Opinion Score. Effective learning from noisy labels necessitates the training of complex models on large-scale datasets. Unfortunately, this is challenging for AIOs as the media quality assessment community lacks extensive datasets that include individual opinion scores. To address the complexity of the task, we first created a dataset comprising two million samples, with synthetic labels derived from human annotation. We then trained a customized network for image quality assessment, named Multi-Distortion ResNet50 (MDResNet50), on this dataset. The weights of the MDResNet50 were subsequently utilized to initialize the learning process of each AIO, thereby avoiding the need to train a complex model from scratch on a small-scale dataset with raw individual opinion scores. Computational experiments show that our approach significantly advances the state-of-the-art in the AIO research. In particular: (i) we demonstrate through a simulation the ability of AIOs to mimic two well-known behavioral characteristics of a subject, i.e., bias and inconsistency, when scoring the media quality; (ii) we train and release DNN-based AIOs that, compared to the state-of-the-art, exhibit a higher performance with a statistical significance in assessing multiple image distortions; (iii) we train AIOs that more accurately mimic the sensitivity of real subjects to noise and color saturation and also better predict the opinion score distribution compared to the state-of-the-art AIOs.
Lohic Fotio Tiotsop, Antonio Servetti, Peter Pocta, Glenn Van Wallendael, Marcus Barkowsky, Enrico Masala
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Training the DNN of a Single Observer by Conducting Individualized Subjective Experiments
abstract
Predicting the quality perception of an individual subject instead of the mean opinion score is a new and very promising research direction. Deep Neural Networks (DNNs) are suitable for such prediction but the training process is particularly data demanding due to the noisy nature of individual opinion scores. We propose a human-in-the-loop training process using multiple cycles of a human voting, DNN training, and inference procedure. Thus, opinion scores on individualized sets of images were progressively collected from each observer to refine the performance of their DNN. The results of computational experiments demonstrate the effectiveness of our approach. For future research and benchmarking, five DNNs trained to mimic five observers are released together with a dataset containing the 1500 opinion scores progressively gathered from each of these observers during our training cycles.
Pavel Majerl, Lohic Fotio Tiotsop, Marcus Barkowsky
QoMEX3
2023 Predicting individual quality ratings of compressed images through deep CNNs-based artificial observers
Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Peter Pocta, Tomas Mizdos, Glenn Van Wallendael, Enrico Masala
Signal Process. Image Commun.3
2022 Regularized Maximum Likelihood Estimation of the Subjective Quality from Noisy Individual Ratings
abstract
Despite several approaches to recover the ground truth subjective quality score from noisy individual ratings in subjective experiments have been explored in the literature, there is still room for improvement, in particular in terms of robustness to noise. This paper proposes a new approach that combines the traditional maximum likelihood estimation framework with a newly proposed regularization term, based on information theory concepts, that is meant to underweight surprising ratings of the quality of a given stimulus, looked at as a noise manifestation, in the final analytical expression of the recovered subjective quality. Computational experiments show the higher robustness to noise of our proposal when compared to three state-of-the-art methods.
Lohic Fotio Tiotsop, Antonio Servetti, Marcus Barkowsky, Enrico Masala
QoMEX3
2022 Mimicking Individual Media Quality Perception with Neural Network based Artificial Observers
abstract
The media quality assessment research community has traditionally been focusing on developing objective algorithms to predict the result of a typical subjective experiment in terms of Mean Opinion Score (MOS) value. However, the MOS, being a single value, is insufficient to model the complexity and diversity of human opinions encountered in an actual subjective experiment. In this work we propose a complementary approach for objective media quality assessment that attempts to more closely model what happens in a subjective experiment in terms of single observers and, at the same time, we perform a qualitative analysis of the proposed approach while highlighting its suitability. More precisely, we propose to model, using neural networks (NNs) , the way single observers perceive media quality. Once trained, these NNs, one for each observer, are expected to mimic the corresponding observer in terms of quality perception. Then, similarly to a subjective experiment, such NNs can be used to simulate the users’ single opinions, which can be later aggregated by means of different statistical indicators such as average, standard deviation, quantiles, etc. Unlike previous approaches that consider subjective experiments as a black box providing reliable ground truth data for training, the proposed approach is able to consider human factors by analyzing and weighting individual observers. Such a model may therefore implicitly account for users’ expectations and tendencies, that have been shown in many studies to significantly correlate with visual quality perception. Furthermore, our proposal also introduces and investigates an index measuring how much inconsistency there would be if an observer was asked to rate many times the same stimulus. Simulation experiments conducted on several datasets demonstrate that the proposed approach can be effectively implemented in practice and thus yielding a more complete objective assessment of end users’ quality of experience.
Lohic Fotio Tiotsop, Tomas Mizdos, Marcus Barkowsky, Peter Pocta, Antonio Servetti, Enrico Masala
ACM Trans. Multim. Comput. Commun. Appl.3
2021 How to Train No Reference Video Quality Measures for New Coding Standards using Existing Annotated Datasets?
abstract
Subjective experiments are important for developing objective Video Quality Measures (VQMs). However, they are time-consuming and resource-demanding. In this context, being able to reuse existing subjective data on previous video coding standards to train models capable of predicting the perceptual quality of video content processed with newer codecs acquires significant importance. This paper investigates the possibility of generating an HEVC encoded Processed Video Sequence (PVS) in such a way that its perceptual quality is as similar as possible to that of an AVC encoded PVS whose quality has already been assessed by human subjects. In this way, the perceptual quality of the newly generated HEVC encoded PVS may be annotated approximately with the Mean Opinion Score (MOS) of the related AVC encoded PVS. To show the effectiveness of our approach, we compared the performance of a simple and low complexity but yet effective no reference hybrid model trained on the data generated with our approach with the same model trained on data collected in the context of a pristine subjective experiment. In addition, we merged seven subjective experiments such that they can be used as one aligned dataset containing either original HEVC bitstreams or the newly generated data explained in our proposed approach. The merging process accounts for the differences in terms of quality scale, chosen assessment method and context influence factors. This yields a large annotated dataset of HEVC sequences that is made publicly available for the design and training of no reference hybrid VQMs for HEVC encoded content.
Lohic Fotio Tiotsop, Tomas Mizdos, Enrico Masala, Marcus Barkowsky, Peter Pocta
MMSP4
2021 How to reuse existing annotated image quality datasets to enlarge available training data with new distortion types
Tomas Mizdos, Marcus Barkowsky, Miroslav Uhrina, Peter Pocta
Multim. Tools Appl.2
2021 Modeling and estimating the subjects' diversity of opinions in video quality assessment: a neural network based approach
abstract
Abstract Subjective experiments are considered the most reliable way to assess the perceived visual quality. However, observers’ opinions are characterized by large diversity: in fact, even the same observer is often not able to exactly repeat his first opinion when rating again a given stimulus. This makes the Mean Opinion Score (MOS) alone, in many cases, not sufficient to get accurate information about the perceived visual quality. To this aim, it is important to have a measure characterizing to what extent the observed or predicted MOS value is reliable and stable. For instance, the Standard deviation of the Opinions of the Subjects (SOS) could be considered as a measure of reliability when evaluating the quality subjectively. However, we are not aware of the existence of models or algorithms that allow to objectively predict how much diversity would be observed in subjects’ opinions in terms of SOS. In this work we observe, on the basis of a statistical analysis made on several subjective experiments, that the disagreement between the quality as measured by means of different objective video quality metrics (VQMs) can provide information on the diversity of the observers’ ratings on a given processed video sequence (PVS). In light of this observation we: i) propose and validate a model for the SOS observed in a subjective experiment; ii) design and train Neural Networks (NNs) that predict the average diversity that would be observed among the subjects’ ratings for a PVS starting from a set of VQMs values computed on such a PVS; iii) give insights into how the same NN based approach can be used to identify potential anomalies in the data collected in subjective experiments.
Lohic Fotio Tiotsop, Tomas Mizdos, Miroslav Uhrina, Marcus Barkowsky, Peter Pocta, Enrico Masala
Multim. Tools Appl.4
2019 Computing Quality-of-Experience Ranges for Video Quality Estimation
abstract
Typically, the measurement of the Quality of Experience for video sequences aims at a single value, in most cases the Mean Opinion Score (MOS). Predicting this value using various algorithms has been widely studied. However, deviation from the MOS is often handled as an unpredictable error. The approach in this contribution estimates intervals of video quality instead of the single valued MOS. Well-known video quality estimators are fused together to output a lower and upper border for the expected video quality, on the basis of a model derived from a well-known subjectively annotated dataset. Results on different datasets provide insight on the suitability of the well-known estimators for this particular approach.
Lohic Fotio Tiotsop, Enrico Masala, Ahmed Aldahdooh, Glenn Van Wallendael, Marcus Barkowsky
QoMEX5
2019 Improving relevant subjective testing for validation: Comparing machine learning algorithms for finding similarities in VQA datasets using objective measures
Ahmed Aldahdooh, Enrico Masala, Glenn Van Wallendael, Peter Lambert, Marcus Barkowsky
Signal Process. Image Commun.5
2019 Improved Performance Measures for Video Quality Assessment Algorithms Using Training and Validation Sets
abstract
The training and performance analysis of objective video quality assessment algorithms is complex due to the huge variety of possible content classes and transmission distortions. Several secondary issues such as free parameters in machine learning algorithms and alignment of subjective datasets put an additional burden on the developer. In this paper, three subsequent steps are presented to address such issues. First, the content and coding parameter space of a large-scale database is used to select dedicated subsets for training objective algorithms. This aims at providing a method for selecting the most significant contents and coding parameters from all imaginable combinations. In the practical case where only a limited set is available, it also helps us to avoid redundancy in the training subset selection. The second step is a discussion on performance measures for algorithms that employ machine-learning methods. The particularity of the performance measures is that the quality of the training and verification datasets is taken into consideration. Common issues that often use existing measures are presented, and improved or complementary methods are proposed. The measures are applied to two examples of no-reference objective assessment algorithms using the aforementioned subsets of the large-scale database. While limited in terms of practical applications, this sandbox approach of objectively predicting an objectively evaluated video sequences allows for eliminating additional influence factors from subjective studies. In the third step, the proposed performance measures are applied to the practical case of training and analyzing assessment algorithms on readily available subjectively annotated image datasets. The presentation method in this part of the paper can also be used as an exemplified recommendation for reporting in-depth information on the performance. Using this presentation method, future publications presenting newly developed quality assessment algorithms may be significantly improved.
Ahmed Aldahdooh, Enrico Masala, Olivier Janssens, Glenn Van Wallendael, Marcus Barkowsky, Patrick Le Callet
IEEE Trans. Multim.5
2018 Proof-of-concept: role of generic content characteristics in optimizing video encoders - Complexity- and content-aware sequence-level encoder parameter decision framework
Ahmed Aldahdooh, Marcus Barkowsky, Patrick Le Callet
Multim. Tools Appl.2
2017 Inpainting-based error concealment for low-delay video communication
abstract
Error concealment (EC) is one of the target applications of inpainting techniques. Some methods combine the estimated lost motion vectors (MVs) with the exemplar-based inpainting technique to recover the lost regions. Due to the erroneous motion vectors that might indicate a moving object as background object and vice versa, these methods are still showing visual artifacts in the recovered regions. In this paper, a concept of motion map that can be easily generated in the decoder side is introduced and it is combined with the exemplar-based inpainting technique. The proposed method introduces an adaptive search window size that trades-off the quality and complexity. Moreover, an optional blending technique is proposed to limit the spatio-temporal artifacts. Experiments show that the proposed method improves the visual quality with 5dB on average relative to the state-of-the-art inpainting-based EC method.
Ahmed Aldahdooh, Marcus Barkowsky, David Bull 0001, Patrick Le Callet
ICASSP2
2016 Effect of content features on short-term video quality in the visual periphery
abstract
The area outside our central field of vision, also referred to as the visual periphery, captures most information in a visual scene, although much less sensitive than the central Fovea. Vision studies in the past have stated that there is reduced sensitivity of texture, color, motion and flicker (temporal harmonic) perception in this area, that bears an interesting application in the domain of quality perception. In this work, we particularly analyze the perceived subjective quality of videos containing H.264/AVC transmission impairments, incident at various degrees of retinal eccentricities of observers. We relate the perceived drop in quality, to five basic types of features that are important from a perceptive standpoint: texture, color, flicker, motion trajectory distortions and also the semantic importance of the underlying regions. We are able to observe that the perceived drop in quality across the visual periphery, is closely related to the Cortical Magnification fall-off characteristics of the V1 cortical region. Additionally, we see that while object importance and low frequency spatial distortions are important indicators of quality in the central foveal region, temporal flicker and color distortions are the most important determinants of quality in the periphery. We therefore conclude that, although users are more forgiving of distortions they viewed peripherally, they are nevertheless not totally blind towards it: the effects of flicker and color distortions being particularly important.
Yashas Rai, Ahmed Aldahdooh, Suiyi Ling, Marcus Barkowsky, Patrick Le Callet
MMSP4
2016 Comparing temporal behavior of fast objective video quality measures on a large-scale database
abstract
In many application scenarios, video quality assessment is required to be fast and reasonably accurate. The characterization of objective algorithms by subjective assessment is well established but limited due to the small number of test samples. Verification using large-scale objectively annotated databases provides a complementary solution. In this contribution, three simple but fast measures are compared regarding their agreement on a large-scale database. In contrast to subjective experiments, not only sequence-wise but also framewise agreement can be analyzed. Insight is gained into the behavior of the measures with respect to 5952 different coding configurations of High Efficiency Video Coding (HEVC). Consistency within a video sequence is analyzed as well as across video sequences. The results show that the occurrence of discrepancies depends mostly on the configured coding structure and the source content. The detailed observations stimulate questions on the combined usage of several video quality measures for encoder optimization.
Ahmed Aldahdooh, Enrico Masala, Glenn Van Wallendael, Marcus Barkowsky
PCS4
2016 Spatio-temporal error concealment technique for high order multiple description coding schemes including subjective assessment
abstract
Error resilience (ER) is an important tool in video coding to maximize the quality of Experience (QoE). The prediction process in video coding became complex which yields an unsatisfying video quality when NALunit packets are lost in error-prone channels. There are different ER techniques and multiple description coding (MDC) is one of the promising technique for this problem. MDC is categorized into different types and, in this paper, we focus on temporal MDC techniques. In this paper, a new temporal MDC scheme is proposed. In the encoding process, the encoded descriptions contain primary frames and secondary frames (redundant representations). The secondary frames represent the MVs that are predicted from previous primary frames such that the residual signal is set to zero and is not part of the rate distortion optimization. In the decoding process of the lost frames, a weighted average error concealment (EC) strategy is proposed to conceal these frames. The proposed scheme is subjectively evaluated along with other schemes and the results show that the proposed scheme is significantly different from most of other temporal MDC schemes.
Ahmed Aldahdooh, Marcus Barkowsky, Patrick Le Callet
QoMEX2
2016 Comparing simple video quality measures for loss-impaired video sequences on a large-scale database
abstract
The performance of objective video quality measures is usually identified by comparing their predictions to subjective assessment results which are regarded as the ground truth. In this work we propose a complementary approach for this performance evaluation by means of a large-scale database of test sequences evaluated with several objective measurement algorithms. Such an approach is expected to detect performance anomalies that could highlight shortcomings in current objective measurement algorithms. Using realistic coding and network transmission conditions, we investigate the consistency of the prediction of different measures as well as how much their behavior can be predicted by content, coding and transmission features, discussing unexpected and peculiar behaviors, and highlighting how a large-scale database can help in identifying anomalies not easily found by means of subjective testing. We expect that this analysis will shed light on directions to pursue in order to overcome some of the limitations of existing reliability assessment methods for objective video quality measures.
Ahmed Aldahdooh, Enrico Masala, Olivier Janssens, Glenn Van Wallendael, Marcus Barkowsky
QoMEX5
2016 Studying user agreement on aesthetic appeal ratings and its relation with technical knowledge
abstract
In this paper, a crowdsourcing experiment was conducted involving different panels of participants. The aim of this study is to evaluate how the preference of one image over another one is related with the knowledge of the participant in photography. In previous work the two discriminant evaluation concepts “presence of a main subject” and “exposure” were found to distinguish group participants with different degrees of knowledge in photography. Each of these groups provided different means of aesthetic appeal ratings when asked to rate on an absolute category scale. The present paper extends previous work by studying preference ratings on a set of image pairs as a function of technical knowledge and more specifically adding a focus on the variance of rating and agreement between participants. The conducted study was composed of two different steps where the participants had to first report their preference of one image over another (paired comparison), and an evaluation of the technical background of the participant using a specific set of images. Based on preference-rating patterns groups of participants were identified. These groups were formed by clustering the participants who saw and shared the same preference rating on images in one group, and the participants with low agreement with other participants in another group. A per-group analysis showed that a high agreement between participants could be observed when participants have technical knowledge. This indicates that higher consistency between participants can be reached when expert users are being recruited, and therefore participants should be carefully selected in image aesthetic appeal evaluation to ensure stable results.
Pierre R. Lebreton, Alexander Raake, Marcus Barkowsky
QoMEX3
2015 Hybrid video quality prediction: reviewing video quality measurement for widening application scope
abstract
A tremendous number of objective video quality measurement algorithms have been developed during the last two decades. Most of them either measure a very limited aspect of the perceived video quality or they measure broad ranges of quality with limited prediction accuracy. This paper lists several perceptual artifacts that may be computationally measured in an isolated algorithm and some of the modeling approaches that have been proposed to predict the resulting quality from those algorithms. These algorithms usually have a very limited application scope but have been verified carefully. The paper continues with a review of some standardized and well-known video quality measurement algorithms that are meant for a wide range of applications, thus have a larger scope. Their individual artifacts prediction accuracy is usually lower but some of them were validated to perform sufficiently well for standardization. Several difficulties and shortcomings in developing a general purpose model with high prediction performance are identified such as a common objective quality scale or the behavior of individual indicators when confronted with stimuli that are out of their prediction scope. The paper concludes with a systematic framework approach to tackle the development of a hybrid video quality measurement in a joint research collaboration.
Marcus Barkowsky, Iñigo Sedano, Kjell Brunnström, Mikolaj Leszczuk, Nicolas Staelens
Multim. Tools Appl.1
2014 Optimizing feature pooling and prediction models of VQA algorithms
abstract
In this paper, we propose a strategy to optimize feature pooling and prediction models of video quality assessment (VQA) algorithms with a much smaller number of parameters than methods based on machine learning, such as neural networks. Based on optimization, the proposed mapping strategy is composed of a global linear model for pooling extracted features, a simple linear model for local alignment in which local factors depend on source videos, and a non-linear model for quality calibration. Also, a reduced-reference VQA algorithm is proposed to predict the local factors from the source video. In the IRCCyN/IVC video database of content influence and the LIVE mobile video database, the performance of VQA algorithms is improved significantly by local alignment. The proposed mapping strategy with prediction of local factors outperforms one no-reference VQA metric and is comparable to one full-reference VQA metric. Thus predicting the local factors in local alignment based on video content will be a promising new approach for VQA.
Kongfeng Zhu, Marcus Barkowsky, Minmin Shen, Patrick Le Callet, Dietmar Saupe
ICIP2
2013 Measurement of Individual Changes in the Performance of Human Stereoscopic Vision for Disparities at the Limits of the Zone of Comfortable Viewing
abstract
3D displays enable immersive visual impressions but the impact on the human perception still is not fully understood. Viewing conditions like the convergence-accommodation (C-A) conflict have an unnatural influence on the visual system and might even lead to visual discomfort. As visual perception is individual we assumed the impact of simulated 3D content on the visual system to be as well. In this study we aimed to analyze the stereoscopic visual performance of 17 subjects for disparities inside and outside the in literature defined zone of comfortable viewing to provide an individual evaluation of the impact of increased disparities on the performance of the visual system. Stereoscopic stimuli were presented in a four-alternative forced choice (4AFC) setup in different disparities. The response times as well as the correct decision rates indicated the performance of stereoscopic vision. The results showed that increased disparities lead to a decline in performance. Further, the impact of the presented disparities is dependent on the difficulty of the task. The decline of performance as well as the deciding disparities for the decline were subject dependent.
Jan Paulus, Georg Michelson, Marcus Barkowsky, Joachim Hornegger, Björn M. Eskofier, Michael Schmidt 0004
3DV3
2013 Open collaboration on hybrid video quality models - VQEG joint effort group hybrid
abstract
Several factors limit the advances on automatizing video quality measurement. Modelling the human visual system requires multi- and interdisciplinary efforts. A joint effort may bridge the large gap between the knowledge required in conducting a psychophysical experiment on isolated visual stimuli to engineering a universal model for video quality estimation under real-time constraints. The verification and validation requires input reaching from professional content production to innovative machine learning algorithms. Our paper aims at highlighting the complex interactions and the multitude of open questions as well as industrial requirements that led to the creation of the Joint Effort Group in the Video Quality Experts Group. The paper will zoom in on the first activity, the creation of a hybrid video quality model.
Marcus Barkowsky, Nicolas Staelens, Lucjan Janowski
MMSP1
2012 Analysis and improvement of a paired comparison method in the application of 3DTV subjective experiment
abstract
Paired comparison is a frequently used method in psychophysical studies. However, with the increase of the number of the stimuli, the number of comparisons increases exponentially. Square design is one of the balanced sub-set paired comparison methods which could reduce the number of comparisons while producing comparably precise results under some assumptions. However, when there are observation errors from observers' attentiveness, the square design would produce large estimation errors. Thus, an improved square design which is robust to observation errors is proposed. Using a Monte Carlo simulation, the proposed method is evaluated and shows improvement in efficiency. The original design is applied in a visual discomfort subjective test of 3DTV. In addition, both of the two designs are studied by utilizing our previous full comparison data. The test results showed that the proposed improved square design is more robust to observation errors. Another important finding is that the influence of the occurrence of some other stimuli on voting is significant. Whether the proposed method could reduce the prediction errors induced by it is still under study.
Jing Li 0026, Marcus Barkowsky, Patrick Le Callet
ICIP2
2011 A Subjective Evaluation of 3D Iptv Broadcasting Implementations Considering Coding and Transmission Degradation
abstract
This paper describes the results of a subjective test to assess current technology used for 3DTV broadcasting. As a first aspect, the performance of the currently deployed coding schemes was compared to state of the art algorithms. Our results show that down sampling and packing 3D stereoscopic videos according to the so called Side-By-Side format gives the highest perceived quality for a given bit rate. The second aspect of the study was to investigate how common 2D error concealment algorithms perform in case of 3D, and how their 3D-related performance compares with the 2D case. The results provide information on whether binocular suppression or binocular rivalries play the most important role for 3D video quality under transmission error. The results indicate that binocular rivalries and related visual discomfort are the dominant factors. Another aspect of the paper is a comparison of the test results with results from different labs to evaluate the repeatability of a subjective experiment in the 3D case, and to compare the employed test methodologies. Here, the study shows the variation between observers when they are rating visual discomfort and illustrates the difficulty to evaluate this new dimension.
Pierre R. Lebreton, Alexander Raake, Marcus Barkowsky, Patrick Le Callet
ISM3
2010 On the perceptual similarity of realistic looking tone mapped High Dynamic Range images
abstract
High Dynamic Range (HDR) images are usually displayed on conventional Low Dynamic Range (LDR) displays because of the limited availability of HDR displays. For the conversion of the large dynamic luminance range into the eight bit quantized values, parameterized Tone Mapping Operators (TMO) are applied. Human observers are able to optimize the parameters in order to get the highest Quality of Experience by judging the displayed LDR images on a realism scale. In the study presented in this paper, two TMOs with three parameters each were evaluated by observers in a subjective experiment. Although the chosen parameter settings vary largely, the chosen images appear to have the same QoE for the observers. In order to assess this similarity objectively, three commonly used image quality measurement algorithms were applied. Their agreement with the preference of the observers was analyzed and it was found that the Visual Difference Predictor (VDP) outperforms the Structural Similarity Index and the Root Mean Square Error. A threshold value for VDP is derived that indicates when two LDR images appear to have the same Quality of Experience.
Marcus Barkowsky, Patrick Le Callet
ICIP1
2010 Video quality assessment: From 2D to 3D - Challenges and future trends
abstract
Three-dimensional (3D) video is gaining a strong momentum both in the cinema and broadcasting industries as it is seen as a technology that will extensively enhance the user's visual experience. One of the major concerns for the wide adoption of such technology is the ability to provide sufficient visual quality, especially if 3D video is to be transmitted over a limited bandwidth for home viewing (i.e. 3DTV). Means to measure perceptual video quality in an accurate and practical way is therefore of highest importance for content providers, service providers, and display manufacturers. This paper discusses recent advances in video quality assessment and the challenges foreseen for 3D video. Both subjective and objective aspects are examined. An outline of ongoing efforts in standards-related bodies is also provided.
Quan Huynh-Thu, Patrick Le Callet, Marcus Barkowsky
ICIP3
2008 Histogram-Based Prefiltering for Luminance and Chrominance Compensation of Multiview Video
abstract
Significant advances have recently been made in the coding of video data recorded with multiple cameras. However, luminance and chrominance variations between the camera views may deteriorate the performance of multiview codecs and image-based rendering algorithms. A histogram matching algorithm can be applied to efficiently compensate for these differences in a prefiltering step. A mapping function is derived which adapts the cumulative histogram of a distorted sequence to the cumulative histogram of a reference sequence. If all camera views of a multiview sequence are adapted to a common reference using histogram matching, the spatial prediction across camera views is improved. The basic algorithm is extended in three ways: a time-constant calculation of the mapping function, RGB color conversion, and the use of global disparity compensation. The best coding results are achieved when time-constant histogram calculation and RGB color conversion are combined. In this case, the usage of histogram matching prior to multiview encoding leads to substantial gains in the coding efficiency of up to 0.7 dB for the luminance component and up to 1.9 dB for the chrominance components. This prefiltering step can be combined with block-based illumination compensation techniques that modify the coder and decoder themselves, especially with the approach implemented in the multiview reference software of the joint video team (JVT). Additional coding gains up to 0.4 dB can be observed when both methods are combined.
Ulrich Fecker, Marcus Barkowsky, André Kaup
IEEE Trans. Circuits Syst. Video Technol.2
2007 Temporal registration using 3D phase correlation and a maximum likelihood approach in the perceptual evaluation of video quality
abstract
The estimation of the video quality is often performed using a full reference approach. One of the most important steps in a video quality measurement algorithm is to find the corresponding frames between the reference and the distorted video sequence. In this paper an algorithm with three steps is proposed. First, an extended version of the phase correlation is used to find candidate images with an arbitrary temporal offset, spatial scaling or spatial shift. Based on the assumption that the spatial scaling and spatial shift does not change during the sequence a set of probable parameters is selected. Finally, a maximum likelihood estimation is applied to select those temporal offsets which support the smoothest playback. A set of video sequences degraded with several distortions which are typical for multimedia scenarios are used to compare the performance to other algorithms.
Marcus Barkowsky, Jens Bialkowski, Roland Bitto, André Kaup
MMSP1
2007 Fast video transcoding from H.263 to H.264/MPEG-4 AVC
Jens Bialkowski, Marcus Barkowsky, André Kaup
Multim. Tools Appl.2
2006 Influence of the Presentation Time on Subjective Votings of Coded Still Images
abstract
The quality of coded images is often assessed by a subjective test. Usually the viewers get as much time as they need to find a stable result. In video sequences however, the viewer has to judge the quality in a shorter time that is defined by the changing content or a following scene cut. Therefore it is desirable to know the influence of a shorter presentation time on the perceptibility of distortions. In this paper we present the results of a suitable subjective test on coded still images. The images were presented for six different durations, ranging from 200 ms to 3 s. Special care was taken to avoid the memorization effect usually present after short presentations. The results show that the viewers tend to avoid extreme votings at short durations. The variance of the votings is also discussed in detail. Based on the result of the voting for the longest presentation time, we propose a prediction model for the voting of the shorter durations using a logistic curve fit. This presentation time model (PTM) is presented and analysed in detail.
Marcus Barkowsky, Björn M. Eskofier, Jens Bialkowski, André Kaup
ICIP1
2006 Low-Complexity Transcoding of Inter Coded Video Frames from H.264 to H.263
abstract
The presented work addresses the reduction of computational complexity for transcoding of interframes from H.264 to H.263 baseline profiles maintaining the quality of a full search approach. This scenario aims to achieve fast backward compatible interoperability inbetween new and existing video coding platforms, e.g. between DVB-H and UMTS. By exploiting side information of the H.264 input bitstream the encoding complexity of the motion estimation is strongly reduced. Due to the possibility to divide a macroblock (MB) into partitions with different motion vectors (MV), one single MV has to be selected for H.263. It will be shown, that this vector is suboptimal for all sequences, even if all existing MVs of a MB of H.264 are compared as candidate. Also motion vector refinement with a fixed ½-pel refinement window as used by transcoders throughout the literature is not sufficient for scenes with fast movement. We propose an algorithm for selecting a suitable vector candidate from the input bitstream and this MV is then refined using an adaptive window. Using this technique, the complexity is still low at nearly optimum rate-distortion results compared to an exhaustive full-search approach.
Jens Bialkowski, Marcus Barkowsky, Florian Leschka, André Kaup
ICIP2
2006 Overview of Low-Complexity Video Transcoding from H.263 to H.264
abstract
With the standardization of H.264/AVC by ITU-T and ISO/IEC and the adaptatation into new hardware, the necessity of transcoding between existing standards and H.264 will arise to achieve interoperability between hardware devices. Because of the many new prediction parameters as well as the pixel-based deblocking filter and the new transform of H.264 this is a difficult task to perform. In our work we propose a fast cascaded pixel-domain transcoder from H.263 to H.264 for both intra- and inter-frame coding. The rate-distortion (RD) performance of the encoded bitstreams is compared to an exhaustive full-search approach. Our approach leads to 9% higher data rate in average, but the computational complexity for the prediction can be reduced by 90% and more. It will be shown that the algorithms proposed for H.263 are applicable for transcoding MPEG-2 to H.264, too
Jens Bialkowski, Marcus Barkowsky, André Kaup
ICME2
2005 On Requantization in Intra-Frame Video Transcoding with Different Transform Block Sizes
abstract
Transcoding is a technique to convert one video bit-stream into another. While homogeneous transcoding is done at the same coding standard, inhomogeneous transcoding converts from one standard format to another standard. Inhomogeneous transcoding between MPEG-2, MPEG-4 or H.263 was performed using the same transform. With the standardisation of H.264 also a new transform basis and different block size was defined. For requantization from block size 8times8 to 4times4 this leads to the effect that the quantization error of one coefficient in a block of size 8times8 is distributed over multiple coefficients in blocks of size 4times4. In our work, we analyze the requantization process for inhomogeneous transcoding with different transforms. The deduced equations result in an expression for the correlation of the error contributions from the coefficients of block size 8times8 at each coefficient of block size 4times4. We then compare the mathematical analysis to simulations on real sequences. The reference to the requantization process is the direct quantization of the undistorted signal. It will be shown that the loss is as high as 3 dB PSNR at equivalent step size for input and output bitstream. Also an equation for the choice of the second quantization step size in dependency of the requantization loss is deduced. The model is then extended from the DCT to the integer-based transform as defined in H.264
Jens Bialkowski, Marcus Barkowsky, André Kaup
MMSP2