EDBT 2026 Demo / reviewers in the wild / expert
Lukas Krasula
dblp:154/3698 · also Lukás Krasula
· DBLP profile ↗
22ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0003-1205-7429ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Estimating the resize parameter in end-to-end learned image compression
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Lukas Krasula, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2024 | "Discriminability-Experimental Cost" Tradeoff in Subjective Video Quality Assessment of Codec: DCR with EVP Rating Scale Versus ACR-HRabstractThis work uses naive observers to compare two subjective studies conducted in a controlled laboratory environment on SDR HD, UHD, and HDR UHD contents. These tests aim to compare the precision and accuracy of a modified Degradation Category Rating (DCR) and Absolute Category Rating with Hidden Reference (ACR-HR) subjective methods for video quality assessment. The modified version of the DCR method includes a repetition of both reference and distorted stimuli; and utilizes an 11-grade rating scale from Expert Viewing Protocol (EVP) of ITU-R BT.500-15 standards. In the second subjective protocol, ACR-HR operates without repetition and with the 5-grade quality scale from ITU standards. We extensively analyze the scale usage and compare Mean Opinion Score (MOS) discriminability in both subjective studies. We show that both methods can retrieve accurate MOS. However, the ACR-HR method achieves better discriminability among MOS than DCR with the EVP rating scale while reducing the experimental effort by a factor of two, i.e., the cost of the experiment. The findings of this work give new insight into how to perform cost-efficient subjective tests for video quality estimation with naive observers and how to retrieve good MOS estimates. Andreas Pastor, Ioannis Katsavounidis, Lukas Krasula, Andrey Norkin, Hassene Tmar, Patrick Le Callet |
PCS | 4 |
| 2024 | The effect of viewing distance and display peak luminance - HDR AV1 video streaming quality datasetabstractWhile it is well recognized that the visibility of distortions is affected by the viewing distance and display peak luminance, very few datasets control those conditions, and also few video quality metrics can account for them. To address this gap, we collected a new video quality dataset, HDR-VDC, which captures the quality degradation of HDR content due to AV1 coding artifacts and the resolution reduction. The quality drop was measured at two viewing distances, corresponding to 60 and 120 pixels per visual degree, and two display mean luminance levels, 51 and 5.6 nits. In contrast to the existing datasets that use direct rating protocol, we employ a highly sensitive pairwise comparison protocol with active sampling and comparisons across viewing distances to ensure possibly accurate quality measurements. We also provide the first publicly available dataset that measures the effect of display peak luminance and includes HDR videos encoded with AV1. Our results indicate that the effect of both viewing distance and display luminance is significant, and it reduces the visibility of coding and upsampling artifacts on dimmer displays or those seen from a further distance. The dataset is available at https://doi.org/10.17863/CAM.107964 and the code at https://github.com/gfxdisp/HDR-VDC. Dounia Hammou, Lukas Krasula, Christos G. Bampis, Zhi Li 0001, Rafal Mantiuk |
QoMEX | 2 |
| 2023 | Recovering Quality Scores in Noisy Pairwise Subjective Experiments Using Negative Log-LikelihoodabstractTo gather larger datasets to train data-angry deep learning quality assessment models, crowdsourcing has become essential to recruit participants. These participants are asked their opinion by directly rating stimuli, e.g., using single or double stimulus methodologies, or indirectly by ranking stimuli or comparing distances as in the Maximum Likelihood Difference Scaling method. In crowdsourcing, participants’ behaviors and environmental distractions are not controlled. So, the researcher must pay attention to the answers’ reliability. Cleaning methods exist for direct annotation subjective methodologies. However, solutions for indirect annotation methods are limited. In this work, we propose a method based on the negative log-likelihood to detect spammers among participants from their answers. To demonstrate its use, we applied it in a quadruplet preference-based scenario. The proposed method requires low computation and can be integrated into active-sampling strategies, where annotations available per comparison are small. We demonstrate that our method is robust to various spammer behaviors and accurate by removing only spammers. It helps reduce the gap between data collected in in-lab conditions (i.e., no spammer) and through crowdsourcing: our method reduces estimated uncertainties around data-points by 50%, and RMSE between estimations from an in-lab experiment and the same experiment performed in crowdsourcing by 1.8. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
ICIP | 2 |
| 2023 | Comparison of Metrics for Predicting Image and Video Quality at Varying Viewing DistancesabstractViewing distance and display resolution have ar-guably a significant impact on perceived image quality; images seen on a mobile phone with high pixel density reveal fewer distortions than the same images seen on a large TV from a close distance. However, only a few image and video quality metrics account for the effect of viewing distance and resolution. Those that do, typically rely on contrast sensitivity functions (CSFs) of the visual system. Other metrics can be potentially adapted to different viewing distances by rescaling input images. In this paper, we investigate the performance of such adapted metrics together with those that natively account for viewing distance. The results for three testing datasets indicate that there is no evidence that the metrics based on the CSF outperform those that rely on rescaled images. Moreover, we found that both methods are not successful to account for the changes in quality introduced by the change in viewing distance. We conclude that accounting for viewing distances requires better models. Dounia Hammou, Lukas Krasula, Christos G. Bampis, Zhi Li 0001, Rafal Mantiuk |
MMSP | 2 |
| 2023 | Measuring and Predicting Perceptions of Video Quality Across Screen Sizes with CrowdsourcingabstractA large-scale crowdsourcing experiment was carried out to study how changes in screen size affect perceptions of video quality. By rescaling video stimuli to different canvas sizes on participants' devices, we collected responses that enabled us to study how encoding artifact visibility is influenced by changes to the screen size. We collected ratings on 1,674 distorted videos from 14,450 participants from four countries and found the data to be reliable and largely consistent. With these data, we evaluated state-of-the-art subjective modeling techniques and benchmarked several objective quality models across screen sizes. Christos G. Bampis, Lukas Krasula, Zhi Li 0001, Omair Akhtar |
QoMEX | 2 |
| 2023 | Predicting local distortions introduced by AV1 using Deep FeaturesabstractSemantics extracted by filters in deep learning networks correlate well with how human eyes perceive distortions. These methods (e.g., LPIPS, PieAPP, etc.) rely on the relative difference in activation between feature maps in pairs of references and distorted patches. However, Deep Feature extraction can be expensive to compute as a difference of latent code between reference and distorted frames. Therefore, it is challenging to integrate them into the decision process of modern video codecs like AV1, making thousands of encoding trials during exhaustive Rate-Distortion Optimization (RDO) searches. In this study, we present a method using deep features to predict the distortion perceived locally by human eyes in AV1-encoded videos. The prediction relies on Deep Features extracted from the reference frame only to weigh the Mean Squared Error (MSE) introduced during encoding. This approach will make integration into video codecs easier as a pre-processing step before starting encoding. We show the superiority of the proposed metric against other Reference-Only metrics on a dataset of local distortions in videos. We achieve comparable performance as state-of-the-art Full-Reference video quality metrics. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
VCIP | 2 |
| 2022 | Improving Maximum Likelihood Difference Scaling Method To Measure Inter Content ScaleabstractThe goal of most subjective studies is to place a set of stimuli on a perceptual scale. This is mostly done directly by rating, e.g. using single or double stimulus methodologies, or indirectly by ranking or pairwise comparison. All these methods estimate the perceptual magnitudes of the stimuli on a scale. However, procedures such as Maximum Likelihood Difference Scaling (MLDS) have shown that considering perceptual distances can bring benefits in terms of discriminatory power, observers’ cognitive load, and the number of trials required. One of the disadvantages of the MLDS method is that the perceptual scales obtained for stimuli created from different source content are generally not comparable. In this paper, we propose an extension of the MLDS method that ensures inter-content comparability of the results and shows its usefulness especially in the presence of observer errors. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
ICASSP | 2 |
| 2022 | Banding vs. Quality: perceptual impact and objective assessmentabstractStaircase-like contours introduced to a video by quantization in flat areas, commonly known as banding, have been a longstanding problem in both video processing and quality assessment communities. The fact that even a relatively small change of the original pixel values can result in a strong impact on perceived quality makes banding especially difficult to be detected by objective quality metrics. In this paper, we study how banding annoyance compares to more commonly studied scaling and compression artifacts with respect to the overall perceptual quality. We further propose a simple combination of VMAF and the recently developed banding index, CAMBI, into a banding-aware video quality metric showing improved correlation with overall perceived quality. Lukas Krasula, Zhi Li 0001, Christos G. Bampis, Mariana Afonso, Nil Fons Miret, Joel Sole |
ICIP | 1 |
| 2022 | On the Accuracy of Open Video Quality Metrics for Local Decision in AV1 Video CodecabstractVMAF is a popular objective quality metric used for video quality evaluation. The power of VMAF has been demonstrated for a wide variety of video scales and encoding processes. However, its ability to evaluate the quality of small video patches has not yet been tested, despite its importance for encoding algorithms. We applied Maximum Likelihood Difference Scaling (MLDS) methodology to estimate supra-threshold perceptual differences in localized sections in videos, also known as tubes, encoded using AV1. We further used the results to assess the performance of VMAF in this scenario and proposed a recalibration of the algorithm to improve its agreement with the subjective data. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
ICIP | 2 |
| 2021 | Enhancing VMAF through New Feature Integration and Model CombinationabstractVMAF is a machine learning based video quality assessment method, originally designed for streaming applications, which combines multiple quality metrics and video features through SVM regression. It offers higher correlation with subjective opinions compared to many conventional quality assessment methods. In this paper we propose enhancements to VMAF through the integration of new video features and alternative quality metrics (selected from a diverse pool) alongside multiple model combination. The proposed combination approach enables training on multiple databases with varying content and distortion characteristics. Our enhanced VMAF method has been evaluated on eight HD video databases, and consistently outperforms the original VMAF model (0.6.1) and other benchmark quality metrics, exhibiting higher correlation with subjective ground truth data. Fan Zhang 0017, Angeliki V. Katsenou, Christos G. Bampis, Lukas Krasula, Zhi Li 0001, David Bull 0001 |
PCS | 4 |
| 2021 | CAMBI: Contrast-aware Multiscale Banding IndexabstractBanding artifacts are artificially-introduced contours arising from the quantization of a smooth region in a video. Despite the advent of recent higher quality video systems with more efficient codecs, these artifacts remain conspicuous, especially on larger displays. In this work, a comprehensive subjective study is performed to understand the dependence of the banding visibility on encoding parameters and dithering. We subsequently develop a simple and intuitive no-reference banding index called CAMBI (Contrast-aware Multiscale Banding Index) which uses insights from Contrast Sensitivity Function in the Human Visual System to predict banding visibility. CAMBI correlates well with subjective perception of banding while using only a few visually-motivated hyperparameters. Pulkit Tandon, Mariana Afonso, Joel Sole, Lukas Krasula |
PCS | 4 |
| 2021 | Ambiguity of objective image quality metrics: A new methodology for performance evaluation
Manri Cheon, Toinon Vigier, Lukas Krasula, Junghyuk Lee, Patrick Le Callet, Jong-Seok Lee |
Signal Process. Image Commun. | 3 |
| 2020 | Training Objective Image and Video Quality Estimators Using Multiple DatabasesabstractMachine learning (ML) is an essential part of recent advances in computer science. To fully exploit its potential, ML-based algorithms require a considerable amount of annotated data to be used for training. This represents a severe limitation in the field of image and video quality assessment since obtaining large-scale annotated databases is time-consuming and expensive. Moreover, the resulting quality estimators are mainly restricted only to the usecases included in the dataset used for their training. This paper proposes a strategy allowing for combination of multiple databases for training of objective image and video quality assessment algorithms. Using this strategy, the algorithms can be trained using all of the existing relevant databases together which allows to increase the amount of data-points and usecases in orders of magnitude. The potential of the proposed method is demonstrated by re-training the combination of features from Video Multimethod Assessment Fusion (VMAF) algorithm resulting in the significant improvement of its performance with respect to 20 video databases. Lukas Krasula, Yoann Baveye, Patrick Le Callet |
IEEE Trans. Multim. | 1 |
| 2020 | FFTMI: Features Fusion for Natural Tone-Mapped Images Quality EvaluationabstractTone-mapping is a crucial step in the task towards displaying high dynamic range (HDR) images on standard displays. Given the number of possible ways to tone-map such images, development of an objective quality criterion, enabling selection of the most suitable tone-mapping operator (TMO) and setting its parameters in order to maximize the quality of the reproduction, is of high interest. In this paper, a new objective metric for natural tone-mapped images is proposed. It is based on a fusion of several perceptually relevant features that have been carefully selected using an appropriate feature selection procedure. The outcome of the selection also provides a valuable insight into the importance of particular perceptual aspects when judging the quality of tone-mapped HDR content. The performance of the resulting combination of features is thoroughly evaluated with respect to three publicly available databases and compared to several relevant state-of-the-art criteria. The proposed approach is shown to significantly outperform the tested metrics and can, therefore, be considered a competitive alternative for tone-mapped images evaluation. Lukas Krasula, Karel Fliegel, Patrick Le Callet |
IEEE Trans. Multim. | 1 |
| 2019 | AccAnn: A New Subjective Assessment Methodology for Measuring Acceptability and Annoyance of Quality of ExperienceabstractUser expectations have a crucial impact on the levels of quality of experience (QoE) that they consider acceptable or satisfying. Measuring acceptability and annoyance has mainly been performed in separate or multi-step experiments without any control over participants' expectations. This paper introduces a simple methodology to obtain the information about both of the entities in a single step and compares several data processing strategies useful for results interpretation. A specifically designed subjective experiment, conducted on compressed videos, has shown that the multi-step procedures could be replaced by our proposed single-step approach, regardless of the viewing conditions, while the novel approach is significantly preferred by observers for its low time requirements and higher intuitiveness. The test has simultaneously proven that user expectations can be altered by the instructions and it is, therefore, possible to simulate different user profiles regardless of the participants' real habits. The acceptability/annoyance experimental results are also used to benchmark the state-of-the-art objective video quality metrics in predicting acceptability/annoyance of QoE. A case study on the determination of the threshold of acceptability/annoyance for objective quality metrics is conducted, which can be served as a guideline for video streaming service providers. Jing Li 0026, Lukas Krasula, Yoann Baveye, Zhi Li 0001, Patrick Le Callet |
IEEE Trans. Multim. | 2 |
| 2018 | Quantifying the Influence of Devices on Quality of Experience for Video StreamingabstractThe Internet streaming is changing the way of watching videos for people. Traditional quality assessment on the cable/satellite broadcasting system mainly focused on the perceptual quality. Nowadays, this concept has been extended to Quality of Experience (QoE) which considers also the contextual factors, such as the environment, the display devices, etc. In this study, we focus on the influence of devices on QoE. A subjective experiment was conducted by using our proposed AccAnn methodology. The observers evaluated the QoE of the video sequences by considering their Acceptance and Annoyance. Two devices were used in this study, TV and Tablet. The experimental results showed that the device was a significant influence factor on QoE. In addition, we found that this influence varied with the QoE of the video sequences. To quantify this influence, the Eliminated-By-Aspects model was used. The results could be used for the training of a device-neutral objective QoE metric. For video streaming providers, the quantification results of the influence from devices could be used to optimize the selection of streaming content. On one hand it could satisfy the QoE expectations of the observers according to the used devices, on the other hand it could help to save the bitrates. Jing Li 0026, Lukas Krasula, Patrick Le Callet, Zhi Li 0001, Yoann Baveye |
PCS | 2 |
| 2018 | Data Analysis in Multimedia Quality Assessment: Revisiting the Statistical TestsabstractAssessment of multimedia quality relies heavily on subjective assessment, and is typically done by human subjects in the form of preferences or continuous ratings. Such data are crucial for analysis of different multimedia-processing algorithms as well as validation of objective (computational) methods for the said purpose. To that end, statistical testing provides a theoretical framework toward drawing meaningful inferences, and making well-grounded conclusions and recommendations. While parametric tests (such as t test, ANOVA, and error estimates like confidence intervals) are popular and widely used in the community, there appears to be a certain degree of confusion in the application of such tests. Specifically, the assumptions of normality and homogeneity of variance are often not well understood, leading to incorrect application and/or interpretation of the statistical test results. Therefore, the main goal of this paper is to present new guidelines toward proper use of statistical tests and, hence, fix some of the issues in multimedia quality assessment. The said guidelines are derived based on theoretical analysis of sampling distribution of test statistics, and consider practical aspects of data analysis in the said domain. Experimental results on both simulated and real data are presented to support the arguments made. Software that implements the said recommendations is also made publicly available, in order to help researchers and practitioners perform correct statistical comparison of models. Manish Narwaria, Lukas Krasula, Patrick Le Callet |
IEEE Trans. Multim. | 2 |
| 2017 | Quality Assessment of Sharpened Images: Challenges, Methodology, and Objective MetricsabstractMost of the effort in image quality assessment (QA) has been so far dedicated to the degradation of the image. However, there are also many algorithms in the image processing chain that can enhance the quality of an input image. These include procedures for contrast enhancement, deblurring, sharpening, up-sampling, denoising, transfer function compensation, etc. In this work, possible strategies for the quality assessment of sharpened images are investigated. This task is not trivial because the sharpening techniques can increase the perceived quality, as well as introduce artifacts leading to the quality drop (over-sharpening). Here, the framework specifically adapted for the quality assessment of sharpened images and objective metrics comparison in this context is introduced. However, the framework can be adopted in other quality assessment areas as well. The problem of selecting the correct procedure for subjective evaluation was addressed and a subjective test on blurred, sharpened, and over-sharpened images was performed in order to demonstrate the use of the framework. The obtained ground-truth data were used for testing the suitability of state-ofthe- art objective quality metrics for the assessment of sharpened images. The comparison was performed by novel procedure using ROC analyses which is found more appropriate for the task than standard methods. Furthermore, seven possible augmentations of the no-reference S3 metric adapted for sharpened images are proposed. The performance of the metric is significantly improved and also superior over the rest of the tested quality criteria with respect to the subjective data. Lukas Krasula, Patrick Le Callet, Karel Fliegel, Milos Klima |
IEEE Trans. Image Process. | 1 |
| 2016 | How to benchmark objective quality metrics from paired comparison data?abstractThe procedures commonly used to evaluate the performance of objective quality metrics rely on ground truth mean opinion scores and associated confidence intervals, which are usually obtained via direct scaling methods. However, indirect scaling methods, such as the paired comparison method, can also be used to collect ground truth preference scores. Indirect scaling methods have a higher discriminatory power and are gaining popularity, for example in crowdsourcing evaluations. In this paper, we present how the classification errors, an existing analysis tool, can also be used with subjective preference scores. Additionally, we propose a new analysis tool based on the receiver operating characteristic analysis. This tool can be used to further assess the performance of objective metrics based on ground truth preference scores. We provide a MATLAB script with an implementation of the proposed tools and we show one example of application of the proposed tools. Philippe Hanhart, Lukas Krasula, Patrick Le Callet, Touradj Ebrahimi |
QoMEX | 2 |
| 2016 | On the accuracy of objective image and video quality models: New methodology for performance evaluationabstractThere are several standard methods for evaluating the performance of models for objective quality assessment with respect to results of subjective tests. However, all of them suffer from one or more of the following drawbacks: They do not consider the uncertainty in the subjective scores, requiring the models to make certain decision where the correct behavior is not known. They are vulnerable to the quality range of the stimuli in the experiments. In order to compare the models, they require a mapping of predicted values to the subjective scores, thus not comparing the models exactly as they are used in the real scenarios. In this paper, new methodology for objective models performance evaluation is proposed. The method is based on determining the classification abilities of the models considering two scenarios inspired by the real applications. It does not suffer from the previously stated drawbacks and enables to easily evaluate the performance on the data from multiple subjective experiments. Moreover, techniques to determine statistical significance of the performance differences are suggested. The proposed framework is tested on several selected metrics and datasets, showing the ability to provide a complementary information about the models' behavior while being in parallel with other state-of-the-art methods. Lukas Krasula, Karel Fliegel, Patrick Le Callet, Milos Klima |
QoMEX | 1 |
| 2014 | Performance evaluation of the emerging JPEG XT image compression standardabstractThe upcoming JPEG XT is under development for High Dynamic Range (HDR) image compression. This standard encodes a Low Dynamic Range (LDR) version of the HDR image generated by a Tone-Mapping Operator (TMO) using the conventional JPEG coding as a base layer and encodes the extra HDR information in a residual layer. This paper studies the performance of the three profiles of JPEG XT (referred to as profiles A, B and C) using a test set of six HDR images. Four TMO techniques were used for the base layer image generation to assess the influence of the TMOs on the performance of JPEG XT profiles. Then, the HDR images were coded with different quality levels for the base layer and for the residual layer. The performance of each profile was evaluated using Signal to Noise Ratio (SNR), Feature SIMilarity Index (FSIM), Root Mean Square Error (RMSE), and CIEDE2000 color difference objective metrics. The evaluation results demonstrate that profiles A and B lead to similar saturation of quality at the higher bit rates, while profile C exhibits no saturation. Profiles B and C appear to be more dependent on TMOs used for the base layer compared to profile A. António M. G. Pinheiro, Karel Fliegel, Pavel Korshunov, Lukas Krasula, Marco V. Bernardo, Maria Pereira, Touradj Ebrahimi |
MMSP | 4 |