EDBT 2026 Demo / reviewers in the wild / expert
Andreas Pastor
dblp:252/8072 · also Andréas Pastor
· DBLP profile ↗
13ranked-venue papers
9as first author
11since 2021 · last 2024
0000-0002-6790-0009ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Comparison of Conditions for Omnidirectional Video with Spatial Audio in Terms of Subjective Quality and Impacts on Objective Metrics Resolving PowerabstractOmnidirectional media formats, particularly 360° videos with spatial audio, provide new immersive experiences and introduce a novel dimension to content consumption.We explore the relationship between subjective data quality and metric performance evaluation in the context of Omnidirectional videos with spatial audio. While methodologies for 360° video quality assessment have been standardized and well-documented, previous efforts primarily focus on video with limited audio conditions, e.g., mono/stereo rendering. Moreover, the experimental test setup and subjective test methodologies impact data quality and the ability to use these data for objective quality metrics performance evaluation. Such a problem is key in the industry and the standardization activities, as codecs and quality models must be compared. Hence, the requirements on the ground truth data quality have to be clarified to allow proper conclusions.In this paper, we compare two setups and three test methodologies to study how experiment discriminability changes with conditions and participant number. Then, we show how discriminability impacts the resolving power of quality metrics. We show that higher-performing metrics require higher-quality data to reveal their full potential. In doing so, we put into relation the experimental cost, data quality, and resolving power. Andreas Pastor, Pierre R. Lebreton, Toinon Vigier, Patrick Le Callet |
ICASSP | 1 |
| 2024 | "Discriminability-Experimental Cost" Tradeoff in Subjective Video Quality Assessment of Codec: DCR with EVP Rating Scale Versus ACR-HRabstractThis work uses naive observers to compare two subjective studies conducted in a controlled laboratory environment on SDR HD, UHD, and HDR UHD contents. These tests aim to compare the precision and accuracy of a modified Degradation Category Rating (DCR) and Absolute Category Rating with Hidden Reference (ACR-HR) subjective methods for video quality assessment. The modified version of the DCR method includes a repetition of both reference and distorted stimuli; and utilizes an 11-grade rating scale from Expert Viewing Protocol (EVP) of ITU-R BT.500-15 standards. In the second subjective protocol, ACR-HR operates without repetition and with the 5-grade quality scale from ITU standards. We extensively analyze the scale usage and compare Mean Opinion Score (MOS) discriminability in both subjective studies. We show that both methods can retrieve accurate MOS. However, the ACR-HR method achieves better discriminability among MOS than DCR with the EVP rating scale while reducing the experimental effort by a factor of two, i.e., the cost of the experiment. The findings of this work give new insight into how to perform cost-efficient subjective tests for video quality estimation with naive observers and how to retrieve good MOS estimates. Andreas Pastor, Ioannis Katsavounidis, Lukas Krasula, Andrey Norkin, Hassene Tmar, Patrick Le Callet |
PCS | 1 |
| 2023 | Recovering Quality Scores in Noisy Pairwise Subjective Experiments Using Negative Log-LikelihoodabstractTo gather larger datasets to train data-angry deep learning quality assessment models, crowdsourcing has become essential to recruit participants. These participants are asked their opinion by directly rating stimuli, e.g., using single or double stimulus methodologies, or indirectly by ranking stimuli or comparing distances as in the Maximum Likelihood Difference Scaling method. In crowdsourcing, participants’ behaviors and environmental distractions are not controlled. So, the researcher must pay attention to the answers’ reliability. Cleaning methods exist for direct annotation subjective methodologies. However, solutions for indirect annotation methods are limited. In this work, we propose a method based on the negative log-likelihood to detect spammers among participants from their answers. To demonstrate its use, we applied it in a quadruplet preference-based scenario. The proposed method requires low computation and can be integrated into active-sampling strategies, where annotations available per comparison are small. We demonstrate that our method is robust to various spammer behaviors and accurate by removing only spammers. It helps reduce the gap between data collected in in-lab conditions (i.e., no spammer) and through crowdsourcing: our method reduces estimated uncertainties around data-points by 50%, and RMSE between estimations from an in-lab experiment and the same experiment performed in crowdsourcing by 1.8. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
ICIP | 1 |
| 2023 | Towards Guidelines for Subjective Haptic Quality Assessment: A Case Study on Quality Assessment of Compressed Haptic SignalsabstractModern systems are multimodal (e.g., video, audio, smell), and haptic feedback provides the user with additional entertainment and sensory immersion. Standard recommendation groups extensively studied and focused on video and audio subjective quality assessment, especially in signal transmission. In that context, subjective quality assessment and Quality of Experience (QoE) of Haptic signals is at its infant age. We propose further analyzing the collected data from a recent subjective quality assessment campaign as part of the MPEG haptic standardization group. In particular, we are addressing the following questions: 1) How the emerging field of haptic signal QoE can benefit from existing efforts of video and audio quality assessment standards? 2) How to detect possible outliers or characterize the rater’s reliability? 3) How does the discriminability of haptic tests increases with the number of raters? Towards this goal, we question if traditional analysis as proposed for audio or video signal are suitable, as well as other state-of-the-art techniques. We also compare the discriminability of the haptics quality assessment tests with other modalities such as audio, video, and immersive content (360° contents). We propose recommendations on the number of raters required to meet the usual discriminability obtained for other perceptual modalities and how to process ratings to remove possible noise and biases. These results could feed future recommendations in standards such as BT500-14 or P.913 but for haptic signals. Andreas Pastor, Patrick Le Callet |
ICME | 1 |
| 2023 | Perceptual annotation of local distortions in videos: tools and datasetsabstractTo assess the quality of multimedia content, create datasets, and train objective quality metrics, one needs to collect subjective opinions from annotators. Different subjective methodologies exist, from direct rating with single or double stimuli to indirect rating with pairwise comparisons. Triplet and quadruplet-based comparisons are a type of indirect rating. From these comparisons and preferences on stimuli, we can place the assessed stimuli on a perceptual scale (e.g., from low to high quality). Maximum Likelihood Difference Scaling (MLDS) solver is one of these algorithms working with triplets and quadruplets. A participant is asked to compare intervals inside pairs of stimuli: (a,b) and (c,d), where a,b,c,d are stimuli forming a quadruplet. However, one limitation is that the perceptual scales retrieved from stimuli of different contents are usually not comparable. We previously offered a solution to measure the inter-content scale of multiple contents. This paper presents an open-source python implementation of the method and demonstrates its use on three datasets collected in an in-lab environment. We compared the accuracy and effectiveness of the method using pairwise, triplet, and quadruplet for intra-content annotations. The code is available here: https://github.com/andreaspastor/MLDS_inter_content_scaling. Andreas Pastor, Patrick Le Callet |
MMSys | 1 |
| 2023 | Predicting local distortions introduced by AV1 using Deep FeaturesabstractSemantics extracted by filters in deep learning networks correlate well with how human eyes perceive distortions. These methods (e.g., LPIPS, PieAPP, etc.) rely on the relative difference in activation between feature maps in pairs of references and distorted patches. However, Deep Feature extraction can be expensive to compute as a difference of latent code between reference and distorted frames. Therefore, it is challenging to integrate them into the decision process of modern video codecs like AV1, making thousands of encoding trials during exhaustive Rate-Distortion Optimization (RDO) searches. In this study, we present a method using deep features to predict the distortion perceived locally by human eyes in AV1-encoded videos. The prediction relies on Deep Features extracted from the reference frame only to weigh the Mean Squared Error (MSE) introduced during encoding. This approach will make integration into video codecs easier as a pre-processing step before starting encoding. We show the superiority of the proposed metric against other Reference-Only metrics on a dataset of local distortions in videos. We achieve comparable performance as state-of-the-art Full-Reference video quality metrics. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
VCIP | 1 |
| 2022 | Considering User Agreement in Learning to Predict the Aesthetic QualityabstractHow to robustly rank the aesthetic quality of given images has been a long-standing ill-posed topic. Such challenge stems mainly from the diverse subjective opinions of different observers about the varied types of content. There is a growing interest in estimating the user agreement by considering the standard deviation (σ) of the scores, instead of only predicting the mean aesthetic opinion score (µ). Nevertheless, when comparing a pair of contents, few studies consider how confident are we regarding the difference in the aesthetic scores. In this paper, we thus propose (1) a re-adapted multi-task attention network to predict both the mean opinion score and the standard deviation in an end-to-end manner; (2) a brand-new confidence interval ranking loss that encourages the model to focus on image-pairs that are less certain about the difference of their aesthetic scores. With such loss, the model is encouraged to learn the uncertainty of the content that is relevant to the diversity of observers’ opinions, i.e., user disagreement. Extensive experiments have demonstrated that the proposed multi-task aesthetic model achieves state-of-the-art performance on two different types of aesthetic datasets, i.e., AVA and TMGA. Suiyi Ling, Andreas Pastor, Junle Wang, Patrick Le Callet |
ICASSP | 2 |
| 2022 | Improving Maximum Likelihood Difference Scaling Method To Measure Inter Content ScaleabstractThe goal of most subjective studies is to place a set of stimuli on a perceptual scale. This is mostly done directly by rating, e.g. using single or double stimulus methodologies, or indirectly by ranking or pairwise comparison. All these methods estimate the perceptual magnitudes of the stimuli on a scale. However, procedures such as Maximum Likelihood Difference Scaling (MLDS) have shown that considering perceptual distances can bring benefits in terms of discriminatory power, observers’ cognitive load, and the number of trials required. One of the disadvantages of the MLDS method is that the perceptual scales obtained for stimuli created from different source content are generally not comparable. In this paper, we propose an extension of the MLDS method that ensures inter-content comparability of the results and shows its usefulness especially in the presence of observer errors. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
ICASSP | 1 |
| 2022 | On the Accuracy of Open Video Quality Metrics for Local Decision in AV1 Video CodecabstractVMAF is a popular objective quality metric used for video quality evaluation. The power of VMAF has been demonstrated for a wide variety of video scales and encoding processes. However, its ability to evaluate the quality of small video patches has not yet been tested, despite its importance for encoding algorithms. We applied Maximum Likelihood Difference Scaling (MLDS) methodology to estimate supra-threshold perceptual differences in localized sections in videos, also known as tubes, encoded using AV1. We further used the results to assess the performance of VMAF in this scenario and proposed a recalibration of the algorithm to improve its agreement with the subjective data. Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet |
ICIP | 1 |
| 2022 | Perception of video quality at a local spatio-temporal horizon: research proposalabstractThis paper contains the research proposal of Andréas Pastor that was presented at the MMSys 2022 doctoral symposium. Encoding video for streaming on Internet has become a major topic to reduce the consumption of bandwidth and latency. At the same time, the human perception of distortions has been explored in multiple research projects, especially for distortions generated by Coder-DECoder (CODEC) algorithms. These algorithms operate in a rate-distortion optimization paradigm to efficiently compress video content. This optimization can be driven by metrics that are most of the time not based on the human perception, and more importantly, not tuned to reflect the local perception of distortions by human eyes. Andreas Pastor, Patrick Le Callet |
MMSys | 1 |
| 2021 | Multi-Modal Aesthetic Assessment for Mobile Gaming ImageabstractWith the proliferation of various gaming technology, services, game styles, and platforms, multi-dimensional aesthetic assessment of the gaming contents is becoming more and more important for the gaming industry. Depending on the diverse needs of diversified game players, game designers, graphical developers, etc. in particular conditions, multi-modal aesthetic assessment is required to consider different aesthetic dimensions/perspectives. Since there are different underlying relationships between different aesthetic dimensions, e.g., between the ‘Colorfulness’ and ‘Color Harmony’, it could be advantageous to leverage effective information attached in multiple relevant dimensions. To this end, we solve this problem via multi-task learning. Our inclination is to seek and learn the correlations between different aesthetic relevant dimensions to further boost the generalization performance in predicting all the aesthetic dimensions. Therefore, the ‘bottleneck’ of obtaining good predictions with limited labeled data for one individual dimension could be unplugged by harnessing complementary sources of other dimensions, i.e., augment the training data indirectly by sharing training information across dimensions. According to experimental results, the proposed model outperforms state-of-the-art aesthetic metrics significantly in predicting four gaming aesthetic dimensions. Yejing Xie, Suiyi Ling, Andreas Pastor, Junle Wang, Junyu Dong, Patrick Le Callet |
MMSP | 4 |
| 2020 | Few-Shot Pill RecognitionabstractPill image recognition is vital for many personal/public health-care applications and should be robust to diverse unconstrained real-world conditions. Most existing pill recognition models are limited in tackling this challenging few-shot learning problem due to the insufficient instances per category. With limited training data, neural network-based models have limitations in discovering most discriminating features, or going deeper. Especially, existing models fail to handle the hard samples taken under less controlled imaging conditions. In this study, a new pill image database, namely CURE, is first developed with more varied imaging conditions and instances for each pill category. Secondly, a W2-net is proposed for better pill segmentation. Thirdly, a Multi-Stream (MS) deep network that captures task-related features along with a novel two-stage training methodology are proposed. Within the proposed framework, a Batch All strategy that considers all the samples is first employed for the sub-streams, and then a Batch Hard strategy that considers only the hard samples mined in the first stage is utilized for the fusion network. By doing so, complex samples that could not be represented by one type of feature could be focused and the model could be forced to exploit other domain-related information more effectively. Experiment results show that the proposed model outperforms state-of-the-art models on both the National Institute of Health (NIH) and our CURE database. Suiyi Ling, Andreas Pastor, Jing Li 0026, Zhaohui Che, Junle Wang, Patrick Le Callet |
CVPR | 2 |
| 2019 | Predicting the Torso Direction from HMD Movements for Walk-in-Place Navigation through Deep LearningabstractIn this paper, we propose to use the deep learning technique to estimate and predict the torso direction from the head movements alone. The prediction allows to implement the walk-in-place navigation interface without additional sensing of the torso direction, and thereby improves the convenience and usability. We created a small dataset and tested our idea by training an LSTM model and obtained a 3-class prediction rate of about 90%, a figure higher than using other conventional machine learning techniques. While preliminary, the results show the possible inter-dependence between the viewing and torso directions, and with richer dataset and more parameters, a more accurate level of prediction seems possible. Andreas Pastor, Jae-In Hwang, Gerard Jounghyun Kim |
VRST | 2 |